7,027 Matching Annotations
  1. Last 7 days
    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Thank you for the positive assessment and for recognizing the novelty of our findings, the complementarity of our experimental approaches, and the breadth of our multi-modal readouts. We will address the weaknesses raised below point by point.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      This is an insightful comment. We fully agree that early functional changes preceding structural degeneration represent an important and exciting avenue, and we will explicitly acknowledge this in the revised manuscript, including the relevant literature on early cholinergic hyperactivity. Examining earlier time points in our model to capture such changes is a compelling perspective that we will discuss as a key direction for future work.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      We acknowledge this limitation and will increase the sample size for the tracing experiments in the revised manuscript. However, we wish to clarify that the tracing was performed in a separate cohort of young animals specifically to characterize baseline MS cholinergic projections independently of any amyloid-related disturbances. While we agree that replicating these experiments in APP mice would be of great interest, this falls outside the scope of the current study and will be highlighted as an important direction for future work.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      We agree with this pertinent suggestion. We will perform and report region-specific analyses of the hippocampal formation in the revised manuscript, in line with our discussion of specific amyloid accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      We respectfully maintain our analytical approach. As the test phase was terminated upon reaching a predefined cumulative exploration time of 20 seconds rather than using a fixed trial duration, computing a discrimination index is not appropriate in this context, as total exploration time is constrained by design. We instead followed the validated protocol described by Leger et al. (2013, Nature Protocols; PMID: 24263092), which controls for inter-individual differences in exploratory motivation by fixing cumulative exploration time, ensuring equivalent sampling conditions across animals. We will clarify this methodological choice in the revised manuscript.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

      We appreciate this constructive suggestion. We will provide additional methodological detail on interictal spike detection, include representative examples, and perform a vigilance state-specific analysis in the revised manuscript. We agree this will provide valuable mechanistic insight and will update the reference accordingly.

      Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Thank you for this positive assessment and for recognizing the rigor of our experimental design, the value of our multi-modal approach, and the mechanistic significance of our cholinergic lesion comparison. We will address the weaknesses below point by point.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

      Regarding amyloid broadcasting, we fully acknowledge that the precise mechanism remains to be elucidated; while this was not a primary objective of the study, it represents a fascinating and unexpected finding that we will discuss more carefully as an open question for future investigation. Regarding the causal relationship between MS<sup>ChAT</sup> cell loss and the observed phenotypes, we agree that the specific circuits involved remain to be fully resolved; however, we would like to emphasize that the convergent evidence from our three complementary models (and in particular the recapitulation of cognitive and REM sleep deficits by selective cholinergic ablation) provides strong causal support for MS<sup>ChAT</sup> neuronal loss as a key driver of these phenotypes, independent of amyloid deposition per se.

      Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Thank you for acknowledging the compelling conceptual premise of our study, the thoughtful and broad experimental framework, and the value of our longitudinal sleep analysis and lesion comparison. We will address all concerns raised below point by point.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      This is a valid point. We agree that the term “prodromal phase” requires more careful framing in the context of animal models, and we will revise the title and relevant statements accordingly. We would like to note, however, that the REM sleep disturbances we report seem to emerge prior to overt cognitive decline in our longitudinal analysis, which we consider to reflect a prodromal-like feature of the model. Nevertheless, we will adopt more precise terminology throughout the manuscript to avoid conflation with clinical staging criteria.

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      We appreciate this comment and will carefully nuance our opening statement to better reflect the current state of evidence. However, we respectfully note that several reports, including Pase et al. (Neurology, 2017; PMID: 28835407) and Ibrahim et al. (Sleep, 2024; PMID: 38001022), have demonstrated that REM sleep loss is associated with increased risk of incident neurodegenerative disorders, particularly Alzheimer's disease, supporting the broader validity of our framing. We will revise the statement to more accurately capture the complexity of this relationship while preserving its scientific relevance.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      We agree and will revise this sentence to better reflect the multi-system complexity of sleep/wake regulation, acknowledging the interplay between multiple neurotransmitters and neuromodulatory systems.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      We agree and will revise this sentence to clarify that ACh is a crucial component of a broader REM sleep-generating circuitry, rather than a sole requirement, while better contextualizing its specific contribution to REM sleep regulation.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      We will revise this statement to more accurately reflect the evidence. We would like to note, however, that while the Pase et al. (2017) cohort included a mixed dementia population, 75% of incident dementia cases (24 out of 32) were consistent with Alzheimer's disease, lending meaningful support to the relevance of REM sleep alterations specifically in the context of AD. We will ensure this nuance is clearly conveyed in the revised manuscript.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      We agree and will revise this section to acknowledge the well-established contribution of tau pathology to BF cholinergic neuronal vulnerability in human AD, including its early accumulation at Braak stages I-II. We wish to clarify, however, that the present study focuses exclusively on amyloid-driven mechanisms, and the introduction will be revised to ensure readers do not infer a purely amyloid-dependent mechanism in the broader human disease context.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      As noted in our response to weakness (1), we will systematically revise the manuscript to replace “prodromal phase” with more precise terminology that clearly distinguishes our animal model findings from clinical disease staging.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      Thank you for raising these important points. Regarding age matching, we acknowledge that chronological ages are not perfectly aligned across groups; however, we wish to emphasize that the duration of amyloid pathology is carefully matched across models. Indeed, AAV injection in MS<sup>ChAT</sup>-AppNL-G-F mice marks the onset of amyloid expression, directly paralleling the onset of pathology from birth in App<sup>NL-G-F/NL-G-F</sup> knock-in mice. We believe pathology duration represents the most biologically relevant variable for comparison in this context, and we will clarify this in the revised manuscript. Regarding the MS<sup>ΔChAT</sup> cohort, animals were culled upon reaching a comparable degree of REM sleep loss, providing a functionally meaningful matching criterion. Finally, regarding sham-operated controls, we acknowledge this limitation; however, based on our experience, surgical procedure alone has negligible effects on the cellular populations under study, and the inclusion of an additional sham group across all experimental cohorts would have required a prohibitive number of animals, raising significant ethical concerns under the 3R principles. We will address these points more explicitly in the revised manuscript.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      Thank you for this observation, we will correct the figure citation accordingly in the revised manuscript.

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      We appreciate this important observation. We will carefully re-examine our tracing data to determine whether MS<sup>ChAT</sup> axons specifically innervate the LHA, and whether amyloid deposition was detectable in this region in our MS<sup>ChAT</sup>-App<sup>NL-G-F</sup> model. We agree that clarifying the potential anatomical relationship between MS<sup>ChAT</sup> projections and LHA-MCH neurons is important to properly position our proposed circuit mechanism relative to this well-established REM sleep-promoting node, and we will address this in the revised manuscript.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      Thank you for catching this inconsistency. We confirm that this is an error in the manuscript: line 125 should read “15-16 month-old” referring to the chronological age of the animals at the time of perfusion, rather than “13-14 months,” which corresponds to the duration of AAV expression. We will correct this in the revised manuscript.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      Thank you for flagging this ambiguity. We will revise the sentence to clearly distinguish between AAV-injected and global knock-in cohorts in the revised manuscript.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      The supplementary information referenced in the main text will be provided in full in the revised manuscript, together with analyzed datasets and analysis scripts, in accordance with eLife's data sharing policy.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      We agree that a comprehensive description of our histological quantification pipeline is essential. We will provide full methodological details in the revised manuscript, including: the metric used for Aβ load quantification, ROI definitions, Fiji/ImageJ thresholding algorithms (accounting for batch effects and autofluorescence normalization), as well as the strategy for identifying and counting ChAT-positive neurons, section spacing, number of sections per animal, hemisphere coverage, normalization, and blinding procedures. All ImageJ scripts will be made available to ensure full transparency and reproducibility.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ<sup>-</sup> is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ<sup>-</sup> and AAV‑Aβ<sup>+</sup> is marked by asterisks (***) placed above AAV‑Aβ<sup>+</sup>. In addition, the single asterisk above KI-Aβ<sup>+</sup> does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ<sup>-</sup> vs WT‑Aβ<sup>-</sup> or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      We acknowledge that the current statistical annotations can be difficult to interpret when multiple experimental groups are displayed within a single plot. While the notation is consistent across figures, we agree that clarity can be improved, and we will revise all figure annotations to explicitly link significance indicators to their corresponding pairwise comparisons, using standardized brackets throughout, with all contrasts clearly specified in the figure legends.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ<sup>+</sup> mice lack a preference for the novel object compared with AAV‑Aβ<sup>-</sup> controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ<sup>-</sup>, and the notation over AAV‑Aβ<sup>+</sup> is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ<sup>+</sup>, KI-Aβ<sup>++</sup>, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      We wish to clarify that in Figure 5D, the absence of novel object preference manifests as equal exploration of both objects (approximately 10 seconds each), such that comparisons against the 10-second chance level reflect this lack of preference. We will revise the annotations and figure legend to make the statistical comparisons explicit and fully consistent with the results text.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      This is a constructive feedback, and we agree that the Discussion would benefit from pruning and recalibration. We will streamline it to focus on the most evidentially supported points, particularly those linking cholinergic loss to REM sleep and cognitive outcomes, while toning down or removing speculative statements regarding disease-stage equivalence, prion-like spreading mechanisms, and direct therapeutic implications.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      Thank you for raising this important point. We will expand the Discussion to address this apparent contradiction with the human literature. Indeed, while increased anxiety is reported in early AD stages, it tends to normalize or decrease at later stages (Botto et al., 2022; PMID: 35461471), which may partly reconcile our findings. In addition, anxiety-like phenotypes are highly inconsistent across AD mouse models, varying with model type, age, sex, and behavioral assay (Pentkowski et al., 2021; PMID: 33979573). Anxiety-related changes in human AD may reflect damage to brain regions beyond the MS cholinergic system, involving additional circuits and mechanisms not captured by our model.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

      We will nuance our wording at line 654 by replacing “comparable durations of amyloid pathology” with “comparable amyloid burden,” acknowledging that while both cohorts were aged for matched durations, the kinetics of amyloid progression may inherently differ between an AAV-driven focal model and a germline knock-in model.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      We thank the reviewer for raising this important point. The emergence of second-site mutations at residue D31 in Vero cells was actually observed in two independent experiments, each initiated by transfection of viral RNAs encoding either the N24A or N24D mutant. In the second experiment, additional timepoints were sampled for sequencing. Comparison of the two experiments revealed that, in the second replicate, only the D31N variant was recovered, whereas D31H was not detected. We will include the results of both independent experiments in the updated Figure 1 and will modify the text accordingly to improve clarity.

      In addition, we will present new experiments performed in other cell types that further confirm the reproducibility of D31 mutation emergence across different cellular contexts. Specifically, we will include results from new viral RNA transfections in BHK-21 mammalian cells and C6/36 mosquito cells, which will be included in the updated Supplementary Figure 1.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      We thank the reviewer for this correction. The sequencing was indeed focused on the region of the macrodomain surrounding the introduced mutations, covering the first 119 amino acids of nsP3. This region was selected to confirm the stability of the introduced mutations at position 24 and to monitor the emergence of potential second-site mutations in its immediate vicinity. We acknowledge that this approach does not rule out the emergence of compensatory mutations elsewhere in the macrodomain or in the full-length nsP3 protein. We will correct the text accordingly to accurately reflect the region that was analyzed.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      This is an excellent point. To address it, we introduced N24A and N24D mutations into an infectious clone of the Indian Ocean strain, a representative of the ECSA lineage, and assessed the emergence of D31 second-site mutations during viral stock production. As observed with the Caribbean strain, the D31N secondary mutation consistently emerged in both N24A and N24D Indian Ocean mutant viruses. These results will be included in the updated Supplementary Figure 1 and in the main text.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      We thank the reviewer for raising this point. We acknowledge that previous studies have reported limited infection of A549 cells by certain CHIKV strains. However, in our hands, with our viral stocks and under our experimental conditions, we observe an increase in viral titers following infection of A549 cells with WT Caribbean strain virus, indicating productive viral replication. We also confirmed that the Indian Ocean strain replicates in A549 cells under the same conditions, and we will include growth curve data for this strain in the updated version of the manuscript. Importantly, we did not observe detectable cytopathic effects in A549 cells infected with either WT or N24 mutant viruses, which is consistent with the notion that CHIKV replicates less efficiently in this cell line compared to other mammalian cell lines. We will include a sentence acknowledging this in the discussion of the revised manuscript.

      Regarding the suggestion to use human fibroblasts, we respectfully note that the primary focus of this study is the role of macrodomain catalytic activity in the mosquito host, and the experiments in A549 cells were performed to confirm the known importance of the macrodomain in interferon-competent mammalian cells. We therefore consider the current data in A549 cells sufficient to support this conclusion within the scope of the manuscript.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      We acknowledge that the full viral genome was not sequenced, and we cannot rule out the presence of additional mutations elsewhere in the genome that could contribute to the observed phenotype. However, we note that the enhanced dissemination phenotype was also observed in independent experiments performed with the Indian Ocean strain mutant viruses in Ae. albopictus, which were generated independently from the Caribbean strain stocks. The consistency of the phenotype across two independently generated sets of mutant viruses from different CHIKV lineages strongly supports the conclusion that the enhanced dissemination phenotype is linked to the macrodomain mutations. These new data will be included in the updated manuscript.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

      We thank the reviewer for pointing this out. The text will be modified accordingly to accurately reflect that we assessed transmission potential, based on viral dissemination to heads, rather than actual transmission.

      Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      We agree that the mutants may indeed have stronger binding for modified substrates than the WT protein; however, given that DSF is not a quantitative measure of binding affinity and that free ADPr is not the relevant ligand (in fact we do not know the relevant ADPr-modified molecule), we have refrained from speculating further than saying in the discussion:

      “The progressive selection of D31H over D31N in the mosquito host further suggests that subtle differences in ADP-ribose substrate recognition may influence viral fitness in the mosquito environment in ways that are not yet understood.” (Lines 368-370)

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

      Regarding the title, we agree with the reviewer that it could be misleading, as both mutants lack catalytic activity yet show distinct phenotypes in mosquitoes. We will therefore modify the title to: "Loss of macrodomain catalytic activity modulates Chikungunya virus dissemination and transmission potential in Aedes mosquitoes", which more accurately reflects that the observed phenotypes arise as a consequence of the loss of catalytic activity at position N24.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      We thank the reviewer for this important comment. Regarding the generation of viral stocks in A549 cells, transfection of N24A and N24D viral RNAs into A549 cells did not yield sufficient viral titers to produce usable stocks. However, as described in our response to Reviewer #1, we confirmed the reproducible emergence of D31 second-site mutations in C6/36 mosquito cells and BHK-21 mammalian cells, which will be included in the updated Supplementary Figure 1. These results demonstrate that D31 mutations spontaneously emerge in stocks generated directly in mosquito cells, addressing the reviewer's concern.

      Regarding the question about insect-specific alphaviruses, we examined the sequence at position 31 using a multiple sequence alignment of 14 alphaviruses with diverse host ranges, including dual-host, insect-specific, and aquatic alphaviruses. We observed that Yada Yada virus (GenBank: QGR15362.1) has an asparagine (N), Tai Forest alphavirus (GenBank: YP_009333615) has an aspartic acid (D), Mwinilunga alphavirus (GenBank: BBC45634.1) has an aspartic acid (D), Eilat virus (GenBank: QBG67155.1) has an aspartic acid (D), and Agua Salud alphavirus (GenBank: QEV83787.1) has a lysine (K) at this position. These results suggest that insect-specific alphaviruses do not share a conserved residue at position 31, and therefore no clear conclusion can be drawn regarding a direct correspondence with the compensatory mutations observed in our study. This sequence alignment with the corresponding text will be included as supplementary data in the revised manuscript.

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      We thank the reviewer for this observation. We acknowledge that the prevalence of WT virus infection shows variability between experiments and timepoints. Based on our experience with infectious blood meal experiments, this variability is sometimes observed between independent experiments even under identical experimental conditions and with the same virus. In this particular case, each set of experiments was performed with independently produced viral stocks. Specifically, for the experiments shown in Figure 3, WT, N24A-D31N, and N24D-D31H/N viral stocks were produced in parallel from transfection of viral RNAs, and the same stocks were used for sequencing, growth curves, and mosquito infections. Subsequently, when we generated single D31H and D31N mutant viruses, a new WT viral stock was produced in parallel with the D31 mutant stocks, and these independently produced stocks were used for the experiments shown in Supplementary Figure 2. Differences in absolute infection rates between experiments are therefore expected, as they reflect both the use of independently produced viral stocks and the inherent variability in the efficiency of midgut infection and dissemination between mosquito cohorts. Importantly, the comparisons between WT and mutant viruses are always made within the same experiment, using stocks produced in parallel.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    1. Author response:

      Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      We appreciate the reviewer’s thoughtful and comprehensive summary of our study.  One thing we would like to clarify is that the non-critical targets used in Experiment 1 received the same type of online continuous feedback as the critical target.Therefore the extent of implicit recalibration should be comparable between the critical and non-critical targets in Experiment 1. In contrast, for Experiment 2, we tightened the control for implicit recalibration at the non-critical targets by: 1) delivering delayed endpoint feedback for all non-critical targets while keeping the online continuous feedback for the critical target, which is known to suppress implicit recalibration, and 2) we widened the spatial gap between the critical and non-critical targets, to minimize spillover 2) As a result, the implicit recalibration was diminished at the non-critical targets, contributing to the shrinkage of the generalization curve around the critical target. 

      One other issue that we would like to make clear is that Experiment 1 was not a straight replication of a previous study, at least to our knowledge. We believe that the reviewer is referring to our previous study (McDougle and Taylor 2019), which found broader generalization for algorithmic strategies (Experiment 4). However, in that study cursor feedback was always delayed. As such, the observed broader generalization was most likely due to the strategy itself and not implicit recalibration. 

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      We are glad to know that our primary finding and conclusion appears reasonable. While we acknowledge that the finding doesn’t appear to be particularly exciting at face value, it does speak to larger questions regarding the independence of different learning systems and how just statistical or surface-level differences in training can result in relatively large differences in apparent behavior that could be easily misinterpreted as the result of system interactions.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      This is a fair concern, as RT is an indirect marker of strategy use and, by itself, cannot establish that participants used different strategies at the Critical target. Prior work, however, provided guidance for the design of our experimental manipulations. Algorithmic strategies, in which an aiming solution is computed online, are associated with longer RTs, whereas retrieval of a previously cached stimulus–response association produces substantially shorter RTs (McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Moreover, caching becomes increasingly difficult as the number of target-specific solutions increases, particularly beyond approximately four targets (Velazquez-Vargas and Taylor, 2024; Bejjanki and Taylor 2026). Our manipulation was designed around these findings: participants in the Algorithmic condition learned the 45° rotation across 10 targets, whereas participants in the Retrieval condition repeatedly encountered the 45° rotation only at the Critical target. As expected with this experimental design, RTs at the Critical target were significantly longer in the Algorithmic condition across all three experiments.

      We agree with the reviewer, however, that this RT difference could in principle reflect a more general carryover of cognitive demands from the Non-Critical targets rather than online computation at the Critical target itself. We can address this possibility more directly by asking whether RT at the Critical target exhibits the parametric signature expected of an algorithmic process. A defining feature of mental rotation is that RT scales with the magnitude of the computed aiming solution (Georgopoulos and Massey, 1987; Bhat and Sanes, 1998; McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Although rotation magnitude was fixed in the present experiments, participants’ actual reach angles varied naturally from trial to trial. Indeed, McDougle and Taylor (2019) originally demonstrated this relationship using actual reach angle rather than imposed rotation magnitude. We can therefore test whether trial-by-trial RT covaries with reach angle at the Critical target in the Algorithmic condition but not in the Retrieval condition. Such a relationship would be difficult to explain as a nonspecific carryover of cognitive load and would instead provide direct evidence that preparation time at the Critical target reflects the computation of the aiming solution.

      We also observe a second, independent difference at the Critical target: reach angles are consistently more variable in the Algorithmic condition than in the Retrieval condition across all three experiments (Figure S6). This pattern is consistent with repeated online computation producing variability in the selected aiming solution, whereas retrieval of a cached stimulus–response association produces a more stable response. Importantly, a general carryover account based solely on increased cognitive demands does not readily explain why movements to the Critical target should also be systematically more variable. Nor is this pattern easily explained by a speed–accuracy tradeoff: the Algorithmic group had more preparation time yet nevertheless exhibited greater variability. Consistent with this notion, Velázquez-Vargas and Taylor (2024) found that retrieving cached solutions produced less variable and more precise movements than movements that are not cached in the memory trace. 

      Third, we can conduct additional analysis to compare RT variability between algorithmic and retrieval groups at the critical target location. According to Logan instance theory (Logan, 1988), the retrieval of cached stimulus-response associations produces a stable RT profile, in contrast, trial-by-trial algorithmic computation can result in more variable trial-by-trial RT differences. 

      While we agree that RT differences alone should not be taken as definitive evidence of distinct strategies, our specific experimental design, the longer and more variable RTs at the Critical target, the greater trial-by-trial variability at that same target, and, if confirmed, a parametric relationship between RT and reach angle provide converging evidence that participants in the two conditions relied on different strategy implementations when preparing movements to the Critical target.

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

      We appreciate the reviewer’s thoughtful comment and agree that this is an important distinction. First, we would like to clarify the relationship among reaching variability, movement-plan variability, and the width of the implicit generalization function. Previous work has shown that when implicit recalibration occurs at a particular target location without an explicit aiming strategy, its generalization across the workspace can be described by a Gaussian-shaped function centered near the trained target location (Morehead et al., 2017). When an explicit aiming strategy is involved, however, implicit recalibration is centered closer to the planned aiming location rather than the visual target itself (McDougle et al., 2017). Thus, implicit recalibration is greatest near the direction in which the movement is planned, and a broader spatial distribution of movement plans can, in principle, produce a broader aggregate implicit generalization function. In the present study, we therefore use trial-to-trial variability in endpoint hand angle as a behavioral proxy for variability in planned movement direction, while recognizing that endpoint variability may also contain contributions from execution-related noise.

      This framework motivated the progression from Experiments 1 to 3. In Experiment 1, participants in the algorithmic condition exhibited substantially greater reaching variability, consistent with the idea that they sampled a wider range of movement plans across trials. Because error feedback associated with these different movement plans can induce implicit recalibration around each planned direction, greater variability in strategy use could contribute to the broader implicit generalization observed in the algorithmic group. In Experiment 2, we imposed stricter controls on spillover from noncritical targets, which reduced the overall breadth of generalization; nevertheless, model fits still suggested a modestly broader implicit generalization function in the algorithmic group, consistent with the remaining difference in reaching variability.

      Experiment 3 was designed to further reduce the direct influence of strategic variability on the induction of implicit recalibration by using a modified error-clamp paradigm. Error-clamp feedback has been shown to elicit implicit recalibration independently of task success and the participant’s explicit strategy (Morehead et al., 2017). We therefore used error-clamp feedback as an incidental signal to induce implicit recalibration while participants implemented either algorithmic or retrieval-based strategies. Importantly, however, Experiment 3 did not completely eliminate between-group differences in reaching variability: the algorithmic group continued to show greater variability around the critical 45-degree location than the retrieval group. We agree with the reviewer that, in principle, comparable generalization widths could arise if this broader distribution of movement plans were offset by a narrower local generalization of implicit recalibration in the algorithmic group. 

      However, not all variability is equivalent—or well described by a Gaussian distribution. Depending on the direction of the skew, variability in aiming can produce different effects on the implicit recalibration function, appearing as either broader generalization or greater amplitude. These effects are difficult to appreciate in Figures 3 and 5 for Experiments 2 and 3, respectively. We therefore sought to illustrate this more clearly in Figure 6, which shows how differences in the underlying aim distributions can shape the resulting implicit recalibration function in directionally complex ways. What complicates matters further is that implicit recalibration can asymptote (Morehead et al., 2017; Kim et al., 2018; Wilterson & Taylor, 2021). As a result, plan-based generalization can distort the implicit recalibration function in different ways depending on which side of the aim the error falls. The block-by-block analysis, suggested by reviewer 2, may shed light on this issue because we can get a sense if implicit recalibration has reached asymptote. 

      At a minimum, in the revised manuscript, we work to make clearer that subtle changes in the reach distribution may have a corresponding impact on the shape of implicit recalibration’s generalization function.  

      Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      We appreciate the reviewer raising this potential confound. We agree that the five pre-exposure trials in the Retrieval condition introduce a small difference in initial task familiarity. However, this account makes a straightforward prediction: if the shorter RTs in the Retrieval condition simply reflect five additional trials of general task experience, then the RT difference should disappear once the Algorithmic group has received a comparable amount of practice.

      We can test this directly by comparing the five pre-exposure trials in the Retrieval condition with the first five trials of the Algorithmic condition. We will also compare these pre-exposure trials with a later five-trial window from the Algorithmic condition to determine whether additional practice substantially reduces Algorithmic RTs. Assuming the observed pattern is as expected, RTs in the Algorithmic condition remain substantially longer even after considerably more than five trials of practice. Thus, the group difference cannot be explained simply by the Retrieval group having five additional trials of task familiarity. This persistent RT difference, together with our prior work showing characteristic RT differences between algorithmic computation and retrieval of cached aiming solutions, supports our interpretation that the groups relied on different strategy implementations.

      We are less certain what the reviewer means by the “early training advantage.” If this refers to angular error, we agree that the Retrieval group shows somewhat better performance very early in training, but this difference is not a central focus of the present study and largely disappears by the second or third training block. We will clarify the text so that we do not overinterpret this transient difference.

      If instead the reviewer is referring to RT, then the matched-trial analysis directly addresses the concern. Five additional familiarization trials could plausibly produce a brief initial advantage, but such an effect should dissipate within a small number of subsequent trials. In contrast, the RT difference between the Algorithmic and Retrieval conditions remains robust throughout training. We therefore do not think that general task familiarity provides a sufficient explanation for the observed RT differences.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      The reviewer raises an interesting possibility. In McDougle and Taylor (2019), however, the transition from algorithmic computation to retrieval occurred in a condition with only two targets in the task set. With repeated practice, participants needed to retain only two target-specific aiming solutions, making it feasible to replace online computation with retrieval of cached stimulus–response associations. By contrast, the Algorithmic condition in the present study contained 10 target locations. Our prior work suggests that caching becomes increasingly difficult once the number of target-specific solutions exceeds approximately four, at least over the timescale of several hundred trials (Velázquez-Vargas and Taylor, 2024; Bejjanki and Taylor, 2026). Thus, our task was designed to maintain pressure toward an algorithmic strategy throughout training.

      The reviewer nevertheless raises a more specific possibility that is not ruled out simply by the size of the target set: participants might selectively cache the aiming solution for the frequently sampled Critical target while continuing to use an algorithmic strategy at the remaining targets. We think the existing behavioral data argue against such a clear transition.

      First, reaction times at the Critical target in the Algorithmic condition remained substantially longer than those in the Retrieval condition throughout training. If participants had cached the aiming solution for the Critical target, we would expect preparation times at that location to be the same as the Retrieval condition. However, the RTs for the Algorithmic and Retrieval conditions are significantly different in the last block of training. 

      Second, within the Algorithmic condition, reaction times at the Critical target remained similar to those at the Non-Critical targets. Selective caching of the Critical target predicts a different pattern: preparation should become faster at the Critical target than at the surrounding locations, where participants would still need to compute the appropriate aiming solution. However, we do not observe a significant difference between RTs at Critical and Non-critical targets for the Algorithmic conditions at the end of the training block. Taken together, these two observations suggest that the Critical target continued to be treated similarly to the other members of the 10-target set rather than becoming a privileged, cached stimulus–response association.

      Third, participants would have to single out the Critical target as being distinct. All targets had the same visual appearance, the Critical target was never presented on consecutive trials, and participants were not informed that it had a special role in the experiment. However, we acknowledge that its higher sampling frequency, its somewhat greater separation from neighboring targets, and the location of the subsequent exclusion trials could nevertheless have made it more salient. Thus, we cannot rule out selective caching solely from the task structure.

      For this reason, we agree that the reviewer’s proposed analysis provides a useful additional test. If the Critical-target strategy progressively transitioned from algorithmic computation to retrieval, one prediction is that the generalization function in the Algorithmic condition should become narrower across successive exclusion blocks and increasingly resemble that of the Retrieval condition. We will therefore attempt to estimate the width of the generalization functions as a function of the training block between the Algorithmic and Retrieval conditions.

      There is, however, an important limitation to interpreting block-by-block generalization functions in this experiment. Implicit recalibration is both plan-based and temporally labile. Generalization is centered around the planned aiming direction (McDougle et al., 2017), and recent work indicates that implicit adaptation can decay over relatively short intervals (Zhou et al 2017; Hadjiosif et al 2023). Consequently, the first trial of an exclusion block provides the cleanest sample of the current state of implicit recalibration. Across later trials in the block, the measured response can be influenced both by temporal decay and by where the sampled target falls relative to the participant’s current aiming direction.

      To minimize systematic sampling bias, the starting exclusion target was randomized across participants. This means that these effects should average out at the group level, but individual exclusion blocks do not provide equally precise samples of the entire generalization function. A fully balanced estimate of every position within each exclusion block would require substantially more participants than were included in the present experiments. We will therefore present the blockwise analysis while interpreting changes in the estimated breadth cautiously.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback). Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

      The reviewer raises an important point. By design, the error-clamp manipulation in Experiment 3 decouples the participant’s planned movement from the visual consequence of that movement. While this gives us precise control over the error driving implicit recalibration, it could reduce agency over the cursor and thereby alter the interaction between explicit strategy and implicit learning in ways that are difficult to rule out completely. In particular, as the reviewer notes, reduced agency could potentially weaken either the influence of the explicit plan on implicit recalibration or the strategies themselves. Because Experiment 3 was intended to test for the absence of a strategy-dependent difference in implicit recalibration, we acknowledge that higher-order interactions of this kind represent an inherent limitation of our study if Experiment 3 is taken in isolation. 

      There are nevertheless several observations that make us think that reduced agency is unlikely to provide the primary explanation for the convergence between groups. First, the progression across Experiments 1–3 is informative. In Experiment 1, where participants retained normal control over the cursor, the broader generalization function in the Algorithmic condition closely mirrored the broader distribution of reach directions. This relationship suggests that the apparent difference in implicit generalization could arise from differences in where participants planned their movements rather than from a direct effect of strategy type on the implicit system. In Experiment 2, we sought to reduce the difference in the distribution of planned movements while preserving normal action–outcome contingencies and, importantly, the generalization functions became correspondingly more similar. Experiment 3 then controlled the error signal itself and again produced similar generalization across strategy conditions. Taken together, this progression favors the interpretation that strategy affects the measured generalization function indirectly, through differences in the distribution of movement plans, rather than directly altering the underlying implicit recalibration process.

      We nevertheless agree that these experiments cannot exclude all possible interactions between explicit and implicit learning systems. Indeed, whether these systems interact directly has been an important and persistent question in the sensorimotor adaptation literature. Several studies have reported evidence consistent with direct interactions (e.g., Albert et al., 2022; Maresch and Donchin 2021; t’Hart and Henriques 2024), whereas our own work has generally pointed toward indirect interactions mediated by factors such as movement planning and the current state of implicit adaptation (Taylor et al 2010; Taylor and Ivry 2011; McDougle et al 2017). Indeed, our recent study was designed specifically to distinguish these possibilities under tighter experimental control (Chen and Taylor, 2026), yet we found that the interaction between explicit and implicit processes is more complex than a simple independent-versus-interacting dichotomy. Going forward, we think it is more cautious to first rule out low-level statistical or distributional differences that could account for apparent effects before invoking higher-order interactions between learning systems.

      Finally, one motivation for the present study was that much of the literature on implicit generalization trains participants at a single target location before measuring generalization across the workspace. Under such conditions, participants have ample opportunity to retrieve a stable target-specific aiming solution. If algorithmic and retrieval strategies fundamentally alter implicit generalization, then many existing estimates of generalization may characterize implicit learning under retrieval-like conditions rather than providing a strategy-independent property of the implicit system. Across the present experiments, we find little evidence for such a fundamental difference once the distribution of movement plans and the experienced error are better controlled. We therefore think the most parsimonious interpretation of the current results is that algorithmic and retrieval strategies primarily influence implicit generalization indirectly through how movements are planned. 

      We plan to revise the manuscript to acknowledge that Experiment 3 cannot completely rule out higher-order effects associated with reduced agency under error-clamp feedback. We will also provide additional validation that participants were implementing distinct strategies, beyond the group-level RT differences, by testing whether RT in the Algorithmic condition scales with the instructed rotation magnitude, and whether RT variability shows group-level difference (Logan, 1988). If present, this relationship would provide stronger evidence that participants continued to engage the intended strategy under the clamp. We agree, however, that confirming distinct strategy use would not by itself rule out the possibility that reduced agency altered how those strategies interacted with implicit recalibration.

      Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      We thank the reviewer for this thoughtful and constructive assessment of the manuscript. We especially appreciate their recognition of the progressive experimental logic across the three experiments and of the broader theoretical motivation for distinguishing algorithmic and retrieval-based strategies. We are also grateful for the reviewer’s comments on the clarity and transparency of the work.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      We agree with the reviewer that our original interpretation of several non-significant effects was too strong, especially without providing some form of equivalence test. Our central hypothesis predicts little or no difference between algorithmic and retrieval-based strategies under conditions in which movement plans and error exposure are controlled, and we therefore face the inherent difficulty of drawing conclusions from an expected null effect. As the reviewer notes, a non-significant conventional hypothesis test does not by itself provide evidence that two conditions are equivalent.

      We therefore plan to supplement the existing analyses with quantitative assessments of the strength of evidence for the null/equivalence, using Bayesian factor analyses to confirm whether two conditions are equivalent. These analyses will allow us to distinguish between effects that are sufficiently small to support our theoretical interpretation and effects for which the data are simply inconclusive. We will also revise the manuscript throughout to avoid treating p > .05 as evidence of equivalence in the absence of such supporting analyses.

      We also now appreciate that the magnitude of implicit recalibration decreases progressively across experiments. This reduction could diminish our sensitivity to differences between the Algorithmic and Retrieval conditions and therefore represents an important qualification on the null result. Because the experiments used similar trial structures and were conducted with the same experimental equipment, the source of this reduction is not immediately clear. The block-by-block analysis suggested by Reviewer 2 may provide useful insight into how implicit recalibration evolves over the course of training and whether this attenuation emerges gradually within experiments.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Based on prior work characterizing implicit generalization in relative isolation from explicit strategy (Morehead et al., 2017; Poh and Taylor, 2019), we expected a relatively narrow generalization function, with a full width at half maximum of approximately 30°. We therefore expected probes spanning −45° to +45° around the Critical target to capture most of the function. At the same time, prior work on plan-based generalization predicts that the function should shift toward the participant’s aiming direction (Day et al., 2016; McDougle et al., 2017; Chen and Taylor, 2026), which complicates the choice of where to center the probes. Expanding the range and density of probe locations is also not cost-free, because additional exclusion trials increase temporal decay (Hajiosif et al., 2023) and begin to overlap with trained locations.

      We nevertheless agree with the reviewer that, when the fitted center approaches the edge of the sampled range, estimates of Gaussian width can become poorly constrained and may partly reflect parameter trade-offs or extrapolation beyond the observed data. We therefore plan to test whether the group differences persist when the Gaussian fits are constrained so that their centers fall within the sampled range. We will also examine complementary nonparametric measures of generalization breadth, such as the area between the group generalization curves across the sampled probe locations. Convergence across these approaches would provide stronger evidence that the reported differences reflect the observed shape of the generalization functions rather than instability in the Gaussian parameter estimates. If the results are not robust across approaches, we will revise the manuscript to qualify the interpretation of the fitted width estimates accordingly.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      This concern is closely related to that raised by Reviewer 2. We are fairly confident that participants in Experiment 3 were nevertheless engaging in the intended strategy manipulation. Participants in the Algorithmic condition showed substantially longer RTs than those in the Retrieval condition, and they were able to accurately generate the instructed angular reach directions across trials. The two conditions also differed in the variability of both RT and reach direction, consistent with online computation of an aiming solution in the Algorithmic condition and retrieval of a more stable cached response in the Retrieval condition. We can provide an additional validation by examining the relationship between RT and reach angle and comparing RT variability across two conditions. If RT scales parametrically with instructed reach angle in the Algorithmic condition but not in the Retrieval condition, this would provide stronger evidence that participants were engaging an online mental-rotation-like computation rather than simply following arbitrary spatial instructions. Moreover, RT yielded by the retrieval strategy would tend to be less variable than the algorithmic strategy.

      We agree, however, that this does not fully address the reviewer’s broader concern. Experiment 3 necessarily changed the context in which the strategy was implemented. In Experiments 1 and 2, mental rotation was used to counteract a visuomotor perturbation and was therefore directly tied to successful control of the cursor. In Experiment 3, the error clamp removed this instrumental relationship: participants still had to compute and execute different angular reach directions, but those computations no longer determined the visual cursor outcome. In that sense, the algorithmic process in Experiment 3 was less naturally embedded in the task and could reasonably be viewed as a somewhat different instantiation of the task.

      We therefore acknowledge that Experiment 3 cannot establish with certainty that the same higher-order cognitive process was engaged in exactly the same way as in Experiments 1 and 2, nor can it rule out the possibility that this change in task relevance altered how explicit strategy interacted with implicit recalibration. At the same time, when considered together with Experiments 1 and 2, we think the overall pattern remains informative. The apparent strategy-dependent difference in generalization was largest when movement plans and error exposure differed most, became smaller when these factors were better controlled while normal action–outcome contingencies were preserved, and was eliminated when the error signal itself was experimentally controlled. This progression is more consistent with an indirect influence of strategy through differences in movement planning and error exposure than with a robust direct effect of strategy type on implicit recalibration.

      Nonetheless, we agree that Experiment 3 should not be interpreted as a definitive test of whether algorithmic strategy, in its more relevant visuomotor adaptation context, can lead to different interactions with implicit recalibration compared to a retrieval strategy. We plan to revise the manuscript to make this limitation explicit.  

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      We agree that the RT difference in Experiment 3, by itself, cannot rule out the possibility that the Algorithmic condition imposed greater spatial precision demands because participants were reaching to invisible locations specified by angular instructions. We can, however, test for a more diagnostic signature of algorithmic computation by examining whether RT scales parametrically with the instructed reach angle. A general cost associated with reaching to invisible targets could increase overall RT, but it would not necessarily predict the characteristic increase in preparation time with the magnitude of the required angular transformation.

      We will therefore examine the relationship between RT and instructed reach angle in Experiment 3. If RT increases systematically with angular displacement in the Algorithmic condition, this would provide additional evidence that the longer RTs reflect online computation of the instructed aiming direction rather than simply the greater difficulty of reaching invisible targets.

      We can also compare this RT–angle relationship across experiments. If participants are engaging the same underlying algorithmic computation in Experiments 1–3, we would expect the slope relating RT to angular displacement to be similar across experiments, even if the overall intercept differs because of differences in task structure and spatial demands. A comparable slope would therefore provide converging evidence that the same computational process was engaged despite the altered task context in Experiment 3.

      We acknowledge, however, that similarity of the slopes would itself require yet another inference from a null difference and should therefore be interpreted cautiously. As with the generalization functions, we will use a Bayes factor analysis to quantify the strength of evidence for the null.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

      Based on the reviewers’ comments and the additional analyses they have suggested, we think we will be able to place our conclusions on a firmer empirical footing while also tightening and narrowing them. In the revised manuscript, we will strengthen the generalization analyses, use more appropriate inferential tools for interpreting null effects, and temper our broader claims about the insensitivity of implicit recalibration to strategy type. We will also more explicitly acknowledge the limitations of the present experiments, especially Experiment 3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high-frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high-frequency events during non-REM or in the hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from downregulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique: I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this comparison, it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs?

      Here, we provide additional analyses to demonstrate that the phasic, theta-modulated PFC activity that we observe during HFOs is specifically tied to the occurrence of HFOs and not a strong phenomenon during non-HFO-associated theta periods. In Figure S5 M (middle and right), we find that aligning PFC multiunit activity to theta periods in phasic REM but temporally distant from HFOs does not elicit the same degree of theta-modulated activity as aligning to HFOs (as in Figure 2A and Figure S5 M, left).

      Additionally, we provide analyses of the theta periods adjacent to HFOs (at different temporal thresholds) and demonstrate that this theta-modulated spiking activity is largely absent (Figure S5; compare to Figure 2A). Unlike Figure S5, this analysis was not restricted to putative phasic REM.

      We have now added Supplementary Figures S5L-N to the revised manuscript.

      What percentage of theta waves contain HFOs, and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      Although theta oscillations are continuously expressed during REM sleep, HFOs occur only intermittently, such that only a small subset of theta cycles contain HFOs. Across all animals and epochs included for analysis, we found that ~7.4% of theta cycles contain HFOs. We present an epoch-level quantification of this in Author response image 1, where the proportions were calculated across all, tonic, and phasic theta cycles. As expected, a higher proportion of putative phasic theta cycles contain HFOs.

      Regarding differential firing rate modulation, we refer reviewer to normalized MUA plots in the manuscript (Figures 2A and 4F). We would like the emphasize that what we show is PFC multiunit activity that is normalized by the mean population firing rate during REM sleep. Thus, these figures, specifically the chain HFO aligned figure, indicate that there are peaks in activity around baseline level in the background of an overall decrease in population activity relative to baseline (i.e. the troughs surrounding the peaks have lower activity compared to baseline, as in Figure 9C). While this suggests that there is an overall decrease in firing rates during HFOs as compared to baseline theta periods without HFOs, this simply provides a qualitative account of this difference. In Figure S5N, we present firing rate comparisons during HFO chains versus theta periods during putative phasic REM bouts at least 4 theta cycles away from HFOs. We find that HFO chain-associated PFC neuron firing rates are lower compared to non-HFO theta periods, supporting our finding of activity suppression during HFOs.

      Author response image 1.

      Proportion of theta cycles with HFOs. (A) Proportion of theta cycles with HFOs across all, tonic, and phasic cycles (***p = 4.90e-05, rank sum test).

      Lastly, we appreciate the reviewer's suggestion that REM HFO quantifications could be framed relative to phasic theta cycles without HFOs. We agree that such comparisons are informative and ensure that the results we present are specific to periods with detected HFOs. In response, we have added additional analyses of theta periods outside of HFOs in Figure S5. Furthermore, while we found that a larger proportion of chain HFOs occurred during bouts of putative phasic REM compared to isolated events (Figure S5 and Figure S3C), most of the analyses that we performed were on events pooled across putative tonic and phasic, since putative phasic REM is relatively scarce (<10%).

      We also refer the Reviewer to our response to Reviewer 2’s major comment #1 below where we reiterate several of our findings that demonstrate the HFO-specificity of the reported dynamics, as well as our extended response to Reviewer 3’s Public Review comment #1 where we show that the dynamics associated with REM HFOs are absent during HFOs detected during awake behavior (comparable theta state) on the W-Track (Figure S11). We hope that the additional control analyses we present as well as our expanded explanations adequately address the Reviewer’s concerns.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods, often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods, often for each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and what gross calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of the overall calculations being done for each panel.

      We have now split and updated the figures in a more logical progression of ideas:

      Figure 1: Prefrontal cortical HFOs in REM sleep using spectral analyses.

      Figure 2: Characteristic spiking modulation in PFC during REM HFOs.

      Figure 3: HFO and gamma distinction in theta cycles, and PFC-CA1 coherence (including chain/ isolated HFOs in phasic and tonic REM stages).

      Figure 4: Differential modulation of PFC spiking activity during REM PFC HFOs vs. NREM PFC ripples.

      Figure 5: Characteristic PFC population activity during REM HFO chains.

      Figure 6: Comparison of PFC reactivation during REM PFC HFOs vs. NREM PFC ripples.

      Figure 7: Differential engagement and excitability modulation of CA1 neurons by REM HFOs.

      Figure 8: REM theta phase shifting CA1 neurons preferentially engaged by REM HFOs.

      Figure 9: (Model) Network model with ACh reproduces spiking modulation during REM PFC HFOs vs NREM PFC ripples.

      Figure 10: (Model) Model reproduces restricted REM coactivity vs. widespread NREM coactivity.

      We have also rearranged the figures to parallel the main figures and results section.

      The authors mention in the discussion section that they see increased functional connectivity between mPFC and CA1, but most data suggesting this seems to be based on LFP rather than spiking. Functional connectivity is best defined by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions.

      We have updated the manuscript accordingly. Specifically, we have modified the text on Page 13, Lines 23-24:

      “These chains are associated with increased measures of oscillatory coupling between PFC and CA1…”

      Reviewer #1 (Recommendations for the authors):

      (1) Please ensure that analytical methods are presented in the same order in the methods section as they are in the results section - panel by panel. That said, the methods section is well-written and presented.

      We apologize for any confusion this may have caused. We have now reorganized the methods section to ensure they presented in the same order as the panels in the main figures.

      (2) Please specify whether the recordings/behaviors occur during the animal's light circadian phase.

      We have added a statement in the methods under the “Behavior” section on Page 34, Lines 9-12 indicating that the experiments took place during the light phase:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Please specify how many tetrodes are in mPFC and how many in CA1?

      We have added this clarification on Page 33, Lines 27-28:

      “Tetrodes were split equally between PFC and CA1 (15, 16, or 32 tetrodes in each region).”

      (4) All mentions of "coherence" should have frequency bands specified. "Theta coherence", for example.

      We have now added this information to all relevant mentions of “coherence”.

      (5) The intuitive logic of the phase slope index (2K) should be briefly explained for maybe half a sentence in the results section. The intuition should be explained better in the methods section devoted to it.

      We have added more clarification on the phase slope index method, both in the legend of Figure 3 on Page 21, Lines 22-23:

      “Phase slope index (PSI), which is a measure of phase lag consistency across different frequencies…” and in the methods section under “Phase slope index” on Page 39, Lines 21-26:

      “In practice, PSI is used to assess the consistency of phase lag relationships between two signals across different frequency bands and is a measure that is weighted by oscillatory coherence. We opted to use PSI to estimate the directional flow of information instead of other methods, such as Granger causality, since it has been demonstrated that PSI is less prone to false positives.”

      (6) "Cofiring" and "coactivity" should be defined clearly as measures - preferably in the results section if possible. They sound similar and are somewhat jargon-y without self-explanatory meaning (or difference from each other). How should readers understand and interpret them?

      We apologize for the confusion regarding these two terms, which are both used throughout the manuscript. Here, we specifically used to term “cofiring” to specify the explicit quantification of coincident activity between pairs of neurons (e.g. Figure 4E, left) as described in the Methods section under “Ripple and HFO co-firing”. There were instances where “coactivity” was used to refer to this quantification, and they have been changed to “cofiring”. We have now added a statement to the manuscript to clarify that “cofiring” is a quantification of coincident activity between neurons on Page 7, Lines 34-35:

      “Overall cortical co-firing, which is a measure of coincident activity between neuron pairs during discrete events”

      Additionally, the term “coactivity” in the manuscript is used when describing neuronal activity in the model or when there are mentions of coincident activity other than the explicit quantification described above.

      (7) The temporal threshold for "cofiring" should be stated in the results to enable interpretation of the results.

      We apologize for the lack of clarity regarding the cofiring metric. Here, the temporal threshold that we are imposing is determined by the ripple/HFO event times (see Methods section “LFP event detection”). If both neurons in a pair emit spikes within the defined window of an event, they are considered to have “cofired” (Cheng and Frank 2008, Singer and Frank 2009, Sosa, Joo et al. 2020).

      We have now clarified this in the results section on Page 7, Lines 34-35:

      “Overall cortical cofiring, which is a measure of coincident activity between neuron pairs during discreet events…”

      We have also added a clarifying statement in the Methods section under “Ripple and HFO cofiring” on Page 40, Lines 24-25:

      “Here, cofiring assesses coincident activity between neuron pairs within the start and end times of events, and thus, no explicit temporal threshold was implemented.”

      (8) Please clarify more systematically in which region the NREM ripples were detected. The natural assumption is the hippocampus, but at times it is mentioned that they are detected in the cortex. Are they cortical in all analyses? Readers could easily get confused about this and misinterpret "ripples". To clarify further, if these are always cortical events, I suggest renaming "ripples" in this text to "cortical NREM HFOs".

      We apologize for the confusion. For all mentions of “ripples” in the text, we are referring to PFC ripples specifically in NREM sleep. When hippocampal ripples are mentioned, we differentiate them by explicitly using “sharp-wave ripples” or “SWRs”. Our decision to use “ripples” for NREM sleep was based on previous studies that investigated NREM high-frequency events in cortex (Khodagholy, Gelinas et al. 2017, Helfrich, Lendner et al. 2019, Vaz, Inati et al. 2019, Aleman-Zapata, Morris et al. 2022, Ghosh, Yang et al. 2022, Shin and Jadhav 2024). Furthermore, we elected to use the “HFO” nomenclature for high-frequency events in REM sleep, since it has been used in previous studies to describe these events (Tort, Scheffer-Teixeira et al. 2013, Bueno-Junior, Ruckstuhl et al. 2023). Thus, to remain consistent with the literature, we decided to use these terms to describe these sleep-state-specific events throughout the manuscript:

      Hippocampal sharp-wave ripples (SWRs) in NREM sleep

      PFC ripples in NREM sleep

      PFC HFOs in REM sleep

      We have now added the following statement on Page 5, Lines 1-3:

      “However, to avoid ambiguity, and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout.”

      (9) The analysis performed for 4B is not explained clearly. It is somewhat better explained in the methods. I believe the reader should understand that each HFO is treated as an event, and spiking participation per unit was measured, and then the similarity of that pattern was assessed between HFOs. I also find the x-axis being quartiles makes understanding this graph particularly difficult.

      Why not label by raw lag and show a correlation plot rather than break down by quartiles? Alternatively, labeling the millisecond values of these quartiles on the x-axis labels may make comprehension much easier.

      We apologize for the lack of clarity regarding this method. We have added points to clarify the analytical procedure used for Figure 4B (Now Figure 5B) on Page 8, Lines 25-28:

      “We represented each HFO as a binary vector of PFC neurons active during the event, and computed the Pearson correlation between the vectors of every pair of consecutive HFOs” As well as in the Figure 5 legend on Page 24, Lines 7-10:

      “Here, the PFC spiking activity during each HFO was binarized across all neurons and the Pearson correlation coefficient was calculated between adjacent events as a measure of pattern similarity. Then the relationship between pattern similarity and IEI was reported.”

      In addition to the quartile labels, we have now added the average inter-event interval (IEI) of each quartile to the x-axis labels of Figure 5B to improve comprehension as well as a statement in the figure legend on Page 24, Line 11:

      “Below each quartile is the average IEI of that quartile in milliseconds.”

      (10) "Rank order analysis" should be again defined in the results in a manner that the reader can follow the point of the analysis and figure. For example, the goal is to assess the spike sequence across HFO events by looking at the regularity of spike timing rank for each neuron in each HFO event. Also, could this correlation be more simply calculated and presented as just a standard deviation around the mean of that cell's rank?

      We apologize for not including an explanation of this analysis in the results section. We have now added more detail about the rank order analysis to the Figure 5 legend on Page 24, Lines 18-21:

      “The mean rank-order correlation from the leave-one-out cross-validation procedure. Each event’s rank was correlated with the averaged rank across all other events. The average across all events compared to a distribution of means generated by jittering (n = 1000) spike times is shown.”

      We also provided a short description of the procedure in the results section on Page 8, Lines 32-36:

      “Furthermore, we examined whether the order in which individual PFC neurons fired during HFOs within a chain was preserved across chains. For each chain, we extracted each cell's first-spike rank order and compared it to a leave-one-out template constructed from the average normalized rank across all other chains.”

      Regarding the Reviewer’s second comment about the correlation, if we presented the result as the standard deviation around the mean of the cell’s rank, the result would be similar to the example rank order in Figure 5C, which we provided as a visualization of the rank-order template procedure we used. However, this would not necessarily demonstrate the consistency of population activity across HFO chains. To demonstrate consistency in activity across HFO chains, we used a leave-one-out cross-validated approach where each chain event was assessed separately. For each event, the neurons firing during that event was ranked and normalized 0-1. Then, the average rank of each neuron across all other events was calculated. Lastly, the correlation between the ranks of the left-out event and the average ranks of the template was taken to assess similarity in sequential activity. We opted to use this method since it has been demonstrated to be effective for evaluating similarities in sequential across events (Stark, Roux et al. 2015, Valero, Viney et al. 2021).

      (11) In Figure 5B and the related results section, I gather that "spatial" relates to place field location rather than anatomical/tetrode location of the neuron? This was not my original understanding and should be stated clearly.

      We apologize for the lack of clarity regarding this result. Yes, the term “spatial” refers to spatial rate map correlation between CA1 neuron pairs, specifically during the behavioral W-Track session prior to the sleep session where cofiring was assessed. We have updated the y-axis label of Figure 5B (Now Figure 7B) to include “rate map” and have updated the results section on Page 9, Lines 32-33:

      “…we observed a higher degree of spatial rate map correlation for high cofiring pairs…”

      (12) Figure 5E and its caption are almost totally unable to be explained since axes aren't explained well in either. It is only by inference from the results text that meaning can be assumed.

      We apologize for the lack of clarity regarding Figure 5E (Now Figure 7E). We have now updated the y-axis labels on both updated figures (Figures 7E and 8B) and have reworded and added more detail to the figure legend on Page 28, Lines 1-5:

      “High REM PFC HFO cofiring CA1 neurons exhibited a greater degree of suppression during NREM PFC ripples. For a description of modulation index, see Methods section Ripple/HFO aligned modulation. Here, since CA1 neurons exhibit a robust decrease in activity in response to NREM PFC ripples, we refer to the modulation as suppressive.”

      (13) Figure 5F is interesting, supporting the concept that HFOs "protect" neurons from downscaling. However, how can a "neuron" be cofiring? Would cofiring not be defined in a pairwise manner, and so each unit of the cofiring measure would be a pair of neurons? This question applies to other panels in this figure. Please clarify this.

      We apologize for the lack of clarity regarding the exact cofiring metric that we use. For Figures 7D-F and 8A, since we wanted to relate cofiring to changes in firing rate and modulation state during NREM PFC ripples, we calculated single cofiring values for each CA1 neurons by averaging across all pairings with PFC neurons. Thus, this metric gives us an estimate of the overall cofiring strength of each CA1 neuron. We have now added a new section in the Methods under “Calculation of a single cofiring metric and separation into populations of high and low cofiring neurons” on Page 44, Lines 34-43:

      “Since we wanted to relate the above cofiring metric to other measures, we needed to obtain a single-value cofiring metric for each neuron. To do this, we averaged the cofiring values across all pairings for a neuron (e.g. 1 CA1 neuron paired with all PFC neurons) and reported it as the cell’s cofiring. Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0 (High cofiring > 0; low cofiring < 0). Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions (High cofiring > mean; low cofiring < mean).”

      We have also clarified this in the Figure 7 legend on Page 27, Lines 25-28:

      “For comparisons between cofiring and other metrics (e.g. firing rate), a single cofiring value was calculated for each neuron by averaging the cofiring metric across all neuron pairings. Additionally, high and low cofiring CA1 neurons were split based on average cofiring values > 0 and < 0, respectively.”

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Weaknesses:

      The primary weakness of the study lies in the lack of a clear distinction between global brain states and the specific events being analyzed. Because the authors compare HFOs across different sleep stages (NREM, tonic REM, and phasic REM) without sufficient controls, it is difficult to determine if the observed differences are intrinsic to the HFOs themselves or simply a reflection of the different physiological states in which they occur.

      We would first like to note in case it was unclear – as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly spiking activity patterns observed in prefrontal cortical-hippocampal circuits during these event-types that occur in the two sleep stages. As noted in response to Reviewer 1’s Comment #8 and Reviewer 3’s Comment #1, to avoid ambiguity and to remain consistent with existing literature, we use the following terms to describe these sleep-state specific events throughout the manuscript: Hippocampal sharp-wave ripples (SWRs) in NREM sleep, PFC ripples in NREM sleep, and PFC HFOs in REM sleep.

      We refer the Reviewer to our Response to Reviewer 1’s first comment where we provide additional analyses of theta periods outside of HFO times (i.e. baseline REM periods), including Figure S5. We also further address these comments as a response to Reviewer 2’s major comment #1 below.

      Furthermore, the evidence for "structured reactivation" is not yet convincing. The temporal alignment of these reactivation events appears inconsistent, with peaks occurring well before the HFO itself, and the analysis does not sufficiently control for pre-existing cellular assembly strengths.

      We have now addressed this as a response to Reviewer 2’s major comments #3 and #8 below, including Figure 6.

      Additionally, some of the sleep architecture presented appears atypical, such as very short REM bouts and direct NREM-to-REM transitions that bypass standard progression, raising questions about the consistency of the sleep detection across animals.

      We have now addressed this as a response to Reviewer 2’s major comment #2 below, including Figure S1 and Figure S3. We expect that addition of these figures will mitigate concerns about our sleep staging procedures and will clarify certain points raised by the Reviewer.

      Finally, the study does not account for potential confounds like baseline firing rates when interpreting the behavior of "high-cofiring" neurons, which may simply be the most active cells in the population.

      We have now addressed this as a response to Reviewer 2’s major comment #6 below.

      Reviewer #2 (Recommendations for the authors):

      In this study, the authors detect activity periods during REM sleep that feature high-frequency oscillations in the prefrontal cortex. They term these REM HFOs and report that they occur in "chains" phase-locked to theta oscillations. They contrast these with HFOs that appear in isolation and with ripples observed during NREM sleep, either in the PFC or in the hippocampus. The data presented is very interesting in places. The authors show that these HFOs are observed primarily during "phasic REM" periods that have higher theta power. It appears that overall firing is substantially lower surrounding these events. Intriguingly, it appears that CA1 cells that co-fire with the PFC during HFOs increase their firing rates over the course of sleep, whereas other neurons show a firing decrease. This seems to be the most striking finding of this study.

      Major:

      (1) The findings are generally intriguing, but the study makes some choices that are hard to understand. They begin comparing HFOs that occur in chains to HFOs that occur in isolation, even though it appears that chained and isolated HFOs likely occur at different times, some during tonic REM, some during phasic REM, and others during NREM. The lack of control for sleep state makes it difficult to determine if the reported differences are intrinsic to the HFO patterns or merely reflections of the underlying global brain state. Overall, I just didn't quite understand the motivation for comparing isolated and chain HFOs, as it seems natural that chains would occur during periods of greater synchrony.

      We appreciate the Reviewer for raising these points and agree that the properties of ripples and HFOs cannot be interpreted independently of the underlying brain state. As we mentioned in the initial response, we do expect that the generation of these ripples/HFOs in NREM and REM sleep are inextricably linked to global brain state (ex., cholinergic tone, as shown in the model in Figures 9-10), which results in differing patterns of activity across sleep states. We noted in the response to the public review comment #1 above that as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly, spiking activity patterns observed in cortical-hippocampal circuits during these event types that occur in the two sleep stages.

      Sleep state: Our primary goal in comparing isolated and chain HFOs (chain vs. isolated HFOs are only compared in REM periods, not NREM periods) was not to suggest that the differences that we observe are state-independent, but rather to provide evidence that temporal clustering of HFOs underlies distinct PFC dynamics, as well as enhanced CA1 engagement. Rather than attempting to dissociate ripple/HFO occurrence and the differential physiology that underlie NREM and REM sleep states, we view them as complementary—both sleep states are permissive to generation of high-frequency oscillations. Similarly, putative tonic and phasic REM substates (low vs high theta power) are differentially permissive for isolated and chain HFOs (Figure S3C). We add additional detail on tonic vs. phasic REM substages in Supplementary Figure S3, showing that the rate of HFOs and HFO chains is significantly elevated in putative phasic REM. While we would have liked to analyze REM HFOs during tonic and phasic states separately to be able to make more concrete conclusions, the scarcity of phasic REM sleep made it difficult to make accurate comparisons between the two, especially for spiking modulation. We have therefore removed the qualifying term “phasic REM” from the abstract.

      Also, we would like to clarify that while we investigated chains of events (ripples) in NREM sleep as a comparison to REM HFOs, we do not refer to them as HFOs in NREM anywhere in the manuscript. In NREM, while there are PFC ripples that are clustered into chains based on our definition (separation of <200 ms), we do not observe a prominent peak in the IEI distribution that would suggest entrainment by other oscillations (e.g. theta or spindles). This analysis was only to show that associated results are specific to REM HFO chains, and not seen during comparable chains of NREM cortical ripple events.

      Chain vs isolated HFOs: Regarding the Reviewers comment about why we chose to compare isolated and chain HFOs, the motivation is clearly demonstrated by differences in these events in spectral properties (Figure 3H-I) and spiking modulation (Figure 4F-G). Our initial motivation to investigate these chains of events in REM sleep came from our observation that PFC population activity aligned to REM HFOs was theta-modulated (multipeaked, suggesting multiple high-frequency events over a short duration). This led us to hypothesize that there may be chaining of events in REM sleep, which in line with previous studies demonstrating the clustering of events, such as spindles in NREM sleep (Darevsky, Kim et al. 2024). In that study (Darevsky, Kim et al. 2024), the authors showed that reactivation of motor patterns was more persistent during trains of spindle events as compared to isolated spindles. Furthermore, trains/chains of multiple hippocampal SWRs have been shown to underlie the replay of extended experience (Davidson, Kloosterman et al. 2009). Thus, the separation of high-frequency events into isolated and chained events has precedence and may have functional significance. Furthermore, since we see that isolated and chain events have a bias for occurring during putative bouts of tonic and phasic REM sleep, respectively (Figure S3C), characterization of both event types is an important step for understanding potential differences in interregional interactions during REM sleep. Furthermore, a recent appreciation for the role of sleep stage sub-states (Chang, Tang et al. 2025) further emphasizes the importance of investigating these events separately. We expect future studies to further dissect the roles of tonic and phasic REM states in memory and cognition, and our finding provides an account of the existence of different events for future reference.

      HFO chains during periods of greater synchrony: While it may seem natural that chaining would occur during periods of high synchrony (theta synchrony here), we show that our result is HFO-specific, especially for spiking activity modulation (Figures 4F-H and Supplementary Figures S5K-N), which makes it novel and important. Previous studies have focused on theta-gamma cross-frequency phase amplitude coupling, which has been proposed as a mechanism where slower theta oscillations temporally organize faster local population activity, thereby synchronizing neural ensembles within and across brain regions (Belluscio, Mizuseki et al. 2012). Here, we present a comparison of HFOs and gamma events and show that while gamma events can also occur in chains (added to Supplementary Figure S4H), possibly due to increased synchrony as the Reviewer stated above, we do not observe theta-modulated population activity aligned to these events (Supplementary Figure S4J), which is a defining property of HFOs that we propose underlies our results. This indicates that our reported results are unique to periods of synchrony associated with HFOs.

      Control analyses for theta periods: Regarding the comment about lack of control for sleep state, we refer the Reviewer to our response to Reviewer 1’s comments. We have performed additional control analyses comparing PFC activity during REM theta periods adjacent to detected HFOs and demonstrate that phasic PFC population activity is largely absent (Figure S5).

      We are aware that it is difficult to dissociate the generation of ripples and HFOs from the underlying brain state, since they are so tightly linked. However, we do expect that our clarifying points in addition to the REM theta state controls provide strong evidence that the results that we present are intrinsic to HFOs and not general reflections of activity during baseline theta activity in REM. Here, we further reiterate a few of the main results that demonstrate this:

      (1) Phasic spiking modulation of PFC population activity associated with HFOs is not present during baseline theta periods (Supplementary Figure S5K).

      (2) Phasic spiking modulation during HFOs is not linked to extracted gamma events (Supplementary Figure S4J).

      (3) Phasic spiking modulation, as well as activity suppression, is strongest during chains of HFOs (Figure 4F, Supplementary Figure S4J).

      (4) PFC-CA1 theta coherence surrounding HFOs increases relative to baseline (Figure 3E, z-scored relative to baseline coherence).

      (5) Assembly peaks are sequentially organized surrounding HFO chains but not during isolated or shuffled (baseline) HFO times (Figure 6C, Supplementary Figure S7E, and Author response image 2).

      (6) A higher proportion of CA1-CA1 pairs are high cofiring during HFOs as compared to baseline periods (Figure 7B), indicating specific CA1 engagement during HFOs (in addition to coherence).

      Lastly, we show that HFOs detected during periods of active behavior (high theta) on the W-Track are not associated with many of the defining features of HFOs in REM sleep, thus demonstrating the specificity of REM HFOs despite similar background theta activity (Figure S11, in response to Reviewer 3 comment #1).

      (2) It is crucial that the study provides REM-specific sleep examples for each of the data sessions, marking tonic and phasic REM and indicating when isolated and chain HFOs are observed. It remains unclear how interspersed these events are. Do some REM episodes just have isolated HFOs and others chains?

      We thank the Reviewer for raising this point and apologize for the lack of clarity. For additional transparency, we now provide example sleep state plots for each animal (Supplementary Figure S1H-I).

      We also provide hypnograms that show bouts of putative phasic REM on top of REM periods (Supplementary Figure S3). Additionally, as a compact way of demonstrating the validity of our separation of putative tonic and phasic bouts, we provide a plot showing the average velocity and spectrogram surrounding putative phasic REM bouts across all putative phasic REM transitions (Supplementary Figure S3A). Regarding the incidence of HFOs in tonic and phasic REM, we refer the Reviewer to Supplementary Figure S3D, where we show that HFO rate is significantly higher during putative bouts of phasic REM, during which they tend to be organized in chains as compared to putative tonic bouts (Supplementary Figure S3E). In line with this, we additionally report that although both isolated and chain HFOs occur in putative tonic and phasic bouts of REM sleep, the proportion of chain HFOs during phasic REM is significantly higher than that of isolated HFOs (Figure S3), indicating a bias for chains to occur in phasic REM sleep, potentially due to stronger theta input. However, since putative phasic REM accounts for <10% of REM sleep in our dataset, consistent with previous reports using similar methods (Mizuseki, Diba et al. 2011), we were unable to restrict spiking analysis to HFOs in phasic REM. We found that chain HFOs during both putative tonic and phasic REM sleep elicited theta modulated population activity in PFC, thus we pooled HFOs across REM states for analysis.

      While a larger proportion of chain events occur in putative bouts of phasic REM sleep as compared to isolated events (Figure S3), chain and isolated events occur in both putative tonic and phasic substates (Proportion of events in putative tonic REM is (1 - proportion in phasic) shown in Figure S3C, right). Thus, the majority of analyses comparing isolated and chain HFOs, especially for spiking data, were pooled across states. Instead of showing example plots for all 22 sleep epochs (22 out of 36 epochs with >5 s of putative phasic REM), we expect that this analysis will be sufficient to illustrate the distributions of isolated and chain HFOs across putative REM sleep substates. We have also removed the qualifying term “phasic REM” from the abstract.

      To clarify this, we have added a statement to the new “Limitations” section of the manuscript on Page 16, Lines 10-20:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states. (Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult. Finally, although our model predicts that distinct cell-type activity profiles shape REM sleep HFO dynamics, we did not record a suDicient number of interneurons to test these predictions directly. Future studies using appropriate behavioral tasks, longitudinal sleep recordings, and cell-type specific opto-tagging will be able to resolve these limitations and further clarify the roles of high-frequency oscillations in REM sleep.”

      We have now added Supplementary Figure S3.

      This is also important because some of the REM sleeps detected and shown in Figure S1H seem unusual. For example, in Animal 1, there are multiple bursts of REM that seem very irregular. In some other sessions, animals occasionally appear to enter REM sleep with very little preceding NREM, which goes against existing literature. It's not clear which ones of these meet the 30 s duration threshold. Are the chain events occurring in these periods?

      In reference to the hypnograms shown in Supplementary Figure S1H (Now Supplementary Figure S1I), these were generated by concatenating all 9 sleep epochs regardless of whether they passed the inclusion criterion of > 30 s of total REM sleep. Furthermore, only a subset of the epochs shown were included for analysis based on a secondary, manual inspection that is performed to confirm inclusion. Sleep state plots (e.g. Supplementary Figure S1H) were visually inspected to further confirm the transition into REM sleep, ensuring absence of noisy T/D ratio or spurious detection due to noisy signals – epochs where microarousals or persistent subthreshold fluctuations in animal movement induced noisy TD ratio increases, and thus inaccurate REM designation, were excluded. We thus used a total of 36 sleep sessions from a possible 90 sleep sessions.

      We apologize for not specifying what portions of data represented by the hypnograms were included. We have now provided updated hypnograms only illustrating the sleep epochs included for analysis (Figure S1I).

      We have now added Supplementary Figure S1.

      Regarding the Reviewer’s point “Are the chain events occurring in these periods?”: Yes, all of the chain events (and isolated events) come from these updated epochs that are now shown. No events in the excluded epochs were included, since we wanted to only analyze the data that came from curated REM epochs that we were confident in.

      (3) The analysis supporting structured reactivation was not generally convincing. Figure 4E does not provide convincing evidence of this. Indeed, reactivation strength is lowest around the time of the event, and appears highest 0.5s before. The example panel seems rather anecdotal. It's also not clear why REM reactivations should be compared to NREM ones here. I could not follow what was done in Figures 4F-H. Why should the first half and second half of an HFO event be correlated?

      We appreciate the Reviewer for raising these concerns. We agree that this section is somewhat dense and at time hard to follow, so we will clarify with further explanations and analyses (see also response to Comment #8 with new reactivation figures).

      In Figure 6B (originally Figure 4E), our intention was to show that if you simply align assembly activation to REM HFOs, a relatively flat response is observed when averaged across all assemblies, in stark contrast to assembly reactivation during NREM cortical ripples. This could suggest a couple of things: 1) There is no real assembly activation in response to REM HFOs or 2) Assembly activation is organized differentially (compared to NREM) surrounding HFOs. What we hypothesized, and then quantified based on observations, is that assembly activity is sequentially organized around HFOs. If this were true, it would suggest a consistent temporal relationship between REM HFOs and assemblies (for example, assembly 1 tends to be active 115 ms after the onset of chains, assembly 2 is active 230 ms after, etc.). Of course, the sequential pattern that we show in Figure 6A may arise trivially, especially since these types of sequential visualization plots can simply arise from noise. Thus, a cross-validation method must be utilized to ensure the sequences are indeed reflective of an underlying computation.

      In the methods section “Assembly sequence detection surrounding HFOs” we explain the splitripple/HFO procedure that we used to compare assembly sequences across two halves of the data, similar to methods used in hippocampal place cell sequence cross-validation (Plitt and Giocomo 2021, Sosa, Plitt et al. 2025). For every split and assembly alignment, a sequence similar to Figure 6A, right is generated based on the first half of aligned data. Then the second half of aligned data is sorted based on the peak indices of the first half of data and the correlation between the peak reactivation indices across all assemblies for the two datasets is calculated. A high correlation indicates high sequence similarity across the two halves of data (not two halves of an HFO as the Reviewer mentioned), suggesting temporal consistency of assembly reactivation. By utilizing this method, we show that assembly reactivation sequences across randomly chosen halves of data are most similar for chained HFOs (Figure 6C and Supplementary Figures S7C-F). We have now provided an additional control analysis that investigates this sequential assembly activity for time-shifted chain events (Author response image 2). As in Figure 6, assembly activity was aligned to the first event in each time-shifted chain. In addition to the analyses in Supplementary Figure S7, this control further indicates that sequential assembly activity is preferentially restricted to HFO chains.

      Author response image 2.

      Assembly sequences surrounding time-shifted HFO chains. (A) Distributions of r values calculated from the Pearson correlation between peak reactivation bins across all assemblies for two randomly chosen halves of the shifted HFO-aligned data. There was no difference between r values for time-shifted chains and shuffled data, indicating no structured assembly activity during periods outside of real HFOs chains.

      To further quantify this, we calculated two additional metrics: 1) the slope difference between fitted lines for the two halves of data for each split and 2) the absolute peak difference between assemblies in the two halves of data. First, for the slope difference metric, the slope of the best-fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data in Figure 6D indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. Secondly, Figure 6E is a quantification of the peak reactivation displacement between the two halves of data. If there is a high probability of small peak differences, as in the real data, this indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data.

      We apologize for the omission of the description of the slope and peak difference metrics that we used in Figures 6D,E. We have added information in the Methods section under “Assembly sequence detection surrounding HFOs” on Page 43, Lines 23-33:

      “Furthermore, the slope and peak differences were calculated as additional metrics of sequence and temporal reactivation consistency between the two halves of data, respectively. For the slope difference metric, the slope of the best fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. For peak reactivation difference, the temporal displacement of the peak reactivation index between the two halves of data was calculated and compared to shuffle. A high probability of small peak differences indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data. Shuffling of assembly strength was carried out as above.”

      (4) The authors argue that the occurrence of HFOs, rather than theta power, is the reason for lower MUA activity, but the analysis for this (Figure S4I) is quite confusing. The left panel actually seems to indicate that MUA is indeed lower when theta power is high.

      We apologize for the confusion regarding this figure and appreciate the Reviewer’s point that periods of high theta power can also appear to be associated with reduced MUA. We agree that the original presentation may not have clearly separated the contributions of theta and HFOs. What we convey with Supplementary Figure S5K (originally Supplementary Figure S4I) and new Supplementary Figures S5L-M is that the observed theta-modulated PFC spiking response (Figure 2A) cannot be solely explained by baseline theta periods (outside of HFOs) in REM sleep. Since theta oscillations are ubiquitous during REM sleep, an important control is to demonstrate that the theta-modulated population activity is not simply a consequence of ongoing theta activity. Thus, we aligned PFC activity to theta oscillations of varied power and show that the fluctuating, theta-modulated population activity is absent, indicating that HFOs associated with the theta oscillation are driving this phasic response.

      We are not solely arguing that the presence of HFOs is the driver of decreased PFC multiunit activity. In both the data and model, we show that the magnitude of theta power detected in PFC (possibly input to PFC) is inversely related to multiunit activity (Figure 9E). Our interpretation is not that theta is unrelated to MUA, but rather that HFO occurrence provides additional explanatory power beyond theta alone. Accordingly, we show directly in Supplementary Figures S5L-M, in response to Reviewer 1’s comment #1, that theta periods in phasic REM not associated with HFOs do not elicit the MUA activity suppression similar to HFOs.

      The phase-alignment performed is hard to follow and is not being applied to HFO periods. If the study is trying to argue that high-theta periods without HFOs in the same recording sessions show lower MUA, then perhaps some sort of shuffle or jitter would be more suitable. For example, in Figure 1B, it seems there are some high-theta periods that don't have HFOs and appear to have higher MUA.

      We expect that the new analyses where we provide additional baseline theta controls for periods adjacent to HFOs in Supplementary Figures S5L-M now clarifies this point. Briefly, alignment to theta phases at different temporal distances from detected HFOs does not exhibit the same fluctuating PFC activity as HFO alignment.

      Regarding the phase alignment procedure that we used for the control analysis, since REM sleep is characterized almost entirely by ongoing theta activity, the control condition was not a separate brain state but rather theta periods outside of HFOs. We therefore needed a systematic way to select comparable theta cycles and phase bins in order to make a valid comparison with HFO-aligned PFC population activity. Since we demonstrated that there is significant phase amplitude coupling between theta and HFOs, thus a theta phase preference of HFOs, we used that specific phase bin across multiple theta cycles to align PFC activity. This phase bin selection was performed separately for each epoch to account for inter epoch and animal variability in phase preference. We reasoned that this procedure would allow for a valid comparison as compared to random alignment, since activity was aligned to similar phases in the baseline theta vs HFO conditions.

      (5) As far as I could tell, the study does not distinguish between putative excitatory and inhibitory neurons in the PFC, but only in the CA1, even though these play very different roles in the model. What is the rationale for not separating these? How are reactivations to be interpreted among interneurons?

      We apologize for the lack of emphasis on this point, which was originally highlighted in Supplementary Figure S2G (now in Supplementary Figure S2F), and for omitting the explanation as to why we did not separately analyze putative excitatory and inhibitory neurons. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with recording primarily from pyramidal neurons (Supplementary Figure S2F, left). We therefore decided to pool and not separate the populations into putative pyramidal cells and interneurons for the spiking analyses presented. To further validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO-aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript, as we did not have enough interneurons to investigate them separately. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      Regarding the interpretation of reactivation in the context of interneurons, we expect reactivation reflects coordinated ensemble activity with excitatory neurons encoding task-relevant information, and inhibitory interneurons shaping timing and neural synchronization. We however did not record enough distinct interneurons to test the predictions of the model, which is now noted in the Limitations on Page 16.

      Relatedly, we find that PFC assemblies detected from pooled data have task relevant representations (Figures 5F-G).

      (6) Are the high-cofiring CA1 neurons generally higher-firing than the other cells? Could this perhaps explain why they behave differently?

      We apologize that this information was not more evident in the manuscript, as it is an important control. In Supplementary Figure S9A, we show that there was no difference in baseline firing rate between low and high cofiring CA1 neurons.

      We have now explicitly referenced this figure in the main text on Page 9, Lines 42-45:

      “Analysis of low and high cofiring CA1 neurons during REM HFOs showed that high cofiring neurons exhibited elevated activity during chained events as compared to low cofiring neurons (Figure 7D), independent of baseline firing rates (Figure S9A).”

      (7) It appears that the decreased firing around HFO's could be a consequence of the stronger firing modulation around these periods, related to time averaging, rather than suppression per se. How does the firing rate compare to other periods with similar modulation that might not have HFOs?

      We thank the Reviewer for raising this important point. We expect that the new analyses, where we provide additional baseline non-HFO-associated theta periods, and theta periods adjacent to HFOs at different temporal distance as controls in also Supplementary Figure S5L-N now clarifies this point. These figures show that the decreased firing rate is specific to HFO chains (Supplementary Figure S5N). Indeed, if the suppression that we observe is related to time averaging, or another analytical artifact, our claims of suppression during HFOs would not be valid. We present a number of results and provide further explanations to support the accuracy of our characterization of PFC population suppression.

      First, event-aligned multiunit activity was quantified as baseline-normalized population firing relative to detected events. For both NREM and REM, activity was normalized by the mean population firing rate during a baseline period within the same sleep state in which events were detected. Values are therefore expressed as deviations from baseline. This normalization allows comparison of relative changes in firing around events within each state. Importantly, values below baseline reflect reductions relative to the state-matched baseline period.

      Second, we show that smoothing activity with a larger gaussian kernel preserves the dip in population activity, consistent with suppression of activity surrounding HFOs (Figure 9C, note that this is a data figure presented in the context of the model). However, as raised by the Reviewer, this normalized measure does not on its own distinguish sustained suppression from transient deviations introduced by event-locked temporal structure in firing.

      Third, to address this, we employed an alternative method to demonstrate that HFO chains tend to occur during periods of PFC suppression (Figure 4G). Briefly, we detected events in PFC where activity fell below a threshold and calculated the probability of HFOs surrounding these “suppression” events. We refer the Reviewer to the Methods section under “Detection of population suppression” where we explain this procedure in more detail. We found that chain ripples, during which the strongest suppression is observed (Figure 4F), are associated with decreases in PFC activity (Figure 4G and Supplementary Figure S6).

      Fourth, we refer the reviewer to Figure S5N, where we compare the firing rates of PFC neurons during HFO chains and theta periods outside of HFOs during putative phasic REM bouts. The observed reduction in firing during true HFO chains compared to non-HFO periods therefore reflects HFO event-specific activity suppression rather than a common occurrence during baseline REM periods.

      Lastly, we performed a control analysis complimentary to Supplementary Figure S5K where we aligned PFC activity to the preferred theta phase of HFOs and investigated how distance from detected HFOs modulates PFC activity (Figure S5). We found that there was no consistent theta-modulated activity aligned to HFO-adjacent theta phases.

      Regarding the final comment, if the Reviewer meant “modulation” as in the theta modulation or suppression observed when PFC population activity is aligned to HFOs (Figures 2A and 4F), we are not aware of any other REM periods where this strong theta modulation or suppression of PFC population activity is present. To our knowledge, we are the first to demonstrate such a modulation of PFC population activity in REM sleep. The closest comparison that we can make is PFC activity aligned to gamma events that are coordinated with HFOs (suppression of PFC), but this is explained only with association with HFOs, as shown in Supplementary Figure S4J.

      (8) The assembly reactivation measure does not control for pre-existing assemblies. The term "activation strength" would therefore be more appropriate.

      We thank the reviewer for this important methodological point. The concern that ICA-based reactivation strength does not, by itself, distinguish behavior-induced reactivation from pre-existing assembly activity/structure is well-taken, and we have implemented several complementary analyses that directly address these concerns.

      First, the interleaved structure of our recordings (8 run epochs interleaved with 9 sleep epochs) allows us to investigate the within-session pre/post assembly strength differences (i.e. each W-Track run session has a preceding (pre) and following (post) sleep session). An increase in assembly strength from pre to post is a hallmark of behaviorally relevant assembly reactivation (Kudrimoti, Barnes et al. 1999, Peyrache, Khamassi et al. 2009). For each run epoch, the same run-derived templates were projected onto the preceding and following sleep epochs, and the pre and post strengths were compared. The distribution of post-minus-pre reactivation differences across all epoch pairs is significantly skewed toward positive values (Figure 6F), indicating that templates from the run epochs are more strongly expressed in the post-sleep epochs. This asymmetry cannot be explained by pre-existing assembly structure, which would predict similar assembly strengths.

      Second, reactivation strength in post-experience sleep increases across the experiment, with templates from later running epochs producing the strongest reactivation in the following post-sleep (Figure 6G). This increase in reactivation strength over time cannot be explained by preexisting assembly structure, which predicts similar assembly expression strength independent of experience.

      Third, the detected assemblies carry behaviorally meaningful structure. Assembly activation maps computed during running exhibit spatially organized "assembly fields" similar to single-cell place fields (Figure 6H), demonstrating that the detected assemblies represent specific spatial locations or task variables rather than behavior-independent states. Pre-existing co-firing structure unrelated to ongoing experience would not be expected to produce spatially tuned, task-locked assembly activation. Furthermore, this spatial tuning of assemblies was verified by comparison with surrogates, where assembly maps were generated using circularly shuffled activation times (1000 shuffles). Assemblies with p < 0.05 (z > 1.65) were considered to have significant spatial structure (Figure 6I).

      These results establish that what we measure is the selective re-expression of behaviorally relevant assemblies in subsequent sleep epochs, consistent with the use of "reactivation" in the established literature (Peyrache, Khamassi et al. 2009, Lopes-dos-Santos, Ribeiro et al. 2013). We have therefore retained the term "reactivation strength" and have added text to the manuscript noting these new results that justify the use of “reactivation” strength.

      We have now added Figure 6.

      We have also added the procedure for the calculation of spatial information to the Methods section under “Spatial information of assembly fields” on Page 44, Lines 8-23.

      (9) Can the study rule out that the rank-ordering in Fig 4C is related to firing rates? Higher-firing rates tend to fire earlier, and lower-firing cells later.

      We thank the Reviewer for raising this interesting point. Here, we assume that the Reviewer meant the baseline firing rates of the neurons, not the intra-HFO firing rates of the neurons. Indeed, when we look at baseline REM firing rates of these PFC neurons, we do find that neurons with higher firing rates tend to fire earlier than low-firing-rate neurons (Author response image 3). This is also true when PFC rank and firing rate are assessed for isolated REM HFOs and NREM PFC ripples (Author response image 3). Similarly, we also observe this relationship in CA1 during SWRs, during which rank order correlation is typically assessed as a method for replay detection. In line with this, a previous study has shown that CA1 neurons with high excitability at animals’ current location tend to initiate replay events (Karlsson and Frank 2009). Furthermore, high-firing-rate, rigid CA1 neurons are more active during SWRs than low-rate, plastic neurons (Grosmark and Buzsaki 2016), and there are distinct populations of neurons in both hippocampus and PFC that are preferentially active during immobility in sleep epochs (Jarosiewicz, McNaughton et al. 2002, Kay, Sosa et al. 2016, Tang, Shin et al. 2017), potentially biasing replay activity during high-frequency events. Similar dynamics may underlie activity during PFC ripples and HFOs in NREM and REM sleep, respectively. The critical point here is in the leave-one-out cross-validation that we implemented to determine sequence similarity—each left out event’s cell rank was correlated with the averaged rank of the template that was generated from all other events. This analysis provides a basis for our claim that there is preserved sequential PFC activity across HFO chains. We did not observe neuron firing consistency during isolated HFOs or during pseudo-HFO chains (coherently shifted chain HFO times), which indicates that REM HFO chains are unique temporal windows during which PFC activity proceeds in a more structured manner.

      Author response image 3.

      Firing rate difference of low and high rank neurons (A) Comparison of baseline firing rates of PFC and CA1 neurons split by average rank across all PFC REM HFOs, PFC NREM ripples, or CA1 SWRs. Baseline rates were calculated separately for NREM and REM sleep.

      Minor:

      (1) It gets confusing that the authors sometimes (but not always) refer to HFOs during NREM as "ripples" but not if they occur during REM. The terminology is inconsistent. When they refer to HFO chains, it seems they now pool between REM and NREM periods, as well as across phasic and tonic REM periods, which is confusing.

      We apologize for the confusion regarding the terminology. In the revised manuscript, we now use NREM ripples exclusively for NREM events and REM HFOs exclusively for REM events. We have removed mixed labels such as “ripple/HFO” except where a collective term is explicitly defined. We also clarified that HFO chains refer to REM events only and revised the relevant text/figure legends to avoid any implication that chain analyses pool NREM and REM events.

      (2) P7 L8: It might be helpful to emphasize "broader temporal distribution".

      We thank the Reviewer for the suggestion. We have updated the text on Page 8, Line 18:

      “Since we observed a broader temporal distribution of activity…”

      (3) P8 L9: What do they mean by spatial? Do they mean the place-fields of these same neurons during a previous task period?

      We apologize for the confusion. The Reviewer is correct. Here, we calculated the spatial rate map correlation between CA1 neurons as a measure of place field similarity during the W-Track session prior to the sleep epoch being assessed.

      For clarification, we have added “rate map” to the text on Page 9, Line 33:

      “…spatial rate map correlation…”

      We have also updated the y-axis label for Figure 7B for clarity.

      (4) P8 L27: What do they mean by "coordinated SWRs"? As opposed to what?

      Here, we are referring to our previous study where we investigated ripples in NREM sleep and showed that ripples and SWRs in PFC and CA1, respectively could either be independent from or coordinated with events in the other region (Shin and Jadhav 2024). A main result in the study showed that CA1 neurons are strongly suppressed during independent PFC ripples and that there was a relationship between activity suppression and reactivation during coordinated SWRs (CA1 SWR-PFC ripple coordination in NREM). We specifically mentioned “coordinated” since these are SWRs that are also coupled with SOs and spindles as compared to SWRs that are independent from PFC ripples (Shin and Jadhav 2024). Overall, we wanted to frame this result in the context of oscillatory coupling and mechanisms of memory consolidation.

      (5) P37 L28 says "we observed a bimodal distribution" but L31 says "unimodal". Which is it?

      We apologize for the confusion. We observed a bimodal distribution for CA1-PFC cofiring in REM sleep, but a unimodal distribution in NREM sleep. Because of these two observations, we decided to split the CA1 population into high and low cofiring neurons based on two different thresholds:

      (1) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) 0.

      (2) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) the mean of the distribution of averaged cofiring values.

      Using two separate thresholds to split high and low cofiring CA1 neurons demonstrates the robustness of the firing rate change result in Figures 7F and Supplementary Figures S9B-D.

      We added a statement that clarifies that the bimodal distribution was seen in REM sleep only on Page 44, Lines 37-43:

      “Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0. Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions.”

      Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which, if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      (1) Distinction from wake HFOs

      The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent a more general theta-associated phenomenon.

      To investigate PFC high-frequency oscillations during running behavior on the W-Track, events were extracted in the same manner as NREM and REM events (Methods). Events during wake were subset by periods where the animals’ velocity was >4 cm/s to provide a comparison of events during periods of high theta. While we were able to detect HFOs during wake that exhibited a similar spectral profile in the high frequency band, we did not observe 1) strong association with gamma or theta oscillations, 2) prominent HFO chaining, 3) HFO aligned theta modulated PFC activity, 4) comparable levels of theta phase amplitude coupling, 5) association with population suppression, or 6) a relationship between peri-event theta power and multiunit activity (Supplementary Figure S11). Many of the defining features of PFC REM HFOs are absent during wake, indicating REM specificity of the results we present.

      (2) Link to memory consolidation

      The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      To address these concerns, we have now added an explicit “Limitations” section in the main text of the manuscript that includes a statement about learning. We have also added Supplementary Figure S1 in the manuscript, which illustrates the performance of all 10 animals on the W-Track task. Finally, we have also included Figures 6F G in the manuscript, showing that PFC assembly reactivation strength during sleep epochs increases during learning.

      Reviewer #3 (Recommendations for the authors):

      Most of my specific comments were related to further clarification that will help the reader's understanding.

      (1) I'd recommend simplifying terminology. Open to debate, but would it not be simpler and clearer to just say NREM HFO vs REM HFO? Obviously, there is a need to mention how NREM HFOs have previously been referred to as cortical ripples, but I'm not sure it is such a helpful terminology to continue for the field, given, as you state, how different cortical ripples are from hippocampal SWRs. If not, I'd at least provide a clearer explanation early in the manuscript, distinguishing NREM ripples from REM HFOs but collectively still calling them 'cortical events'.

      We appreciate the Reviewer’s suggestion regarding terminology and agree that it would be simpler and clearer to use a single term, ripple or HFO, to describe these events. Initially, we had used a unified term (ripples across both states) but ultimately decided to switch to state-specific terminology due to previous comments we received on the manuscript and to emphasize the distinctions between NREM and REM events. Ultimately, we decided on calling them ripples in NREM and HFOs in REM since there is precedence for both terms in each respective sleep state (Khodagholy, Gelinas et al. 2017, Vaz, Inati et al. 2019, Bueno-Junior, Ruckstuhl et al. 2023, Shin and Jadhav 2024), but we do agree that this distinction can be confusing if not clearly stated. Thus, we have added an additional statement in the manuscript on Page 4, Line 44 to Page 5, Lines 1-3 for clarity:

      “Similar criteria were used to detect cortical high-frequency events in NREM and REM states; however, to avoid ambiguity and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout”

      (2) I assume experiments occurred during the light phase, but it would be good if this could be stated explicitly.

      We have now added text specifying that these experiments took place in the light phase on Page 34, Lines 9-12 of the Methods section under “Behavior”:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Page 2 Line 14: Rephrase to make clearer, e.g. 'that have a shift in...'.

      We have rephrased the sentence for clarification on Page 2, Lines 13-15:

      “REM HFO chains also preferentially engage CA1 neuronal populations that demonstrate a shift in their preferred theta-phase from behavior to REM sleep.”

      (4) Page 4 Line 36 - 'and find coherent shifts in TD', this is self-fulfilling. I would rephrase to something like 'resulting in...'.

      We have rephrased the sentence on Page 4, Lines 36-38:

      “We separated NREM and REM sleep stages based on theta-to-delta (TD) ratio in CA1, which revealed coherent shifts in TD ratio across CA1 and PFC at the onset and offset of REM sleep…”

      (5) Figure 1H left - clarify how many animals or multiunits this is based on.

      We have now updated Figure 1 legend to specify the number of animals and epochs included on Page 19, Line 4:

      “REM HFO aligned multiunit activity (MUA) in PFC (n = 10 animals, 36 epochs)…”

      (6) Figure 1I (right), it would be useful to see the x-axis frequency start from 0, since you are cutting the peak in power.

      We thank the Reviewer for this suggestion. We had initially set the frequency limits to 4 and 12 to specifically illustrate the absence of theta-modulated activity during NREM ripples. However, as suggested by the Reviewer, it is informative to expand the frequency range to ascertain the location of the peak frequency for NREM. Indeed, the peak frequency of NREM ripple-aligned PFC activity tends to be lower than 4 Hz, which is consistent with a single peak of activity that lasts <1 s.

      We have now updated Figure 2C with these new panels.

      (7) Figure 5 E/I - use of ** is confusing, it looks like a significance comparing e.g. quartile 2 to 1, but I think this is the correlation significance. I'd move ** to the top right corner and ideally include rho values.

      We thank the Reviewer for this suggestion. We have now added the r values for the correlation to Figures 7E and 8B to resolve any ambiguity. In addition, we updated Figure 7F, right and Figure 5B to maintain consistency across main figures.

      (8) Figure 7D - Why are stimulated and non-stimulated cells so different at baseline? This is not the case for the NREM results.

      The y-axes in Figures 10C,D (Formerly Figure 7) show the fraction of cells with at least one spike per 10ms bin, either for all pyramidal cells or all interneurons. We report in the figure legend that the stimulated cells represent only 30% of the network for any given stimulus (Page 32, Line 18), so at baseline in both the NREM and REM simulations, there are roughly 3x as many non-stimulated cells with a spike per 10ms bin compared to stimulated cells, as these populations are active at roughly equal firing rates per neuron outside of stimulation periods.

      (9) Methods - there is limited info on spike sorting procedure - can you provide a reference with further details (I couldn't find)? Is this method equally valid for identifying PFC units?

      Matclust is a MATLAB-based spike sorting graphical user interface that allows for manual curation of neuron clusters through the visualization of spike waveform amplitude, peak-to-trough, and principal components. Polygons or boxes are drawn around spike data points, and single unit clusters are resolved through refinement in multiple dimensions. It was developed by Mattias Karlsson and was first used in a publication reporting replay of remote experiences in the hippocampus (Karlsson and Frank 2009). Although there is no formal reference, it can be found at https://bitbucket.org/mkarlsso/matclust/src/master/. Other labs have used the software for clustering neurons from cortical areas (Yu, Liu et al. 2018, Proskurin, Manakov et al. 2023), which demonstrates its robustness across multiple brain areas.

      Additionally, examples of clustered neurons in PFC over the course of the experimental paradigm used here can be found in our previous publication (Shin, Tang et al. 2019). In addition to the aforementioned references, we show that PFC neurons can be accurately clustered and that neurons are stable over time, according to a number of cluster metrics.

      (10) Why were there no further analyses of pyramidal cells and interneurons beyond Figure S2?

      We thank the Reviewer for bringing up this important point, which is similar to Reviewer 2, comment #5 above. We repeat our response here. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with primarily recording from pyramidal neurons (Supplementary Figure S2F, left). All results are similar if we exclude putative interneurons. To validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      (11) Looking at Figure S1B, the second to last main block of NREM sleep shown has a clear peak passing TD threshold, but oddly not classed as REM - I can only assume this is due to the duration limits on your classification?

      Yes, this is due to the REM duration threshold that we implement in our sleep scoring algorithm. We used a minimum REM bout threshold criterion of 10 s for inclusion, as in previous reports (Rothschild, Eban et al. 2017, Zhang, Zhang et al. 2020). Additionally, we have included example sleep plots for all 10 animals in Supplementary Figure S1.

      (12) Include details of how head speed was calculated - just based on the 30fps video?

      Yes, the head speed of the animal was determined by tracking the animals’ position and calculating the speed based on cm/pixel values. We have now added more detail on this in the Methods section under “Surgical implant and electrophysiology” on Page 33, Line 44 to Page 34, Lines 1-2:

      “Additionally, the animals’ speed was calculated based on predetermined cm/pixel values and the position displacement between frames captured at 30 fps.”

      (13) Page 31 line 7 - With the reference you cite, they didn't really show tonic and phasic REM can be segregated based on theta frequency - they just defined it as such. You've done it for some, but I'd ensure all references to phasic REM are defined as putative - mostly missed within the discussion - I'd also add this as a brief limitation (without recording of eye movements or PGO waves).

      We thank the Reviewer for these suggestions. We have now ensured that all mentions of “phasic REM” are qualified with “putative” and have added the requested limitation in the Limitations section on Page 16, Lines 10-15:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states.(Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult.”

      We have also modified the wording in the Methods section under “Theta inter-peak intervals during bouts of high and low theta power” on Page 37, Lines 14-15 to indicate that the referenced article simply used the theta frequency-based method to define tonic and phasic REM – not to explicitly separate the two states:

      “Since previous studies have segregated putative tonic and phasic substages of REM sleep based on CA1 theta frequency…”

      References:

      Abdou, K., M. Nomoto, M. H. Aly, A. Z. Ibrahim, K. Choko, R. Okubo-Suzuki, S. I. Muramatsu and K. Inokuchi (2024). "Prefrontal coding of learned and inferred knowledge during REM and NREM sleep." Nat Commun 15(1): 4566.

      Aleman-Zapata, A., R. G. M. Morris and L. Genzel (2022). "Sleep deprivation and hippocampal ripple disruption after one-session learning eliminate memory expression the next day." Proc Natl Acad Sci U S A 119(44): e2123424119.

      Belluscio, M. A., K. Mizuseki, R. Schmidt, R. Kempter and G. Buzsaki (2012). "Cross-frequency phase coupling between theta and gamma oscillations in the hippocampus." J Neurosci 32(2): 423–435.

      Bueno-Junior, L. S., M. S. Ruckstuhl, M. M. Lim and B. O. Watson (2023). "The temporal structure of REM sleep shows minute-scale fluctuations across brain and body in mice and humans." Proc Natl Acad Sci U S A 120(18): e2213438120.

      Cairney, S. A., S. J. Durrant, R. Power and P. A. Lewis (2015). "Complementary roles of slow-wave sleep and rapid eye movement sleep in emotional memory consolidation." Cereb Cortex 25(6): 1565–1575. Chang, H., W. Tang, A. M. Wulf, T. Nyasulu, M. E. Wolf, A. Fernandez-Ruiz and A. Oliva (2025). "Sleep microstructure organizes memory replay." Nature 637(8048): 1161–1169.

      Cheng, S. and L. M. Frank (2008). "New experiences enhance coordinated neural activity in the hippocampus." Neuron 57(2): 303–313.

      Darevsky, D., J. Kim and K. Ganguly (2024). "Coupling of Slow Oscillations in the Prefrontal and Motor Cortex Predicts Onset of Spindle Trains and Persistent Memory Reactivations." J Neurosci 44(43). Davidson, T. J., F. Kloosterman and M. A. Wilson (2009). "Hippocampal replay of extended experience." Neuron 63(4): 497–507.

      Ellenbogen, J. M., P. T. Hu, J. D. Payne, D. Titone and M. P. Walker (2007). "Human relational memory requires time and sleep." Proc Natl Acad Sci U S A 104(18): 7723–7728.

      Foster, D. J. and M. A. Wilson (2006). "Reverse replay of behavioural sequences in hippocampal place cells during the awake state." Nature 440(7084): 680–683.

      Ghosh, M., F. C. Yang, S. P. Rice, V. Hetrick, A. L. Gonzalez, D. Siu, E. K. W. Brennan, T. T. John, A. M. Ahrens and O. J. Ahmed (2022). "Running speed and REM sleep control two distinct modes of rapid interhemispheric communication." Cell Rep 40(1): 111028.

      Grosmark, A. D. and G. Buzsaki (2016). "Diversity in neural firing dynamics supports both rigid and learned hippocampal sequences." Science 351(6280): 1440–1443.

      Helfrich, R. F., J. D. Lendner, B. A. Mander, H. Guillen, M. PaD, L. Mnatsakanyan, S. Vadera, M. P. Walker, J. J. Lin and R. T. Knight (2019). "Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans." Nat Commun 10(1): 3572.

      Jarosiewicz, B., B. L. McNaughton and W. E. Skaggs (2002). "Hippocampal population activity during the small-amplitude irregular activity state in the rat." J Neurosci 22(4): 1373–1384.

      Ji, D. and M. A. Wilson (2007). "Coordinated memory replay in the visual cortex and hippocampus during sleep." Nat Neurosci 10(1): 100–107.

      Karlsson, M. P. and L. M. Frank (2009). "Awake replay of remote experiences in the hippocampus." Nat Neurosci 12(7): 913–918.

      Kay, K., M. Sosa, J. E. Chung, M. P. Karlsson, M. C. Larkin and L. M. Frank (2016). "A hippocampal network for spatial coding during immobility and sleep." Nature 531(7593): 185–190.

      Khodagholy, D., J. N. Gelinas and G. Buzsaki (2017). "Learning-enhanced coupling between ripple oscillations in association cortices and hippocampus." Science 358(6361): 369–372.

      Kudrimoti, H. S., C. A. Barnes and B. L. McNaughton (1999). "Reactivation of hippocampal cell assemblies: effects of behavioral state, experience, and EEG dynamics." J Neurosci 19(10): 4090–4101. Lee, A. K. and M. A. Wilson (2002). "Memory of sequential experience in the hippocampus during slow wave sleep." Neuron 36(6): 1183–1194.

      Leemburg, S., V. V. Vyazovskiy, U. Olcese, C. L. Bassetti, G. Tononi and C. Cirelli (2010). "Sleep homeostasis in the rat is preserved during chronic sleep restriction." Proc Natl Acad Sci U S A 107(36): 15939–15944.

      Lopes-dos-Santos, V., S. Ribeiro and A. B. Tort (2013). "Detecting cell assemblies in large neuronal populations." J Neurosci Methods 220(2): 149–166.

      Mizuseki, K., K. Diba, E. Pastalkova and G. Buzsaki (2011). "Hippocampal CA1 pyramidal cells form functionally distinct sublayers." Nat Neurosci 14(9): 1174–1181.

      Nitsche, M. A., M. Jakoubkova, N. Thirugnanasambandam, L. Schmalfuss, S. Hullemann, K. Sonka, W. Paulus, C. Trenkwalder and S. Happe (2010). "Contribution of the premotor cortex to consolidation of motor sequence learning in humans during sleep." J Neurophysiol 104(5): 2603–2614.

      Peyrache, A., M. Khamassi, K. Benchenane, S. I. Wiener and F. P. Battaglia (2009). "Replay of rule-learning related neural patterns in the prefrontal cortex during sleep." Nat Neurosci 12(7): 919–926.

      Plitt, M. H. and L. M. Giocomo (2021). "Experience-dependent contextual codes in the hippocampus." Nat Neurosci 24(5): 705–714.

      Proskurin, M., M. Manakov and A. Karpova (2023). "ACC neural ensemble dynamics are structured by strategy prevalence." Elife 12.

      Rothschild, G., E. Eban and L. M. Frank (2017). "A cortical-hippocampal-cortical loop of information processing during memory consolidation." Nat Neurosci 20(2): 251–259.

      Shin, J. D. and S. P. Jadhav (2024). "Prefrontal cortical ripples mediate top-down suppression of hippocampal reactivation during sleep memory consolidation." Curr Biol 34(13): 2801–2811 e2809.

      Shin, J. D., W. Tang and S. P. Jadhav (2019). "Dynamics of Awake Hippocampal-Prefrontal Replay for Spatial Learning and Memory-Guided Decision Making." Neuron 104(6): 1110–1125 e1117.

      Siapas, A. G. and M. A. Wilson (1998). "Coordinated interactions between hippocampal ripples and cortical spindles during slow-wave sleep." Neuron 21(5): 1123–1128.

      Simor, P., G. van der Wijk, L. Nobili and P. Peigneux (2020). "The microstructure of REM sleep: Why phasic and tonic?" Sleep Med Rev 52: 101305.

      Singer, A. C. and L. M. Frank (2009). "Rewarded outcomes enhance reactivation of experience in the hippocampus." Neuron 64(6): 910–921.

      Sosa, M., H. R. Joo and L. M. Frank (2020). "Dorsal and Ventral Hippocampal Sharp-Wave Ripples Activate Distinct Nucleus Accumbens Networks." Neuron 105(4): 725–741 e728.

      Sosa, M., M. H. Plitt and L. M. Giocomo (2025). "A flexible hippocampal population code for experience relative to reward." Nat Neurosci 28(7): 1497–1509.

      Stark, E., L. Roux, R. Eichler and G. Buzsaki (2015). "Local generation of multineuronal spike sequences in the hippocampal CA1 region." Proc Natl Acad Sci U S A 112(33): 10521–10526.

      Tang, W., J. D. Shin, L. M. Frank and S. P. Jadhav (2017). "Hippocampal-Prefrontal Reactivation during Learning Is Stronger in Awake Compared with Sleep States." J Neurosci 37(49): 11789–11805.

      Tort, A. B., R. Scheffer-Teixeira, B. C. Souza, A. Draguhn and J. Brankack (2013). "Theta-associated high-frequency oscillations (110-160Hz) in the hippocampus and neocortex." Prog Neurobiol 100: 1–14.

      Valero, M., T. J. Viney, R. Machold, S. Mederos, I. Zutshi, B. Schuman, Y. Senzai, B. Rudy and G. Buzsaki (2021). "Sleep down state-active ID2/Nkx2.1 interneurons in the neocortex." Nat Neurosci 24(3): 401–411.

      van de Ven, G. M., S. Trouche, C. G. McNamara, K. Allen and D. Dupret (2016). "Hippocampal Offline Reactivation Consolidates Recently Formed Cell Assembly Patterns during Sharp Wave-Ripples." Neuron 92(5): 968–974.

      van der Helm, E. and M. P. Walker (2011). "Sleep and Emotional Memory Processing." Sleep Med Clin 6(1): 31–43.

      Vaz, A. P., S. K. Inati, N. Brunel and K. A. Zaghloul (2019). "Coupled ripple oscillations between the medial temporal lobe and neocortex retrieve human memory." Science 363(6430): 975–978.

      Wilson, M. A. and B. L. McNaughton (1994). "Reactivation of hippocampal ensemble memories during sleep." Science 265(5172): 676–679.

      Yang, S. R., H. Sun, Z. L. Huang, M. H. Yao and W. M. Qu (2012). "Repeated sleep restriction in adolescent rats altered sleep patterns and impaired spatial learning/memory ability." Sleep 35(6): 849–859.

      Yu, J. Y., D. F. Liu, A. Loback, I. Grossrubatscher and L. M. Frank (2018). "Specific hippocampal representations are linked to generalized cortical representations in memory." Nat Commun 9(1): 2209.

      Zhang, L. B., J. Zhang, M. J. Sun, H. Chen, J. Yan, F. L. Luo, Z. X. Yao, Y. M. Wu and B. Hu (2020). "Neuronal Activity in the Cerebellum During the Sleep-Wakefulness Transition in Mice." Neurosci Bull 36(8): 919– 931.

    1. Author response:

      We read the Assessment as identifying two decisive gaps: (i) claims of group differences are supported by within-group tests rather than by a test of the group X condition interaction, and (ii) the drift-diffusion modeling is not yet validated or fit in a framework that permits group comparison. We agree with both, and we do not defend the current versions of these analyses. The revision will rebuild them rather than supplement them. We also agree with the reviewers that several of our conclusions about “internal beliefs” are currently stated more strongly than the analyses support, and these will be scaled back to what the modeling can carry.

      Below we first note factual errors and internal inconsistencies in the manuscript that the editors asked us to flag promptly, then summarize the planned revisions, then respond to each public review comment in turn.

      (1) Corrections and clarifications for the record

      Reviewer #2 identified one outright error in our text, and in re-checking the manuscript we found several further inconsistencies. We list them here so that they are on the record alongside the first version of the Reviewed Preprint. In each case the reviewers' reading of the manuscript is correct and the manuscript is at fault.

      (a) Aggressive investors and pupil modulation (Reviewer #2, 4.3). Our Results text states that “aggressive investors showed no significant pupil modulation by ambiguity,”. This is incorrect and contradicts our own Fig. 2d, which reports a significant, ambiguous vs. non-ambiguous difference in all three groups, including aggressive investors (T(31) = 2.74, P = 0.0304, Bonferroni-corrected). The figure and its statistics are correct; the text is wrong. The interpretive claim built on it - that aggressive investors show “blunted” arousal - is therefore unsupported and will be removed. The revised manuscript will state the correct within-group result and will test group differences directly rather than by contrasting significant against non-significant within-group tests.

      (b) Within-group versus between-group claims. Relatedly, our Results state that ambiguity-related pupil differences under the subjective model are “no longer significant relative to zero,” whereas the Discussion describes them as no longer differing across groups. These are different claims, and only the first was tested. This conflation runs through several of our physiological conclusions and is the same problem the Assessment identifies. It will be resolved by replacing these statements with explicit between-group and interaction tests.

      (c) Sample description (Reviewer #1, 5). The numbers are not contradictory but are certainly underspecified, and we accept that as written they cannot be reconciled by a reader. To state them plainly: 57 unique individuals each completed one to three sessions, yielding 108 sessions; 7 sessions were incomplete and 5 showed no response variability, leaving 96 sessions with usable behavior; of these, 95 had usable pupil data and 79 usable EEG. The three strategy groups are tertiles of these 96 sessions (32 each), and the smaller Ns in Figs. 2-4 (N = 31, 29, 25) reflect modality-specific exclusions within each tertile. The revision will include a participant/session flow diagram and a table reporting how many individuals contributed one, two, or three sessions, together with per-figure Ns.

      (d) DDM fitting procedure (Reviewer #1, 11). One clarification: the DDMs were fit separately for each participant, not separately per group (Methods, Section 4.6); group comparisons were then performed on the participant-level parameter estimates. We note this only for accuracy of the record. The reviewer's substantive objection is unaffected and we accept it: comparing parameters estimated in independent per-participant fits does not constitute a test of group differences within a common statistical framework, and it cannot test interactions.

      (e) Inconsistent specification of the EEG regressor. The trial-wise EEG covariate is described in one paragraph of Methods 4.6 as 8-13 Hz power averaged over 0-0.5 s, and in the text following Eq. 3 as a 1315 Hz difference over 0.1-0.4 s. The former corresponds to the analysis actually performed. We will correct Eq. 3's description and report the frequency band and window once, unambiguously.

      (f) Time-locking and figure/caption errors. As Reviewer #1 notes (comment 14), Methods describe stimulus-locked epochs (-0.25 to 1 s) while Figs. 2-3 are labeled relative to decision onset over -1 to 2 s. We will state for each analysis whether it is stimulus- or response-locked and harmonize axis labels accordingly. In addition: the Fig. 3 caption reads "N = 25 for ideal and aggressive; N = 29 for aggressive," where the second instance should read conservative; and Section 2.6 cites Fig. 1b and 1c for the leadership and team-performance results, which are Fig. 4e and 4f.

      (g) Delta/theta claims in the Discussion. Our Discussion attributes early delta- and theta-band enhancements to ideal investors. No such effects appear in our own time-frequency results, which report frontal beta-range and parietal alpha/low-beta clusters. These Discussion statements are not supported by the data presented and will be deleted. This also bears on Reviewer #1's comment 13, since it removes the only interpretation that depended on the low-frequency edge of our filter passband.

      (2) Summary of planned revisions

      (2.1) A single statistical framework with explicit interaction tests. All group comparisons will be replaced by unified models that include group, condition, and their interaction, with sessions nested within participants. Choice will be modeled with a generalized linear mixed model of the form invest ~ ambiguity ⨉ group + trial + (1 + ambiguity | participant/session); decision time with the corresponding linear mixed model including random slopes for ambiguity (Reviewer #1, 15). Time-resolved pupil and time-frequency EEG effects will be evaluated using cluster-based permutation tests on the group-by-condition interaction statistic rather than by aggregating within-group tests. Where our claim is that an effect is absent, most importantly, that ambiguity-related pupil differences vanish under subjective valuation, we will support it with equivalence testing and Bayes factors rather than with a non-significant P value, since a null result is not evidence for the null.

      (2.2) Continuous analyses as primary; grouping justified or abandoned. We accept that trichotomizing a continuous decision tendency requires justification that we did not provide. The revision will (i) define a continuous EV-consistency index and re-run every key analysis with it as a continuous predictor, establishing that the principal conclusions do not depend on the split. The “ideal” label implies a normative optimality we have not demonstrated and will be replaced with descriptive labels (EV-consistent, EV-exceeding, EV-shortfall).

      (2.3) Full specification, validation, and reliability of the belief parameter k. We agree the current description of k is inadequate. The revision will give the generative choice model in full, and explicitly state the estimation procedure and parameter bounds.

      (2.4) Rebuilt and validated drift-diffusion modeling. The boundary will be freed and estimated per participant/session, so that variance is no longer forced into the drift term. We will report whether the group effect on baseline drift survives.

      (2.5) Report why repeated sessions are treated as independent samples. We will conduct additional behavioral analyses to show repeated sessions from the same individual are sufficiently independent to be analyzed as unique samples. In particular, we will examine whether within-participant similarity across sessions is greater than cross-participant similarity. These analyses will provide direct evidence for whether sessions can reasonably be treated as separate samples rather than requiring sessions to be nested within participants.

      (2.6) Sharper framing and appropriately scaled claims. We will remove the framing that positions the field as treating ambiguity aversion as uniform and instead situate the work within literature that already documents heterogeneity in ambiguity attitudes. The distinction between ambiguity aversion and the internal representation of ambiguity will be made explicit, with the latter identified as our actual question. Claims about “internal belief models” will be restated in terms of what is estimated. Physiological hypotheses will be stated as specific, directional, falsifiable predictions rather than by appeal to the broad range of processes EEG has been linked to.

      (2.7) Quality control for pupillometry, EEG, and the VR context. We will report blink rates and counts, proportion of interpolated samples, trials and channels rejected, and ICA components removed. We will also report validate the luminance GLM.

      (2.8) The collaborative task. The Apollo Distributed Control Task preceded the Lottery Choice Task by design, so it must be reported as part of the protocol regardless of the leadership analysis. However, we agree that the leadership and team-performance analyses are not sufficiently motivated to carry the interpretive weight currently given them. They will be moved to a clearly labeled exploratory section, removed from the Abstract and framing, and presented without causal or trait-level interpretation.

      (3) Timeline and next steps

      The revisions above require refitting the behavioral, physiological, and computational analyses rather than adding to them, including hierarchical model fitting, parameter recovery, and new quality-control analyses. We therefore anticipate submitting the revised manuscript within approximately three months, and would welcome guidance if the editors would prefer a different schedule. We are content for the first version of the Reviewed Preprint to be published with this provisional response attached.

      We are grateful to both reviewers for the time invested in this manuscript. Several of the problems they identify are ones we should have caught ourselves, and the paper will be considerably stronger for their having been raised now rather than after publication.

    1. Author response:

      General Statements.

      We thank all reviewers for their careful evaluation and constructive comments.

      Fidel Serrano recently completed a related study on the role of CK1δ in the circadian clock which is published in eLife (https://doi.org/10.7554/eLife.110786.1). During these studies, we became interested in the physiological role of CK1δ autophosphorylation, whose functional significance has remained unclear.

      In the present manuscript, Fidel discovered that the autophosphorylated, auto-inhibited form of CK1δ accumulates specifically during mitosis. These findings provide a physiological context for CK1δ autophosphorylation that has remained elusive for many years. They constitute the central foundation of our model, which further integrates previous findings on APC/C function and activity with the data presented in our eLife study.

      Several suggestions and questions raised by the reviewers concern the regulation of CK1δ in cultured cells, which are predominantly in G1. Many of the corresponding experiments and analyses are already included in our eLife paper. We apologize that this overlap may complicate the review process. However, we are unable to publish identical datasets in both manuscripts.

      We therefore briefly summarize here the findings from the eLife study that are most relevant to the present work.

      - Overexpressed CK1δ is unstable, whereas CK1δ-K38R (catalytically inactive) and CK1δ-R178Q (altered specificity/activity) are comparatively stable (eLife Figs. 3A-E and 5B).

      - Assembly with PER2 stabilizes overexpressed CK1δ (eLife Figs. 4A-D).

      - Treatment with PF670462 similarly stabilizes the kinase (eLife Figs. 3F and 5D).

      - CK1δ binding sites in the centrosomal/Golgi area are present in excess even relative to overexpressed kinase (eLife Fig. 5C).

      The reviewers also raised questions about the PER2-CRY1 nuclear foci.

      Thermodynamically stable/persisting nuclear foci form upon coexpression of PER2 and CRY1. They were preliminarily characterized in the eLife study (eLife Fig. 2). For the purposes of the present work, they constitute a serendipitous and convenient tool for studying interactions of PER with CK1δ. To induce foci formation, we used a stable cell line expressing DOX-inducible mK2-CRY1. mK2-CRY1 accumulates only at relatively low levels because the protein without a binding partner is intrinsically unstable, as confirmed by Western blot analysis (eLife Fig. 4C). Coexpression of PER2 stabilizes mK2-CRY1 and promotes the formation of nuclear PER2-mK2-CRY1 foci.

      These foci contain elevated levels of endogenous CK1δ (eLife Fig. 2E and EV2), which accumulate gradually over a 24-hour period. This observation indicates that, at steady state, a fraction of endogenous CK1δ is continuously degraded in the absence of overexpressed PER2. Because endogenous CK1δ is synthesized at a relatively low rate, the unstable pool is small at any given time and therefore not readily detectable in conventional cycloheximide chase assays against the kinetically stabilized steady-state background of CK1δ.

      In our point-by-point response, we therefore refer the reviewers to the corresponding datasets and analyses presented in the eLife manuscript.

      Point-by-point description of the revisions.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      Involved in various cellular pathways and processes, Ser/Thr kinase CK1δ is thought to be constitutively active and its tail autophosphorylation is suggested as a putative inhibitory mechanism of CK1δ kinase activity. Here, the authors investigate CK1δ's phosphorylation status, in relation to its location and dynamics throughout the cell cycle. Immunofluorescence and biochemistry studies showed that the subcellular distribution of CK1δ is dynamic and in equilibrium between centrosomal and nuclear pools. The authors argue that CK1δ phosphorylation protects it from degradation. Finally, the authors show differences in CK1δ phosphorylation status and location throughout the cell cycle. Combining their findings with the existing literature, they propose a model of CK1δ phosphorylation status, location and dynamics throughout the cell cycle.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Lack of a positive control for centrosomal staining, eg. PLK1 (Fig.1A,C, 2B, 6A-B). This would also allow a colocalisation analyses between CK1δ and a centrosomal marker, which would build a more convincing body of evidence in favour of centrosomal recruitment of CK1δ in baseline conditions. A lack of a negative control for CK1δ IF? Most CK1 antibodies also cross-react with other isoforms.

      As suggested, we show centrosomal staining with a well-studied centrosomal marker pericentrin (PCNT; revised Fig. 1). Consistent with what has been well-described in the literature that CK1δ localizes to the centrosome (Sillibourne et al., 2002; Greer & Rubin, 2011; Greer et al., 2014) and with our observations, we confirm that CK1δ is indeed localized at the pericentral region.

      The CK1δ antibody does not crossreact, this is shown in Fig EV3A of the eLife paper.

      b. Lack of proper image quantification (Fig.1A,C, 2B, 6A-B). To establish a centrosomal recruitment of the CK1δ pools, colocalisation of CK1δ and a centrosomal marker must be quantified using a standard colocalisation quantification method and shown appropriately. Same comments regarding the colocalisation of CK1δ with PER2/mK2-CRY1-positive foci.

      Fig. 1: Colocalization analysis with pericentrin (PCNT) as centrosomal marker is provided.

      Fig. 2: The colocalization analysis of CK1δ (green) and mK-CRY1 (magenta) shows that in all PER2-untransfected cells (50 cells evaluated), CK1δ is concentrated at a single, mK2-CRY-negative spot, corresponding to the pericentrosomal region, which we have established in Figure 1. In contrast, in all PER2-transfected cells (15 cells evaluated) CK1δ localizes to mK2-CRY1 positive nuclear foci. The fields-of-view we have provided are representative of the typical phenotype of CK1δ localization in the presence or absence of PER2/mK2-CRY1-positive foci. Here it should be noted that CRY1 foci formation is strictly dependent of co-expression with PER2, (eLife paper).

      Fig. 6: This figure presents an analysis of fixed cells. Because only a small fraction of asynchronously growing cells is in mitosis at any given time, the number of cells assigned to individual mitotic stages is necessarily low (typically fewer than 10 cells per stage). The purpose of this figure is therefore primarily illustrative: to document and confirm the known subcellular localization of CK1δ during the different stages of mitosis rather than to provide a comprehensive quantitative analysis.

      c. Lack of proper statistical analyses (Fig.1, 2, 6). As immunofluorescence constrains one to only show few images per condition at best, statistical analyses on broader image analysis data are essential to measure the significance of the changes observed on the representative images provided.

      See answer to previous question.

      d. Lack of biological replicate numbers (Fig.1, 2, 6). Was it from 3 independent biological replicates?

      Fig.1 and 2: three independent biological replicates.

      Fig.6 is just illustrative to show the already known localization of CK1δ across mitosis. The localization shown in the four panels is seen in all cells at the respective cell cycle stages.

      e. NOTE: if available, using a confocal microscope would be best to provide optimal Z-axis resolution. This would provide you with more accurate colocalisation data, like CK1δ recruitment to the centrosome or PER2/mK2-CRY1-positive foci.

      Characterization of PER-CRY foci is published in the eLife paper in Fig. 2 and EV2.

      (2) Shortcomings in biochemistry experiments

      Lack of loading controls in Fig.3A-C and 4A-B. Until they are added, no conclusions can be safely interpreted from the experiments. Myc is a good turnover control for CHX across all experiments. Enrichment of phospho-proteins upon CalA treatment?

      Loading controls are provided in this revision. Enrichment of phospho-proteins following CalA treatment has been demonstrated and published in numerous previous studies (Cegielska et al., 1998; Ishihara et al., 1989; Rivers et al., 1998).

      (3) The authors infer from the literature that overexpressed CK1δ is unassembled, without checking it in any experiment. It could be that CK1δ overexpression drives overexpression of its binding partners, in which case most of overexpressed kinase could actually be assembled. Since CK1δ assembly is of great importance to the study's conclusions, it should be confirmed experimentally.

      In our eLife study, we show that overexpressed CK1δ is not fully assembled with stabilizing binding partners such as PER2. In fact, centrosomal/Golgi binding partners remain always available in excess even relative to overexpressed CK1δ, as demonstrated by the MG132-induced increase in centrosomal accumulation (eLife Fig. 5C). However, binding is dynamic and the affinities and concentrations are such that a substantial fraction of overexpressed CK1δ remains unbound and is therefore subject to degradation.

      (4) The authors infer from the literature that phosphorylated CK1δ is inactive - which is not a given. All Western blotting experiments should include confirmed CK1δ substrates like PER2 or DVL3 to confirm that phosphorylated CK1δ is indeed inhibited. As added benefit, blotting for known CK1δ substrates will act as confirmation that the FLAG-tag on overexpressed CK1δ/ε does not impact its kinase function, and that CK1δ/ε inhibitor treatments like PHF670 have indeed worked. Furthermore, because CalA and OA inhibit many phosphatases, inferences on CK1δ/ε activity upon these treatments should be taken catiously.

      The activity of phosphorylated CK1δ is severely attenuated by autoinhibition, although the kinase is not completely inactive. This has been demonstrated repeatedly in the literature and does not require re-establishment in the present study. We previously showed this directly in a PNAS study by Marzoll et al. (https://doi.org/10.1073/pnas.2118286119). At some point, one must rely on established published findings unless there is compelling evidence supporting an alternative interpretation.

      (5) Figure 1 requires further biochemical analyses to support immunofluorescence data. Western blotting analyses showing fluctuating CK1δ levels in nuclear vs. Centrosomal (https://pmc.ncbi.nlm.nih.gov/articles/PMC7618310/) fractions in the different treatment conditions would help illustrate the redistribution of CK1δ pools between nuclear and centrosomal areas. Additionally, adding a cytoplasmic fraction in each treatment condition would help visualise the amounts of unrecruited CK1δ in overexpressed conditions.

      Microscopy is a well-established and widely accepted approach for assessing subcellular localization. In many cases, it is superior to biochemical fractionation assays, which rarely yield completely pure nuclear or centrosomal fractions. The fractionation experiments suggested by the reviewer could provide additional supportive evidence, but they are not required to substantiate the conclusions presented here.

      (6) Figure 2 requires further biochemical analyses to support immunofluorescence data showing colocalisation of CK1δ with PER2, using for example immunoprecipitation to show that pulling down PER2 also pulls down CK1δ, and vice versa. Optimally, an additional Western blotting experiment showing total / phospho-PER2 and CK1δ levels for each condition in cytoplasmic vs. nuclear fractions would consolidate the evidence from immunofluorescence and immunoprecipitation.

      Association of CK1δ with PER2 has been demonstrated extensively in numerous previous studies (Aryal et al., 2017; Narasimamurthy et al., 2018; Philpott et al., 2020; Cao et al., 2021; Marzoll et al., 2022; An et al., 2022) including our own eLife publication.

      (7) The authors consistently interpret data showing CK1δ phosphorylation as CK1δ tail phosphorylation (p.8-11). A band shift in CK1δ signal on a Western blot does not in any way show the tail specifically is phosphorylated. Indeed, although CK1δ tail phosphorylation is widely recognised, several identified sites on other CK1δ domains, e.g.: the kinase domain, can be phosphorylated by CK1δ and other kinases. The authors' conclusions must therefore not mention tail-specific phosphorylation unless domain-specific phosphorylation is established. This could be done by comparing signal from antibodies specifically recognising known phospho-sites on the CK1δ tail against total CK1δ signal in experiments of Fig.3, 4, and 5.

      The electrophoretic mobility shift is caused by phosphorylation of the CK1δ tail. This has been demonstrated in numerous previous studies. For example, we generated a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines. In this mutant, CK1δA, inhibition of phosphatases by CalA no longer induces an electrophoretic shift (Figure 3B).

      (8) The use of anti-FLAG antibody in all figures looking at overexpressed CK1δ is not the most appropriate choice, as the anti-CK1δ antibody works perfectly for Western blotting and immunofluorescence studies. This adds more variables and hinders comparisons with endogenous CK1δ. If CK1δ is successfully overexpressed, wild-type levels should be negligible compared to overexpressed CK1δ levels, so the need for a FLAG antibody is not justified. Using an anti-CK1δ antibody in all experiments instead of anti-FLAG would confer them greater solidity.

      The amount of endogenous CK1δ is limited, and we therefore use endogenous detection only when it is scientifically necessary. FLAG-tagged constructs, in contrast, provide a robust and convenient experimental system. In the absence of evidence that the FLAG tag introduces artifacts or alters the observed behavior of the kinase, we do not consider it necessary to repeat these experiments using endogenous detection alone.

      (9) The authors base themselves off a correlation between CK1δ phosphorylation status and its observed stability to establish a causal relationship between both ('phosphorylation protects the overexpressed kinase from degradation' p. 7; 'tail phosphorylation protects the overexpressed kinase from rapid degradation' p.9). However, no experiments performed suggest there is a causal relationship occurring. This could be done by looking at specific known CK1δ tail phospho-sites using phosphosite-specific antibodies. The authors could assess whether CK1δ stability is impacted when overexpressing wild-type vs. phospho-dead mutant CK1δ.

      In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable.

      (10) The authors consistently link APC/CCDH1 with CK1δ's putative degradation, when none of their study touched on APC/CCDH1 activity or its involvement in CK1δ dynamics. To support these conclusions, fig.5 requires further biochemistry studies. To show APC/CCDH1's interaction with CK1δ, I suggest the authors (1) pull down CK1δ by immunoprecipitation and check for APC/CCDH1, phosphorylated CK1δ, and ubiquitin levels in enriched samples for each cell cycle stage in untreated vs. MG132-treated conditions. To consolidate this, the authors could (2) pull down CDH1 and check for phosphorylated and total CK1δ levels in pulled-down samples. Finally, they should investigate whether reducing CDH1 (3) expression (siRNA knock-down) and (4) CDH1 activity (specific E3 ligase inhibitor) have any impact on CK1δ levels at each cell cycle stage.

      These data were previously published in Penas et al. (doi: 10.1016/j.celrep.2015.03.016). For example, Fig. 5C shows that siRNA-mediated depletion of Cdh1, the G1-specific cofactor of APC/C, stabilizes overexpressed CK1δ, as well as the established APC/C-Cdh1 substrate Cyclin B1.

      (11) Methods section lacks a subsection detailing image acquisition, including the type of microscope used (confocal vs. widefield), magnification used, whether acquisition settings were kept consistent throughout conditions / technical/biological replicates.

      Is provided in the revised manuscript.

      (12) Methods section must include a subsection detailing immunofluorescence image analysis parameters and statistical analyses, e.g.: criteria for centrosomal localisation categories in Fig.1, criteria for categorising cells in different stages of mitosis (Fig.6), overall threshold / criteria stringency.

      Is provided in the revised manuscript.

      Minor comments:

      The Reviewer has raised about 50 minor points, counting the individual remarks and sub-point and sub-sub points. We have addressed a number of these comments where they were scientifically relevant or helpful for improving clarity. Most of the questions raised concern published and generally accepted data. We therefore do not believe that a point-by-point response to every individual minor remark is constructive or necessary.

      (1) PF670 was shown to selectively inhibit CK1δ/ε over 42 common kinases (TOCRIS), however it is not well-characterised regarding the remaining 474 kinases encoded by the human genome. There is thus a possibility for PF670 inhibition overlap between CK1δ/ε and other kinases, which should be mentioned, and the authors' conclusions should he more nuanced as a result.

      (2) CHX is a protein synthesis inhibitor, which means it does not selectively target CK1δ expression, but that of every protein within the cell. This should be addressed and controlled for, if possible. A potential way to go about it would be to knock-down CK1δ using siRNA and compare untransfected controls with select post-transfection timepoints to assess the impact of inhibiting CK1δ expression on total CK1δ levels.

      (3) All unshown Western blotting replicates should be included in the supplementary materials.

      (4) Figure 1:

      a. A-B: experiment is missing PF670 alone and CHX alone conditions to control for direct effects of either compound, versus combined.

      b. A-B: in the text, please address the fact that >50% cells show unclear or no centrosomal pattern in untreated conditions.

      c. C: the authors claim that CK1δ-FLAG levels are decreased in PF670-treated conditions because PF670 inhibits excess CK1δ autophosphorylation, inducing its subsequent degradation. Before this claim is made, a proteasome inhibitor like MG132 should be included within the experiment, to show that this decrease in CK1δ expression under PF670 treatment can be rescued with MG132 treatment. This would plead in favour of excess CK1δ degradation and exclude the possibility that 4h of PF670 treatment simply reduces the rate of CK1δ expression. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ degradation occurs via APC/CCDH1-mediated ubiquitination of CK1δ.

      d. C-D: experiment is missing pre-DOX induction control to confirm that CK1δ-FLAG is indeed being overexpressed.

      (5). Figure 2:

      a. A: should include non-transfected control panels to confirm the success of PER2 overexpression.

      b. B: Although observed in a preprint from the same team, there is no peer-reviewed evidence that establishes mK2-CRY1 as a reliable indicator of PER2 overexpression. 2A shows a correlation between both but does not exclude the fact that PER2 must be included in the imaging of the 2B panels, instead of using mK2-CRY1 as a proxy readout of PER2 expression and location. Appropriate analyses would then be required to show colocalisation of CK1δ with PER2 in the highlighted puncta.

      c. 'PER2-dependent foci': cannot be said of the data unless PER2 dependency has been validated in those images. Please see above point to resolve this.

      (6) Figure 3:

      a. B: the use of kinase-dead CK1δ mutant does not allow to fully separate direct autophosphorylation from phosphorylation by other kinases, unlike what the authors mention: CK1δ kinase activity may be required for the phosphorylation of certain sites by other kinases. In this case, the decrease of phosphorylation in the kinase-dead CK1δ mutant would not only result from inhibited autophosphorylation, but also from reduced phosphorylation by other kinases. Please adjust your conclusions accordingly (p.8).

      b. C: poor visualisation of CK1δ overexpressed condition, especially showing critically reduced signal at the 6min CHX timepoint, compared to its kinase-dead homolog. If this change is present in every replicate performed, please address it in the text. If not, perhaps you may have to display another representative replicate in the figure.

      c. The authors claim that the increased levels of overexpressed with CalA treatment suggest 'that full or partial phosphorylation of the CK1δ tail stabilizes both active and inactive forms of the kinase' (p.8). Importantly, tail phosphorylation was never shown, so this should be corrected. Additionally, it could be that CalA treatment considerably increases the rate of CK1δ expression - which would also match data in fig.4A, since CalA and CalA+PF670 treatments alone drastically increased CK1δ levels. This should be checked by pre-treating cells with CHX before applying CalN, or by treating cells with CHX and CalN simultaneously, and interpreted accordingly.

      (7) Figure 4:

      a. A: CK1δ levels in CHX-free controls from both 1h pre-treatments look much higher compared to untreated controls, which the authors interpret as an indication that 'most of the newly synthetised overexpressed kinase was degraded in untreated cells' (p.9). However other explanations are not explored: since loading controls are not provided, it may be that sample loading in the gel is simply off. Importantly, it is also possible that pre-treatments increased CK1δ expression before CHX application. Please make sure you touch on each

      b. A-B: the authors claim that CK1δ-FLAG and CK1ε-FLAG levels are decreasing with CHX treatment because they are being degraded. Adding a panel with a proteasome inhibitor like MG132 would solidify this argument. Rescue of CK1δ/ε degradation under CHX treatment would show that the loss is indeed mediated by the UPS. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ/ε degradation occurs via its ubiquitination by APC/CCDH1.

      c. B, D: blot in B does not match the quantification trends in D. E.g.: 60min CHX + CalA + PF670 condition shows clearly lower CK1ε signal compared to its 0min CHX control. Please ensure the biological replicate you display on the figure is indeed representative of your results.

      d. A, C: 'Hyperphosphorylated CK1δ remained stable throughout the CHX chase [...], indicating that tail phosphorylation protects the overexpressed kinase from rapid degradation' (p.9). Meanwhile this is true, unphosphorylated CK1δ in the CalA+PF670 treatment condition was also stabilised, showing that CK1δ phosphorylation may not be required for kinase stabilisation. This is an important point and should be addressed in the data interpretation. On another note - and as mentioned above -, tail phosphorylation specifically is not shown and cannot be inferred unless domain-specific phosphorylation is investigated.

      e. B, D: CK1ε-FLAG levels decrease with CHX treatment compared to its baseline in the CalA+PF670 condition, which is not the case for CK1δ-FLAG (A). Thus, the data shown does not support the authors' conclusions 'CK1ε turnover is regulated in a similar manner to CK1δ'. Please address this issue and adjust your conclusions accordingly.

      f. C-D: performing statistical analyses on protein band intensity in different conditions would be interesting to establish the significance of those changes.

      (8) Figure 5

      a. B: non-arrested controls mentioned in the text are missing from the figure.

      b. B: 'the accumulated CK1δ remained dephosphorylated' (p.10). The experiment is missing important positive / negative controls of CK1δ phosphorylation status to conclude whether CK1δ remained phosphorylated or unphosphorylated across conditions. One or the other cannot be concluded from the blot as it is. Using phospho-specific antibodies may also help to visualise phosphorylation status.

      c. B, D-E: blots should include CDH1 phosphorylation levels (hyperphosphorylated CDH1 is inactive), CK1δ substrates (indicators of CK1δ activity), and relevant phosphatase substrates to create a cohesive picture of changes in mitosis and support fig.7.

      d. D-E: please address the differences in CK1δ profile between overexpressed and endogenous CK1δ in G2/M phase.

      (9) Figure 6

      The authors say: 'similar results were observed when using U2OStx cells and staining for endogenous CK1δ'. However, CK1δ's subcellular distribution pattern is different in endogenous vs. overexpressed conditions in the telophase-cytokinesis and post-mitosis stages. In the telophase-cytokinesis stage, endogenous CK1δ seem to form nuclear hotspots, while overexpressed CK1δ is more diffuse. In the post-mitosis stage, overexpressed CK1δ shows a clear polar pattern in the nuclear periphery, while endogenous CK1δ shows a diffuse pattern similar to that of overexpressed CK1δ described in fig.1 as 'unassembled' by the authors. Please address this and adjust your conclusions appropriately.

      (10) Figure 7

      a. The authors infer APC/CCDH1's activity levels or relationship to CK1δ from existing literature only. Since it is a central mechanism of their study, literature-only components are insufficient for a summary figure. To include these elements in the figure, the authors must include investigations of APC/CCDH1's activity levels and involvement in CK1δ degradation at different stages of the cell cycle in their study. Please refer to point 10.

      b. Similar comment for CK1δ assembly status and activity levels. Please refer to point 4.

      c. Similar comment for phosphatase activity levels in different stages of the cel cycle. By observing phosphorylation status in known substrates of established CK1δ phosphatases, one can easily confirm phosphatase activity levels.

      d. Does not consider the fact that some results were different in endogenous vs. overexpressed CK1δ models. Please nuance your claims.

      (11) Reference missing p.12 paragraph 1: 'Yet, overexpression of CK1δ consistently accelerates the circadian clock, implying that kinase availability can influence clock speed. This finding suggests that CK1δ activity may be regulated not only by catalytic mechanisms but also by its spatial availability within the cell'. The facts stated there are not covered in the study's findings and is not referenced with a published study.

      (12) Grammar mistakes / typos to report:

      p.3 paragraph2: 'the kinases undergo futile cycles of phosphorylation and dephosphorylation'

      p.26 Fig.1A legend: 'Endogenous CK1δ was detected by IF to'? Unfinished formulation

      p.8: 'PP1 was previously suggested as a major PPase of CK1δ/ε'

      p.8: 'fewer phosphorylation sites are targeted by other kinases'

      Figure 4C legend: 'Overexpressed unphosphorylated CK1δ [instead of CK1ε] is degraded with a half-life of about 15 min'

      Figure 4E: Ponceau staining

      (13) Nomenclature inconsistencies to report:

      p.4 paragraph2: 'protein PER2', then p.4 paragraph3 'PERIOD2'.

      'CK1δ' used throughout the article's body text, but 'CSNK1D' is used in IF panel legends. The gene name was never introduced in the main text, nor has it been explained in the figure legends, so perhaps go for CK1δ for all mentions, including in figures.

      'FOV' nomenclature is unclear, please define in the figure legend.

      Figure 2B: what are the arrowheads pointing to? - please clarify in the figure legend.

      Mislabelled figures: figure 3 in the text refers to figure 4 in the figure section, and figure 4 in the text to figure 3 in the figure section.

      Figure 6A-B: abbreviated mitosis stages 'Pro' and 'Meta-Ana' should either be defined in the figure legend or put in full writing within the figure.

      Reviewer #1 (Significance):

      The paper will appeal to those working on CK1 biology, including cell cycle and circadian rhythms.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      In this study, authors aimed to address functional links between CK1 activation, subcellular localization and protein stability, which is an important biological question. This manuscript is a follow up of a recent study by the same team, which is currently deposited at Biorxiv and a fraction of data seems to overlap, which is somewhat confusing. Overall, the concept that the stability of CK1d is dynamically controlled across the cell cycle is interesting. On the other hand, how this is functionally connected to the circadian cycle described in the previous study remains unclear.

      The data presented in this study are highly preliminary and lack a number of essential controls, which weakens an otherwise interesting concept.

      Major issues:

      One of the main conclusions of the study is that CK1d is degraded predominantly in the nucleus while the centrosomal pool is protected from the degradation. This concept is interesting but unfortunately, the experimental evidence supporting this model is very limited. Authors previously showed that neither inhibition of proteasome or treatment of cells with PF670462 significantly influenced levels of endogenous CK1d but both treatments promoted accumulation of the tagged and overexpressed FLAG-CK1d. Absence of the phenotype at the level of endogenous protein clearly raises question whether this may be just an artifact of overexpression, tagging or both combined.

      As shown in our previously published eLife study, CK1δ is synthesized at a relatively low rate from its endogenous locus. Free CK1δ continuously shuttles between the cytosol and the nucleus. Although association of CK1δ with the centrosomal/Golgi area is dynamic, nuclear export followed by binding to centrosomal/Golgi structures constitutes the major sink for the kinase at steady state. Consequently, the fraction of unbound CK1δ that is targeted for degradation in the nucleus at any given time is very small and cannot be detected in cycloheximide chase assays against the much larger background of kinase stabilized by association with centrosomal/Golgi binding sites. To reveal that CK1δ is subject to degradation at all, we had to overexpress the kinase. Under these conditions, the fraction of unbound CK1δ increases substantially, allowing nuclear degradation to be detected experimentally.

      The statement that degradation occurs in the nucleus is based on weak data with CK1d-NES construct that was claimed to accumulate at higher levels compared to the wild type CK1d. However, that experiment used quantification of microscopic data in which two distinct regions (nucleus, cytosol) were compared, which is technically challenging. Including a reporter for normalizing for the transfection efficiency would strengthen conclusions of that experiment. Overall, quantification by immunoblotting may be more accurate.

      As suggested, we performed CHX time-course experiments with NES- and NLS-tagged CK1δ constructs followed by immunoblot analysis (shown in revised Fig. 4E).

      Specificity of the microscopy staining was not validated Fig. 1. Confirmation of the staining specificity by RNAi or KO approaches is essential. This antibody from Abcam has been discontinued which makes it impossible to reproduce this experiment.

      The authors should therefore attempt to demonstrate localization using other available antibodies against CK1d.

      The antibody is available through Thermo Fisher in the United Kingdom: https://www.fishersci.co.uk/shop/products/100-ul-mouse-monoclonal-af12g4-casein-kinase/13070202#

      It is specific for CK1δ and does not recognize CK1ε, as demonstrated in our eLife study.

      Furthermore, FLAG-tagged CK1δ detected with anti-FLAG antibodies, as well as endogenous CK1δ detected with the commercial antibody, localize to the pericentrosomal region and relocalize upon PER2 overexpression to mCRY1-containing nuclear foci. Together, we consider these data compelling evidence for the specificity of the antibody staining.

      The authors also failed to demonstrate that the dots represent centrosomes. To do so, they must perform co-staining with a robust centrosome marker.

      As requested, Pericentrin (PCNT) was used as a centrosomal marker and is now provided in the revised Fig. 1.

      This would help classify approximately 50% of the cases that are currently assessed as "potential" but are in fact inconclusive.

      Centrosomal colocalization with PCNT improved the robustness of the quantification.

      The quantification in the experiment is incorrect. Since the percentages are reported, both columns should add up to 100%. In Panel 1B, approximately 10% of the cells are missing, while in Panel 1D, there appear to be about 20% more cells.

      As indicated in our previous figure, the y-axis represents the number of cells, not percentages. We analyzed approximately 100 cells, which may have led to the misunderstanding that the values refer to percentages. As well, we have replaced this figure with a revised Figure 1 which shows that CK1δ co-localizes with PCNT as a centrosomal marker.

      Fig. 2B shows four different fields based on which authors come to conclusion that co-expression of CRY and PER2 promotes re-localisation of CK1d to nuclear foci. This is an interesting possibility but it is hard to conclude without any quantification. What was the fraction of cells that expressed CRY and PER2 that showed this phenotype? What was the fraction of cells that did not show this phenotype although both of these proteins were expressed?

      It would help to label cells expressing PER2 either by expressing it as a fusion protein or by co-expressing a marker protein ideally from the same plasmid.

      All cells in this stable cell line express mK2-CRY1. The cells were transiently transfected with PER2, and therefore only a fraction of the cells received the PER2 expression construct. In our eLife paper, we showed that all PER2-expressing cells stabilized mK2-CRY1 and formed nuclear foci. We demonstrate here that every cell containing such nuclear foci also showed accumulation of endogenous CK1δ within these structures. We did not observe a single cell with PER2-induced nuclear foci that lacked endogenous CK1δ accumulation.

      Cells that did not express PER2 were identified by the absence of nuclear foci and by low levels of mK2-CRY1, which was homogeneously distributed throughout the nucleus, as also shown in the eLife manuscript. In these cells, endogenous CK1δ was concentrated at a single discrete structure that, based on data shown in Fig. 1, corresponds to the pericentrosomal region. Under these conditions, neither CK1δ nor mK2-CRY1 was detected in nuclear foci.

      Fig. 5 suggests that massive phosphorylation of CK1d, which is responsible for its mobility shift on SDS-PAGE, is most likely linked with mitosis. Authors should use established markers to estimate a fraction of mitotic cells in their G2/M fraction. In principle, there are two possible explanations for the doublet observed with CK1d staining. Ether CK1d exists in two pools with different phosphorylation states in mitosis, or perhaps more likely, this fraction contains G2 cells where CK1d is not yet modified and mitotic cells where CK1d is fully phosphorylated. Performing a shake-off experiment yielding a pure fraction of mitotic cells could help to distinguish between these two options.

      As described in the main text and the methods section, the fractions were prepared by mitotic shake-off to further enrich our sample for rounded cells that are loosely attached during mitosis.

      Degradation of CK1d in telophase/cytokinesis when APC/Cdh1 becomes active is not apparent in Fig 6. The signal at mitotic spindle is missing, but there is still plenty of signal remaining in the cells. It is possible that the signal is just redistributed in the cell and the data shown do not support degradation of the protein. Authors could film cells expressing fluorescently labeled CK1d and quantify the signal during progression through mitosis and mitotic exit. The statement that "Following nuclear envelope reformation and mitotic exit, CK1d localized primarily to the single centrosome in each daughter cell" is incorrect. First, authors cannot deduce from the fixed cells whether they have just formed the nuclear envelope and exited mitosis.

      The reviewer is, of course, absolutely correct in the points raised.

      First, we cannot deduce from fixed cells whether they have only recently exited mitosis. This was not our intention. We merely selected cells in G1 and referred to them as “post-mitotic,” without intending to imply that these cells had just exited mitosis. To clarify this point, we changed the previous statement:

      “Following nuclear envelope reformation and mitotic exit, CK1δ localized primarily to the single centrosome in each daughter cell (Fig. 6A, 4th column)”

      to:

      “In G1, CK1δ localized primarily to the single centrosome (Fig. 6A, 4th column),” and replaced in column 4 of Fig. 6A and B the label “post-mitosis” with “G1.”

      Furthermore, we cannot deduce from fixed cells whether, or to what extent, CK1δ is degraded upon mitotic exit. This was neither the intention nor the conclusion drawn from Fig. 6. Rather, we show in Fig. 3 that phosphorylated CK1δ is not degraded, and in Fig. 5 that CK1δ is predominantly hyperphosphorylated during mitosis, leading us to conclude that this phosphorylated pool of CK1δ is stable. In G1, CK1δ is dephosphorylated and unassembled kinase is degraded. We currently have no data regarding the kinetics of CK1δ dephosphorylation, assembly with centrosomal/Golgi structures, versus degradation of unassembled dephosphorylated CK1δ.

      The data shown in Fig. 6 serve merely to illustrate the subcellular distribution of CK1δ, which is consistent with previous reports. In Fig. 7, we present a model that attempts to integrate the new findings reported here together with the data from our recent eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation. Of course, we do not claim that this model does by no means represents a final verdict, and many important questions remain open. However, we believe that the model provides plausible novel concepts and mechanistic ideas that have not been proposed in previous publications and therefore merit publication, as they provide a basis for further investigation and discussion.

      Live-cell imaging:

      Live-cell imaging of fluorescently tagged CK1δ throughout mitotic progression and mitotic exit could, in principle, provide additional insight, but such experiments are technically extremely challenging. Moreover, the central idea of our model is that as much CK1δ as possible is preserved throughout the cell cycle, whereas degradation selectively targets unassembled and potentially harmful kinase. When CK1δ is expressed at physiological levels, which would require tagging the endogenous locus, the fraction of kinase degraded upon mitotic exit is expected to be very small and therefore likely below the threshold for reliable quantification by fluorescence microscopy. Similarly, although overexpressed CK1δ undergoes substantial degradation in G1, the kinase is simultaneously synthesized at a high rate. Hence, quantitative interpretation of overexpressed CK1δ levels during mitotic exit by microscopy (without CHX) would still be difficult.

      Second, the images of interphase cells constantly show multiple dots (probably surrounding the centrosome), which is a pattern that likely corresponds to Golgi rather than a single centrosome.

      The reviewer is correct. Indeed, CK1δ localization to the Golgi apparatus is well established in the literature. In our original wording, we did not explicitly distinguish between Golgi and centrosomal localization, which may have been somewhat misleading. In the revised version, we therefore refer more cautiously to the “pericentrosomal region” rather than strictly to the centrosome.

      The data shown in Fig. 6 serve to illustrate the known subcellular distribution of CK1δ. In the model presented in Fig. 7, we attempt to integrate the new findings reported here together with the data from our eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation.

      It is unclear to which figure points the paragraph "Tail phosphorylation protects CK1δ/ε from degradation". I assume that one figure is missing.

      Figs. 3 and 4 were accidentally swapped, and we apologize for this error. The paragraph in question refers to Fig. 4, which shows the cycloheximide-induced degradation kinetics of CK1δ/ε.

      Minor points:

      CK1 kinase inhibitor PF670462 should not be named as PF670 as this causes confusion. Authors should either use the full name of the compound or just call it as CK1 inhibitor with providing details in the methods.

      PF670 has been changed to PF670462.

      Fig. 3B is discrepant with the figure legend. Figure shows CK1e but legend says kinase dead CK1D-K38R

      The captions to Figs. 3 and 4, as well as the references to these figures in the text, are correct. However, the actual Figs. 3 and 4 were inadvertently swapped during figure assembly. We apologize for this mix-up.

      The authors` interpretation of CK1 involvement in checkpoint is incorrect. The authors state that CK1 activity decreases p53 function promoting recovery, but Inuzuka et al (ref. 51) showed that inhibition of CK1 leads to this outcome.

      We thank the Reviewer for noting this mistake. We are no experts in p53 regulation, which is rather complex. CK1 decreases MDM2 stability and hence enhances p53 function.

      We corrected the statement and placed it in the right context: “CK1 phosphorylation triggers β-TrCP-mediated degradation of MDM2 and activates p53, thereby enhancing p53-dependent responses involved in checkpoint signaling and DNA repair (Inuzuka et al., 2010; Winter et al., 2004). After DNA repair, CK1δ has…”

      Reviewer #2 (Significance):

      It is generally assumed that CK1 is constitutively active, which is likely an oversimplified view; in a physiological context, some degree of regulation can be expected. Demonstrating that there are several pools of CK1 that are differently regulated at the level of protein stability during the cell cycle would be a significant advance in our understanding of CK1 functions.

      Reviewer #3 (Evidence, reproducibility and clarity):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U2OS cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin), and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      Overall, they propose that the activity and abundance of CK1 are regulated during the cell cycle. However, this claim would require several experiments to support it.

      Major comments:

      - As presented, some of the data are inconclusive. Co-staining with a centrosomal marker is required to determine whether or not Ck1 localises to the centrosomes. A large proportion of cells exhibit "potential" (their term) centrosomal staining, so a centrosomal marker is essential before any conclusions can really be drawn.

      In our revised Figure 1, we confirm this centrosomal staining using an antibody against pericentrin (PCNT).

      - Figures 3 and 4 have no loading controls, and these two figures have been mixed up in the text.

      The reviewer is right, we have corrected the Fig. 3 and Fig. 4 mix-up.

      Loading controls are now also provided.

      - A mobility shift on SDS-PAGE does not prove that a protein is phosphorylated. The authors should provide experimental evidence that the mobility shift is really due to phosphorylation. As they are inactivating phosphatases using CalA, it is likely the case, but they should prove it. Furthermore, the authors did not map any phosphorylation sites in this study, so they do not know whether CK1 phosphorylation occurs in the tail (as they assert) or elsewhere.

      We provide data in Fig. 3B showing that a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines does not undergo an electrophoretic mobility shift upon CalA treatment. These results demonstrate that the CalA-induced mobility shift is caused by phosphorylation of the CK1δ C-terminal tail.

      - The figure legends in general are limited and lack crucial information. For instance, in Figures 3C and 3D, how was the half-life of CK1 determined?

      We have adapted the figure caption:

      (C) Densitometric quantification of n=3 Western blots (see A) shown as mean ± SD. Overexpressed unphosphorylated CK1δ is degraded with a half-life of about 15 min. Both CalA and CalA + PF670462 treatments stabilize the kinase. (D) Densitometric quantification of n=3 Western blots (see B) shown as mean ± SD.

      - The CK1 regulatory model presented in Figure 7 is not supported by the data. What experimental evidence, for instance, shows that CK1 is inactive during mitosis? To make this claim the authors should directly assay its activity.

      We show that the majority of CK1δ is phosphorylated and therefore auto-inhibited during mitosis. In the original version, we referred to this fraction as inactive. In the revised manuscript, we refer to the phosphorylated kinase as auto-inhibited.

      Reviewer #3 (Significance):

      This study may be of interest to researchers working on cell cycle regulation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are largely supported by strong genetic and biochemical evidence, although some claims regarding the dispensability of fatty acid synthesis pathways remain incomplete. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We kindly ask that the editorial assessment be revised to focus on the major conclusions. Alternatively, we would respectfully suggest revising the second sentence to something akin to “The main conclusions are largely supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve.”

      Public Reviews:

      Reviewer #1 (Public review):

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

      We thank the reviewer for these positive comments.

      Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      To clarify, we conclude that the essential function of ACP in blood-stage P. falciparum parasites includes a critical stabilizing interaction with pyruvate kinase II. Apicoplast ACP presumably still plays a central biochemical role in FASII pathway function, but that role in FASII is dispensable for blood-stage parasites.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages. However, there are currently some important experimental issues/flaws, missing experiments that induced wrong interpretations and thus do not support some important claims of the study, notably for the role of FASII and the interaction between ACP and PKII.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We elaborate on these points and address the reviewer’s critiques below.

      Therefore, at this point, the study is only partial and would require major additions and/or important text edits/revisions before being considered for acceptance.

      We note that the manuscript has already been accepted for publication in accordance with the current eLife publishing model.

      Major points:

      From the graph of P. falciparum growth, we can see that in the lipid-rich condition, where both FabH KO and ACP KO can survive, the addition of mevalonate was essential for the growth of ACP KO. Along with the other evidence (PKII association, DNA levels...), we therefore agree that PfACP is involved in the mevalonate pathway.

      To clarify, our model is that ACP supports IPP synthesis by the apicoplast nonmevalonate/MEP pathway indirectly by stabilizing and thus supporting function by pyruvate kinase II that supplies the pyruvate and NTPs required for IPP synthesis by the MEP pathway.

      The authors claim that the FASII pathway is inactive/not essential in the P. falciparum blood stage. However, the authors have not shown any evidence on whether ACP is or not involved in the FASII pathway during the asexual blood stage.

      To clarify, there is overwhelming data in the prior published literature that we cite (including refs. 13, 14, and 32) to establish that FASII is dispensable for blood-stage Plasmodium growth in vivo in rodent parasites and in vitro culture in human parasites. Prior studies also strongly support a role for apicoplast ACP as the central scaffold for FASII-mediated acyl chain synthesis. However, our and prior studies support the conclusion that essential ACP function in blood-stage parasites is independent of its role in FASII.

      As currently designed, the experiments presented cannot conclude on that point for several reasons. Indeed, it was previously shown that (i) the expression of the protein from the FASII pathway are all present in blood stages and are significantly upregulated in patients that are under under "nutrient starvation" (Daily et al. Nature 2007), (ii) that, growing parasites under similar low lipid conditions in vitro induces an activation/upregulation of FASII, which can be measured by stable isotope precursor labelling and lipidomics (Botté et al. 2013).

      We are aware of these prior studies and cite and discuss the Botté et al. 2013 reference in our manuscript, which provided isotope-labeling evidence to support FASII activity in low-lipid growth conditions for P. falciparum. We note that neither study addresses whether FASII activity is required for growth in low-lipid conditions.

      (iii) that growing the PfFabI KO line under deprived lipid conditions leads to parasite death (Amiar et al. 2020), indicating that the FASII pathway can become critical, if not essential, depending on the host nutritionnal content together correlating patients' data and metabolic adaptation for the same reasons in the related parastie Toxoplasma gondii (Amiar et al. 2020, Krishnan et al. 2020, Liang et al. 2020, Primo et al. 2021, Charital et al. 2024, Dass et al. 2024, Bitew et al. 2025).

      All of the studies cited by the reviewer focus primarily or exclusively on Toxoplasma gondii parasites. We agree with the reviewer that these and other studies provide strong evidence that FASII activity contributes to growth of T. gondii parasites, including roles for apicoplast ACP that appear to differ from what we have unveiled for P. falciparum malaria parasites. We acknowledge and discuss these differences from T. gondii in the final section of the Discussion section and think that exploring these differences will be a fascinating area for future study.

      The Amiar et al. 2020 paper cited by the reviewer is the only study we are aware of that has directly tested the ability of a ∆FASII parasite (in this case, ∆FabI) to grow in low-lipid conditions. We acknowledge that they observed little to no growth of ∆FabI parasites in these conditions. Our growth assays with ∆ACP and ∆FabD parasites indicated a different outcome in which both WT and ∆FASII parasites grew similarly in low-lipid conditions. Our results thus contrast with the prior study. As noted below, the minimal lipid growth conditions explicitly reported in the methods section of the Amiar et al. paper are identical to those used in our study: fatty acid-free BSA, 30 µM palmitic acid, and 45 µM oleic acid (all sourced from Sigma) with daily media changes. Thus, the basis for these differences is unclear and additional follow-up work will be needed to explore and resolve these differences.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

      Here, the authors are expecting to show that FabH (and thus the FASII pathway) is not essential in an experiment that is not designed to be in low lipid conditions but rather in lipid rich conditions: Such high lipid conditions of culture in this study is granted by daily feedings with high fatty acid supplement (30-90 uM palmitic acid and 30-60 uM oleic acid). These fatty acid concentrations were used previously by Mitamura et al. (2005) and Miichi et al.(2007) to replace non-determined supplements such as Serum or Albumax supplement to grant similar growth by a completely controlled culture medium.

      This means the concentrations above do not represent limited fatty acid concentrations, especially not with daily feeding (representing an excess supplied amount of lipids, unlike regular 48h feedings) that allowed the authors to easily reach very high non-physiological parasitaemia of more than 20%!! Amiar et al. previously showed essentiality of FabI in P. falciparum in the limited fatty acid culture at a lower concentration (<30uM 16:0, <45um 18:1), than the Mi-Ichi et al. controlled medium with regular 48 h culture feeding. Therefore, with the current experimental settings, the FAH KO is placed in high lipid conditions, thus preventing any conclusion on its essentiality under low lipid conditions.

      The basis for the reviewer’s statements here is unclear, as this critique and the conditions it describes do not conform to the published conditions reported in the Amiar et al. 2020 paper. The methods section of that study for “Plasmodium falciparum growth assays” explicitly states (page e7):

      “Media was replaced daily, sub-culturing were performed every 48 h when required, and parasitemia monitored by Giemsa-stained blood smears. Growth assays in lipid-depleted media were performed by synchronizing parasites before transferring trophozoites to lipid-depleted media as previously reported (Botte ´ et al., 2013; Shears et al., 2017). Briefly, lipid-rich AlbuMAX II was replaced by complementing culture media with an equivalent amount of fatty acid-free bovine serum albumin (Sigma), 30 µM palmitic acid (C16:0; Sigma) and 45 µM oleic acid (C18:1; Sigma).”

      We used identical culture conditions to those described above: fatty acid-free BSA in place of lipid-rich AlbuMAX, 30 µM palmitic acid, and 45 µM oleic acid. We thus obtained growth results that contrast with the prior study and suggest that additional, future studies will be required to understand and resolve these differences.

      Furthermore, it is too uncertain to conclude that ACP is only essential for the mevalonate pathway.

      Please see our response above that clarifies our model for ACP function in supporting pyruvate kinase II and the many apicoplast pathways that appear to depend on PKII.

      This would be a similar discussion to the Yeh et al. 2011 and the Swift et al., where induced Apicoplast knockout caused parasites to require IPP to survive, but there were always remnant apicoplast vesicles and thus the putative presence of an active FASII in the parasite, where de novo fatty acid synthesis could be maintained.

      It is extremely unlikely that FASII remains active upon apicoplast disruption and loss of the apicoplast genome. The apicoplast-encoded SufB is lost upon apicoplast disruption and can no longer participate in making Fe-S clusters. Without Fe-S synthesis, the apicoplast lipoate synthase (LipA) cannot make lipoate to activate pyruvate dehydrogenase (E2 subunit) and produce the acetyl-CoA needed for FASII activity. There is no experimental evidence that FASII remains active upon apicoplast disruption and loss of the apicoplast genome.

      Amiar et al. (2020) and Krishnan et al. (2020) showed that disruption of FASII and absence of de novo FA synthesis in T. gondii could be compensated by the exogenous supplementation of myristic acid, C14:0.

      As explained above, we acknowledge that FASII contributes to Toxoplasma gondii growth and includes functions that appear to differ from P. falciparum.

      Here, high fatty acid supplementation using commercially available fatty acids may include unexpected fatty acid species such as myristic acid in palmitic acid or oleic acid, since all commercially available fatty acids guarantee only >99% but not 100%. If P. falciparum requires a very, very low amount of myristic acid to survive, the amount of possible contamination, like 1 nM, may be sufficient to maintain their survival. Thus, ACP and FabH might be very important to generate de novo fatty acids within parasites, but this was not shown by the authors.

      As noted above, we used identical culture conditions and commercial sources of defined fatty acids to those reported in the Amiar et al. study. We do not see a basis for the reviewer’s critique that the two studies utilized differing culture conditions. Nevertheless, we agree that future studies are needed to understand and resolve these differences.

      Therefore, the manuscript currently contains incorrect conclusions on the potential essentiality/use of FASII, against current experimental evidence.

      As explained above, we do not see a basis for the reviewer’s critique here or for viewing one study as more or less definitive than the other, as identical culture conditions were used yet contrasting results were obtained for reasons that remain uncertain. Future studies beyond the scope of the present manuscript will be required to fully understand and resolve these differences.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      We either request more solid experimental evidence showing the absence of fatty acid synthesis at low fatty acid conditions by re-doing the growth assay in the lower fatty acid feeding conditions without daily feeding to clarify if the ACP and FabH are essential in the blood stage growth, or not; as well as showing the absence of fatty acid synthesis at low fatty acid conditions using isotope labelled precursor. Without these, the authors cannot conclude on this important point. Alternatively, toning down the text to acknowledge the possibility of FASII being active and critical under certain conditions would be acceptable.

      We note that the Amiar et. al 2020 study cited by the reviewer reported growth assays for WT and ∆FabI NF54 P. falciparum in low-lipid conditions that similarly lacked direct tests of FASII activity by isotope labeling.

      We agree that our results, which utilized distinct ∆ACP and ∆FabD NF54 (PfMev) lines, contrast with the results and conclusions of the Amiar et al. study. We fully agree with the reviewer that future studies, utilizing tandem growth assays and isotope-labeling metabolic flux assays (e.g., mass spectrometry), will be required to fully test and understand the dependence of FASII activity on the lipid content of the growth medium and the functional dependence of P. falciparum growth on FASII activity in low-lipid conditions.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the wellestablished role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      We thank the reviewer for this important suggestion. To address this possibility, we examined epicardial cells in primordial CM-ablated hearts using the epicardial marker tcf21. Surprisingly, we found that epicardial cells rapidly expanded following primordial CM ablation, with an increase already detected at 5 days post-treatment. This increase persisted through at least 30 dpt, when epicardial cell abundance remained higher than that in control hearts. These findings suggest that the coronary vessel defects are unlikely to result from a reduction in epicardial cell number. In addition, we investigated whether loss of the primordial CM layer permits abnormal inward migration of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation. While we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype, our results indicate that the vascular defects are not attributable to reduced epicardial cell abundance or abnormal epicardial invasion. We have included these experiments in Page 10 and 11, and Fig. 3I-3J of revised manuscript.

      “Because epicardial cells play critical roles in coronary vessel development, we next asked whether the vascular defects observed following primordial CM ablation were secondary to alterations in the epicardium. We found that tcf21+ epicardial cells revealed an increase rather than a decrease in epicardial cell abundance in primordial CM-ablated hearts (Fig. 3I3J). These findings suggest that the impaired coronary vessel organization is unlikely to result from a reduction in epicardial cell number, although we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype.”

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      We appreciate the reviewer for this insightful suggestion. Because primordial cardiomyocytes form a continuous single-cell-thick layer at the ventricular surface, we examined whether their ablation affects the spatial distribution or promotes inward migration of epicardial cells or coronary endothelial cells. Using tcf21 and deltaC reporters, we carefully analyzed the localization of these cell populations within the myocardium. We did not observe any evidence of abnormal inward migration or ectopic localization of either epicardial cells or coronary endothelial cells following primordial CM ablation (Fig. 3J–3K). These results indicate that loss of the primordial CM layer does not disrupt tissue compartmentalization or lead to inappropriate cellular invasion into the myocardial interior. We have included these experiments in Page 11, and Fig. 3J-3K of revised manuscript.

      “In addition, because primordial cardiomyocytes form a continuous layer at the ventricular surface, we investigated whether their ablation permits abnormal invasion of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation (Fig. 3J-3K). Thus, loss of the primordial CM layer does not appear to disrupt tissue compartmentalization or permit abnormal cellular invasion into the myocardium.”

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4<sup>+</sup> cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      We thank the reviewer for this important suggestion. To further address the relationship between primordial cardiomyocytes and gata4<sup>+</sup> cardiomyocytes during heart development, we examined their spatial and cellular relationship in juvenile zebrafish hearts. Consistent with our observations during regeneration, we did not detect any overlap between phlda2<sup>+</sup> cardiomyocytes and gata4<sup>+</sup> cardiomyocytes in the juvenile heart (7-8 wpf). These results indicate that primordial CMs and gata4<sup>+</sup> proliferative CMs represent distinct cardiomyocyte populations during both heart development and regeneration. We have included these experiments in Page 13, and Fig.5E of revised manuscript.

      “Consistently, no overlap between phlda2<sup>+</sup> and gata4<sup>+</sup> cardiomyocytes was observed in the wpf juvenile zebrafish heart (Fig. 5E).”

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

      We appreciate the reviewer for this important suggestion. To determine whether ablation of primordial cardiomyocytes alone is sufficient to activate regenerative signaling, we treated adult phlda2:mCherry-NTR;gata4:EGFP fish and gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. Under these conditions, we did not observe any induction of gata4 expression following primordial CM ablation (Fig. S5). These results indicate that loss of phlda2<sup>+</sup> cardiomyocytes alone is not sufficient to trigger regenerative gata4 activation in the absence of injury, further supporting that primordial CMs are dispensable for activation of the regenerative response. We have included these experiments in Page 11 and Fig. S5 of revised manuscript.

      “To determine whether ablation of primordial cardiomyocytes is sufficient to activate regenerative signaling, we first treated adult phlda2:mCherry-NTR;gata4:EGFP fish and control gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. We did not observe any induction of gata4:EGFP expression following primordial CM ablation (Fig. S5), indicating that loss of phlda2<sup>+</sup> cardiomyocytes is not sufficient to trigger regenerative gata4 activation in the absence of injury.”

      Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive?

      We thank the reviewer for this important question. We found that efficient ablation of phlda2<sup>+</sup> primordial cardiomyocytes does not affect overall survival of zebrafish under standard laboratory conditions. Although these animals exhibit clear defects in cardiac structure and coronary vascular organization, they remain viable during the experimental period, indicating that the primordial CM layer is not essential for survival. While we did not assess detailed physiological parameters such as cardiac function, swimming behavior, or long-term fitness in this study, the observed structural abnormalities suggest that subtle functional consequences may exist. These aspects will be important directions for future investigation. We have included the description on page 14, paragraph 2.

      “Although primordial CM ablation does not affect survival under laboratory conditions, the observed structural defects may have functional consequences on cardiac performance and overall physiological fitness, which warrant further investigation.”

      Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days posttreatment when they were detected?

      We appreciate the reviewer for this important question. To determine whether primordial cardiomyocytes regenerate following ablation during development, we performed Mtz-mediated ablation in juvenile zebrafish (7–8 wpf) and examined the hearts at extended time points after treatment. We found that phlda2<sup>+</sup> cardiomyocytes did not recover even at 90 days post-treatment, indicating a persistent loss of this population following developmental-stage ablation (Fig. S6). Importantly, we further assessed whether the cardiac defects observed at earlier time points resolve over time. We found that the abnormalities in trabecular and compact myocardium, as well as coronary vessel organization, persisted at 90 days post-treatment and did not show evidence of recovery (Fig. S4C). These findings demonstrate that the observed defects are not transient developmental delays but represent long-lasting structural alterations of the heart following primordial CM ablation.

      Major Comments:

      (1) Figure 1: A more detailed characterization of the three CM populations would be helpful in the text as well as a new Supplemental Excel Data Sheet with the top unique genes expressed in each.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have expanded the description of the three cardiomyocyte populations in the Results section to provide a more detailed characterization. We have revised the description on page 8. In addition, we have added a new Supplemental Excel Data Sheet (Table S1) listing the top differentially expressed genes for each CM cluster.

      “Notably, Cluster 2 showed additional enrichment for pathways involved in mitochondrial respiratory chain assembly, ATP synthesis, TCA cycle, and ribosome biogenesis, suggesting a relatively higher metabolic and biosynthetic activity state. In contrast, Cluster 1 was enriched for GO terms associated with mitochondrial stress responses, protein degradation, and cytoprotective pathways, indicating a stress-adapted cardiomyocyte state. Cluster 3 showed reduced enrichment of metabolic pathways, consistent with an immature metabolic profile, and was further enriched for genes involved in muscle development and epithelial morphogenesis, suggesting a role in cardiac morphogenesis and tissue organization.”

      (2) Figure 3: Lower magnification views of control and MTZ-treated hearts are needed for all reporters shown (cmlc2, gata4, deltaC) at 10 days post-treatment. These "whole heart" views will enable the reader to get a gross sense of how disrupted heart development is following ablation of the primordial layer.

      We appreciate the reviewer for this helpful suggestion. We have added low-magnification whole-heart images for all reported markers (cmlc2, gata4, and deltaC) to better illustrate the overall cardiac morphology following ablation of the primordial layer. We have included these data in Fig. S3 of revised manuscript.

      (3) It should also be stated clearly in the Results section text what day post-fertilization Mtx treatment began. A section on Mtx treatment, including timing (what day was it applied and what day was it washed out) and dose, should be added to the methods section.

      We thank the reviewer for this helpful suggestion. In response, we have now clearly stated the timing of Mtz treatment in the Results section in page 9, 10 and 11 of revised manuscript.

      “To address the role of phlda2<sup>+</sup> cells during heart development, we performed the following experiments using a standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to 7-8 wpf juvenile zebrafish.”

      “To address the role of primordial cells during heart regeneration, we perform the below experiments using standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to adult zebrafish (4–6 months post-fertilization).”

      In addition, we have added a Mtz treatment section in the Methods, which now includes detailed information on dosage, duration, and washout schedule to ensure full reproducibility of the experiments.

      “Mtz treatment

      For conditional ablation of phlda2<sup>+</sup> cardiomyocytes, zebrafish expressing phlda2:mCherryNTR were treated with 10 mM metronidazole (Mtz) for 12 hours per day for three consecutive days. Fish were washed out and maintained in fresh system water after each daily treatment. For developmental analyses, juvenile zebrafish (7–8 weeks post-fertilization) were subjected to Mtz treatment as described above. Following completion of Mtz exposure, fish were maintained under standard conditions. Hearts were collected at 10 days post-treatment for assessment of gata4 activation, and at 30 days post-treatment for analysis of myocardial structure and coronary vessel development. For regeneration experiments, adult zebrafish (4–6 months old) were similarly treated with Mtz for three consecutive days with daily washout. Three days after the final Mtz treatment, ventricular apex resection was performed. Hearts were harvested at 7 days post-amputation (dpa) for analysis of gata4 activation and early regenerative responses, and at 30 dpa for evaluation of myocardial regeneration and coronary vessel revascularization.”

      (4) What happens to the heart 2 and 6 months post-treatment? Are there long-term consequences to primordial layer ablation or do the defects seen at 10 days post-treatment eventually resolve?

      We thank the reviewer for this important question. To determine whether the developmental defects observed following primordial CM ablation are transient or persist long term, we performed additional analyses at later time points after Mtz washout. First, we found that phlda2<sup>+</sup> cardiomyocytes failed to recover following Mtz-mediated ablation. In adult phlda2:mCherry-NTR fish, phlda2<sup>+</sup> cells remained absent 30 days after Mtz washout (Fig. 5B). Similarly, when juvenile fish (7–8 wpf) were treated with Mtz and subsequently allowed to recover, phlda2<sup>+</sup> cells were still not restored at 90 days post-treatment (Fig. S6). Second, juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated primordial CM ablation and analyzed 90 days after Mtz washout. We found that vascular abnormalities persisted long after ablation of primordial CMs (Fig. S4C). Coronary vessels remained disorganized and fragmented, indicating that the vascular phenotype does not resolve over time. Together, these findings demonstrate that primordial CM ablation causes long-lasting defects and that the abnormalities observed are not transient developmental delays. Instead, loss of primordial CMs results in persistent cellular and vascular defects that remain evident months after Mtz treatment. We have included these experiments in Fig. S4C and Fig. S6 of revised manuscript.

      “Notably, these vascular abnormalities persisted at 90 days post-treatment, indicating that the defects do not resolve during subsequent cardiac growth and maturation (Fig. S4C).”

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (5) Also, does the primordial layer come back in these animals where the lineage is ablated during development? Or is it permanently lost as shown in Figure 5 when it is ablated during adulthood?

      We appreciate the reviewer for raising this important question. To determine whether the primordial layer can be reestablished following ablation during development, we treated juvenile phlda2:mCherry-NTR fish (7–8 wpf) with Mtz and examined hearts 90 days after treatment. We found that phlda2<sup>+</sup> cardiomyocytes remained absent at this late time point (Fig. S6), indicating that the primordial layer does not recover following developmental-stage ablation. We have included these experiments in Page 12, and Fig. S6 of revised manuscript.

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (6) Figure 4: Need to show that phlda2 reporter fluorescence in lost/reduced following Mtz treatment during adulthood before apex amputation.

      We thank the reviewer for this important suggestion. We agree that confirming efficient ablation of phlda2<sup>+</sup> cardiomyocytes in adult fish prior to regeneration analysis is essential.

      In our study, we have already demonstrated in Fig.5B that Mtz treatment in adult phlda2:mCherry-NTR fish results in efficient and sustained loss of phlda2<sup>+</sup> cells, with no detectable recovery at 7 and 30 days post-treatment. These data confirm robust ablation of the primordial CM population following Mtz treatment in adults. Therefore, additional redundant imaging prior to apex resection was not performed in Fig 4.

      (7) Figure 5: It is interesting that primordial CMs do not regenerate following apex amputation or genetic ablation. This result suggests that primordial CMs are only important during development and dispensable during adulthood? This result also makes me question whether primordial CMs are actually required for heart development, which is why it is important to address whether the fish recovers.

      We appreciate the reviewer for this insightful comment. Our data indicate that primordial cardiomyocytes are essential for proper heart development, as their ablation during juvenile stages leads to significant structural and vascular abnormalities. Importantly, we further examined long-term outcomes and found that these defects do not resolve over time. Juvenile zebrafish subjected to primordial CM ablation failed to recover phlda2<sup>+</sup> cardiomyocytes even at 90 days post-treatment, and coronary vascular abnormalities also persisted at this late stage (Fig. S4C). These findings indicate that the observed developmental defects are not transient delays but instead reflect long-lasting structural alterations of the heart. In addition, we found that adult zebrafish similarly fail to regenerate primordial cardiomyocytes following genetic ablation (Fig. 5B), further supporting the limited regenerative capacity of this population. Together, these data demonstrate that primordial cardiomyocytes are required for proper cardiac development, and their loss leads to persistent defects that are not reversed during subsequent growth or regeneration.

      Minor:

      Line 234: Did the authors mean to write Cluster 3 (instead of Cluster 2)?

      We thank the reviewer for pointing out this error. We confirm that this was a labeling mistake, and “Cluster 3” is correct. The text has been corrected in the revised manuscript.

      Line 265: There is a typo of some sort in the phrase, "54.7% reduction closed to the ventricular wall".

      We thank the reviewer for pointing out this error. We have changed the description on page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      Is there a corollary lineage in mammals? This should be addressed in the Introduction or Discussion.

      We thank the reviewer for this insightful suggestion. At present, a direct corollary lineage to zebrafish phlda2<sup>+</sup> primordial cardiomyocytes have not been clearly defined in mammals. However, mammalian hearts also contain heterogeneous cardiomyocyte populations with distinct developmental states and metabolic profiles, including immature or embryonic-like cardiomyocytes that persist in specific regions during development and early postnatal stages. These populations may share functional similarities with the zebrafish primordial CMs in terms of developmental organization and maturation roles. We have now discussed this point in the Discussion and emphasized that whether a comparable lineage exists in mammals remains an important open question for future studies. We have included the discussion on page 15 of the revised manuscript.

      “Although a direct corollary of phlda2<sup>+</sup> primordial cardiomyocytes has not yet been identified in mammals, mammalian hearts contain heterogeneous cardiomyocyte populations with immature states. Whether these populations represent a functional equivalent of zebrafish primordial CMs remains an open question and needs further investigation.”

      Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

      We appreciate the reviewer for this important suggestion. In the revised manuscript, we have added supplemental data to support the scRNA-seq analysis, including gene expression tables for each cardiomyocyte cluster (Table S1), and full GO-term enrichment results (Table S2). In addition, we have revised the Results section to improve the clarity and description of the scRNA-seq findings. Please see detailed responses below for point-by-point clarification.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors did not provide enough data to demonstrate the quality of their scRNAseq. They only mentioned that they obtained "high-quality" transcriptomics. Specific parameters such as how many total reads and reads per cell should be provided.

      We thank the reviewer for this important suggestion. In the revised manuscript, we have added detailed sequencing quality metrics to the Methods section in Page 6 of the revised manuscript. The dataset contains 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.

      “The newly generated scRNA-seq data yielded 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.”

      (2) The authors utilized cmlc2:EGFP fish to perform scRNASeq. It will be helpful to include feature plots of cmlc2 and EGFP transcripts.

      We thank the reviewer for this suggestion. We have now included the description in Page 8, and feature plots of cmlc2 transcripts and EGFP reporter expression in the scRNA-seq dataset as a supplementary figure (Fig. S1A and S1B).

      “The expression of cmlc2 transcripts and EGFP reporter signal in the single-cell dataset further confirmed the enrichment of cardiomyocytes (Fig. S1A and S1B)”

      (3) The authors show that notch 3 is in cluster 3 of cardiomyocytes and suggest that this reflects elevated NOTCH signaling activity. The authors might consider using RNAScope to further validate that Notch 3 is expressed in cardiomyocytes. It will be also helpful to confirm phlda2 expression patterns during zebrafish heart development and regeneration by RNAScope.

      We appreciate the reviewer for this helpful suggestion. To further examine the spatial expression patterns of these genes, we analyzed publicly available spatial transcriptomic data from zebrafish hearts. We found that phlda2 is enriched in the outer region of the heart during both uninjured and regenerating conditions. Similarly, notch3 and actn1 were also predominantly localized to the outer heart region in the uninjured heart. These findings are consistent with our scRNA-seq results and support the spatially restricted signature of Cluster 3 cardiomyocytes. We have included these data in Page 8, and Fig. S2A-S2C of revised manuscript.

      “To further validate their spatial distribution, analysis of previously published spatial transcriptomic data revealed that phlda2, notch3, and actn1 were predominantly expressed in the outer region of the heart (Fig.S2A-S2C).”

      (4) The authors did not provide any data as supplemental tables to support their analyses of scRNAseq and GO-term analysis.

      We thank the reviewer for this suggestion. In the revised manuscript, we have added new supplemental tables providing full support for the scRNA-seq and GO-term analyses (Table S1 and S2), including lists of differentially expressed genes for each cardiomyocyte cluster and the corresponding GO enrichment results.

      (5) The description of the phenotype in Fig. 3A and B is not clear, especially for the sentence "trabecular muscle formation was severely impaired showing an approximately 54.7% reduction close to the ventricular wall.". The authors might consider using a bracket to show the compact muscle and trabecular muscle and the distance to the ventricular wall.

      We appreciate the reviewer for this suggestion. We have added brackets in Fig. 3A to label the compact and trabecular myocardium for improved clarity. Regarding “distance to the ventricular wall,” we found that this measurement varies substantially across different regions within the same heart, making a single distance-based metric unreliable. Therefore, we quantified trabecular muscle using the trabecular area fraction (trabecular area/total ventricular area) within a defined region of interest as a robust and unbiased indicator. The reported 54.7% reduction refers to this area fraction, and we have revised the text accordingly for clarity in Page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      (6) It is not clear how the authors quantify the vessel "length"/ventricular area and found that there is no difference (Fig. 3G). The vessel length is significantly shorter in the images (Fig. 3F) as the author also indicated that the vessels are fragmented.

      We thank the reviewer for this important comment. Coronary vessel “length” was quantified by selecting a fixed region of interest (ROI) within the ventricular area, followed by skeletonization of deltaC:EGFP<sup>+</sup> vessels using ImageJ. The total vessel length within the ROI was measured and normalized to the ROI area to obtain vessel length density. Although the representative images (Fig. 3F) show a more fragmented vascular pattern, the total summed vessel length within the defined region was not reduced. This indicates that primordial CM ablation primarily affects vascular organization rather than overall vessel length within the ventricular area. We have updated the figure legend for clarity of the revised manuscript

      “Vessel length was measured within a fixed region of interest (ROI) after skeletonization of deltaC:EGFP<sup>+</sup> vessels in ImageJ.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      We sincerely thank the reviewer for the meticulous evaluation of our revised manuscript and for the constructive and insightful comments.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      We appreciate the reviewer’s suggestion. In the revised manuscript, we have clarified that PRRT2-mediated regulation of Nav1.2, together with its regulation of Nav1.6, may contribute to neuronal and network excitability in the cortex. Please refer to Page 15, Line 439-440.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      We thank the reviewer for this comment. We have corrected the timescale description in the Discussion accordingly. Please refer to Page 13, Line 382.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".

      PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      We appreciate the reviewer’s constructive suggestion. We agree that PRRT2 is unlikely to be the sole regulator of neuronal excitability or Nav channel slow inactivation. In the revised Discussion, we have clarified that PRRT2-dependent regulation of Nav channel slow inactivation represents one mechanism among several that fine-tune neuronal excitability. Other mechanisms, including regulation of potassium channels such as Kv7.2/Kv7.3, and other ion channel- or signaling- dependent pathways, may also contribute to excitability control in both PRRT2-positive and PRRT2-negative neurons. We have revised relevant paragraph to provide a more balanced discussion of alternative mechanisms. Please refer to Page 15, Line 430-438.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last compound APs in panels B and C.

      We appreciate the reviewer’s valuable suggestion. We have now added representative traces of the first and last compound action potentials, corresponding to the 1st and the 100th responses, respectively, to panels B and C of Figure 7-figure supplement 2. Please refer to Page 51, Figure 7-figure supplement 2.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Page 19, line 546: a sampling rate of 10kHz was used for recording Nav currents. Because the Nav channel isoforms examined in this manuscript (Nav1.1, Nav1.2, Nav1.4, Nav1.5, and Nav1.6) exhibit extremely rapid activation and inactivation kinetics, the temporal resolution provided by a 10 kHz sampling rate may not be optimal for detailed kinetic analysis. Although I am not requesting additional experiments, the authors may wish to choose a higher sampling rate (e.g., 50 kHz) in future Nav channel studies.

      We thank the reviewer for this helpful suggestion regarding the sampling rate for Nav current recordings. We will adopt higher sampling rates for rapid kinetic analyses of sodium currents in our future work.

      (2) Typo: Page 24, line 712-713: should "a Digidata (Molecular Devices, 1332A)" be "... 1322A"?

      We thank the reviewer for pointing out this typo. We have corrected “Digidata 1332A” to “Digidata 1322A” in the relevant Method section of revised manuscript. Please refer to Page 25, Line 722.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We again thank the editor and reviewers for their detailed attention to our work. In our previous revisions we endeavored to address the principal concerns raised by reviewers that we were capable of addressing. We recognize that asymptomatic pertussis transmission represents a particularly thorny area of epidemiology and public health, where multiple (and sometimes overlapping) mechanisms have been proffered even as empirical evidence remains thin, particularly in low-resource settings such as sub-Saharan Africa. A key finding of our work is that prospective surveillance in such a low-resource setting revealed abundant evidence of otherwise unobserved asymptomatic incidence. Moreover, as we note in our prior revisions, this finding is supported by recent work in South Africa and elsewhere (Kayina et al., 2015; Moosa et al., 2019, 2025). As such, we believe that further prospective surveillance in similar settings would be highly informative, a point that we have sought to emphasize in our present manuscript (and associated commentary).

      While we broadly agree with many of the concerns raised by the reviewers, we believe that we are unable to significantly strengthen the present work through further revisions. We do, however, wish to respond to several points raised in these reviews. Of note, a reviewer raises the possibility of a "breakthrough strain" without reference to existing literature. We agree that we cannot test this hypothesis, and though it is not incompatible with our own findings, it is, however, not consistent with recent molecular surveillance in South Africa (Moosa et al., 2023). The reviewer also raises the potential of low adult vaccination coupled with recent reintroduction. This hypothesis relies on our investigation looking "at the right place and the right time", and further does not explain how immunologically naive adults would have escaped morbidity. We have adopted what we believe is a more parsimonious interpretation of our results (i.e., that asymptomatic infection represents evidence of previous immune exposure), though we agree that a more thorough exploration of this particular issue is warranted, particularly in light of our persistently colonized mothers. In addition, we have noted similar studies in sub-Saharan Africa that also found widespread evidence of asymptomatic pertussis, which we believe is inconsistent with a “right time, right place” interpretation.

      A reviewer also pointed to the work of Warfel et al. (2014) and Althouse and Scarpino (2015). We are familiar with both of these studies and agree with their broad relevance to the field (our apologies for omitting Warfel et al.). The reviewer states that, "If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research." Critically, we believe that our prospective field study of human patients in a real-world public health system complements previous research, including animal trials and simulation studies. Simply put, given that our study was unable to establish the prior vaccination or exposure status of participants, we do not claim to have shown evidence for transmission despite wP vaccination or prior infection.

      Regarding Althouse & Scarpino (2015), we believe that, for the majority of readers, the most compelling analysis in their paper was the examination of genome sequences that pointed to substantial asymptomatic transmission in the US. This conclusion emerged from their population model, which required that “births” (representing transmission events) exceeded “deaths” (representing recovery of infectious individuals) in order to be consistent with the sequence data. Unfortunately, this paper does not provide a detailed explanation of their methods and data sources, nor is this work directly reproducible through, for example, an open-access code/data repository. We have explored the availability of US genome sequences over the time period of their study and were able to find only 36 sequences: 2 from the pre-vaccine era, 8 from the wP vaccine era, and 26 from the aP vaccine era. Given this notable imbalance in the number of sequences (and thus sequence diversity) that was biased in favour of the most recent time period, is it then surprising that the “birth rate” in their model had to exceed the “death rate” in order to match the genetic diversity in the data? Based on a careful inspection of this work, we do not consider its conclusions to represent a gold standard against which all subsequent studies should be judged. We also note that genomic surveillance and analysis of pertussis remains sparse relative to other fields, though recent works have added dramatically to the corpus of available sequences (Bridel et al., 2022).

      Finally, we note that our previous revisions addressed several concerns raised in the present reviews. For example, we previously sought to address reviewers' about our presentation of the strength of our evidence. In this regard, we broadly agree with the reviewers, and we now state that our results "suggest that pertussis transmission occurs between minimally symptomatic mothers and their newborn infants." We believe this largely addresses a present reviewer's concern that 'the mother-to-infant transmission pathway should be framed as "highly suggestive" rather than "confirmed"'. We also note that our results examine three different threshold Ct values (survival analysis, Fig 4), a point that we believe partially addresses a reviewer's suggestion to "including a sensitivity analysis using a stricter cut-off" and concerns about "the decision to use a Ct<45 threshold, as this is higher than standard clinical cut-offs". Indeed, we discuss the issue of clinical cut-offs (and their appropriateness) at some length in the section, "Test reliability, disease surveillance, and public health where we state, "We recognize that such weak and potentially ambiguous signals may not be appropriate for clinical diagnosis. However, our results demonstrate that they nonetheless contain valuable information about pathogen presence and infection intensity that can (and should) be leveraged for disease surveillance." We have also included in the present work a detailed discussion of qPCR sensitivity and efficiency that we believe should interest others working in pertussis surveillance.

      We do not view our own research as the "last word" in this rather controversial subject. In this spirit, we have attempted to present our work transparently, state our claims carefully, and underscore future activities that we believe would benefit the pertussis research community going forward. For example, we agree that further attention to shared exposure and functional immune data among low-resource communities could provide valuable insights into epidemiology and ecology of pertussis. However, we also believe the trade-offs of including one set of activities over another should be clearly acknowledged by researchers, clinicians, and public health officials. To simply state that we must measure more fails to account for the very real resource constraints that we all face.

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC Medicine, 1–12. https://doi.org/10.1186/s12916-015-0382-8

      Bridel, S., Bouchez, V., Brancotte, B., Hauck, S., Armatys, N., Landier, A., Mühle, E., Guillot, S., Toubiana, J., Maiden, M. C. J., Jolley, K. A., & Brisse, S. (2022). A comprehensive resource for Bordetella genomic epidemiology and biodiversity studies. Nature Communications, 13(1), 3807. https://doi.org/10.1038/s41467-022-31517-8

      Kayina, V., Kyobe, S., Katabazi, F. A., Kigozi, E., Okee, M., Odongkara, B., Babikako, H. M., Whalen, C. C., Joloba, M. L., Musoke, P. M., & others. (2015). Pertussis prevalence and its determinants among children with persistent cough in urban Uganda. PLoS One, 10(4), e0123240.

      Moosa, F., du Plessis, M., Weigand, M. R., Peng, Y., Mogale, D., de Gouveia, L., Nunes, M. C., Madhi, S. A., Zar, H. J., Reubenson, G., & others. (2023). Genomic characterization of Bordetella pertussis in South Africa, 2015–2019. Microbial Genomics, 9(12), 001162.

      Moosa, F., du Plessis, M., Wolter, N., Carrim, M., Cohen, C., von Mollendorf, C., Walaza, S., Tempia, S., Dawood, H., Variava, E., & others. (2019). Challenges and clinical relevance of molecular detection of Bordetella pertussis in South Africa. BMC Infectious Diseases, 19, 1–11.

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018: Findings from the PHIRST study. Journal of Infection, 106550.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787–792. https://doi.org/10.1073/pnas.1314688110


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants, in particular. The authors used a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka, Zambia, and they utilized qPCR-based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants, thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high-burden low-resource settings, and they advocated for enhanced molecular surveillance strategies to capture full pertussis infection, including those that might have gone undetected.

      Strengths:

      The strengths are the use of innovative study design, especially the longitudinal approach and routine sampling, rather than symptom-driven testing that minimizes bias in the study. The methodology was also rigorous and transparent by evaluating the IS481 signal strength to classify pertussis detection and conducting retesting to assess qPCR reliability. There were also important epidemiological insights, and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low-resource settings.

      Weaknesses:

      These include reliability on qPCR-based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanations for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there is limited generalizability as the study was done in a single urban site in Zambia. There is also a lack of functional immune data.

      Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Weaknesses:

      While I am quite enthusiastic about the work, I am concerned that a number of likely relevant confounders were not discussed and that the broader implications of their findings were not well grounded in the existing literature. For example, I could not find information on the vaccination status of the mothers in the study. Given the conclusions about asymptomatic transmission and the durability of immunity, it is important to know the vaccination status of the mothers. Moreover, did the authors have other metadata on the mother/infant dyads, e.g., household size, vaccination status of household members, etc.? Given the potential implications of more widespread asymptomatic transmission associated with pertussis infection, I believe the authors should better couch their results in the context of the broader debate around asymptomatic transmission.

      We appreciate the reviewers' detailed feedback. We provide an overview of our responses here and we address specific recommendations below. In light of reviewers’ comments, we have revised our manuscript in order to improve the clarity of our presentation and to better situate our results within the context of the existing literature. Unfortunately, as the field study has been concluded, many of the reviewers’ recommendations are not possible. These include additional testing (i.e., culture or serology) or sequencing. We have updated the manuscript to more clearly indicate our knowledge regarding maternal vaccine status and immunological immunity of study participants. We have also provided a more comprehensive overview of existing pertussis studies, including genomic surveillance and details regarding sub-Saharan Africa and Zambia in particular. Finally, we have revised the formatting of Figure 4 (survival analysis) to more clearly highlight differences between mothers and infants and to better align with the text, and note that the underlying results are unchanged.

      A particular concern raised in the reviews that we wish to address is the recommendation of culture- or serology-based tests as "confirmatory". We have revised the manuscript in light of this feedback to better reflect our own position on this matter. We believe these recommendations do not adequately account for important trade-offs between testing sensitivity and specificity that are widely recognized in both clinical practice and epidemiology (Enøe et al., 2000; Florkowski, 2008; Swift et al., 2020). When the results of different testing methodology disagree, rarely is one method, a priori, correct. Rather, the disagreement may point to specific test limitations or important biological questions about the study system.

      In the case of pertussis detection, cell culture is recognized for its very low sensitivity, while serological detection is complicated by debate around appropriate threshold levels and time horizons for seroconversion and subsequent decay (Lee et al., 2018; van der Zee et al., 2015). Furthermore, while anti-PT antibodies are a common target of serological detection, these are not reliable correlates of protection (Mills, 2001; Wilk et al., 2019), nor are they reliably generated in response to colonization (de Cellès & Rohani, 2024; Graaf et al., 2020). Overall, the detailed relationship between exposure, carriage, transmissible infection, and the dynamics of anti-PT serology remains poorly characterized (Craig et al., 2020; de Cellès et al., 2025).

      While we agree that these are important questions in epidemiology and public health, we nonetheless wish to highlight that there is no “free lunch": each additional test and protocol comes with additional cost and complexity that should be evaluated based on the specific goals of the intended surveillance. In our case, the repeated sampling of longitudinal surveillance serves as a low-cost "confirmatory" testing regime. We disagree that cell culture would have added value to the present study and would not recommend its addition to future studies (primarily due to low sensitivity). While we agree that before-and-after serology of mothers would have added important context to the present study, we nonetheless expect that significant ambiguity would have surrounded any such results (e.g., Moosa et al. (2025)).

      One area that we strongly agree warrants further attention is the household dynamics in pertussis transmission, particularly in low-resource settings where crowding is common. In our study we were not able to rule out environmental and/or shared transmission events, though our survival analysis did demonstrate a greater impact of mothers on infants than vice versa, results which suggest a causal mechanistic role. In previous studies we detailed the demographics of household size, number of children, and mothers' age (Gill et al., 2021; Gunning et al., 2020), though we have not conducted formal analyses of these important covariates here. We also note that the code and data are freely available, allowing for others to build on our work.

      We believe that future prospective studies are an invaluable tool for directly tracking pertussis disease transmission, including both community and household studies. We have argued here for the value of qPCR-based community surveillance, which could integrate into existing public health activities. Regarding household studies, we note that a key challenge in implementing these studies is selecting an appropriate sampling interval and duration to best capture epidemiological linkages. Our results suggest that qPCR-based real-time population-level surveillance could be used to initiate such a prospective household study during a pertussis outbreak so that a higher sampling frequency (e.g., weekly swabs) could be gainfully employed over a shorter time period.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Enhance Validation of qPCR Findings:

      We address these comments above at greater length. We also note that the term “false positive” is rather ambiguous here, as we lack a clear distinction between carriage and transmissible infection. We note that our manuscript includes a considerable discussion of qPCR validation, including negative controls and sample retesting. We agree that test sensitivity and specificity remains an important question, and have endeavored to clearly indicate these concerns throughout the manuscript.

      (2) Clarify Transmission Dynamics:

      While we agree that this is an important question, we lack the relevant sequence data to test it. Notably, we suspect that any such phylodynamic linkage would require genomic sequencing due to the relatively low genetic diversity observed in pertussis (see, e.g., population-level estimates of time to most common ancestor in Lefranc et al., (2022). However, sequencing pertussis genomes remains resource-intensive, and we expect that the deployment of such sequencing at scale would be cost-prohibitive in low-resource settings.

      We have revised our manuscript to underscore uncertainties around shared/household exposure. We also direct the reviewer’s attention to our survival analysis, where a notable asymmetry exists between mothers and infants. Here, mothers’ prior qPCR signals exhibit a larger impact on their infants than infants on their mothers (Fig 4). If shared exposure was the principal cause of the observed increase in hazard in ego from alter, then we would expect no such asymmetry between mothers and infants.

      (3) Expand Discussion on Public Health Implications:

      As noted above, we have revised our introduction and expanded our discussion to better account for the existing literature. And, while we are hesitant to put forward specific recommendations based on our study, we feel confident in stating that active pertussis surveillance in low-resource settings is A) almost entirely absent B) possible to achieve, and C) necessary to resolve long-standing questions about pertussis epidemiology at regional and national levels. We have endeavored to clarify these points, particularly within the discussion.

      (4) Address the Role of Immunity More Directly:

      While we lack such immune data, we have pointed to recent work from the notable PHIRST study in South Africa, as well as highlighted ambiguities surrounding these data.

      Reviewer #1 (Recommendations for the authors):

      (1) Do we know the vaccination status of the mothers in the study? … If these data are not available, I think that the paper must be re-framed to acknowledge that all the conclusions are statistical in nature, based on publicly available vaccine coverage data from Zambia.

      We do not have information on the immunization status of mothers, though we cite national rates for Zambia across the relevant time period. We have clarified this point in the revised manuscript. We have endeavored to clearly acknowledge that many of our conclusions are statistical in nature and to clearly quantify the strength of evidence.

      We strongly disagree that anomalously low vaccination rates amongst mothers (i.e., relative to national averages) would materially alter the interpretation of our findings.

      Overall, our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity. Indeed, some have argued that the preponderance of mild/asymptomatic infections in mothers is, of itself, evidence of prior immunological exposure (Fine & Clarkson, 1982).

      (2) Do we know anything about rates of pertussis in Zambia, especially in the study site?

      We address this important question in the discussion. In particular, we state that: “As a populous, middle-income and primarily urban country, Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009”.

      (3) I couldn't find information in the paper related to the severity of infection in the infants. It's mentioned in the section describing results in Figure 5, but I only saw analyses with symptoms (as opposed to severe symptoms). Do you have outcome data from infants testing positive?

      This question was addressed in more detail in our previous work (Gill et al., 2021), which we briefly summarize in the Introduction. We also show the frequency of severe symptoms in Fig 5 (bottom panel), and detail mild versus serious symptoms in our subgroup analysis (Fig 7D).

      (4) Do we know anything about vaccine-resistant strains of pertussis in Zambia?

      We are not aware of any such work. As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa.

      (5) While I believe sequencing is beyond the scope of the current study, the authors should comment on the potential utility of sequencing elements of the pertussis genome and use that to demonstrate causality and direction of transmission more strongly.

      We believe that existing literature has addressed the potential of sequence data and phylodynamics to infer transmission, particularly for pathogens with high mutation rates such as RNA viruses. To date, research into the phylodynamics of pertussis has focused exclusively on population-level dynamics (Lefrancq et al., 2022), where estimates of time to most recent common ancestor (TMRCA) are long, indicating low genomic variability at the scale of countries and years. To our knowledge, no work on pertussis has directly inferred transmission chains from sequence data. Given the existing evidence, we expect that any such work would require genome-level sequencing, which would likely be cost-prohibitive in low-resource settings.

      (6) … However, it would be helpful to understand more about how your results fit into the broader story around pertussis resurgence. … if the infant cases were all mild, they might never have been captured in surveillance data sets.

      We believe that a key result of our study is the remarkable mismatch between country-level symptoms-based surveillance and prospective surveillance, which demonstrates that such mild cases have almost certainly not been captured. These findings are mirrored by recent work in South Africa (now addressed in our Discussion, see Moosa et al. (2025)). We believe that prospective surveillance, particularly in under-surveilled regions, is critical to understanding pertussis transmission writ large, which we have attempted to communicate throughout our discussion.

      (7) Relatedly, if there are still high rates of asymptomatic mother-to-infant transmission with whole cell vaccination, then why is there an observed drop in infant pertussis following whole vaccination in most countries?

      In previous work, we demonstrated that some infants in this cohort exhibited asymptomatic infection (Gill et al., 2021). We note that a drop in pertussis incidence amongst infants after the roll-out of the whole-cell vaccine is not contradictory with our findings. We want to clarify that our results, and evidence that mother-to-infant transmission can occur, does not imply that the whole-cell vaccine fails to protect against transmission.

      We have previously used epidemiological evidence to infer the population-level impacts following the roll-out of whole-cell pertussis infant immunization. For example, we observed an increase in the inter-epidemic period that, together with the drop in infant cases, are consistent with a reduction in transmissible infections (Broutin et al., 2010; Rohani et al., 2000).

      References

      Broutin, H., Viboud, C., Grenfell, B. T., Miller, M. A., & Rohani, P. (2010). Impact of vaccination and birth rate on the epidemiology of pertussis: A comparative study in 64 countries. Proceedings of the Royal Society B: Biological Sciences, 277(1698), 3239–3245.  https://doi.org/10.1098/rspb.2010.0994  

      Craig, R., Kunkel, E., Crowcroft, N. S., Fitzpatrick, M. C., Melker, H. de, Althouse, B. M., Merkel, T., Scarpino, S. V., Koelle, K., Friedman, L., Arnold, C., & Bolotin, S. (2020). Asymptomatic Infection and Transmission of Pertussis in Households: A Systematic Review. Clinical Infectious Diseases, 70(1), 152–161. https://doi.org/10.1093/cid/ciz531

      de Cellès, M. D., & Rohani, P. (2024). Pertussis vaccines, epidemiology and evolution. Nature Reviews Microbiology, 1–14. https://doi.org/10.1038/s41579-024-01064-8

      de Cellès, M. D., Wong, A., Dalby, T., & Rohani, P. (2025). Natural immune boosting biases pertussis infection estimates in seroprevalence studies. Nature Communications, 16(1), 8883. 

      Enøe, C., Georgiadis, M. P., & Johnson, W. O. (2000). Estimation of sensitivity and specificity of diagnostic tests and disease prevalence when the true disease state is unknown.  Preventive Veterinary Medicine, 45(1–2), 61–81.

      Fine, P. E. M., & Clarkson, JacquelineA. (1982). The recurrence of whooping cough: Possible implications for assessment of vaccine efficacy. The Lancet, 319(8273), 666–669.  https://doi.org/10.1016/S0140-6736(82)92214-0 

      Florkowski, C. M. (2008). Sensitivity, specificity, receiver-operating characteristic (ROC) curves and likelihood ratios: Communicating the performance of diagnostic tests. The Clinical Biochemist Reviews, 29(Suppl 1), S83.

      Gill, C. J., Gunning, C. E., MacLeod, W. B., Mwananyanda, L., Thea, D. M., Pieciak, R. C., Kwenda, G., Mupila, Z., & Rohani, P. (2021). Asymptomatic Bordetella pertussis infections in a longitudinal cohort of young African infants and their mothers. eLife, 10, e65663. https://doi.org/10.7554/elife.65663

      Graaf, H. de, Ibrahim, M., Hill, A. R., Gbesemete, D., Vaughan, A. T., Gorringe, A., Preston, A.,  Buisman, A. M., Faust, S. N., Kester, K. E., Berbers, G. A. M., Diavatopoulos, D. A., & Read, R. C. (2020). Controlled Human Infection With Bordetella pertussis Induces Asymptomatic, Immunizing Colonization. Clinical Infectious Diseases: An Official Publication of the Infectious Diseases Society of America, 71(2), 403–411.  https://doi.org/10.1093/cid/ciz840

      Gunning, C. E., Mwananyanda, L., MacLeod, W. B., Mwale, M., Thea, D. M., Pieciak, R. C., Rohani, P., & Gill, C. J. (2020). Implementation and adherence of routine pertussis vaccination (DTP) in a low-resource urban birth cohort. BMJ Open, 10(12), e041198.

      Lee, A. D., Cassiday, P. K., Pawloski, L. C., Tatti, K. M., Martin, M. D., Briere, E. C., Tondella, M. L., Martin, S. W., & Group, C. V. S. (2018). Clinical evaluation and validation of laboratory methods for the diagnosis of Bordetella pertussis infection: Culture, polymerase chain reaction (PCR) and anti-pertussis toxin IgG serology (IgG-PT). PLoS One, 13(4), e0195979.

      Lefrancq, N., Bouchez, V., Fernandes, N., Barkoff, A.-M., Bosch, T., Dalby, T., Åkerlund, T.,  Darenberg, J., Fabianova, K., Vestrheim, D. F., Fry, N. K., González-López, J. J.,  Gullsby, K., Habington, A., He, Q., Litt, D., Martini, H., Piérard, D., Stefanelli, P., … Brisse, S. (2022). Global spatial dynamics and vaccine-induced fitness changes of Bordetella pertussis. Science Translational Medicine, 14(642), eabn3253.  https://doi.org/10.1126/scitranslmed.abn3253 

      Mills, K. H. G. (2001). Immunity to Bordetella pertussis. Microbes and Infection, 3(8), 655–677. https://doi.org/10.1016/s1286-4579(01)01421-6 

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018:  Findings from the PHIRST study. Journal of Infection, 106550. 

      Rohani, P., Earn, D. J., & Grenfell, B. T. (2000). Impact of immunisation on pertussis transmission in England and Wales. The Lancet, 355(9200), 285–286. https://doi.org/10.1016/S0140-6736(99)04482-7 

      Swift, A., Heale, R., & Twycross, A. (2020). What are sensitivity and specificity?  Evidence-Based Nursing, 23(1), 2–4.

      van der Zee, A., Schellekens, J. F., & Mooi, F. R. (2015). Laboratory diagnosis of pertussis.  Clinical Microbiology Reviews, 28(4), 1005–1026.

      Wilk, M. M., Allen, A. C., Misiak, A., Borkner, L., & Mills, K. H. G. (2019). The immunology of Bordetella pertussis infection and vaccination. In Pertussis: Epidemiology, Immunology & Evolution. Oxford University Press.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself, and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper, showing the membrane-bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Thank you for the positive feedback.

      Major concerns:

      My biggest concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors' main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is, of course, the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this, any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help Figure 1 is an evolutionary conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This will also help orient the reader to the identity of the mutated residues assayed in Figure 3.

      We agree that crosslinking can capture non-physiologically relevant interfaces. However, because we do not observe any crosslinking between Gβγ and a PLCβ3 variant that retains a cysteine in the X–Y linker or between PLCβ3 and any other cysteine in the Gβγ heterodimer, we believe it is site-specific.

      The question about sequence conservation in the Gβγ–PLCb3 interfaces is interesting and we have included this information in Figures S8 and S9.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent Tesmer structure with PI3K.

      We agree that the orientation of Gβ in the crosslinked structure is different. We include a comparison of this reconstruction to other published Gβγ–effector complexes as Figure S6.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated, as of course that can affect any of the three surfaces. My main question is that, from the way Figure 2A is oriented, the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      Thank you for pointing this out, and we expanded our analysis to include this residue. The R199A and R199E mutations had no defects in basal or Ga<sub>q</sub>-stimulated activities. However, R199A had 3-fold lower activation by Gβγ, while R199E was not activated in this assay. The R199E mutation also had significantly decreased agonist-dependent BRET with Gβγ and decreased PI(4,5)P2 hydrolysis, while retaining robust recruitment to the plasma membrane by Gα<sub>q</sub>. These data is included in the main text and Figures 2-4.

      (4) To help the reader's interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Thank for the suggestion. In revised Figure S3, we show the Gβγ–PLCb3 D892-PH<sub>cys</sub> complexes determined in this study at different contour levels.

      Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify and map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on the membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification. The manuscript is generally strong, but some issues need to be addressed.

      Thank you for the positive comments.

      Major comments:

      (1) BMOE/BM(PEG)2 crosslinking may enforce a non-native docking geometry, potentially compromising the physiological relevance and precision of the Gβγ-PLCβ3 interface as described. Although a >50% 1:1 crosslinked complex is formed and remains active, the solution maps show lower local resolution for Gβγ, consistent with a dynamic, potentially heterogeneous, interface. One interface is captured via a single engineered cysteine pair (PLCβ3 E60C-Gβ C271), which could potentially bias the pose. It would be helpful if the authors could provide additional orthogonal support (e.g., alternative crosslinked sites) and bolster the clarification of its uniqueness and relevance.

      We did attempt to isolate other crosslinked complexes. PLCβ3-D892 self-crosslinked under all reaction conditions, while PLCβ3-D892 XY<sub>Cys</sub>, which retains an endogenous cysteine within the X–Y linker (C516), did not result in any crosslinked product when incubated with Gβγ. Only the PLCβ3-D892 E60C crosslinked to Gβγ. With the exception the C68S mutation at the C-terminus of Gg to eliminate its prenylation site, all endogenous cysteines were retained in both Gβ and Gγ. Indeed, Gβ contains two solvent-exposed cysteines in its canonical effector binding surface (C204 and C271), but we did not observe any crosslinker density involving C204. While we cannot exclude the possibility that crosslinking occurred between PLCβ3-D892 E60C and other residues in Gβγ, we were unable to identify any 2D classes corresponding to these alternative conformations. These observations, together with the high efficiency of crosslinking, are consistent with a stable and persistent interaction.

      (2) In the crosslinked structure, the authors report that GβD228 interacts with PLCβ3 R199 and K183. In Figure 2A, R199 appears closer to Gβ D228 than K183, yet only K183 is functionally tested. Testing R199 (e.g., R199E/R199A) would strengthen the structure-guided validation of this interface.

      We agree, and functional analysis of PLCb3 R199E is included in the revised manuscript (see Figures 2-4).

      (3) The mutagenesis strategy appears inconsistent across figures/assays, which makes it difficult to interpret phenotypes and directly link the functional data to the proposed interfaces. For example, in Figure 2E, we see R185L but R215E, while residue L40 is mutated to Gly in the IP accumulation assays but to Glu/Lys (L40E/K) in the BRET assays (Figures 3B/3D/3F). The authors should (i) clearly justify the rationale for each substitution (conservative vs charge-reversal, interface disruption, etc.) and (ii), where possible, test the same mutants across assays (or provide evidence that alternative substitutions yield consistent conclusions).

      Mutagenesis experiments were initially carried out independently in the Lambert and Lyon Labs. As the study progressed, additional mutants were identified and/or designed based on results from both groups. The residues subject to mutagenesis are overall consistent across the different assays, with differences in the identity of the mutation varying in some cases. The L40G mutation is one such example, where given its modest impact on Gβγ-mediated activation in the IP accumulation assay, more impactful changes were made (L40E and L40K) for the BRET and signaling assays. In the revision, we now state that mutations were designed to maximally disrupt the three observed interfaces, such as by changing the size of the side chain and/or introducing charge reversal mutants.

      Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solutions and cryoEM studies of PLCβ3 / Gβγ, trying to understand the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate, although this is not completely convincing. Although these sites may not naturally be operative, the authors might want to develop the potential role of these sites.

      The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation, in accord with other PLCβ enzymes, but not for PLCβ3, and again, the authors might want to develop this point further.

      Thank you for the suggestions. We are investigating whether the other PLCb isoforms contain multiple Gβγ binding sites and the relative importance of preactivation by Ga<sub>q</sub> for a manuscript in preparation.

      Weaknesses:

      (1) I'm confused as to why the authors feel that their mechanism is distinct from the two-state enzyme, the synergistic activation proposed by Ross in 2011, using a primarily thermodynamic argument. As written, the authors appear to be very reliant on structural and BRET studies that do not give the details that would disprove this interpretation. The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits.

      The reconstitution experiments are under extremely artificial conditions, using nM-µM of purified proteins and liposomes that contain up to 30% PI(4,5)P2. Under these conditions, we think the increased activity is due to interfacial activation promoted by Gβγ binding to the lipase once it is associated with the liposome surface. This would be sufficient to account for the dose-dependent increase in both PLCb2 and PLCb3 activity as a function of Gβγ concentration. Given the higher basal activity of PLCβ2 and its decreased sensitivity to activation by Ga<sub>q</sub>, one possible explanation is that this isoform differs in its autoinhibition and/or structure of its proximal CTD that Ha2’ displacement is not a prerequisite for activation. In addition, Gβγ may also be a direct activator of PLCβ2. Further studies, ideally in cell-based systems, are needed to answer these questions.

      (2) In a recent study, McKinnon presents a model showing that Gαq and Gβγ activate PLCβ3 by two distinct pathways and that activation by Gβγ occurs through membrane recruitment. It is not surprising that the authors find that this is not true since the pelleting method used by McKinnon is subject to error. The authors should directly address the limitations of this previous work and the changes in proteoliposomes with sedimentation that alter partition coefficients. Although the inability of Gβγ to drive membrane binding is in accord with the quantitative studies of Scarlata, showing that the affinity of PLCβ3 to Gβγ is fairly weak as compared to the intrinsic membrane partition coefficient.

      We have added some of the limitations of proteoliposome sedimentation experiments to the discussion.

      (3) It was proposed many years ago that in signaling complexes Gαq - Gβγ may not have to fully dissociate when binding PLCβ, but rather shift their relative orientation when binding to PLCβ to allow activation. Is their model consistent with this? Is it possible that PLCβ3 keeps Gβγ from diffusing to enhance the rate of Gq / Gβγ re-association?

      Our crosslinked complex is compatible with simultaneous binding of a Gα<sub>q</sub>-Gβγ heterotrimer to the PLCb3, without disrupting the observed interface. If Gαq were to interact with the Gβγ molecules bound to the PH or EF hands, the interaction would be mediated by the N-terminal helix of Gα<sub>q</sub>. It is possible Gβγ–PLCβ3 interactions may slow heterotrimer reassociation, but this may be complicated by the intrinsic GAP activity of the lipase.

      (4) The authors find that Gβγ binds multiple sites, and it is clear that the PH domain site is the primary one in accord with previous work. Could these weaker sites be an artifact of the elevated concentrations used in cryoEM and BRET assays?

      While more studies have focused on the PH domain as a Gβγ binding site, our data does confirm the EF hands are also functionally relevant. To our knowledge, the role of the EF hands has not been investigated in this capacity until very recently, and so we hesitate to label them primary or secondary. It is possible the EF hands may be a lower-affinity site for Gβγ and the protein concentrations needed in cryo-EM drive complex formation. However, it is also possible the concentration of free Gβγ adjacent to an activated receptor may be high enough to saturate the PH and EF hand binding sites.

      (5) Although their assays infer differences in binding affinities, it would strengthen the paper if the authors could estimate the association energies of these different binding sites. This estimation would also address the concern stated above.

      We appreciate this suggestion and quantifying the affinities of the Gβγ–PLCβ3 interactions is the subject of future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please correct PIP2 to the correct PIP2 (many examples throughout).

      These have been corrected.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Figure S1B: The lane-condition labels above the third gel appear to be incorrect, as both lanes are marked identically (Gβγ +/+, PLCβ variant +/+, BMOE +/+) despite clearly different banding patterns. Please confirm.

      We have confirmed the markings above the gels are correct.

      (2) In Figure S2, the authors show three fitted models, but the helical density for Gβγ cannot be seen in two of them at the displayed contour level. The authors should provide views of the map at different contour levels (thresholds) to better support the model fitting. Otherwise, it is difficult to assess whether the Gβγ subunit could adopt alternative orientations (i.e., whether it may be rotated) within the density.

      We have included a new figure (Figure S3) that provides images of the maps at different contour levels.

      (3) Page 5: "where PLCβ3 is increased by the overexpression of either Gβγ or Gαq" should be revised to "where PLCβ3 activity is ...".

      This sentence has been corrected.

      (4) Figure 2: Please label residue R215 in Figure 2A/2B (or the relevant structural panel), since R215E is tested in 2E but the position is not shown.

      R215 is now included in Figure 2C.

      (5) Page 19, Figure 2 legend: "Changes ... Figure S3" should be "Changes ... Figure S5".

      We have corrected this figure call.

      Reviewer #3 (Recommendations for the authors):

      The studies seem well carried out, although more details regarding the BRET controls and the significance of the values should be included.

      We have revised the captions to provide more details about the experimental controls and a brief description of significance. Individual p-values are included in the supplemental tables.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      (1) The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of Western blot data and further discussion of mechanistic questions would strengthen the study.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      (2) The findings are likely to have a broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      We thank the reviewer’s comment. We have discussed the potential implications of these findings for cancer treatment in the “Discussion” section.

      Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing cooverexpression and positive correlation of NINJ2 with ferritin genes in iron-addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:

      NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture, is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc</sup>-</sup> pathways, is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

      Areas for improvement:

      While the study is strong, several issues should be addressed for mechanistic depth and general relevance.

      (1) Although NINJ2 is shown to interact with LAMP1 and LAMP1 knockdown rescues ferritin levels, it remains unclear whether the NINJ2-LAMP1 interaction is required for lysosomal protection. The authors could: a) Map the NINJ2 domain required for LAMP1 interaction and test whether an interaction-deficient mutant fails to protect against LMP and ferroptosis. b) Rescue NINJ2 KO cells with wild-type versus mutant NINJ2 to establish causality.

      We thank the reviewer’s comments. Ongoing work in our laboratory is focused on elucidating the molecular mechanism by which the NINJ2-LAMP1 interaction regulates lysosomal membrane integrity and ferroptosis, and we anticipate reporting these findings in a future publication.

      (2) The conclusion that NINJ2 suppresses ferroptosis relies primarily on RSL3 and Erastin sensitivity. A direct assessment of ferroptosis would hence the study, such as:

      (a) Include ferroptosis rescue experiments using ferrostatin 1 or liproxstatin 1.

      (b) Assess lipid peroxidation directly (e.g., C11 BODIPY staining) to strengthen the ferroptosis claim.

      We thank the reviewer for this thoughtful comment. We agree that ferrostatin-1 or liproxstatin-1 rescue experiments, together with direct analysis of lipid peroxidation, would provide complementary evidence for ferroptosis. We will incorporate these additional experiments in future studies to further strengthen the mechanistic basis of our findings.

      (3) The manuscript discusses lysosomal ferritin degradation but does not directly examine NCOA4, a central mediator of ferritinophagy. It would be good to: a) Test whether NCOA4 knockdown rescues ferritin loss and ferroptosis sensitivity in NINJ2 KO cells. b) This would clarify whether NINJ2 acts upstream of canonical ferritinophagy pathways or via an alternative mechanism.

      We appreciate the reviewer's thoughtful suggestion. Defining the contribution of NCOA4 to NINJ2-mediated ferritin degradation is an important question that could further clarify the underlying mechanism. Addressing this issue will require a comprehensive set of additional experiments, which will be addressed in the future studies.

      (4) The study is entirely cell-based, despite references to inflammatory and tumor phenotypes in Ninj2-deficient mice. While not strictly required, even limited in vivo validation (e.g., ferroptosis markers or iron accumulation in existing Ninj2 KO tissues) would substantially strengthen the manuscript.

      We thank the reviewer for this insightful suggestion. We agree that in vivo validation of ferroptosis markers and iron accumulation in Ninj2-deficient tissues would further strengthen our conclusions. However, these experiments will require substantial additional investigation, which will be pursued in the future studies.

      (5) Finally, most imaging data (e.g., Galectin 3/LAMP1 colocalization, PLA signals) and immunoblot data are presented qualitatively. The authors should provide the qualifications of Western blots and other measurements.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) What mechanisms might underlie the regulation of LAMP1 transcript levels by NINJ2?

      A clear mechanism by which NINJ2 regulates LAMP1 transcripts has not been elucidated and warrants further investigation. Nevertheless, several possibilities can be considered. First, LAMP1 transcription is known to be regulated by TFEB (transcription factor EB), a master regulator of the lysosomal–autophagy pathway. Upon lysosomal membrane permeabilization (LMP), TFEB translocates to the nucleus and activates a broad set of lysosome-related genes, including LAMP1. Notably, phosphorylation of TFEB by mTORC1 at Ser211 inhibits its activity by preventing nuclear translocation. Thus, it would be of interest to determine whether NINJ2 modulates TFEB phosphorylation status and subcellular localization. In addition, the tumour suppressor p53 has been reported to engage in complex crosstalk with TFEB in regulating basal autophagy. Interestingly, p53 expression is increased in NINJ2-KO cells. It is therefore plausible that NINJ2 regulates TFEB activity through p53, or alternatively modulates the p53–TFEB signaling axis more broadly to maintain lysosomal integrity.

      (2) Does Ninjurin1 play a similar role in regulating lysosomal membrane permeabilization (LMP)?

      At this moment, it remains unclear whether NINJ1 plays a role similar to that of NINJ2 in regulating LMP. In fact, our previous studies demonstrated that NINJ2 physically interacts with NINJ1 and may antagonize NINJ1-mediated pyroptosis. Furthermore, NINJ1 has recently been reported to promote ferroptosis by interacting with the xCT cystine/glutamate antiporter (PMID: 38464226), a function that contrasts with the protective role of NINJ2 against ferroptosis identified in the present study. These findings suggest that NINJ1 and NINJ2 may have distinct, or even opposing, functions in regulating cell death pathways. Nevertheless, further studies are required to determine whether NINJ1 also participates in the regulation of LMP and to define its relationship with NINJ2 in maintaining lysosomal membrane integrity.

      (3) What are the potential clinical implications of these findings, particularly in the context of cancer progression or therapeutic targeting?

      Targeting NINJ2 may have important clinical implications in cancer therapy. Given its role in maintaining lysosomal membrane integrity, inhibition or loss of NINJ2 could promote lysosomal membrane permeabilization (LMP), thereby sensitizing cancer cells to ferroptosis through increased intracellular labile iron accumulation and disruption of redox homeostasis, ultimately enhancing tumor cell killing. As such, NINJ2 inhibition may represent a strategy to selectively destabilize lysosomal function in cancer cells and improve responsiveness to ferroptosis-inducing agents or other combination therapies that exploit oxidative stress vulnerabilities. Indeed, we previously developed a peptide that targets NINJ2. Whether this NINJ2-targeting peptide can sensitize cancer cells to ferroptosis therefore warrants further investigation.

      (4) The authors should provide quantification of all Western blot data throughout the manuscript to enhance data robustness and reproducibility.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Reviewer #2 (Recommendations for the authors):

      (1) Controls for knockdown efficiency of NINJ2 (Figure 2D) should be shown.

      NINJ2-KO MCF7 cells were generated previously and published in the article (PMID: 38325550) along with sequencing confirmation.

      (2) In Figure 2A legends, the concentration of LLOMe is 1mM or 1µM - need to be clarified?

      The concentration for LLOME is 1µM. This typo has been corrected.

      (3) In the figure legends section, "Figure 4" is missing.

      We thank the reviewer’s comment. Figure 4 has been added to the Figure legends.

      (4) Some description of NINJ2 ko cells generation should be included in the materials section.

      In the Materials and Methods section, we briefly described how these cell lines were generated and cited the original publication (PMID: 38325550)

      (5) The manuscript would benefit from a schematic model figure summarizing the proposed NINJ2-LAMP1-iron-ferroptosis axis.

      In the revised manuscript, we provided a model to elucidate the role of NINJ2 in modulating lysosomal membrane integrity and iron homeostasis.

      (6) Some sections of the Introduction are lengthy and could be streamlined to focus more directly on lysosomes and ferroptosis.

      We have streamlined the introduction.

      (7) Statistical methods should clarify whether data meet assumptions for Student's t-test and whether multiple comparisons were corrected where applicable.

      Statistical analysis has been added to the figure legends when applicable.

    1. Author response:

      We thank the editors and three reviewers for their careful evaluation and constructive feedback. We are pleased that our identification of IR20a as a multimodal tuning receptor required for both low-salt and arginine sensing in a distinct gustatory neuron population was recognized as a valuable contribution to sensory coding. We agree with the major points raised and outline our planned revisions below, organized thematically.

      (1) Receptor assembly and integration model

      Our genetic, heterologous, and calcium imaging data show functional cooperation among IR20a, IR25a, and IR76b but do not demonstrate physical association. In the revision we will replace terms like “distinct subunit assemblies” and “peripheral integration” with more cautious language such as “functional receptor combinations.” We will state explicitly that our data cannot resolve whether the three IRs form a single heteromeric complex, and we will discuss the alternative possibility that IR76b and the IR20a/IR25a pair function as separate receptors within the same neuron. This interpretation better accounts for the response patterns observed upon co-expression. Throughout the text and in a revised summary figure, we will clearly differentiate elements directly supported by data from those that remain inferential.

      (2) Tarsal calcium imaging versus behavior

      The mismatch between tarsal calcium imaging and behavioral arginine responses will be addressed directly. We will explain that tarsal recordings sample only a small subset of IR20a neurons, whereas proboscis extension and feeding assays predominantly engage the more numerous labellar sensilla, whose neurons may carry different receptor compositions. We will generate labellar imaging where feasible; if additional functional data cannot be obtained, we will acknowledge this limitation rather than overinterpreting the tarsal results.

      (3) Synergy versus additivity

      We will discuss behavioral and cellular data separately. Recognizing that true synergy is difficult to demonstrate in feeding and proboscis extension assays, we will adopt conservative terminology when interpreting those experiments. For cellular data, where mechanistic insight is stronger, we will present the evidence for functional synergy and discuss why the outcomes may differ between the cellular and organismal levels.

      (4) Contextualizing prior IR76b and IR20a literatures

      We will expand the Introduction and Discussion to cite more fully the works establishing IR76b as a low-salt sensor and IR20a’s role in amino-acid sensing, including the earlier report that IR20a overexpression can inhibit IR76b-dependent salt responses. We will clarify how our single-cell imaging and loss-of-function data obtained in the native context refine models derived from ectopic expression. In its endogenous setting, IR20a marks neurons narrowly tuned to amino acids such as arginine, and IR20a is strictly required for low-salt detection. These findings contrast with the earlier view that IR20a functions broadly as an amino-acid sensor or as a salt-response blocker.

      (5) State-dependent modulation and IR56b

      Our data do not identify IR56b as the direct molecular sensor of internal state. We will reframe IR56b as a necessary component for state-dependent modulation of low-salt preference and retract any claim that it is the sensor itself. We will also clarify that IR20a and IR56b define genetically separable peripheral pathways that may converge on downstream circuits.

      (6) Methodological and presentation issues

      We will address the following points raised across reviews:

      (1) Co-localization: Higher-resolution confocal images and co-localization analysis will be provided.

      (2) Summary model figure: A new figure will illustrate the distinct functions of IRs in low-salt and amino-acid taste, clearly indicating which aspects are directly supported and which are inferential.

      (3) Feeding-assay control: We will either include an isosmotic sucrose control to avoid the water confound or explicitly discuss this limitation and temper the interpretation of feeding-preference results.

      (4) S2 cell quantification: Complete details on response criteria, responder fractions, and statistical reporting will be added.

      (5) Figure and supplementary corrections: All noted errors in figure legends, scale bars, citations, and supplementary-file mismatches will be fixed.

      We are confident these revisions will bring our mechanistic claims into close alignment with the evidence and substantially improve the manuscript. We again thank the editors and reviewers for their detailed and helpful comments.

    1. Author response:

      We thank the editors and reviewers for recognizing the originality and potential value of the proposed framework. The reviews rightly ask us to distinguish three things more sharply: what is established directly in human breast tissue, what is inferred from mammary and aging studies, and what remains hypothesis. We agree, and the revision will make that distinction explicit throughout.

      We will define the proposed reserve state operationally and specify the findings that would distinguish active niche maintenance from passive persistence. Throughout, we will treat passive persistence as a legitimate competing hypothesis rather than a settled question. Heterogeneity and immune or inflammatory associations will be presented as consistent with active maintenance and causal directionality, not as establishing them. We will also integrate the mixed epidemiologic evidence more centrally into the argument, rather than treating it as a caveat.

      We will reassess the evidence for senescence in the aging breast and describe it more precisely, correcting or narrowing statements that outrun the data. The figures will be revised so that observed inflammatory and immune-regulatory features are clearly separated from proposed senescence- and SASP-mediated mechanisms. We will also clarify the limits of our analogies: postpartum involution and cross-tissue reserve systems will be presented as sources of candidate mechanisms and testable predictions, not as direct evidence for the proposed mechanism in human ARLI. Finally, we will frame the translational implications as contingent. They depend on first demonstrating that senescent cells are enriched near persistent lobules, identifying the relevant cell types and immune states, and establishing causal relevance.

      We appreciate the reviewers' constructive suggestions. We believe these revisions will preserve the conceptual contribution of the model while making its evidentiary status, the alternative explanations, and its falsifiable predictions substantially clearer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

      Thank you for your critical review and insightful comments.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      (b) Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.

      - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function. Thank you.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

      Thank you.

      Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

      Thank you for your critical review of our manuscript and for the very constructive comments and suggestions to improve the manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (There is some redundancy below with my comments in the public review, but I've included here for clarity and further explanation.)

      The authors have addressed my primary concern, as treatment with PLK inhibitor reduces phosphorylation of T301, while also providing some comment on relative impact of PLK-mediated KIN-G phosphorylation.

      It is notable that phosphorylation of T301 was reduced by only ~27%, while phosphorylation of S569 was reduced by ~100% in the presence of PLK inhibitor. The authors note that this might be explained by slower dephosphorylation of T301. In the absence of a phenotype, and with cell doubling continuing unabated in presence of the inhibitor, it is intriguing that more loss is not observed. An alternative explanation is that an alternate kinase might also be able to phosphorylate T301 in the absence of PLK activity, and the authors should consider that possibility.

      We added a sentence in the main text to suggest an alternative explanation.

      The model shown in figure 8 should include a dephosphorylation step, per the authors comments in the text regarding the small fraction of T301 that is phosphorylated and proposal of a phosphorylation/dephosphorylation cycle. The Fig 8 legend needs to have some commentary on the role of phosphorylation, as phosphorylation is the center point of this paper.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function.

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Fig 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Fig 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces motility of KIN-G. These statements are contradictory.

      (3) p. 8, and Fig 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      (4) Fig 7B. Please explain labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      (5) Fig 4, 5, and 7: "% Cells" is reported. Please indicate what number of cells total were examined.

      These minor comments have already been addressed in the previous revision.

    1. Author response:

      Reviewer #1 (Public review):

      R1-Q1: The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      We thank the reviewer for raising this important and thoughtful point. We agree that the relationship between spontaneous gamma oscillations and pain perception remains a matter of active debate, and we will moderate these statements throughout the manuscript. We will acknowledge the ongoing debate and present the gamma-pain relationship as an active area of investigation rather than settled fact.

      R1-Q2: The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as 'responders' based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      We appreciate this careful critique. We wish to clarify the rationale behind our analytical approach and address the concern.

      A well-established finding in the neurofeedback literature is that a substantial proportion of participants are "non-learners" — individuals who, despite receiving real feedback, fail to achieve effective control over the targeted neural activity. This is not a failure of the intervention, but reflects individual differences in neurofeedback learning capacity. Our core research question is therefore: "Among individuals who can successfully learn to upregulate gamma oscillations, does this upregulation reduce pain perception and nociceptive brain responses?"

      To address the circularity concern and improve transparency, we will make the following revisions:

      - We will reframe the wording from "NFB increases gamma and reduces pain" to "Successful gamma upregulation via NFB is associated with reduced pain in those who achieve it." All causal language will be replaced with appropriately cautious, correlation-based terminology.

      - We will report full-sample results for completeness.

      R1-Q3: The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      We thank the reviewer for requesting this clarification. We will specify the exact criterion in the revised manuscript, i.e., a participant was classified as a responder if their post-intervention gamma power minus pre-intervention gamma power was positive (i.e., an increase in gamma power following the neurofeedback intervention).

      R1-Q4: Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      We thank the reviewer for these constructive suggestions. We will supplement and refine the methodological details in the revised manuscript, and we will make the analysis code and the experimental program (including the custom neurofeedback software) publicly available via an open repository.

      R1-Q5: Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      We will discuss the limitations of the discontinuous visual feedback and the viewing distance in the revised manuscript. We acknowledge these as valid methodological concerns and will address them as limitations in the Discussion.

      R1-Q6: The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

      We thank the reviewer for pointing out that the description of the muscle-confound analysis was insufficiently detailed. In the revised manuscript, we will clarify the EMG analysis procedures and explicitly report: (a) the exact number of participants contributing to the EMG analysis; (b) the cohort from which they were drawn; and (c) whether responder selection was performed before or after restricting the sample for EMG analysis.

      Reviewer #2 (Public review):

      R2-Q1: The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increase gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      As detailed in our response to R1-Q2, the responder analysis reflects a conceptually motivated subgroup defined by successful neurofeedback learning — a well-documented challenge in NFB research where many participants are non-learners. Our central question is whether successful gamma upregulation (among those capable of achieving it) is associated with pain reduction. We will make this rationale explicit in the revised manuscript. We will also: (a) transparently report full-sample results; (b) discuss potential biases introduced by subgroup selection (e.g., differences in attention, self-regulation); and (c) acknowledge the lack of preregistration for the responder analysis.

      R2-Q2: Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group → gamma change → pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      We will substantially revise the language throughout the manuscript to avoid causal claims. We also plan to conduct a formal mediation analysis (NFB group → gamma change → pain change) to test whether changes in gamma activity statistically mediate the observed pain reduction.

      R2-Q3: For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      We thank the reviewer for raising this point, and we will clarify both points in the revised manuscript. (a) Because group assignment was randomized, the first participant could in principle have been assigned to the sham group. To prepare for this possibility, we collected EEG data from one participant in advance (equivalent to pilot data) to serve as the sham feedback signal, had the first participant been assigned to the sham group. In the actual experiment, however, the first participant was randomly assigned to the active group, so this pre-collected dataset was never used. (b) We will also compare the discrepancy between actual gamma power and the sham feedback signal in the sham group, and report this result in the revised manuscript.

      R2-Q4: Was baseline gamma power comparable between groups?

      We thank the reviewer for this suggestion. We will report and compare baseline gamma power between the active and sham groups in the revised manuscript.

      Reviewer #3 (Public review):

      R3-Q1: Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      We thank the reviewer for this important suggestion. We will revise the manuscript to discuss more explicitly the inherent difficulty of recording pure gamma-band oscillations with scalp EEG. We will acknowledge that scalp gamma is not reliably observable in all participants, is vulnerable to contamination from multiple muscle sources, and that a single posterior neck EMG channel provides only limited control. We will also discuss the possibility that changes in posture, facial tension, breathing, and arousal may contribute to high-frequency scalp activity.

      R3-Q2: The evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      We appreciate the reviewer's careful consideration of this point. We will temper our conclusions, presenting the current evidence as demonstrating the feasibility of gamma-band neurofeedback training in a subset of participants and an association between successful training and pain reduction, rather than a demonstrated causal relationship. We will discuss individual differences (attention, suggestibility, relaxation ability, neurofeedback learning capacity) as potential confounds that cannot be fully disentangled from gamma-specific effects.

      R3-Q3: It is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      We will clarify that the study employed a single-blind design: participants were unaware of their group assignment. We will state this clearly in the revised manuscript.

      R3-Q4: The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      We will provide a stronger rationale for the selection of the two neurofeedback video scenarios. We will also place greater emphasis on the pain reduction observed in both groups, acknowledging the substantial nonspecific analgesic effects associated with the procedure.

      R3-Q5: The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      We thank the reviewer for this insightful comment. We will expand the discussion of why neurofeedback may produce effects beyond those achieved by gamma-frequency tACS. In particular, we will elaborate on the potential contributions of volitional control, attentional engagement, immersion, expectation, and self-regulatory processes, and discuss how these factors may be central to the observed pain reduction rather than merely incidental to the gamma modulation.

      R3-Q6: Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

      We thank the reviewer for this balanced assessment. We fully agree that the conclusions should be tempered, and we will revise the manuscript accordingly to reflect that the current evidence supports feasibility and association rather than established causal efficacy.

      Summary

      In summary, the planned revisions include:

      (i) full-sample results reported;

      (ii) moderating causal language throughout and adding a formal mediation analysis;

      (iii) clearly specifying the responder criterion and adding this to the preregistration;

      (iv) providing complete methodological documentation and publicly releasing all analysis code;

      (v) expanding the Discussion to address limitations regarding scalp gamma measurement, EMG control, visual feedback, viewing distance, single-blind design, NFB scenario rationale, and nonspecific effects;

      (vi) adding analyses on baseline gamma comparability and sham-feedback discrepancy.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank you and the reviewers for the thoughtful evaluation of our manuscript and for the constructive comments and important suggestions. We are encouraged by the recognition that the study is valuable and of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution. We are also grateful that the reviewers acknowledged the integrative approach of our work, including fecal metabarcoding, behavioral assays, electrophysiological recordings, chemical analyses, and field observations.

      We have carefully considered all comments and have revised the manuscript accordingly. In particular, we have made the following major revisions:

      (1) We clarified the biological origin of limonene and expanded the discussion of possible bat-associated sources, including snout secretions and microbial contributions, while avoiding overinterpretation of limonene as an exclusively endogenous mammalian compound.

      (2) We added more detailed descriptions of contamination controls, including instrument cleaning procedures, materials used for housing and handling, and blank-control results, to address the concern that limonene could have originated from human-associated or environmental contamination.

      (3) We added individual-level odor data and revised the presentation of terpenoid profiles to show more clearly where limonene was detected and how its relative contribution varied among samples.

      (4) We clarified the rationale for the concentration choices used in the electrophysiological and field experiments, emphasizing that these assays were designed to test physiological detectability and functional sufficiency rather than to reproduce exact natural emission concentrations or determine response thresholds.

      (5) We revised our interpretation of limonene more cautiously. We now state that limonene is sufficient to trigger avoidance responses, while we acknowledge that its natural ecological specificity, concentration dynamics, and interactions with other bat odor components require further investigation.

      (6) We corrected the terminology related to limonene enantiomers. Because our GC-MS and GC-EAD analyses did not use an enantioselective column, we replaced “(-)-limonene” with “limonene” throughout the manuscript, figures, legends, and supplementary materials.

      (7) We improved the figures and supporting materials by moving the figure illustrating predator-prey relationship into the main text, adding electrophysiological traces from all tested crickets, revising figure legends, clarifying the rationale for comparing bat body odor with air controls, and providing additional chemical-identification details in the supplementary materials.

      (8) We checked and clarified statistical annotations, including the exact adjusted P-values for relevant comparisons.

      We believe these revisions have substantially improved the clarity, rigor, and balance of the manuscript. We hope that the revised manuscript and the detailed point-by-point responses satisfactorily address all concerns raised.

      We are confident that our study represents an important contribution, as it fundamentally expands the traditional acoustic-centred view of bat–insect interactions by demonstrating that crickets can use olfaction as a complementary sensory modality to detect bat odors and initiate avoidance behavior. The multidisciplinary evidence, integrating behavioral assays, electrophysiology, chemical profiling, and field validation, provides convincing support for this novel olfactory pathway.

      eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

      We sincerely thank the editors for the positive and constructive assessment of our work. We greatly appreciate the recognition of our study’s value and its potential interest to multiple research communities.

      We have carefully considered all comments and have revised the manuscript accordingly. The revisions include clarifications of limonene’s biological origin and ecological specificity, strengthened contamination controls and individual-level odor data, clearer rationales for experimental concentration choices, and more cautious interpretation throughout the manuscript. All changes are addressed in detail in our point-by-point responses below.

      Importantly, the central conclusion of our study—that crickets can detect and avoid bat odor through olfaction, and that limonene is sufficient to trigger avoidance responses in both laboratory and field settings—remains robustly supported by the multidisciplinary evidence we present. We believe the revised manuscript now provides a clearer, more balanced, and scientifically rigorous account of our findings, and we hope it meets the standards of eLife.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are delighted that you found the study interesting and enjoyable. We also appreciate your kind remarks about the clarity of the manuscript, the logical flow of the experiments, and the quality of the figures.

      We are especially grateful that you recognized the value of our integrative approach. Combining behavioral assays, electrophysiology, chemical analysis, and field observations was central to our study design, and we are pleased that this approach resonated with you.

      You provided an accurate and comprehensive summary of our work. You confirmed that our main narrative is clear and logically coherent. Specifically, you followed our progression from establishing the predator–prey relationship, to demonstrating olfactory avoidance, to identifying limonene as an active compound, and finally to validating its behavioral effects in both laboratory and field settings.

      We have carefully considered all your constructive comments and suggestions. We address them in detail in our point-by-point responses below. Your feedback has been extremely helpful, and we believe the revised manuscript is substantially stronger as a result.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

      We sincerely thank you for your critical and constructive comments. Your questions regarding the identity, biological origin, and ecological specificity of limonene are insightful and have helped us substantially strengthen the manuscript. We address each of your points below.

      On the biological origin of limonene and whether mammals synthesize it de novo

      You raised an important point that limonene seems unusual for a mammal to emit endogenously. We fully agree. We agree that direct evidence for de novo synthesis of limonene in mammals is currently limited.

      However, limonene in bat body odor could still have biological origins. First, it may originate from skin- or gland-associated microbiota. Recent work has shown that skin-associated microorganisms can substantially shape bat volatile odor profiles (Sun et al., 2026, BMC Biology), and some microbes possess enzymes capable of terpene biosynthesis. Second, previous studies have independently reported limonene in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.). This suggests that its presence in bats is not unique to our study. Third, our own analyses detected limonene in hair and snout secretions, but not in faeces or blank controls. This pattern is consistent with a biological source associated with the body surface, rather than diet or environmental deposition.

      We have now expanded our Discussion to cover these possibilities more explicitly. We also emphasize that the exact source, i.e., endogenous, microbial, or otherwise, remains an open question that warrants future investigation. Please see lines 258–273 of the clean version of the revised manuscript, or see the excerpt below:

      “Although limonene reliably induced avoidance behaviour in crickets, two related questions still merit careful consideration. One question is whether the limonene we identified genuinely originates from bats or reflects contamination during sampling. Limonene is common in plants and numerous consumer products (Boncan et al., 2020; Schuman, 2023), making its endogenous production by a mammal seem unusual. Nevertheless, multiple lines of evidence militate against contamination. First, we adhered to rigorous protocols. For example, all instruments were cleaned with ethanol and oven-dried before each use; bats were housed in stainless-steel cages, and cloth bags had been rinsed with purified water. Second, limonene was absent from all blank controls, including empty-chamber air samples and clean swabs, and was not detected in bat faecal samples. In contrast, it was consistently identified in hair and snout-secretion samples from bats. Third, independent studies have similarly identified limonene in the secretions of other bat species (Faulkes et al., 2019; Zhang et al., 2022). Furthermore, emerging evidence indicates skin-associated microbes may contribute to bat volatile profiles, with some taxa possessing enzymes involved in terpene biosynthesis (Sun et al., 2026). Taken together, these observations point towards an endogenous or microbe-mediated source, although the exact biosynthetic pathway remains to be determined.”

      On contamination from human-associated sources

      You asked how we excluded the possibility that limonene entered our samples through handling, equipment, cleaning products, or other human-associated sources. We appreciate this concern and have addressed it in detail.

      We believe contamination is highly unlikely for several reasons. First, we followed strict protocols throughout. All instruments were cleaned with ethanol and oven-dried before and after each use. We used stainless-steel cages and cloth bags made of degreased bleached cotton washed with purified water. These materials are not sources of terpenes. Second, we ran multiple blanks. Limonene was not detected in any empty-chamber air controls or in blank cotton swabs. In contrast, it was consistently found in multiple bat snout-secretion samples. This clear difference between samples and blanks strongly argues against contamination. Third, we now report individual-level odor data (see new Figure 3; Supplementary Table 5.xlsx). These data show that limonene was consistently present across individual bats. It did not appear sporadically, as one would expect from accidental contamination.

      In the revised manuscript, we have added detailed descriptions of our contamination controls in the Methods section (Please see lines 405–407, 443–444, 471–477 of the clean version of the revised manuscript, or see the excerpt below). We have also included the blank-control results (Supplementary Table 5.xlsx), as you suggested.

      Lines 405–407: “Prior to sampling, all glassware was thoroughly rinsed with ethanol and dried in an oven at 120°C, and volatile odor collection was conducted in a dedicated odor-free room to minimize environmental contamination.”

      Lines 443–444: “The empty-chamber controls were used to account for potential background signals from the experimental system and to provide a baseline for comparison with bat odor extracts.”

      Lines 471–477: “Upon capture, bats were placed in clean stainless-steel cages and kept in groups consistent with their natural social associations during the brief interval prior to immediate odor sampling. Hair samples (10 mg per individual) were clipped from dorsal and ventral regions. Snout secretions were collected using sterile cotton swabs (CS15-005, Shenzhen SihuaBo Technology Co., Ltd., China), with two blank swabs as controls. These blank swab controls were included to account for potential volatile contamination from ambient air or the swab material itself (Supplementary Table 5).”

      On how crickets distinguish bat-associated limonene from environmental sources

      You raised a thoughtful question about ecological specificity. Given that limonene is abundant in mint, citrus peel, pine, and other non-threatening plants, how would crickets use it as a reliable indicator of bat presence?

      We agree with you completely. We do not claim that limonene alone serves as an unambiguous bat-specific signal. Instead, our interpretation is more nuanced. We argue that elemental perception represents one effective strategy within a broader olfactory toolkit. It does not exclude the importance of other cues.

      In our revised manuscript, we have softened our interpretation accordingly. We now state explicitly that limonene is sufficient to trigger avoidance under our experimental conditions, but we do not interpret it as a uniquely bat-specific keystone kairomone. We also discuss mechanisms that could help crickets reduce false alarms under natural conditions. These include concentration differences, temporal patterning (bats are active at night), spatial context (specific foraging habitats), co-occurrence with other bat-specific volatiles, and possibly enantiomeric composition. Please see lines 274–292 of the clean version of the revised manuscript, or see the excerpt below.

      Lines 274–292: “The second question is how crickets might distinguish bat-derived limonene from environmental sources of this compound, given its prevalence in mint, citrus peel, pine and other non-threatening plants (Boncan et al., 2020; Schuman, 2023). It seems implausible that crickets could simply rely on limonene per se to differentiate a bat from a leaf. Two non-exclusive mechanisms could help resolve this issue. First, limonene need not be the only olfactory cue mediating risk perception. Our findings establish the sufficiency of limonene as an avoidance trigger, but do not preclude a role for other odor components. The crickets’ antennal responses to other bat volatiles in our GC–EAD analyses suggest more complex peripheral perception. Additional compounds, either alone or in synergistic blends, may modulate the full behavioral response in nature. Therefore, elemental perception via limonene likely represents one effective strategy within a broader olfactory toolkit available to insects. Second, crickets may discriminate bat-derived limonene through context-specific cues (e.g., temporal and spatial patterning, co-occurrence with other bat-specific compounds) to minimize false alarms. Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays. Critically, our field data confirm that limonene exposure in nature robustly triggers an adaptive anti-predator response, irrespective of the precise discrimination mechanism.”

      We acknowledge that fully testing these ideas would require substantial additional work. We have therefore framed this as an important direction for future research, rather than as a resolved issue in the present study.

      We thank you again for these insightful comments. Your feedback has helped us present a more rigorous, balanced, and transparent account of our work.

      Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are especially gratified that you recognized our study as having a "complete chain of evidence" and as a "highly praised study." This recognition means a great deal to us, given the multidisciplinary nature of our approach and the effort required to integrate behavioral, electrophysiological, chemical, and field data into a coherent narrative.

      We also appreciate your accurate summary of our work. You correctly highlighted the broader context that while acoustic co-evolution between bats and insects is well documented, whether crickets can use bat scent as an early warning cue has remained unknown. Your summary confirms that our main findings are clear: the body odor of S. kuhlii triggers robust avoidance and electrophysiological responses in L. equestris, and that limonene alone is sufficient to elicit avoidance in the laboratory and suppress calling in the field.

      We are particularly grateful that you acknowledged the multi-sensory nature of cricket predator recognition, with keen olfactory, tactile, and auditory senses. We agree that crickets are an excellent model for studying multimodal predator detection, and we hope our study encourages further exploration of olfaction in this classic predator–prey system.

      We have carefully considered all your specific comments and suggestions. These include questions about the novelty framing of olfactory eavesdropping, the rationale for our concentration choices, and the presentation of electrophysiological comparisons. We address each of these points in detail in our point-by-point responses below. Your thoughtful feedback has been extremely helpful in improving the clarity and rigor of the manuscript.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      Thank you for this comment. You are absolutely right that olfactory eavesdropping across the vertebrate–invertebrate divide is not a new concept in itself, and we acknowledge that this phenomenon has been well documented in previous studies.

      However, we would like to clarify our intended emphasis. In the Introduction and Discussion, we have already stated that empirical examples combining chemical identification, electrophysiological validation, behavioral assays, and field confirmation within a direct predator–prey context remain relatively limited. This is especially true for the bat–insect system, where research has historically focused on acoustic interactions rather than olfaction.

      We did not intend to overstate the novelty of olfactory eavesdropping per se. Instead, our emphasis was on providing a complete chain of evidence in a vertebrate–invertebrate predator–prey system that has traditionally been viewed through an acoustic lens. In that sense, we believe our study adds a complementary olfactory perspective to this classic system, rather than claiming to have discovered olfactory eavesdropping as a novel phenomenon. We hope this clarifies our position, as already stated in the original manuscript (lines 71–78, lines 88–90, lines 235–241 of the clean version of the revised manuscript, or see the excerpt below):

      Lines 71–78: “However, a fundamental gap exists in understanding whether such olfactory eavesdropping can operate across the vast phylogenetic divide separating vertebrate predators and invertebrate prey (Apfelbach et al., 2015; Dicke and Grostal, 2001; Schoeppner and Relyea, 2005). Although olfactory interactions across broad taxonomic boundaries are widespread in nature, such as mosquitoes using host odors to blood-feed, elephants and moths sharing pheromonal components, and aroids chemically mimicking carrion to attract pollinating flies (Kang et al., 2023; Zaremska et al., 2022; Zhao et al., 2022), these interactions are primarily shaped by selective pressures tied to foraging, reproduction, or mutualisms, rather than by predation-related selection.”

      Lines 88–90: “The bat–insect system presents an ideal model to address these questions. Despite the clear importance of olfaction to both taxa and its established role in predator–prey ecology, whether it plays any functional role in the iconic bat–insect arms race remains unexplored.”

      Lines 235–241: “Beyond the specific bat–insect model, our work addresses a central question in sensory ecology: how chemical eavesdropping operates within predator–prey systems between phylogenetically distant taxa with fundamentally divergent olfactory systems (Adams et al., 2020; Emerson and Johnson, 2024; Kaupp, 2010). While intraphyletic kairomone detection is well-established (e.g., rodents avoiding carnivore odors, aphids responding to ladybug chemicals), compelling experimental evidence for such olfaction-mediated recognition across broad phylogenetic divides has been limited (Apfelbach et al., 2005; Ferrari et al., 2007; Tanis et al., 2018).”

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      Thank you for this question. We fully agree that knowing the natural concentrations and relative content of limonene in bat odor would be valuable. However, our experimental aims were not to mimic natural emission levels precisely. Instead, they were designed to answer two distinct questions: First, whether cricket antennae are physiologically capable of detecting limonene at all; and second, whether limonene alone is sufficient to trigger behavioral responses under controlled and semi-natural conditions.

      For the EAG experiments, we selected a concentration gradient (0.001%, 0.01%, 0.1%, 1%, and 10% v/v) following standard practices in insect chemical ecology and referencing a previous study (Tang et al., 2024). The goal was to establish dose-dependent antennal sensitivity, not to match a specific natural concentration. Our data clearly show that cricket antennae respond across a range of concentrations, with stronger responses at higher doses.

      For the field experiment, we used a 10% v/v limonene spray over 25 m<sup>2</sup> plots. We acknowledge that this concentration does not quantitatively reflect natural bat emissions. Natural odor plumes are highly dynamic and depend on airflow, turbulence, temperature, humidity, vegetation structure, and distance from the source. Accurately reconstructing these natural dynamics would require detailed quantitative measurements of bat odor release rates and plume modeling, which were beyond the scope of the present study. Instead, our field experiment was designed for a functional purpose: to test whether limonene could alter cricket calling behavior under semi-natural conditions, using a concentration sufficient to produce a detectable odor stimulus in the field.

      We also note that the field-applied concentration is comparable to what has been used in other chemical ecology studies testing the behavioral effects of single volatile compounds under natural or semi-natural conditions. In that context, our positive result supports the ecological relevance of limonene as an avoidance cue, without requiring that the exact applied concentration matches natural bat emissions.

      We agree that quantitative characterization of natural bat odor composition, limonene release rates, and ambient exposure concentrations is an important direction for future research. According to your comments, we have added this point to the revised Discussion as a clear future direction (please see lines 286–290 of the clean version of the revised manuscript, or see the excerpt below). We have also clarified in the Methods that our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly (please see lines 515–522, 556–558 of the clean version of the revised manuscript, or see the excerpt below).

      Lines 286–290: “Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays.”

      Lines 515–522: “For each antenna, a hexane control was first presented to establish baseline antennal activity. For the initial screening, limonene, undecane, pentadecane, and hexadecane were diluted to 10% (v/v) in hexane and delivered individually in a randomized order. To assess dose-dependent responses, limonene was further tested at five concentrations (0.001%, 0.01%, 0.1%, 1%, and 10%, v/v in n-hexane), following the concentration gradient used in a previous study (Tang et al., 2024). Following the initial hexane control, the five limonene concentrations were tested in a randomized order across trials. Each stimulus lasted 0.5 s, with an inter-stimulus interval of 1 min to allow full recovery of antennal responses.”

      Lines 556–558: “Our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly.”

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

      Thank you for this suggestion. We understand your point that comparing GC-EAD responses between bat body odor and snout secretions would be a more direct way to identify the anatomical source of active compounds.

      However, we would like to explain why we did not include this comparison in the current study.

      First, the purpose of Figures 1C and 1D (i.e., Figure 2C and 2D in the revised manuscript) was to answer a more fundamental question: whether bat body odor, as a whole, contains volatile compounds that are detectable by cricket antennae. Comparing with an odor-free air control was therefore the appropriate first step. It established the basic phenomenon of olfactory detection before we moved on to source attribution.

      Second, we did conduct chemical profiling of snout secretions, hair, and faeces using HS-SPME-GC-MS (presented in Figure 3). These analyses showed that limonene was consistently present in hair and snout secretions, but absent from faeces and blanks. This allowed us to identify snout secretions as the most likely source of limonene, without requiring GC-EAD recordings from secretion samples themselves.

      Third, we did not perform GC-EAD on snout secretions for practical reasons. The secretion samples were collected in very small amounts. They were almost entirely consumed during the HS-SPME-GC-MS chemical analyses, leaving insufficient material for additional GC-EAD testing.

      We agree with you that directly comparing GC-EAD responses to snout secretions versus whole-body odor would be an excellent experiment. It would further strengthen the source attribution and provide more direct evidence. We have noted this as a valuable direction for future studies in the revised Discussion.

      We hope this clarifies our rationale. Thank you again for your thoughtful suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I would suggest moving Supplementary Figure 1 into the main figures. It contains important information about the predator-prey relationship and the ecological relevance of L. equestris, so it seems too central to be placed only in the supplement.

      We agree with your suggestion. The predator–prey relationship and the ecological relevance of L. equestris are indeed central to the biological context of this study. We have therefore moved the original Supplementary Figure 1 into the main text as Figure 1. We have renumbered the remaining figures accordingly, revised the figure legends, and updated the Results text to better highlight this ecological context (revised manuscript, Figure 1 legend, lines 856–864).

      (2) Lines 134 to 136: Only one representative EAD trace is shown. I suggest showing all five traces, either in the main figure or as a supplementary figure, to better illustrate the reproducibility of the antennal responses across individuals.

      We agree. To better illustrate reproducibility across individuals, we have added EAD traces from all five tested crickets as Supplementary Figure 1. The figure legend now describes the sample sizes for both the bat odor treatment and the odor-free control (revised manuscript, Supplementary Figure 1 legend, lines 900–906).

      (3) Lines 144 to 153: It would be helpful if the authors reported which VOC collections contained limonene and in what amounts. Showing individual-level data for the bat odor samples, rather than only pooled or summarized profiles, would strengthen the conclusion that limonene is consistently associated with S. kuhlii body odor.

      We agree. To better show the consistency of limonene detection across individuals, we have added Figure 3B, which displays an individual-level terpenoid profile. Each stacked bar represents one VOC collection, and the limonene segment indicates its presence and relative contribution. We have also revised the Results section to state explicitly that limonene was detected in hair and snout secretions but absent from feces (lines 144–149 in the revised manuscript).

      Lines 144–149: “To identify the biological sources of bat body odor, we analyzed VOCs from hair, faeces, and snout (pararhinal gland) secretions of nine bats using headspace solid–phase microextraction coupled with gas chromatography–mass spectrometry (HS–SPME–GC–MS). Snout secretions and hair shared similar hydrocarbon-rich VOC profiles, whereas faecal volatiles were distinct (Figure 3A). Individual-level terpenoid profiles further showed that limonene was detected in hair and snout secretion VOC collections but was absent from faeces (Figure 3B and Supplementary Table 2).”

      (4) Figure 2A: I recommend adding representative chromatograms or VOC traces for feces, hair, and snout secretions. This would make the source comparison more transparent and would help readers assess the underlying chemical profiles behind the heatmap.

      We agree that representative chromatograms would make the source comparison more transparent. However, due to the analytical workflow, individual chromatograms were not retained in the dataset we received. As an alternative, we added Figure 3B, which shows individual-level terpenoid profiles. This allows readers to assess which samples contained limonene and how its relative contribution varied among sample types and individuals. In addition, we have provided the NIST retention index, quantitative ion, qualitative ion, and molecular weight for each terpenoid compound in Supplementary Table 2.

      (5) The chemical identification of the key compounds would benefit from more detail. The authors state that compound identities were confirmed by matching retention times and mass spectra to authentic standards, including a mixed standard injection. It would be useful to provide the retention times, match and reverse-match scores, blank traces, and, if available, sample-plus-standard co-injection data showing peak augmentation without the appearance of new peaks. For (−)-limonene specifically, the enantiomeric assignment would require an enantioselective method, such as chiral gas chromatography, unless this was already performed and not described.

      We fully agree that comprehensive identification evidence is important for transparency and reproducibility.

      Regarding the identification data: we have already confirmed compound identities by matching retention times and mass spectra to authentic standards, including a mixed standard injection. In the revised version, we will deposit the raw chromatographic data and identification details (retention times, match scores, and blank traces) in Figshare as supporting information.

      Regarding co-injection: we acknowledge that sample-plus-standard co-injection would provide even stronger confirmation. However, our bat odor samples were difficult to obtain and were almost entirely consumed during the GC–EAD and GC–MS analyses. We therefore could not perform additional co-injection validation. We have noted this limitation in the revised Materials and Methods (revised manuscript, lines 463–464).

      Regarding the enantiomer assignment of limonene: you are correct that determining the specific enantiomer requires a chiral GC column, which was not available in this study. To avoid overinterpretation, we have replaced “(-)-limonene” with “limonene” throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) This is a typical study in the field of chemical ecology. As long as the source of the active substances is determined and the biological activity has been detected, this research is complete, so Figure 2 seems to be unnecessary.

      We agree that identifying the source of the active compound and confirming its biological activity are central to this study. However, we believe Figure 2 serves an important purpose. Bat body odor could originate from multiple sources, i.e., hair, faeces, or snout secretions, and comparing VOC profiles across these sources helps us determine which source most likely contributes to the odor cues detected by crickets. This is especially important for limonene, because it is a plant-associated terpenoid and not a typical animal-derived volatile, as the reviewer #2 pointed out. Our analysis showed that hair and snout secretions shared similar VOC profiles, while faecal VOCs were distinct and lacked limonene. This supports snout secretions as the likely primary source. To make this purpose clearer, we have revised the Methods section to state that this analysis was conducted to investigate potential biological sources of bat body odor (revised manuscript, lines 494–495, 567–569).”

      Line 494–495: “To identify the biological source of the characteristic body odor, we compared the VOC profiles from hair, faeces, and snout secretions.”

      Line 567–569: “Principal component analysis (PCA) based on a binary (presence/absence) matrix was performed using the vegan package in R to compare profiles from hair, feces, snout secretions, and bat body odor.”

      (2) Figure 3B seems to be incorrect. The significance of n-hexane and 1% limonene is ***P < 0.001, while for 10% limonene, why is it only two stars, **P < 0.01? Please check it.

      We have rechecked the statistical analysis and confirmed that the original annotation was correct. The significance levels in original Figure 3B (Figure 4B in the revised manuscript) were based on Bonferroni-corrected paired t-tests comparing each limonene concentration with the n-hexane control. The adjusted P value was 0.0006 for 1% limonene (P < 0.001) and 0.00485 for 10% limonene (P < 0.01). The higher adjusted P value for 10% limonene reflects greater among-individual variation at this concentration, which may be due to differential sensitivity of individual antennae at higher doses. We have now added the exact adjusted P values in the revised manuscript, lines 163–167.

      Lines 163–167: “EAG responses to limonene were concentration-dependent (repeated-measures ANOVA, F(5, 25) = 24.95, P < 0.001, η<sup>2</sup>p = 0.83; Figure 4B), with both 1% and 10% limonene solutions eliciting significantly stronger responses than the hexane control (Bonferroni-corrected paired t-tests, 1%: t(5) = 10.77, P < 0.001, Hedges' g = 3.82; 10%: t(5) = 6.92, P = 0.005, Hedges' g = 2.46).”

      Finally, we would like to express our sincere gratitude to the editors and the reviewers for the thoughtful feedback. Your comments have significantly improved the quality and clarity of our manuscript. We hope the revised version satisfactorily addresses all concerns raised.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on smaller studies to more comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound and broadly applicable to large-scale short-read datasets for assessing copy number variation and genomic repeat content. While convincing in its scope and novelty, the findings would be further strengthened with exploratory analyses of datasets from other species with more or fewer repeats in their genomes.

      We agree with the assessment and appreciate the Editor’s handling of our manuscript and careful consideration of our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      While it would be interesting to further investigate the mechanisms underlying cis-QTLs, we do not believe this dataset can distinguish between the two proposed models of cis-variation. As argued in the manuscript, the observed cis-association signals are linked to repeat-associated SNPs and therefore are most likely driven by variation in repeat copy number, supporting the latter hypothesis.

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      As a compromise, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      Our decision to use 12-mers to profile genome content in A. thaliana was based on an empirical evaluation of the sensitivity and specificity of different K-mer lengths for this application. Because our goal was to characterize intraspecific variation in genome content, we did not extend this analysis beyond A. thaliana. Nevertheless, we agree that alternative K-mer lengths may capture additional and potentially useful information and that using the method in other species would be interesting.

      Although we found generating a 14-mer dataset to be computationally infeasible, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      To our knowledge there is not a published gapless A. thaliana assembly, although several of the recent assemblies are pretty close. Although we continue to use the 1001 Genomes SNP dataset that relies on TAIR10 for the GWAS analyses, we now present Figure 2A and Figure S9 using the Col-PEK assembly (Hou, Wang, Cheng, Wang & Jiao 2022 Molecular Plant).

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      We concede that A. thaliana does not possess a typical plant genome and thus have moderated our claims throughout the paper.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

      We agree that the conclusions of our work are correlational, as they are based on GWAS associations and have not been validated through functional experiments. We would love to see tests of the hypotheses we presented, but we believe this is beyond the scope of the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The Snakemake workflow Kmer-it had a few minor bugs that prevented use with modern Snakemake versions, which I fixed as part of the review (submitted as a pull request). Please be sure to validate that the whole workflow works with the latest Snakemake version.

      We appreciate the Reviewer’s testing of our software and have updated it accordingly.

      (2) L162: Since you already map samples to a reference genome to filter organellar DNA, consider adding some form of filter or correction for duplication rate in paired-end data, which I have found to be one major source of batch effects as observed here between sequencing centers.

      We appreciate this suggestion. In earlier versions of the analysis, we removed duplicated reads, but doing so substantially weakened the relationship shown in Figure 1C. Because it is difficult to distinguish duplicate reads arising from genuine repeat copy number variation from those resulting from technical artifacts, we ultimately chose not to include duplicate removal in the final pipeline. However, we did remove the effect of the sequencing center using a linear model.

      (3) Have you used this approach with raw long-read data? Do you see any difference between Illumina and low-error-rate long read data (e.g., HiFi, R10 Nanopore with SUP calling)? This could be compared in a directly paired manner using some of the recent papers with HiFi/ONT data on 1001G accessions, e.g., Lian et al, 2024, Wlodzimierz et al, 2023, Teasdale et al, 2025. (I do not believe such a comparison is required to prove the utility of this method, but it could be of interest as such sequencing technologies become increasingly cost-competitive).

      We have not evaluated this approach using raw long-read sequencing data, although we agree that it would be an interesting direction for future work. As noted in the manuscript, we detected significant batch effects attributable to sequencing center, even among datasets generated with the same sequencing technology. Given these observations, we are cautious about comparing K-mer frequencies across sequencing platforms that have distinct error profiles, as technical differences could confound biological signals.

      Reviewer #2 (Recommendations for the authors):

      (1) I would strongly suggest modulating some of the claims made in the paper, especially with regard to "genome evolution in plants", given that the paper focuses on A. thaliana. Alternatively, the authors could test the method in a different dataset from a different species. This would alleviate some concerns regarding the choice of k-mer length and demonstrate robustness.

      We agree that A. thaliana is not representative of most plant genomes and have revised the manuscript to better reflect this limitation. Since the primary goal of this study was to characterize intraspecific variation in genome content within A. thaliana, we believe that extending the analysis to an additional species falls beyond the scope of the present work. We do not believe that our choice of K-mer length undermines the robustness of the approach. To evaluate this concern, we repeated the GWAS analyses using 10-mers and found broadly consistent results, which are presented in Supplementary Figure S19-21. These findings suggest that the major conclusions are not strongly dependent on the specific K-mer length selected.

      (2) A k-mer length of 12 is quite small. There are 4^12 possible 12-mers (~17 million), so you would expect that all 12-mers would occur in the A. thaliana genome by chance at least once. While larger k-mers may be more computationally expensive to compute, it would be worth repeating the analysis with a larger k-mer length.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      (3) Line 689 on page 31 should include the figure number.

      Thanks! Fixed.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Marconcini et al. report results of an ambitious study on the genetic mechanisms that contribute to resistance of Drosophila flies to the toxin octanoic acid (OA). This study was motivated by two observations: first, Drosophila sechellia, a close relative of D. melanogaster, has evolved specialized feeding on fruits of Morinda citrifolia, which contain high concentrations of OA and second, that artificial selection on Drosophila simulans, a sister species of D. melanogaster, can generate higher resistance to OA. Previous studies had performed genetic mapping studies between D. simulans and D. sechellia that implicated certain genomic regions in resistance to OA and, in particular, implicated several Osiris gene paralogs as contributing to resistance, though the molecular mechanisms of resistance remain unclear. In this study, Marconcini et al. performed two major experiments. First, they performed evolution-and-resequence on Drosophila simulans populations exposed to OA for 50 generations and identified candidate regions with excessive shifts in allele frequencies as candidate regions containing OA resistance genes in D. simulans. Second, they performed a CRISPR knock-out screen in a D. melanogaster cell line to identify genes that contribute to OA resistance and susceptibility.

      Evolve-and-resequence yielded many candidate genomic regions with extreme allele frequency shifts, which may be regions containing OA resistance genes, or linked genes, or regions that happen to show a strong shift in all replicate populations by chance. As the authors note, detecting significant shifts in allele frequencies is a challenging problem, and the authors use two measures of allele frequency shifts (the Cochran-Mantel-Haenszel method and Bait-ER) and perform simulations under neutrality to estimate a reasonable significance threshold. I am not entirely convinced by this method of estimating significance levels, because the simulations involve assumptions that may not be met by the real populations. I would think that a permutation test would provide an assumption-free method of estimating significance levels. I have tried to think whether there is something about the design of these experiments that would preclude the use of permutation tests (which are used widely for genome-wide studies, such as QTL), but I can't think of one. Perhaps the authors are aware of a reason permutation tests would be invalid here, and if so, they should state this reason.

      Significance thresholds have been estimated using a variety of approaches in the literature, including false discovery rate (FDR) control, simulations under neutral drift with predefined cut-offs, arbitrary significance thresholds, haplotype-based analyses, permutation tests, amongst other methods. Our choice was guided by a review (doi:10.1186/s13059-019-1770-8), which compared the performance of several of these approaches and found that the relatively simple assumptions underlying the Cochran–Mantel–Haenszel (CMH) test often performed as well as, or better than, more complex alternatives. As with most significance thresholds used in genome-wide analyses, the CMH threshold applied here is ultimately based on a degree of arbitrariness, although we aimed to be as stringent as possible. As a complementary method, we used Bait-ER; here the threshold used followed the recommendations provided in the original publication describing this method (doi:10.1111/jeb.14134).

      One reason that permutation-based approaches may be less widely adopted than theoretical null models is the extensive linkage disequilibrium among SNPs, which results in large blocks of correlated variants that cannot be considered independent observations (and therefore not shuffled). In principle, permutations could be performed at the haplotype-block level; however, defining haplotype blocks itself requires selecting thresholds or criteria that are often user defined or based on significance cut-offs. Although several tools are available for this purpose, our experience has been that the resulting block definitions remain sensitive to these choices and therefore introduce a comparable degree of arbitrariness. In reality, there is not yet a standard method in the field that has emerged as the “best” one.

      There is overlap between regions detected by the two methods, but the methods disagree for many regions. The authors state that a "majority of prominent peaks were found by both methods," but I am unclear on what "prominent" means here. It would be more helpful to be more quantitative about the extent of overlap.

      To quantify the agreement between the two methods, we calculated the overlap in genomic coverage (base pairs) between CMH and Baiter candidate regions. For G25, the two methods shared 1.90 Mb of candidate regions. The overlap encompassed 65.3% of the genomic span identified by CMH and 64.1% of the span identified by Bait-ER. For G50, the methods shared 4.77 Mb of candidate regions. The overlap encompassed 63.6% of the CMH candidate span and 99.9% of the Bait-ER candidate span. We have now replaced the admittedly qualitative statement the reviewer highlighted ("majority of prominent peaks were found by both methods”) with this information.

      The authors hypothesized that the response would be at similar genomic loci in all populations (line 222). It seems at least possible that epistatic interactions would lead to different combinations of alleles evolving in each population. I wonder if it would be possible to test whether there is heterogeneity in the responses across the replicate populations.

      We agree that epistatic interactions could lead to different allelic combinations being favored in different replicate populations. However, both BaitER and the CMH test are designed to detect parallel evolutionary responses across replicates, which was the focus of our study. One approach for testing population-specific responses is the LRT-2 test (doi:10.1534/genetics.118.301824). We did not pursue this analysis because it would likely generate many additional candidate loci, making interpretation more challenging, while providing limited additional insight into the repeatable genomic responses that were the primary focus of this work.

      The evolve-and-resequence method yielded many possible regions contributing to OA resistance in D. simulans, but perhaps too many regions to test directly or even to build sensible hypotheses about the genes involved. Thus, the authors performed a second experiment to try to narrow down the list of possible candidate genes. They performed a CRISPR knockout screen in a D. melanogaster cell line for genes that contribute to resistance or susceptibility to OA. The authors identify several limitations of this experiment, but they nonetheless identified several genes where knockouts contribute to OA susceptibility or resistance. Intersecting top hits with regions that experienced selection identified two "resistance" genes: kraken and Alkbh7. The selection hit at kraken is quite compelling, whereas the evidence at Alkbh7 is less strong because only two SNPs were marginally significant. Further functional assays, including gene knockouts in D. melanogaster and D. sechellia, provide some support for the claim that both of these genes can contribute to resistance to OA in flies.

      Beyond the few issues raised above, I do not have significant questions about methodology or the results. I do think, however, that the authors should be more conservative about the implications and significance of their results. For example, on line 139, the authors claim that this intersection approach provides a "powerful paradigm to investigate ecotoxicology." I am not sure I agree that the identification of two genes that may contribute to OA resistance, after a seemingly heroic selection experiment and CRISPR screen, suggests that this method is all that powerful. It seems that most of the genes that contribute to the selection response remain unidentified.

      We agree with the reviewer and have modified the sentence accordingly. While we believe that integrating the approaches discussed in this paper can provide valuable insights into the genetic basis of ecotoxicological traits, these approaches are not a panacea for traits with highly complex genetic architectures, such as OA resistance. Nevertheless, the identification of two candidate genes with some evidence of contributing to the trait represents a meaningful advance toward understanding its underlying genetic basis.

      Finally, given that one motivation of this project was to identify genes that contribute to evolved resistance to OA, I am surprised that the authors did not generate CRISPR alleles of kraken and Alkbh7 in D. simulans and then use these together with the existing alleles in D. sechellia to perform reciprocal hemizygosity tests to determine if these two genes actually contribute to evolved resistance in D. sechellia. This test is simpler to perform and may be more sensitive than the allelic replacement that the authors propose (lines 446-449).

      While generating null alleles for the candidate genes in D. simulans is beyond the scope of this revision, we note that we did attempt reciprocal hemizygosity tests using the mutants available in D. melanogaster. However, the resulting hybrids were recovered in low numbers, and these animals were rather weak, making them unsuitable for the severe OA exposure conditions employed in our assays. We agree that, where feasible, future studies should incorporate reciprocal hemizygosity tests, as they represent a powerful approach for validating the contribution of candidate genes to the trait of interest. We have revised the final sentence of the Results section accordingly.

      Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit, in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting, and the idea of narrowing down a large list of candidate loci by CRISPR-based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow-up experiments to validate the results.

      Weaknesses:

      The reviewer is not convinced of the conceptual idea behind their approach: the intersection of the two approaches implicitly assumes that null alleles (or at least compromised alleles) should be selected during experimental evolution. The reviewer considers this unlikely, and the authors made no attempt to test this implicit hypothesis in their data.

      We respectfully disagree with the reviewer’s interpretation of the conceptual idea behind our approach. Our strategy did not assume that experimental evolution selects for null alleles, but rather for any type of variant that could contribute to the trait being selected for (i.e., increases in OA resistance), pointing to candidate genes contributing to the trait. Like many evolve-and-resequence experiments, this approach identified hundreds of candidate genes. This is why we took an orthogonal, genome-wide CRISPR screening approach, where loss-of-function mutations could lead to increases or decreases in OA tolerance of cultured cells. However, we stress that the naturally selected alleles – which could be gain or loss of function – may have much subtler phenotypic effects than the null alleles used for functional validation.

      Along the same lines, it is not clear how to reconcile an upregulation of candidate genes in resistant flies with the knockout experiments.

      We respectfully disagree that these findings are difficult to reconcile. The observed upregulation of the candidate genes in the selected, resistant D. simulans is consistent with a role in promoting OA resistance, while the knockout experiments (whether in cultured cells or in whole animals) demonstrate that loss of gene function reduces resistance. These observations are complementary: increased expression is associated with enhanced resistance, whereas complete loss of function impairs it. The knockout experiments of kraken and Alkbh7 were intended as functional validation of gene involvement and do not imply that the alleles selected during experimental evolution are loss-of-function alleles (we rather hypothesize that the selected alleles are gain-of-function through some, as yet undetermined, mechanism).

      The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This is correct. As our results suggest that the phenotype is shaped by multiple genes, such that the effect of any individual gene is likely modest compared to their combined contribution. Consequently, detecting and validating the effect of a single gene requires more stringent OA conditions than those needed to observe the overall phenotypic response over the course of several generations. In addition, the shorter-term plate assay was more practical for higher temporal resolution of the analysis of mortality in the presence of OA.

      The statistical analysis suffers from some problems and an insufficient description of the analyses performed.

      Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      We agree that comparing the experimental evolution and GWAS results (from our previous work, doi:10.1093/g3journal/jkag032) is of considerable interest. We have now expanded the Discussion to explicitly discuss the relationship between the two datasets. Overall, the overlap between the approaches was limited, although two GWAS candidate genes, bez and CG13003, fall within genomic regions exhibiting significant CMH signals at generation 25 (but not generation 50) of the evolve-and-resequence experiment. (We note that these genes could not have been identified in our CRISPR screen, as they are not expressed in S2R+ cells). More broadly, differences between the GWAS and evolve-and-resequence results likely reflect the distinct evolutionary processes captured by each approach: GWAS maps standing phenotypic variation among isofemale lines, whereas experimental evolution tracks allele frequency changes under sustained selection. Understanding why some signals are shared whereas others are not – whether due to effect size, genetic background, epistasis, pleiotropic costs, or the contribution of initially rare variants – remain important open questions.

      The reviewer would have liked to see more connection between the experimental evolution and the GWAS data. As some D. simulans genotypes have similar resistance to D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      While D. simulans genotypes displayed a range of OA resistance levels, none approached D. sechellia levels of resistance (see Figure 2e from our previous work, doi:10.1093/g3journal/jkag032) (The reviewer might have conflated the data from our GWAS of D. melanogaster strain, shown in Figure 2b of that paper, where some lines of that species exhibit comparable resistance to D. sechellia under the conditions of that assay). Regardless, our experimental evolution data do not provide sufficient resolution to identify the specific favorable alleles underlying the response to selection. Instead, we detect genomic regions containing many linked variants whose frequencies change under selection. Consequently, a direct comparison between evolved genotypes and GWAS-associated genotypes is currently difficult. We note, however, that expression of the D. sechellia kraken allele in D. melanogaster did not produce a significant effect on resistance, suggesting that even the most promising candidate alleles might not have strong effects in isolation.

      At several places, the authors discuss the challenge of studying a polygenic trait, but at the same time, they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      Our results do not imply that kraken and Alkbh7 are major-effect loci or that they explain a substantial proportion of the phenotypic variation. Rather, our data indicate that these genes make measurable contributions to OA resistance, consistent with the expectation that complex traits are influenced by many loci of individually modest effect. The absence of these genes among the top GWAS candidates does not preclude their involvement, as GWAS and experimental evolution interrogate different aspects of the genetic architecture and differ in their power to detect loci of varying effect sizes and allele frequencies.

      It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      We agree that the significant peaks identified in the evolve-andre sequence experiment represent promising targets for future investigation. However, these peaks typically span large genomic regions containing tens to hundreds of genes, making it difficult to prioritize individual candidates based on the experimental evolution data alone. In this work, we focused our functional validation on genes independently supported by the CRISPR screen, which provided gene-level resolution. We fully acknowledge that additional causal genes are likely to reside within the selected regions and remain to be functionally characterized.

      Impact:

      Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid, a trait that is rather easy to evaluate.

      The reviewer appears to imply that a trait being straightforward to phenotype necessarily implies that its genetic basis should also be straightforward to resolve. Many classic complex traits, such as human height, are simple to measure yet have an extraordinarily complex, highly polygenic genetic architecture. We believe that OA resistance represents a similar challenge: while the phenotype is readily assayed, differences in assay conditions, genetic backgrounds, and the contribution of many loci of individually modest effect make its genetic basis difficult to dissect. We have strived to be cautious in our conclusions, in particular the evolutionary interpretations; nevertheless, to our knowledge, this is the first study to provide functional evidence supporting the contribution of specific genes to OA resistance in D. sechellia, combining both loss-of-function phenotypes and expression data. As emphasized by the title of our manuscript, we view the principal contribution of this work as demonstrating how complementary experimental approaches (both of which are fairly novel for study of toxin susceptibility/resistance genetics) can be integrated to prioritize and functionally evaluate candidate genes underlying complex adaptive traits.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers propose to include additional statistical and quantitative measures to strengthen the results. Furthermore, certain parts of the manuscript could be rephrased to make sure that the readers can understand more clearly the implications and significance of the results. To test whether the genes identified in the present manuscript (as contributing to octanoic acid resistance) are also involved in the evolution of the resistance, reciprocal hemizygosity tests using CRISPR alleles of kraken and Alkbh7 in D. simulans would be a plus, but such experiments are not required because the main focus of the paper is on the genetic basis of octanoic acid resistance, and not on evolution.

      We thank the Reviewing Editor for these constructive comments. We have revised the manuscript to clarify the interpretation and significance of our findings and have addressed the reviewers’ comments throughout. Regarding reciprocal hemizygosity tests, we agree that they would provide a valuable means of assessing the evolutionary contribution of candidate genes. However, generating the necessary reagents in D. simulans represents a substantial undertaking beyond the scope of the present study, particularly given the expected modest effects of individual loci underlying this highly polygenic trait. We have nevertheless revised the end of the Results section to mention reciprocal hemizygosity tests as an important direction for future work.

      Reviewer #2 (Recommendations for the authors):

      (1) Provide more details about the selection tests: which sequences were used? Please report p-values. It would also be important to discuss the possibility of false positives caused by a bottleneck in D. sechellia. A genome-wide analysis could help to see if the bottleneck increased the signal of positive selection.

      We are not entirely sure what additional analyses are being suggested. The sequences and methods used for the selection analyses are described in the Methods, and the statistical support for these analyses is reported in the Supplementary Material (MK test p-values and FUBAR posterior probabilities). We are also unclear as to how a genome-wide analysis would address the interpretation of the gene-specific selection analyses presented here. If the reviewer intended a different analysis, we would appreciate further clarification.

      (2) The significance level of the CMH test needs to be determined with the effective population size, not with the census size, as done by the authors. This is important, as it is not clear if the candidate genes remain significant after significance adjustment based on the effective population size.

      We thank the reviewer for this comment. In our analyses, effective population sizes were explicitly incorporated by Bait-ER to model the effects of genetic drift. The CMH significance thresholds, however, were obtained from neutral forward simulations following the recommended workflow for the method, which requires census population sizes rather than effective population sizes as input. To make these simulations as realistic as possible, we therefore used the observed census population sizes at each generation of the experimental evolution. Estimating generation-specific effective population sizes for use in an alternative simulation framework would require substantially more temporal data than are available here (e.g., sequencing many additional time points) and is beyond the scope of the present study.

      (3) Include the allele frequency trajectory across time for the candidate genes.

      We thank the reviewer for this suggestion. However, plotting allele frequency trajectories for the candidate genes is not straightforward because the evolve-and-resequence analysis identified broad linked genomic regions rather than individual causal variants. Each candidate region contains numerous SNPs spanning several kilobases, with each SNP exhibiting its own allele frequency trajectory across the 10 replicate populations. Consequently, there is no single representative trajectory for a given candidate gene, and plotting all SNPs within each region would be difficult to interpret. For kraken, however, we identified seven candidate regulatory SNPs and have plotted their individual allele frequency trajectories, which we include in Author response image 1 to illustrate the diverse patterns observed:

      Author response image 1.

      (4) Estimate the effect of candidate loci in the GWAS data (independent of significance).

      We thank the reviewer for this suggestion. However, we are not entirely sure what analysis is being proposed. In particular, it is unclear which variants the reviewer is referring to, as our candidate genes are associated with multiple linked variants rather than a single causal SNP. Consequently, we are unsure how the effect of a candidate locus should be estimated in the GWAS dataset. Moreover, there is no guarantee that the same variants are represented in both datasets. For example, variants filtered out during the GWAS because of low allele frequency may subsequently have increased in frequency during experimental evolution and therefore contributed to the evolve-and-resequence signals. We would appreciate further clarification of the analysis the reviewer has in mind.

      (5) Discuss the challenge of false positives (see: 10.1016/j.cub.2020.12.023).

      We thank the reviewer for this suggestion. We agree that false positives (as well as false negatives) are an important consideration when studying complex polygenic traits, both through the initial “screening” efforts (e.g., GWAS, experimental evolution) and follow-up functional validation (e.g., RNAi, mutant, overexpression analyses). We believe this issue is already addressed in the manuscript through our discussion of the limitations of the individual approaches, the polygenic nature of OA resistance, and our cautious interpretation of the functional validation results. Our conclusions are limited to identifying candidate genes that contribute to OA resistance, rather than claiming to have identified all of the loci or variants underlying the evolution of this trait. We are therefore not sure what additional discussion the reviewer has in mind and would appreciate further clarification if a specific point from the cited study is intended.

      (6) The figures with the Manhattan plots should be improved to indicate the overlapping genes in the Manhattan plots, rather than in the circle figures below. By using different colors this should be quite easy and clean.

      We thank the reviewer for this suggestion. However, we believe Author response image 2 more appropriately illustrate the overlap between the CMH and BaitER analyses. Although the Manhattan plots display individual SNPs, our candidate loci are defined by broader genomic regions comprising blocks of linked significant SNPs rather than by single variants. Simply highlighting SNPs within overlapping regions would therefore add visual complexity without providing additional biological insight beyond that already captured by the circos plots. As an illustration, we provide here an example of the generation 50 Manhattan plot with overlapping SNPs highlighted in red:

      Author response image 2.

      (7) A more focused discussion of the assumption that functional data from D. melanogaster can explain resistance in D. simulans or D. sechellia. At some places, epistatic interactions are mentioned, but the reviewer feels that the entire screen is based on the idea that similar effects are found across species, hence, this needs to be adequately reflected in the discussion.

      We thank the reviewer for raising this important point. We would like to emphasize that our study was not based on the assumption that functional effects identified in D. melanogaster or D. simulans necessarily explain the evolution of OA resistance in D. sechellia. Rather, our motivation stemmed from the longstanding difficulty of identifying individual genes underlying this highly polygenic trait using mapping approaches alone. We therefore sought to combine orthogonal experimental approaches to identify genes contributing to OA susceptibility (of D. melanogaster and D. simulans) and resistance (D. sechellia). We hoped, but did not assume, that genes supported by multiple independent lines of evidence would be informative for understanding the natural evolution of OA resistance in D. sechellia. Indeed, we acknowledge in the manuscript that the cell-based CRISPR screen, experimental evolution, and the natural evolution of D. sechellia occurred under very different selective contexts and timescales. As reflected in the title of our manuscript, our conclusions are intentionally framed around the identification of novel toxin resistance loci, while remaining cautious about their evolutionary interpretation.

      (8) Discuss that the controls in the RNAi test were quite variable. Could this reflect some problems with the assay?

      We do not believe that the variability among the control lines reflects a problem with the assay. Rather, the different controls (Gal4, UAS-RNAi etc.) represent distinct genetic backgrounds, each of which may exhibit a different baseline level of OA resistance. Indeed, in our recent GWAS of OA resistance (doi:10.1093/g3journal/jkag032), we observed substantial natural variation in OA resistance among D. melanogaster and D. simulans lines. Importantly, each RNAi line was compared with its corresponding genetic background control, so differences among control lines do not affect the interpretation of the individual RNAi experiments.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      We understand the more severe chromosome missegregation and DNA damage phenotypes in the npp-3 mdf-1 or npp-3 mdf-2 double RNAi may suggest that npp-3 RNAi is more sensitized to the loss of spindle assembly checkpoint (SAC) component MDF-1 or MDF-2 for chromosome protection, but may not clearly show that the localization of chromosomes to the nuclear periphery serves a chromosome protective function. Based on our current phenotypic analyses, we will tone down the title to “Prophase Chromosomes Relocalization to Nuclear Periphery in NPP-3/NUP205 Depletion Depends on the Spindle Assembly Checkpoint and Inner Kinetochore Proteins”.

      We have thought about artificially tethering the chromosomes to the nuclear periphery. However, without the same stimulus/defect and activation of the response pathway, the effects could be different. In future, to further clarify the functions of chromosome relocalization, we could analyze the chromosomal missegregation and DNA damage phenotypes in npp-3 lem-2 RNAi, which has a partial chromosomal nuclear periphery phenotype, to see if there is any quantitative relationship between chromosome nuclear periphery location and DNA protection.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      We have not performed partial npp-3 RNAi yet. To explore whether the level of nuclear envelope rupture correlates with chromosome localization to the periphery or if a minimal level of nuclear rupture is required, we have performed the lacI::GFP reporter assay in different NPP RNAi to indicate the nuclear permeability defects (in Fig. S1A and B). npp-2, npp4 or npp-5 RNAi also causes increased permeability, at a comparable level as npp-3, npp-7 and npp-13 RNAi, yet only npp-3, npp-7 and npp-13 RNAi causes chromosome relocalization. This may explain the peripheral chromosome relocalization response may be more closely linked to the specific NPC subcomplex’s (inner ring and nuclear basket) or individual NPP’s function, rather than the general permeability defect or rupture size.

      However, we have also attempted to use other methods to induce targeted nuclear rupture with specific rupture size by laser ablation (Author response image 1, 355 nm UV pulsed laser marked by the white rectangle), and we could occasionally observe all chromosomes relocalizing to all over the nuclear periphery after 660 s upon laser ablation in different strains. However, the chromosome relocalization phenotype is not very consistent (~33%, n =12). Importantly, the laser also introduces DNA damage at the rupture sites, as marked by HUS-1 (Author response image 2), complicating the interpretation, so we did not include this attempt in the manuscript.

      Author response image 1.

      Time-lapse imaging of embryos expressing LEM-2::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce nuclear envelope rupture (white rectangle). 0s is the time applying laser ablation. Scale bar 5 μm.

      Author response image 2.

      Time-lapse imaging of embryos expressing HUS-1::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce DNA damage (white rectangle). 0 s is the time applying laser ablation. Scale bar 5 μm.

      Another attempt to achieve different nuclear rupture sizes is by observing nucleus in different stages of embryos. By npp-3 feeding RNAi approach, we observed an increase of the proportion of the nuclear circumference with nuclear rupture during embryonic development (Fig. S3D). Yet, the chromosome periphery phenotype is displayed in all embryo stages, suggesting that the localization of chromosomes to the nuclear periphery occurs across a range of rupture sizes.

      Given these observations, it appears that the extent of nuclear rupture or permeability increase may not be the sole determinant of chromosome relocalization. Instead, disruption of certain NUP subcomplexes or NPPs might stimulate this process.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      We observed that chromosomes localize to the nuclear periphery in all cells during embryonic development until the embryo dies, as demonstrated by snapshots and time-lapse imaging (Fig. S1C and Fig. S1E).

      We focused on the P1 blastomere for our analyses because its cell cycle timing and division orientation is well characterized, which facilitates precise examination of chromosome localization and cell cycle dynamics under different perturbation conditions. We will add a sentence at the beginning of that result section to clarify this choice: "The chromosome nuclear periphery localization in npp-3 RNAi is consistent in all cells at different embryonic stages (Fig. S1C and Fig. S1E). For analysis in distinct perturbation conditions, we specifically examined the P1 cells, whose cell cycle timing and division orientation is well characterized and suitable for such studies.”

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (see response to point 2) We observed that npp-2 or npp-4 RNAi can cause nuclear rupture with certain permeability defect compared to wildtype (Fig. S1A and S1B), but not chromosomal nuclear periphery localization, suggesting that such chromosome localization to the nuclear periphery does not just depend on the nuclear rupture but could be a more specific phenotype related to individual NUP subcomplex’s (inner ring and nuclear basket) or NPP’s function. 

      For Y complex, we also now tested NPP-5 depletion, which shows permeability defect but did not show nuclear rupture marked by NPP-1:GFP or chromosomal periphery phenotype (Fig. S1A and S1B), which is consistent to the paper cited [1]. 

      To confirm the efficiency of NPP-2 depletion, we have observed the presence of smaller nuclei (as described in phenobank) by NPP-1::GFP marker (Fig. S1A) and a marked reduction in NPP2::GFP signal in npp-2 RNAi-treated embryos (unpublished data). These evidences support that NPP-2 depletion was effective. 

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      We agree with this suggestion and have decided to remove the section on AIR-1 in the manuscript text to avoid confusion. To clarify, air-1(RNAi) causes multiple centrosomes in only a small proportion of embryos (approximately 13%) [2] , primarily by affecting centrosome positioning [3]. and does not routinely lead to significant centrosome amplification.

      Our interest in AIR-1 was inspired by previous findings by Hachet et al., showing that AIR1/Aurora A localizes at sites without NPP-3 during mitotic entry—these sites are believed to correspond to centrosome locations [4]. This relationship prompted us to explore how AIR-1 depletion might influence NPP-3 localization at the nuclear envelope and its potential effects on chromosome positioning. Interestingly, we discovered that following air-1(RNAi) treatment, nuclei exhibit discontinuous NPP-3 localization on the nuclear envelope. Thereby, we could investigate the interplay between NPP-3 and chromosomal dynamics. Chromosomes localize to nuclear periphery specifically without NPP-3 (Author response image 3). This suggests a potential negative correlation of NPP-3 localization with the chromosomes. We did not imply any regulation by AIR-1.

      Author response image 3.

      Condensed chromosomes tend to localize at nuclear envelope sites without NPP-3 or NPP-7. (A) (C) Selected representative confocal images of all chromosomes localizing at the nuclear periphery in the control and air-1(RNAi) embryos expressing H2B::GFP and mCherry::NPP-3 (A) and GFP::NPP-7 and mCherry::H2B (C). The upper right image is the zoom-in view of the nucleus. Scale bar 10 μm. (B) (D) The intensity of GFP and mCherry is normalized to the average intensity along the nuclear envelope and plotted in the control and air- 1(RNAi) nucleus from (A) and (C).

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      Based on our time-lapse data in Figure 1D, most chromosomes appear to condense at or near the nuclear periphery, but there are still some chromosomes inside the nuclear space at approximately -600 seconds before NEBD, and then subsequently moving to the periphery during condensation (with full condensation at 100 s past NEBD). This suggests that initial condensation may occur both within the nucleus and at the periphery. Over time, chromosomes condense and cluster at the nuclear envelope. However, we did not separately measure the condensation level of individual chromosomes at the nuclear periphery versus in the middle of the nucleus, which could be challenging in live cells. So far, we cannot separate the chromosome nuclear periphery phenotype and the condensation.

      To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      We recognize that npp-3 RNAi embryos exhibit broad developmental defects, and therefore global transcriptional changes. Our RNA-seq analysis revealed that differentially expressed genes did not display positional bias within the genome. NPP-3 depletion downregulates many pathways, including pathways related to RNA polymerase II activity and cell cycle regulation, consistent with a global transcriptional downregulation. Thus, while the transcriptomic data are broad, it is consistent with the chromosome condensation phenotype.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      In our co-depletion experiments, the images for npp-3(RNAi) and X(RNAi) (Fig. 4) are displayed side-by-side under the same exposure conditions and scaled the same way with the same intensity thresholds to facilitate comparison. Despite that, some differences in the background intensity can be observed across samples. Thus, the normalised mCherry::NPP-3 signal intensity (subtracting the background) the single and double RNAi samples will be quantified and added to supplementary figure S4.

      To ensure consistency of double RNAi, we used ligation-based RNAi constructs designed to simultaneously target both NPP-3 and X, aiming to achieve comparable knockdown efficiencies. Nonetheless, variability in RNAi efficiency is a recognized limitation, and we interpret our data within this context. To further validate the knockdown, we will also assess the efficiency of the other gene X using the corresponding fluorescent reporter (see response to Reviewer 3, point 2).

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      Inactivation of MDF-1 or MDF-2 suppresses both the chromosome nuclear periphery localization and extended duration of interphase to prophase and prometaphase in npp-3 inactivation. We have performed the analysis to assess the timing and extent of chromosome condensation in the single and double depletion conditions (Author response image 4). The chromosome condensation dynamics is similar in the mdf-1 RNAi and npp-3 mdf-1 double RNAi, as well as the control group, indicating that MDF-1 is also involved the premature chromosome condensation phenotype caused by NPP-3 depletion.

      Author response image 4.

      The dynamic changes of chromosome condensation parameter, in which 30% of pixels in the ROI analyzed is below the threshold scaled intensity (<77), in the different groups. The sample size is 5. Error bars show mean ± SEM.

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      We have aligned the nuclear envelope reassembly (NER) time and highlighted the time point in the images (Fig. 5A). Our data show that in control embryos, this duration is approximately 780 seconds, while in npp-3(RNAi) embryos, it extends to about 930 seconds. This difference is statistically significant, as determined by one-way ANOVA (or appropriate nonparametric/mixed tests). The representative image is consistent with the quantification presented in Fig. 5B. We could add the corresponding videos to the supplementary information.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

      It is correct that the average embryonic lethality observed in npp-3(RNAi) embryos reachs 95% (Fig. S6). Given the broad developmental defects and such high baseline lethality, the genetic interaction with mdf-1 is modest.

      Nevertheless, our findings highlight that in the npp-3 mdf-1 double RNAi condition, we observe significantly increased rates of lagging chromosomes (Fig. 5D), micronuclei formation (Fig. 7A), and elevated DNA damage (Fig. 7B). These effects support the idea that MDF-1-, MDF-2dependent chromosome anchoring to the nuclear periphery (and condensation) plays a positive role in NPP-3 depleted cells.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP- 3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically codepleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Codepletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

      The DNA damage evidence is based on lagging chromosomes in 2-cell stages, the HUS-1 reporter, and micronuclei in embryos at the 20-30 cell stage. It is noted that npp-3 RNAi is pleiotropic and also causes DNA damage. There is additional DNA damage caused by loss of MDF-1 and MDF-2 in npp-3 RNAi, but we agree that it is difficult to say whether the effect is additive or not, complicating the interpretation. Thus, we will tone down in our title to describe the dependency of the chromosomal nuclear periphery phenotype (see response to Reviewer 1 point 1).

      As for the premature chromosome condensation phenotype in NPP-3 depletion, we hypothesize it may result from accumulation of factors such as BAF-1 at the nuclear periphery, which could facilitate chromatin condensation. To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization (also see response to Reviewer 1 point 6).

      Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP- 3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Codepletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. , they unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SACdependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      We agree that the current data are largely correlative. In the early C. elegans embryos, single depletion of SAC components MDF-1 or MDF-2 does not affect the mitosis duration or chromosome segregation. When spindles are defective, the functional SAC delays progression through mitosis [5]. Depleting SAC components such as MDF-1 in npp-3 RNAi indeed impacts multiple processes, including prophase and prometaphase duration, chromosome repositioning and condensation, making it challenging to disentangle effects specifically to related chromosome positioning.

      To address this, future experiments involving NPP-3 LEM-2 co-depletion, which has been shown to partially impair peripheral chromosome positioning, will be utilized to assess whether disruption of partial peripheral localization results in increased DNA damage, micronuclei formation, or HUS-1 foci accumulation, independent of SAC function (see response to Reviewer 1 point 1). 

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      We acknowledge that in our double RNAi experiments, the knockdown efficiency of the npp-3 is checked by imaging (see response to Reviewer 1 point 8), whereas that of the second gene was not validated in each condition. To address this, we confirmed the effectiveness of certain partner gene depletions, e.g. HCP-3, by examining the levels of the respective proteins using available GFP-marked strains or immunofluorescence (Author response image 5). Additionally, for genes like HCP-3, KNL-1, and BUB-1, we also assessed the functional consequences on chromosome segregation, where severe defects observed (Fig. S4C) can support effective depletion. However, for some strains, we do not have GFP makers and will need to check the RNA levels. 

      Author response image 5.

      Representative confocal image of GFP::HCP-3 and mCherry::H2B at the NEBD time point of P1 cell in the control, single and double RNAi. Scale bar, 5 μm.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      We agree that these chromatin modifications and transcriptional alterations could be related to the nuclear transport failure and developmental delay in npp-3 disruption. We will discuss this possibility in results and discussion. 

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      We agree that the evidence for SAC activation during prophase is indirect. The prophase extension (100 s) in npp-3 RNAi is a functional assay to support SAC activation, and the suppression in npp-3 mdf-1 double RNAi suggests dependency. The changes in MDF1/MDF-2 intensities are modest. Biochemical analyses of MCC assembly in C. elegans mixed cell cycle stage embryos is challenging to demonstrate SAC activity in prophase. 

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf- 1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF- 1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

      We did not include the BAF-1/CENP-C bridge hypothesis in the abstract or the model, and we will discuss this as a speculative mechanism rather than an established dependency.

      References:

      (1) Rodenas, E., Gonzalez-Aguilera, C., Ayuso, C. & Askjaer, P. Dissection of the NUP107 nuclear pore subcomplex reveals a novel interaction with spindle assembly checkpoint protein MAD1 in Caenorhabditis elegans. Mol Biol Cell 23, 930-944 (2012).

      (2) Schumacher, J.M., Ashcroft, N., Donovan, P.J. & Golden, A. A highly conserved centrosomal kinase, AIR-1, is required for accurate cell cycle progression and segregation of developmental factors in Caenorhabditis elegans embryos. Development 125, 4391-4402 (1998).

      (3) Kotak, S., Afshar, K., Busso, C. & Gonczy, P. Aurora A kinase regulates proper spindle positioning in C. elegans and in human cells. J Cell Sci 129, 3015-3025 (2016).

      (4) Hachet, V. et al. The nucleoporin Nup205/NPP-3 is lost near centrosomes at mitotic onset and can modulate the timing of this process in Caenorhabditis elegans embryos. Molecular Biology of the Cell 23, 3111-3121 (2012).

      (5) Encalada, S.E., Willis, J., Lyczak, R. & Bowerman, B. A spindle checkpoint functions during mitosis in the early Caenorhabditis elegans embryo. Molecular Biology of the Cell 16, 1056-1070 (2005).

    1. Author response:

      Reviewer #1:

      We thank Reviewer #1 for the thoughtful critique. We have revised the manuscript to clarify the physiological rationale for the experimental system, the potential relevance of mitochondrial STX–VDAC signaling to neurodegeneration, and the appropriate scope of our conclusions.

      (1) Lack of Rationale

      The study provides no justification for investigating sex-specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in POMC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      We thank the reviewer for raising this important point. We agree that the original manuscript did not sufficiently explain the rationale for using POMC neurons.

      Our rationale is primarily physiological rather than disease-specific. POMC neurons are a well-characterized, metabolically sensitive, and estrogen-responsive neuronal population in which membrane-initiated estrogen signaling and the actions of STX have been extensively characterized. They therefore provide a physiologically relevant neuronal system for identifying the molecular mechanisms through which STX regulates mitochondrial function.

      The potential relevance to neurodegeneration is supported by evidence that hypothalamic and POMC neuronal function can be disrupted in neurodegenerative disease models. Do and colleagues (2018) reported hypothalamic neurodegeneration, increased inflammatory and apoptotic markers, and reduced POMC neuronal populations in 3xTg-AD mice and further showed that exercise attenuated hypothalamic apoptosis and restored POMC neuronal populations (Do, Laing et al. 2018). In addition, Shen and colleagues (2016) demonstrated disruption of POMC/MC4R signaling in APP/PS1 mice and showed that restoration of this pathway improved synaptic function (Shen, Tian et al. 2016).

      These studies do not establish POMC neurons as a primary site of AD pathology, but they demonstrate that POMC-related neuronal systems can be vulnerable to neurodegenerative processes. This provides a broader biological context for investigating mitochondrial mechanisms in this neuronal population.

      There is also a strong physiological rationale for examining estrogen-sensitive mechanisms in POMC neurons. These neurons are established targets of 17β-estradiol and are highly responsive to metabolic and mitochondrial state. STX is a non-steroidal estrogenic compound that activates membrane-initiated estrogen signaling, and our previous studies demonstrated neuroprotective and mitochondrial effects of STX in a neurodegenerative disease model.

      Thus, POMC neurons were used because they provide a well-defined estrogen-responsive and metabolically sensitive neuronal population in which STX signaling can be mechanistically investigated—not because we consider them an initiating site of AD pathology.

      We have revised the Introduction and Discussion accordingly. The revised manuscript now emphasizes the physiological significance of STX–VDAC signaling for neuronal mitochondrial function and presents its potential relevance to neurodegeneration as an important area for future investigation.

      (2) Weak Link to AD Pathogenesis

      Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      We agree that the present experiments do not establish VDAC2 as an etiological or pathological driver of AD. This was not the objective of the present study.

      The primary goal was to identify the molecular target(s) through which STX influences mitochondrial function. Using BF-STX chemoproteomic capture and competition experiments together with single-cell gene-expression analysis, planar lipid membrane electrophysiology, and mitochondrial metabolic analyses, we identify VDAC proteins as mitochondrial targets of STX and demonstrate functional effects of STX on VDAC channel properties and mitochondrial bioenergetics.

      Our previous studies demonstrated neuroprotective and mitochondrial effects of STX in the 5xFAD model (Lee, Bostick et al. 2025), providing a neurodegenerative context that motivated the present mechanistic investigation. However, the current findings should not be interpreted as establishing VDAC2 as an AD pathogenic mechanism.

      We have revised the manuscript accordingly. The STX–VDAC interaction is now presented principally as a mitochondrial mechanism with potential relevance to neuronal physiology and neurodegeneration. Whether this pathway contributes to the neuroprotective actions of STX in disease models will require direct experimental testing.

      (3) Unclear Relevance to AD Contexts

      While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in POMC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex-specific AD phenotypes.

      We agree that the present experiments do not directly establish the relevance of the STX–VDAC interaction to mitochondrial dysfunction in AD.

      The current study was designed to identify and characterize the molecular mechanism underlying the mitochondrial actions of STX, rather than to test this pathway in a specific neurodegenerative disease model. Our findings demonstrate that STX interacts with VDAC proteins, modifies VDAC channel properties, and alters mitochondrial bioenergetics.

      The study builds on our previous findings in the 5xFAD model, in which STX reduced amyloid-β-associated pathology and affected mitochondrial function (Lee, Bostick et al. 2025). The identification of VDAC proteins as STX targets therefore provides a mechanistic foundation for future studies examining whether this pathway contributes to neuronal protection under neurodegenerative conditions.

      We have revised the Discussion to acknowledge the absence of direct disease-model validation. Future studies using conditional or neuron-specific manipulation of VDAC isoforms will be important for determining the physiological and neuroprotective significance of STX–VDAC signaling in vivo.

      Accordingly, our conclusions now emphasize what is directly supported by the present experiments: STX interacts with VDAC proteins and modulates VDAC channel function and mitochondrial bioenergetics. The relevance of this mechanism to neurodegenerative disease remains to be established.

      (4) Interpretation of Competitive Binding Data

      The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

      We thank the reviewer for highlighting this important point. We agree that the original manuscript placed too much emphasis on VDAC2 based on the chemoproteomic data.

      Our chemoproteomic experiments identified VDAC1, VDAC2, and VDAC3 as STX-interacting proteins. Competition with unlabeled STX produced a particularly clear reduction in VDAC3 labeling, including complete loss of the VDAC3 signal at three molar equivalents of unlabeled STX. We therefore agree that VDAC3 represents an important candidate STX target.

      We have revised the Results and Discussion so that the competition experiment is no longer interpreted as demonstrating preferential or exclusive binding to VDAC2.

      However, the persistence of VDAC1 and VDAC2 labeling may also be influenced by properties of the BF-STX photoprobe as alkyl diazirine probes can preferentially photolabel membrane proteins (Kleiner, Heydenreuter et al. 2017). We now present this only as a possible technical consideration rather than an explanation established by our data.

      Our subsequent emphasis on VDAC2 was based on the integration of several observations. Single-cell qPCR demonstrated that Vdac2 is the predominant transcript in the native POMC neurons examined (revised Figure 3B), with an approximate expression hierarchy of Vdac2 > Vdac3 >> Vdac1. In addition on a technical note, recombinant VDAC3 is more difficult to reconstitute reliably into artificial membranes because of its lower stability in detergent.

      The revised manuscript therefore recognizes all three VDAC isoforms as candidate STX targets but does not claim that VDAC2 is the exclusive or preferential target. Direct comparisons of STX binding and functional modulation among the three isoforms will be required to establish isoform selectivity.

      Reviewer #2:

      We thank Reviewer #2 for the positive assessment of our chemoproteomic strategy and multidisciplinary characterization of the STX–VDAC interaction. We also appreciate the reviewer’s identification of important limitations and directions for future investigation.

      (1) Physiological Relevance of Cell-Line Experiments

      Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance.

      We agree that the use of immortalized neuronal cell models represents an important limitation.

      However, these models provided the cellular material, reproducibility, and experimental control required for chemoproteomic target capture and Seahorse metabolic measurements. Importantly, we complemented these studies with single-cell analysis of native POMC neurons and electrophysiological characterization of recombinant VDAC channels.

      Nevertheless, these approaches do not substitute for direct demonstration of STX–VDAC signaling in primary POMC neurons or in vivo. We have therefore revised the Discussion to emphasize that establishing the physiological significance of this mitochondrial pathway will require validation in primary neuronal preparations and whole-animal models.

      (2) Structural Binding Site of STX on VDAC

      The exact structural binding site of STX on VDAC remains unresolved.

      We agree. BF-STX chemoproteomics identifies proteins interacting with STX in a cellular environment but does not resolve the amino acid residues or structural pocket responsible for binding.

      We have clarified this limitation in the Discussion. Determining the STX-binding site will require complementary approaches such as targeted mutagenesis, direct binding measurements with purified VDAC proteins, and structural studies in membrane-like environments, including lipid nanodiscs. Such studies should also determine whether STX recognizes a conserved feature among VDAC isoforms or exhibits isoform selectivity.

      (3) Lack of VDAC2 Loss-of-Function Experiments

      No loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects.

      We agree that VDAC loss-of-function experiments would provide an important additional test of causality.

      Our present evidence is convergent: chemoproteomics identifies VDAC proteins as STX-interacting targets; single-cell analyses demonstrate VDAC expression in POMC neurons; electrophysiological studies show that STX modifies VDAC channel properties; and metabolic analyses demonstrate STX-dependent changes in mitochondrial bioenergetics. Together, these findings support a STX–VDAC mechanism but do not establish that VDAC2 alone is necessary for the mitochondrial or neuroprotective actions of STX.

      We have revised the manuscript accordingly and explicitly identify the absence of loss-of-function experiments as a limitation.

      Because whole-body VDAC2 loss-of-function is associated with severe developmental consequences, future studies will require conditional or neuron-specific approaches. Parallel evaluation of VDAC1 and VDAC3 will also be important because all three isoforms were identified by chemoproteomics and potential functional redundancy may complicate single-isoform manipulations.

      (4) Non-linear Dose-Response at Higher STX Concentrations

      The non-linear dose-response at higher STX concentrations also requires further investigation.

      We agree that the non-linear concentration-response warrants further investigation and have revised the Discussion to avoid interpreting the STX response as a simple monotonic concentration-response relationship.

      One possibility is that STX engages multiple mitochondrial targets with different apparent affinities and opposing effects on respiration. At lower concentrations, STX may preferentially engage a higher-affinity target, potentially VDAC, whereas higher concentrations may recruit lower-affinity targets that constrain this response. Consistent with this possibility, our BF-STX dataset (Supplemental Tables) identified several mitochondrial proteins involved in oxidative phosphorylation and metabolite transport, including ATP5F1C, NNT, SLC25A4, SLC25A5 and SLC25A3. ATP5F1C is required for efficient mitochondrial ATP production (Fiorillo, Scatena et al. 2021), and estrogenic regulation of ATP synthase has been reported (Massart, Paolini et al. 2002, Moreno, Moreira et al. 2013). NNT, SLC25A4, SLC25A5 and SLC25A3 also regulate mitochondrial respiration, redox balance, and ATP production (Mayr, Merkel et al. 2007, Lopert and Patel 2014). However, BF-STX enrichment does not establish direct STX binding, relative affinity, or functional modulation of these proteins. We therefore present the multi-target explanation only as a hypothesis. Direct binding and concentration-dependent target-engagement studies will be required to determine whether the non-linear response reflects recruitment of a lower-affinity mitochondrial target, concentration-dependent effects on VDAC itself, or downstream mitochondrial feedback.

      We thank the editors and reviewers again for their constructive comments. The revised manuscript more clearly defines the physiological rationale for the experimental system, establishes the scope of the STX–VDAC mitochondrial mechanism supported by our data, and places its potential relevance to neurodegeneration in an appropriately forward-looking context.

      References cited in the response

      Do, K., B. T. Laing, T. Landry, W. Bunner, N. Mersaud, T. Matsubara, P. Li, Y. Yuan, Q. Lu and H. Huang (2018). "The effects of exercise on hypothalamic neurodegeneration of Alzheimer's disease mouse model." PLoS One 13(1): e0190205.

      Fiorillo, M., C. Scatena, A. G. Naccarato, F. Sotgia and M. P. Lisanti (2021). "Bedaquiline, an FDA-approved drug, inhibits mitochondrial ATP production and metastasis in vivo, by targeting the gamma subunit (ATP5F1C) of the ATP synthase." Cell Death Differ 28(9): 2797-2817.

      Kleiner, P., W. Heydenreuter, M. Stahl, V. S. Korotkov and S. A. Sieber (2017). "A Whole Proteome Inventory of Background Photocrosslinker Binding." Angew Chem Int Ed Engl 56(5): 1396-1401.

      Lee, H.-J., Z. Bostick, J. Doherty, T. L. Swanson, M. J. Kelly, J. F. Quinn, N. E. Gray and P. F. Copenhaver (2025). "Neuroprotection against beta-amyloid toxicity by the novel estrogen receptor modulator STX requires convergent signaling pathways." Frontiers in Molecular Neuroscience Volume 18 - 2025.

      Lopert, P. and M. Patel (2014). "Nicotinamide nucleotide transhydrogenase (Nnt) links the substrate requirement in brain mitochondria for hydrogen peroxide removal to the thioredoxin/peroxiredoxin (Trx/Prx) system." J Biol Chem 289(22): 15611-15620.

      Massart, F., S. Paolini, E. Piscitelli, M. L. Brandi and G. Solaini (2002). "Dose-dependent inhibition of mitochondrial ATP synthase by 17 beta-estradiol." Gynecol Endocrinol 16(5): 373-377.

      Mayr, J. A., O. Merkel, S. D. Kohlwein, B. R. Gebhardt, H. Böhles, U. Fötschl, J. Koch, M. Jaksch, H. Lochmüller, R. Horváth, P. Freisinger and W. Sperl (2007). "Mitochondrial phosphate-carrier deficiency: a novel disorder of oxidative phosphorylation." Am J Hum Genet 80(3): 478-484.

      Moreno, A. J., P. I. Moreira, J. B. Custódio and M. S. Santos (2013). "Mechanism of inhibition of mitochondrial ATP synthase by 17β-estradiol." J Bioenerg Biomembr 45(3): 261-270.

      Shen, Y., M. Tian, Y. Zheng, F. Gong, A. K. Y. Fu and N. Y. Ip (2016). "Stimulation of the Hippocampal POMC/MC4R Circuit Alleviates Synaptic Plasticity Impairment in an Alzheimer's Disease Model." Cell Rep 17(7): 1819-1831.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.

      The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      We agree with the reviewer that the TDP-43 2KQ mutant is a non-physiological variant. To address this concern, we have removed all statements that extrapolate our findings to disease pathogenesis. As reflected in the revised title, we now present this study as an investigation of factors that modulate TDP-43 phase transition and aggregation using a sensitized model system, rather than as a direct disease model. The Abstract and Discussion have also been revised accordingly.

      The choice of the 2KQ mutant was dictated by the requirements of the screening strategy. To identify modulators of TDP-43 phase transition, it was necessary to use a TDP-43 variant that reliably undergoes phase separation in cells within an experimentally practical timeframe. In this context, we believe the 2KQ mutant provides a suitable and justified experimental tool. We have revised the manuscript to clearly distinguish observations made with this engineered construct from conclusions regarding physiological or disease-associated TDP-43.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.

      The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      We agree with the reviewer and have removed all speculative statements regarding the relative contributions of cytoplasmic aggregation and RNA splicing defects to disease pathogenesis. The organoid section has also been revised to focus solely on the experimental findings. Specifically, we only present evidence that inhibition of nuclear export in a sensitized organoid model promotes the accumulation of cytoplasmic phosphorylated TDP-43 without making broader claims regarding its role in disease pathogenesis.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.

      The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      We have now added some sentences on page 11 to acknowledge the limitation of our experiments. It reads as “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.”

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.

      The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      We thank the reviewer for this point. We have now explicitly mentioned in the discussion that the effect of XPO-1 on anisosome dynamics is likely mediated by an indirect mechanism. We also acknowledge that we do not fully understand why over-expression of XPO-1 causes TDP-43 to accumulate in gel-like structures in the cytoplasm. Although we did not check whether overexpression of XPO1 increases cargo export, we cited studies showing that over-expressed XPO1 disrupts the normal distribution of cargos between nucleus and cytoplasm on page 7. To avoid confusion, we also revised the result part on page 7, emphasizing on the difference rather than the similar increase in puncta size by opposing manipulations.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.

      The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      We removed the RNA-seq analysis because both the reviewers and editors agreed that the modest transcriptional rescue should not be overinterpreted. Upon reconsideration, we believe that reinstating these data would not address the reviewer's principal concern, namely whether nuclear export inhibition reduces TDP-43 aggregation in organoids. We have therefore chosen to limit our conclusions to the direct observation supported by the current data, namely a reduction in cytoplasmic phosphorylated TDP-43 immunoreactivity. We also add a sentence to acknowledge that “Whether the reduction in pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined” on page 10.

      (6) No functional neuronal readout has been provided for the organoid model.

      The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      We thank the reviewer for this helpful suggestion. We agree that assessments of neuronal survival or function would be important if the manuscript were claiming that nuclear export inhibition improves neuronal health or rescues disease phenotypes in the organoid model. However, in response to the reviewers' comments regarding the physiological relevance of the homozygous K181E organoids, we have substantially revised both the framing and interpretation of this section.

      Specifically, we have revised the statement to read, "These results imply that maintaining TDP-43 in the nuclear demixed liquid state might diminish p-TDP-43 accumulation but whether the reduction of pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined," thereby limiting our conclusion to the direct experimental observation. We no longer make claims regarding disease modification or functional rescue in the organoid model. Given this revised scope, we believe that additional measurements of neuronal survival or function, while certainly of interest, are not essential to support the conclusions presented in this study. We have also revised the conclusion to explicitly acknowledge that the functional consequences of reducing p-TDP-43-positive puncta (whether this can be translated to reduced aggregation) remain to be determined by future studies (page 10).

      (7) The abstract and title continue to overstate the mechanistic conclusions.

      Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      We have revised the title of the paper, reframing it as a screen that reveals modulators of TDP-43 phase separation. The last sentence of the abstract is also revised accordingly. It now reads as “These findings identify multiple modulators of TDP-43 phase transitions in a sensitized model system and establish a framework for further dissecting the link between nuclear transport and TDP-43 phase dynamics.” We also tone down our conclusions and discussions.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.

      The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      While we agree with the reviewer that studies relying on siRNA should provide sufficient information regarding knockdown efficiency and specificity, we respectfully disagree that this should be a major concern in the present study. As explained in the manuscript, we deliberately chose not to pursue mechanistic studies using chronic XPO1 knockdown because prolonged depletion of this essential nuclear export factor is likely to produce secondary effects that could complicate data interpretation. Instead, we employed multiple chemically distinct XPO1 inhibitors to achieve acute inhibition, thereby minimizing indirect consequences while providing a more appropriate approach for mechanistic analysis.

      We agree that assessing knockdown efficiency is technically straightforward. However, because our mechanistic conclusions are based primarily on acute pharmacological inhibition rather than siRNA-mediated depletion, we prioritized experiments that directly addressed the central mechanistic questions raised by the reviewers, particularly the semi-permeabilized cell assay. Moreover, the XPO1 inhibitors used in this study are well-characterized, highly specific compounds that have been extensively validated and widely used in the literature. We therefore believe that our experimental strategy provides a reliable basis for the conclusions presented.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.

      The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      We thank the reviewer for this helpful suggestion. We have now added a sentence on page 11, which state that “our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires validation.”

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

      We thank the reviewer for his/her appreciation of our previous revision. We hope that the new changes now satisfactorily address the remaining concerns.

      Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      We thank the reviewer for acknowledging the strength and the potential significance of our study.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:

      (a) What is the smear in Figure S1 after VLX treatment?

      We thank the reviewers for the positive assessment. We do not know why VLX treatment causes a fraction of TDP-43 to migrate slowly. We suspect that it may form detergent-insoluble aggregates. However, we cannot be sure whether this occurred during drug treatment or sample preparation. We now add a sentence in the figure legend to clarify this point.

      (b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.

      We reasoned that the reduction in TDP-43 was probably not caused by protein elimination because the puncta could be reformed when permeabilized cells were incubated with exogenously added cytosol and ATP/GTP. We have revised the text to avoid this confusion. The revision on page 8 reads as “The reduction in TDP-43 signal probably resulted from a shift of TDP-43 from a phase-separated high fluorescent state into a soluble state with reduced fluorescence intensity (Zhang et al., 2026). We attributed this phenotype to the depletion of cytosolic factors and ATP during cell permeabilization because it is known that anisosome formation and maintenance require HSP70, a cytosolic ATPase (Yu et al., 2021).”.

      (c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.

      Since TDP-43 protein was apparently still in the nucleus after cell permeabilization (see above) but became invisible, the best interpretation is that the protein is shifted into a soluble state, which reduces the fluorescence intensity substantially. We have revised the text to clarify this point. We also cited a recent study showing that EGFP-alpha-synuclein oligomerization/aggregation enhances its fluorescence intensity in cells.

      (d) Why did the authors choose cow lover cytosol for this study?

      The main reason is because we have access to a large amount of cow liver cytosol that is known to have activities in in vitro reconstitution assays. We now cite several papers from us that reported the use of the same cytosol in other in vitro assays in the method (page 14). 

      (e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.

      We now revise the schematic in Figure 6A and include more details in the method and figure legend.

      (f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      The total TDP-43 level was shown by immunostaining in green in Figure 7. This was used as a control to show that the increase in p-TDP-43 was not simply caused by an overall increase in its protein level. We have added the quantification to Figure 7C. We also discuss the reduced nuclear TDP-43 in organoids bearing the disease mutation in the main text.  

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      We agree that the structures reformed after incubating permeabilized cells with cytosol and ATP/GTP are not fully characterized. Due to their small size, we could not see the typical ring-shaped anisosome morphology. FRAP experiment is also tricky. Due to these issues, we have revised the text to acknowledge that we do not know the exact identity of these structures. We speculate that they are anisosome-related because like anisosome formation, it depends on cytosolic factor and energy (page 9). It is worth noting that whether these structures are anisosomes is not the main conclusion of this experiment. We conclude from this experiment that TDP-43 was still in the nucleus after cell permeabilization (not degraded). The fact that we could not see the protein likely because the protein was in a low-fluorescence soluble state.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      Since the formation of anisosome requires HSP70, a cytosolic chaperone that likely needs to be imported into the nucleus, we included cytosol and ARS/GTP in our in vitro reaction. We revise the description in the result part to improve clarity (page 8-9).

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      As discussed in Yu H et al., Science 2021, proteins in anisosomes cannot be stained by antibodies due to an antibody accessibility issue. It was mentioned in our paper as “since antibody staining could not conclusively demonstrate the sequestration of endogenous XPO-1 in anisosomes due to an antibody penetration barrier {Yu, 2021 #937}.” We now revise this section completely to better clarify this point. We could see overexpressed XPO1 in anisosome because it has a mCherry tag.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      The full characterization of the iPSC-derived organoids is presented in a second paper that is posted in BioRxiv (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v2), which is cited. This manuscript reports not only the accumulation of p-TDP43, but also other ALS-related phenotypes including cell death, gene transcriptional changes, cryptic exon inclusion etc. in mutant organoids.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      The longer treatment (24 h) was used in the chemical genetic screen in which different drugs may act with different efficiency. To maximize our chance of detecting more drug effect, we used a longer treatment scheme. For later follow-up experiments involving Spuatin-1, Bortezomib, TRP, because the phenotype appears quickly. To avoid secondary effects from long treatment, we shortened the treatment to 3-5 hours. For LMB treatment, we used long treatment to reveal the steady-state phenotype (anisosome enlargement in size and reduction in number has reached maximum). This time point was determined in Figure 5A-C. In contrast, shorter treatment (5 h) was to reveal early changes that might be causal to the end-point phenotypes (e.g. the anisosome fusion phenotype could be detected as early as 5 h post-treatment). We have added some explanations in the result section to make this point clear.

      (6) Several figure legends require clarification:

      We thank the reviewer for pointing out the inconsistencies. We have corrected the outstanding issues, as explained below.

      (a) In the section stating “Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.

      Thanks for pointing out this error. This is now corrected.

      (b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.

      Due to the large sample size, each concentration was analyzed once. We have added other information to the figure legend. We also remove the redundant inaccurate information from the figure legend. The treatment time shown in the figure is correct as it was also indicated in the method.

      (c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.

      We thank the reviewer for noticing the discrepancy and apologize for the error. We have corrected the figure legend to indicate that 4-6 anisosomes were analyzed for each condition. In Figure 2G, each data point represents the initial fluorescence loss rate averaged from the first 10 sec after reverse photobleaching.

      (d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."

      In Figure 3B, we used Z score >2 as the threshold. This is now mentioned in the legend and defined in the method. There is no Figure 3G.

      (e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.

      We have changed the figure legend throughout the paper accordingly.

      (f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.

      In Figure 5B, the graph reflects data collected from 3 independent replicates. To ensure reliable baseline measurement, for each experiment, two independent control samples were included, which is why it has 6 data points. In Figure 5H, each dot represents a randomly selected imaging field. We now mention the total number of fields analyzed.

      (g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."

      These are all fixed. Thank you for pointing this out.

      (h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R<sup>2</sup>, p-value).

      We now include the cell number in the legend and R<sup>2</sup> and p-value in the figure.

      (i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.

      We now revise the experimental scheme in Figure 6A to better explain the experiment and the sequence of different events. For Figure 6D, permeabilized cells were incubated with cytosol with or without ARS/GTP for 40 min. This information is added to the figure legend.

      (j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

      To improve clarity, we change mean spot volume to p-TDP-43 puncta mean volume. This refers to the average volume of segmented phosphorylated TDP-43-positive puncta.  

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

      We thank the reviewer for this helpful suggestion. We have added a few sentences in the discussion (page 10) to acknowledge the limitation of the semi-permeabilized cell assay. Specifically, we mentioned that “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.” We also revise our manuscript throughout to avoid over-interpretation.  

      Recommendations for the authors:

      Editor's notes:

      The value of the work is clear.

      We also recognise that it may not be possible/practical to get around the 'incomplete' appellation attached to this body of work, by further experiments. However, there may be scope here to retreat from the less well supported mechanistic claims-by editing the title, abstract and discussion and thus earn a 'solid' descriptor on a revised paper that remains a useful addition.

      The authors are best placed to consider creatively how to achieve this.

      We thank the editors for this helpful suggestion. We have revised the manuscript extensively to address every single concerns of the reviewers.

      Reviewer #1 (Recommendations for the authors):

      This revision addresses some minor concerns and adds one mechanistic experiment of genuine value. However, the major deficiencies of the original submission persist: the exclusive reliance on non-physiological TDP-43 model systems without validation in more disease-relevant contexts, the unresolved asymmetry between the two XPO1 perturbation phenotypes, the thin organoid section that now has fewer readouts than before the revision, and overstatement of the mechanistic conclusions in the abstract and Discussion. The manuscript in its current form still does not provide methods, data, and analyses that sufficiently support the primary claim that nuclear export is an established key regulator of TDP-43 phase transitions with mechanistic and disease relevance.

      As mentioned before, we have clarified the interpretation of the data, tempered conclusions where appropriate, and revised the text to explicitly acknowledge the limitations of the current study. We hope that these changes satisfactorily addressed the reviewer’s concern.

      Reviewer #2 (Recommendations for the authors):

      I would suggest adjusting the title to match the data, which shows so much more than just an effect of nuclear export.

      We thank the reviewer for this suggestion. We have changed the title to “Cellular modifiers of TDP-43 phase transition and cytoplasmic aggregation”

      When introducing the semi-permeabilized cell-based in vitro assay, it would be helpful to add a short statement describing what this model resembles, and what advantage it can bring to use this system in the context of the study.

      We have revised this section extensively and hope that improves the clarity. See marked text in page 8-9.

    1. Author response:

      The following is the authors’ response to the previous reviews

      The revisions in this version are minor and primarily include the addition of RT-qPCR validation experiments. In addition, the benchmarking data against commercial systems have now been incorporated into the main manuscript.

      We also sincerely appreciate Reviewer 2 and Reviewer 3 for their highly encouraging evaluations and recognition of our system's robustness. To fully address the remaining mechanistic queries from Reviewer 1 and the benchmarking concerns from the editors, we have performed quantitative RT-qPCR to directly measure transcript levels and have integrated our commercial benchmarking data into the revised manuscript.

      (1) The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

      To directly validate our claims regarding transcriptional efficiency, we performed quantitative RT-qPCR to determine the transcription levels of the reporter gene in both systems.

      First, we established no-reverse-transcriptase (no-RT) controls to verify complete DNA template removal. The Ct values for these controls remained above 34, confirming the absence of plasmid DNA contamination in our RNA samples.

      Second, transcript levels were calculated using the comparative 2<sup>-ΔΔCt</sup> method, normalized to the standard initial system (100 ng/μL T7) at 30 min. The optimized system achieved a 16.56-fold increase (P < 0.001) in transcript levels. In contrast, supplementing the initial system with high concentrations of T7 RNA polymerase (400 ng/μL) only yielded a 2.87-fold increase (P < 0.01)—which is nearly 6-fold lower than our optimized system.

      These findings perfectly mirror our protein-level titration assays (Figure S3C). Supplementing the initial system with excess T7 RNA polymerase fails to rescue either transcript accumulation or protein expression. This mutual validation confirms that transcription is severely bottlenecked in traditional systems due to rapid nucleotide degradation or inhibitory reaction environments. By streamlining the reaction buffer to seven core components and omitting runoff/dialysis, our system successfully relieves these systemic bottlenecks. We have incorporated these new qPCR findings into Figure 3B, the Methods, and the Results sections of the revised manuscript.

      (2) Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems...

      We appreciate the editor’s emphasis on establishing standard performance benchmarks. To address this important point, we would first like to highlight that our manuscript already contains extensive, rigorous benchmarking against typical cell-free platforms widely utilized in the literature. This includes detailed head-to-head comparisons with both our 35-component "initial" system and the classical, widely established "PEP-based" system across multiple expression kinetics and western blot analyses (as shown in Figure 4 and Figures S3–S4).

      To fully embrace the editor's valuable recommendations regarding standard commercial performance, we are very pleased to formally integrate our commercial benchmarking data into the revised manuscript as Figure S3C.

      To maintain technical neutrality, we have omitted specific brand names, presenting it generically as "a high-end commercial cell-free system." The data demonstrate that our optimized system significantly outperforms this commercial alternative in both expression speed and final absolute yield, reaching an absolute productivity of 0.46 mg/mL compared to approximately 0.21 mg/mL for the commercial kit.

      We are grateful for the guidance from the editors and reviewers, which has significantly strengthened the scientific rigor of our work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include: Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states.

      We thank the reviewer for their constructive feedback. We previously validated our 882 models with whole genome sequencing and teratoma formation upon mouse fat pad injection, while the parental human embryonic stem cell line (WA01 hESCs) used for P904L variant knock-in was validated by our Genome Engineering Stem Cell (GESC) core upon derivation of that variant knock-in model. We have now added both karyotyping and pluripotency staining (SOX2/OCT4) for all other hPSC lines as (new) Supplementary Figure S17 and included further description in our Methods section under “hPSC Model Generation and Culture” (pg. 19, lines 6-7).

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      We thank the reviewer for their insightful suggestions and have detailed our responses in the “recommendations to the authors” section.

      (3) Also, some errors are noted and detailed in the recommendation section.

      We thank the reviewer for catching these errors and have since corrected them, with detailed responses below.

      Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A, which cause Tatton-Brown-Rahman Syndrome (TBRS), disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation, where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability, potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with a rigorous experimental design and novel findings with a potential impact on a better understanding of the OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types, and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed, and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors partially address this tension by noting that DNMT3A LOF alone is insufficient to initiate neuronal differentiation, i.e., V-NPCs upregulate neuronal and synaptic genes while retaining progenitor identity, implying that transcriptomic priming and commitment to differentiation are decoupled. However, the relationship between the proliferative phenotype and the epigenetic priming phenotype remains mechanistically unresolved. The manuscript documents mTOR pathway upregulation at the protein level and identifies shared DEGs that include proliferative regulators, but it does not establish whether mTOR-driven proliferation and mCG-loss-driven neuronal gene de-repression/enhanced differentiation are causally linked or represent two independent consequences of DNMT3A LOF.

      We thank the reviewer for their comment and agree that this phenotype, whereby progenitors exhibited both increased proliferation and hallmarks of gene expression associated with neuronal differentiation is striking and interesting, given that these are typically antagonistic paradigms during normal development.

      We documented that these phenotypes involve upregulated expression of both neuronal/synaptic and proliferative genes in V-NPCs (Figure 2d), with concomitant loss of repressive DNA methylation at regulatory elements associated with these genes (Figure 2f, Supplementary Data 5). In this work, DNMT3A mutation had a more prominent role in de-repressing neuronal and synaptic gene expression to promote hallmarks of neuron differentiation, while playing a relatively less central role in direct regulation of proliferation genes, as seen from the relative prominence of neuronal/synaptic- versus proliferation-related GO terms in our Supplementary Data 5 table (pg. 6, lines 16-19).

      To examine the mechanisms underlying increased V-NPC proliferation in our TBRS models, we assessed a potential relationship with the PIK3/AKT/mTOR pathway, as this is implicated in increased proliferation resulting from DNMT3A-associated mutation in myeloid leukemia (Dai et al., 2017, PMID: 28461508). In our work, DNMT3A mutation increased the expression and/or phosphorylation of mTOR signaling pathway targets specifically in V-NPCs (Figure 1q-r, Supplementary Figure S3a-d). However, while TBRS mutation directly affected repressive DNA methylation at a suite of cell proliferation-related genes, these did not include the PIK3/AKT/mTOR pathway genes themselves, suggesting an indirect relationship between altered DNA methylation and increased mTOR signaling.

      We have since incorporated discussion of how DNMT3A-mediated gene repression and levels of PIK3/AKT/mTOR pathway signaling may be interacting, providing a framework for future studies to identify how these related OGID gene mutations may converge mechanistically (pg. 5, lines 19-21; pg. 16, lines 8-10).

      (2) Relatedly, the rapamycin rescue experiment is a valuable proof-of-concept for the PIK3/AKT/mTOR convergence but is limited to a single dose in a single model (882) with a single readout (Ki67+ proliferation). Given the prominence of mTOR pathway convergence in the manuscript as a potential shared therapeutic avenue across OGIDs, the data supporting this claim are somewhat preliminary. It remains unknown whether mTOR inhibition rescues downstream phenotypes (neurogenesis, gene expression, neuronal maturation) or whether less severe TBRS models respond similarly. This might also help tackle the first comment above. e.g., if mTOR inhibition rescued proliferation but not the transcriptomic priming, that would support two independent mechanisms.

      We thank the reviewer for their comment. We explored both the overall levels and phosphorylation of proteins involved in PIK3/AKT/mTOR signaling in the 882, 904, Del1, Del2, and KO V-NPC models (Figure 1q-r, Supplementary Figure S3a-d), finding specific increases of all proteins. We showed that rapamycin addition reversed the increased proportion of KI67+ proliferating cell nuclei resulting from 882 mutation in V-NPCs in main Figure 1s, while demonstrating that rapamycin also reduced the proportion of KI67+ nuclei observed in both less severe 904 and Del1 V-NPC models (Supplementary Figure S3e-f).

      We agree that understanding whether rapamycin treatment can rescue TBRS neuronal phenotypes would be very interesting, as previous work on Tuberous Sclerosis Complex has utilized rapamycin and other mTOR inhibitors to effectively reverse TSC-related alterations of neuronal morphology and neuronal hyperexcitability (Buttermore et al., 2025, PMID: 40792287). Future studies examining convergent mechanisms and therapeutics for OGIDs should examine how similarly targeting this and related pathways rescues altered neuronal morphology, maturation, and function, as we have demonstrated that TBRS mutation has subsequent consequences for V-IN differentiation, maturation, and function. This point has been detailed in the discussion section on pages 15-16.

      (3) The claim that H3K27me3 compensates for mCG loss is an important mechanistic point, but the current data do not distinguish between active compensation, in which EZH2 is recruited in response to methylation loss, and functional redundancy, in which H3K27me3 is independently established and becomes the dominant repressive mark once DNA methylation is reduced. The EZH2 knockdown/inhibition experiments show that H3K27me3 is sufficient to maintain repression at hypo-DMR sites, but they do not establish that H3K27me3 gain is itself a response to methylation loss. Because H3K27me3 profiling was performed only in the severe 882 model, it is also unclear whether H3K27me3 gain scales with DNMT3A LOF severity, as a compensatory model would predict. Finally, the EZH2 overexpression rescue is performed in V-NPCs, whereas the compensation model is developed primarily in D-NPCs, making it difficult to assess whether the same mechanism operates in the lineage where it was originally inferred.

      We thank the reviewer for the opportunity to clarify our findings and experimental reasoning. A previous study using a conditional Dnmt3a knockout mouse model (Li et al., 2022, PMID: 35604009) demonstrated increased expression of multiple PRC2 components following the loss of Dnmt3a. This study demonstrated that sites which lost DNA methylation gained H3K27me3 in postnatal neurons upon Dnmt3a loss. Therefore, we hypothesize that the gain of H3K27me3 likely occurs in response to loss of DNMT3A methylation.

      While we did not perform CUT&Tag for H3K27me3 in our less severe models, we did validate gene expression changes following EZH2 knockdown and inhibition in both the R882H (Figure 4g-h) and P904L (Supplementary Figure S8b) models, finding that gene expression was unchanged in the model with the less severe DNMT3A mutation (P904L). Based upon these findings, we hypothesized that compensatory H3K27me3 may occur only upon severe DNMT3A loss, as seen in the dominant-negative R882H model. Furthermore, as H3K27me3 compensation was more prominent in D-NPCs, we hypothesized that this might be sufficient to prevent de-repression and aberrant neuronal gene repression upon loss of DNMT3A-mediated repression in D-NPCs. However, since TBRS mutation caused the most prominent de-repression of neuronal gene expression in V-NPCs, we also tested whether EZH2 overexpression could reverse this, finding that it partially suppressed this dysregulated neuronal gene expression. To better clarify this logic and the findings, we have made text edits to this results section and referenced Li et al., 2022 in both the results (pg. 9, lines 16-21; pg. 10, lines 4-7,10-12) and discussion (pg. 16, lines 12-15).

      (4) The narrative framing of dorsal neuron development as unaffected by DNMT3A LOF is somewhat at odds with the data presented. The 882 D-NPCs show substantial DNA methylation changes, and TBRS D-INs exhibit what the authors describe as "substantive transcriptomic differences" involving persistent expression of pluripotency and progenitor genes, which seems to be a distinct but potentially significant phenotype. The impact of DNMT3A loss between ventral and dorsal lineages might be more accurately framed as divergent in nature rather than specific to a certain population.

      We thank the reviewer for their comment. While TBRS mutations appear to have a significantly stronger effect on V-NPCs and subsequently V-INs, both transcriptomic and methylation alterations do also occur upon TBRS mutation in D-NPCs and D-INs, as noted in Supplemental Figure S4d, S11, and Supplemental Data 2. However, we observed substantially greater molecular alterations in V-NPCs/V-INs, a lack of overt cellular phenotypes in D-NPCs where assayed, and a lack of functional consequences in matured D-INs, suggesting a more significant requirement for DNMT3A in regulating the differentiation and subsequent maturation of cortical inhibitory interneurons during embryonic and early pre-natal development, the developmental periods that we can readily model in hPSC-derived neurons.

      It should also be noted that these hPSC differentiation models do not recapitulate post-natal deposition of non-CpG (mCA) DNA methylation, a mechanism disrupted postnatally by TBRS-associated mutations in our prior work in murine models (Harrison Gabel; e.g. Beard et al., 2023, PMID: 37952155), which we have now added in the results section (pg. 7, lines 8-11). Therefore, we hypothesize that if we could sufficiently mature D-INs to a state that modeled postnatal development and recapitulated this non-CpG methylation, we might be able to detect cellular and functional phenotypes in later stage D-INs. To avoid misinterpretation, we have altered the language in the results section to confirm that there are both transcriptomic and methylation changes in our D-NPCs/D-INs, but that these are not accompanied by cellular phenotypes or neuronal dysfunction (pg. 6, lines 1-2; pg. 7 line 3; pg. 7, lines 20-23; pg. 8, lines 1-2; pg. 8, lines 18-23; pg. 9, lines 1-4).

      (5) SST stainings are not entirely convincing. They appear mostly nuclear, and some instances localized to rosettes in organoids, whereas the protein is largely confined to processes and is expected to be found outside progenitor-rich zones like rosettes.

      We agree that the perinuclear SST staining detected in these young ventral telencephalic-patterned organoids at day 30 differs somewhat from the more process-localized and cytosolic signal seen in later stage organoids in other studies. This may be related to the use of different commercial SST antibodies across studies but also likely reflects SST immunoreactivity in newborn neurons near the onset of SST expression. For example, immature SST-immunoreactive neurons in the early postnatal rat cortex exhibit predominant SST staining in perinuclear cytoplasm and short processes (e.g. Fig. 3 in Lee et al, PMID: 9664223) while acquiring more cytosolic and process-localized staining as postnatal neuron maturation occurs. Evaluation of immunopositivity for other markers of neurogenesis (ASCL1) and immature neurons (TUJ1) is also congruent with these findings for SST, with TBRS-associated mutations increasing in the fraction of cells in V-NPCs/V-ORGs that express these three markers.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Weaknesses:

      (1) Li et al., 2022 (referred to in the manuscript) seems to already show the interplay between H3K27me3 and Dnmt3a discussed in this study i.e., that in the absence of DNA methylation, there is an expansion of polycomb-like repression. These data should be better acknowledged in the paragraph 'Repressive H3K27me3 compensates for severe loss of DNA methylation' (page 9), given it supports the data presented in this manuscript and suggests this as a common mechanism in the interplay between these two repressive marks, as it is well established in the literature.

      We thank the reviewer for this suggestion. We have now added Li et al., 2022 to both the results section (pg. 9, lines 16-20) and our discussion section (pg. 16, lines 12-13).

      (2) The authors should acknowledge that the omics data come from a mixed population of cells.

      We thank the reviewer for their comment. We have validated that the established 2-D differentiation methods we used in this study generate cell populations with >85-90% enrichment for the desired progenitor and neuronal cell type, based upon marker expression, but acknowledge that these are bulk -omics data obtained from cells that may represent a mixed population and have now detailed this in the methods section under “Sequencing” (pg. 21, lines 16-18).

      (3) The authors are encouraged to further discuss whether the overgrowth observed in ventral GABAergic cultures or organoids compares to the overgrowth observed in diseased patients. One expects MRIs to have been performed in patients and that these could be harnessed to discern if overgrowth occurs in the cortex or ventral regions of the brain.

      We thank the reviewer for their suggestion and do note that at least one published study documents increased cortical thickness in the MRIs of TBRS patients (Jiménez de la Peña et al., 2024, PMID: 37795572); however, to our knowledge studies have not examined regional or cell type-selective overgrowth of cortical tissue in TBRS patients. Future clinical studies examining the nature of the neuronal progenitor overgrowth and resulting consequences for patient brain imaging would be of interest to better understand TBRS-associated etiology of brain overgrowth and its manifestations.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A potential explanation for the ventral organoid's sensitivity compared to the dorsal can be provided. I understand that the PRC2-mediated methylation loss was more pronounced in the ventral than in the dorsal. However, this raises the question: why is this so? Are the PRC2 components more highly expressed in ventral organoids or in GABAergic neurons than in glutamatergic neurons?

      We thank the reviewer for this question. Upon examining RNA-sequencing data for control V-NPCs and D-NPCs, we found that expression of the PRC2 subunits SUZ12, EED, and EZH2 is significantly elevated in D-NPCs relative to V-NPCs (see Author response image 1). This could contribute to restoring epigenetic repression in D-NPCs in the context of DNMT3A mutation, to yield the more limited transcriptomic changes and greater gain of H3K27me3 we observed in TBRS D-NPCs, despite their comparable levels of global mCG loss). This could also account for our finding that overexpression of EZH2 in V-NPCs could normalize some gene expression changes observed in the TBRS models.

      Author response image 1.

      (2) DNMT3A is a major enzyme that mediates non-CpG methylation during synaptogenesis. There is no discussion of this phenomenon or any related analysis. I suppose the brain organoids may be too immature to detect non-CpG methylation. Even then, this can be described in the results sections.

      We thank the reviewer for this comment. While prior work has examined an important role for DNMT3A in postnatal deposition of non-CpG methylation (Christian et al. 2020; Beard et al. 2023), we found that our hPSC models, as the reviewer suggests, were too immature to detect appreciable levels of non-CpG methylation. Accordingly, we have focused our work here on neurodevelopmental requirements for DNMT3A; this work demonstrated that many of the functional alterations of neuronal network activity originate from altered neurogenesis and neuronal maturation of cortical GABAergic interneurons. We have now included a description of the important role of non CpG methylation by DNMT3A during postnatal development and have indicated that the immaturity of hPSC-derived neurons precludes our ability to study non-CpG methylation using these models in the results (pg. 7, lines 8-11) with present language in the discussion (pg. 15, lines 8-11).

      (3) Given the Figure 4 results showing H3K27 methylation and EZH2 knockdown rescue the DNMT3A mutation, the authors argue that EZH2 and DNMT3A regulate a similar set of genes and that the two diseases are related. To claim this, there should be data showing that gene misregulation upon loss of EZH2 and DNMT3A is similar.

      Given our data, it is premature to draw direct relationships with the dysregulated genes we detected in TBRS versus the molecular basis of Weaver Syndrome; therefore, we have attempted to better reflect the potential but currently unproven relationship between these OGIDs and altered the respective language in the results section (pg. 10, lines 10-12) to reflect this. Understanding convergent mechanisms of OGIDs involving both DNMT3A (TBRS) and EZH2 (Weaver Syndrome) gene mutations will be a compelling topic for future studies, as our work here suggests that they may regulate similar gene suites.

      (4) Stronger phenotypes were observed in the R882H mutant compared to the P904L mutant. However, this cannot be translated to the functionality of the mutant protein or disease severity because one is on iPSCs and another is made in hESCs. The limitation should be noted to avoid misleading.

      We agree that the severity of each pathogenic mutation can be influenced by its presence on a different iPSC vs hESC background; accordingly, our study design indicated the hPSC background for each model and employed paired isogenic controls to clearly define the consequences of each TBRS mutation relative to a model with an identical genetic background but lacking the mutation. Our finding that the R882H mutation resulted in more severe epigenomic and transcriptomic consequences than the P904L mutation is congruent with findings made in prior work (Beard et al., 2023 PMID: 37952155; Russler-Germain et al., 2014 PMID: 24656771).

      (5) It is unclear what measurements were done for DNMT3A level quantification shown in Figure 1e- f.

      Protein quantification for DNMT3A models was performed by western blot, shown in Supplemental Fig. S16. We have now included reference to whole blots in the methods section under “Cellular Phenotyping” (page 20 line 13).

      (6) The results section of the paper for Figures 6 and 7 cites the wrong figure numbers and panels.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      (7) Some typos for PI3K (meaning PIK3) in several places, including the Abstract.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      Reviewer #2 (Recommendations for the authors):

      (1) There is a figure numbering discrepancy in the manuscript - the text references six main figures, but the figure pages include seven, with the MEA network data apparently mislabeled.

      We thank the reviewer for catching these errors and have corrected them in the new manuscript draft.

      (2) The nomenclature "D-IN" for dorsal immature neurons is potentially misleading, as "IN" conventionally denotes interneurons, which these glutamatergic cells are not.

      While we agree that IN could be interpreted as interneurons, this nomenclature is defined at an early point in the manuscript and was used to allow readers to easily identify compare findings made in D-NPCs and D-INs (and, as a counterpart, findings made in V-NPCs versus V-INs).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1.1) The manuscript “Realistic coupling enables flexible macroscopic traveling waves in the mouse cortex” by Sun, Forger, and colleagues presents a novel computational framework for studying macroscopic traveling waves in the mouse cortex by integrating realistic brain connectivity data with large-scale neural simulations.

      The key contributions include: (1) developing an algorithm that combines spatial transcriptomic data (providing detailed neuron positions and molecular properties) with voxelized connectivity data from the Allen Brain Atlas to construct neuron-to-neuron connections across 300,000 cortical neurons; (2) building a GPU-accelerated simulation platform capable of modeling this large-scale network with both excitatory and inhibitory HodgkinHuxley neurons; (3) extending phase-based analysis methods from 2D to 3D to quantify traveling wave activity in the realistic brain geometry; and (4) demonstrating that realistic Allen connectivity generates significantly higher levels of macroscopic traveling waves compared to simplified local or uniform connectivity patterns.

      The study reveals that wave activity depends non-monotonically on coupling strength and that slow oscillations (0.5-4 Hz) are particularly conducive to large-scale wave propagation, providing new insights into how anatomical connectivity enables flexible spatiotemporal dynamics across the cortex.

      The authors leverage two existing dense datasets of spatial transcriptomic data and connection strength between pairwise voxels in the mouse cortex in a novel way, allowing for the computational model to capture molecular and functional properties of neurons as determined by their neurotransmitter profiles, rather than making arbitrary assignments of excitatory/inhibitory roles. Additionally, the author’s expansion of 2D phase dynamics to 3D phase gradient analysis methods is important and can be widely applied to calcium imaging, LFP recordings, and likely other electrophysiological recordings.

      Thank you for the accurate summary of our manuscript and list of strengths.

      (1.2) The model’s Allen connectivity approach overlooks critical aspects of real cortical dynamics. Most importantly, it excludes subcortical structures, especially the thalamus, which drives cortical traveling waves through thalamocortical interactions. The authors’ method of electrically stimulating all layer 4 neurons simultaneously to initiate waves is artificially crude and bears little resemblance to natural wave generation mechanisms.

      We agree that excluding subcortical structures, especially the thalamus, is an important limitation of the current model. Because adding these structures would substantially expand the scope and complexity of the model, we now state this limitation more explicitly in the Discussion and leave it as a future extension:

      “For simulations, we choose to randomly stimulate the total population of layer 4 neurons as a way to mimic subcortical input and generate traveling waves, which can be unrealistic. Subcortical structures, such as the thalamus, are vital to cortical dynamics like slow-wave activity [1] and are known to regulate traveling waves [2]. Therefore, a direct and important future improvement would be adding subcortical structures to the model.”

      We also agree that the original constant-current stimulation was too artificial. We therefore replaced it with a 10Hz Poisson spike train delivered to layer-4 excitatory neurons across the isocortex, which more closely mimics the stochastic input that cortex receives from subcortical regions such as the thalamus. The revised stimulation protocol is described in the Results:

      “Therefore, to mimic the stochastic drive that cortex receives from subcortical regions like thalamus, we deliver a 10Hz Poisson spike train to layer-4 excitatory neurons (Figure 2a), since layer 4 is the canonical thalamocortical input layer [3]; each Poisson event applies a fixed voltage bump V<sub>stim</sub>, and no other external input is applied.”

      as well as in the Methods 4.5 Simulation.

      The revised stimulation protocol allowed us to rerun the parameter sweep under stochastic drive. A direct comparison of alternative wave-generating mechanisms remains an important direction for future work.

      (1.3) The model handles voxel-to-voxel connections crudely when neurons have mixed excitatory/inhibitory properties and varying synaptic strengths. Real connectivity differs dramatically between neuron types (pyramidal cells vs. interneurons, across cortical layers), but the model only distinguishes excitatory and inhibitory neurons. Additionally, uniform synaptic weights ignore natural variations in connection strength based on neuron type, distance, and functional role. Integrating the updated thalamocortical dataset mentioned by the authors, even at regional resolution, would substantially improve the model.

      We thank the reviewer for raising these important points regarding cell-type-specific connectivity and heterogeneity of synaptic weights. We agree that cell-type-specific connectivity and heterogeneous synaptic weights are important limitations. Because the current voxelized projectome is not cell-type specific, we now state these limitations explicitly and outline how future versions of the model could incorporate improved density and synaptic weight assumptions in the Discussion (Construction of the Allen connectivity). Specifically, we now write:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.

      Also, our model uses uniform synaptic weights within each synapse type: every AMPA synapse has conductance g<sub>AMPA</sub> and every GABA synapse has conductance g<sub>GABA</sub>. In reality, synaptic strength varies with presynaptic and postsynaptic cell type, anatomical distance, and the distribution of synaptic weights is typically heavy-tailed (e.g. log-normal [5]). Extending our algorithm to sample synaptic weights from realistic distributions would be a natural next step and is likely necessary for quantitative comparisons with electrophysiological recordings.”

      These additions clarify which aspects of the current model are constrained by the available voxelized projectome and which extensions would require cell-type-resolved connectivity data, region-specific density information, or more detailed synaptic-weight estimates.

      (1.4) While the authors bridge microscopic (single neuron) and mesoscopic (regional connectivity) data to study macroscopic (whole-cortex) waves, they don’t integrate the distinct mechanisms operating at each scale. The framework demonstrates that realistic connectivity enables macroscopic waves but fails to connect how wave dynamics emerge and interact across spatial scales systematically.

      We thank the reviewer for this insightful comment. In the revision, we added the Kuramoto synchrony analysis as a first step toward connecting scales: changes in the microscopic coupling parameter alter network synchrony, and this synchrony measure is closely associated with the macroscopic PGD observable (new Fig. 5). We agree that a full separation of layer-specific microcircuit mechanisms, region-specific connectivity motifs, and whole-cortex wave propagation remains beyond the scope of the present study. The revised text now frames the synchrony-to-wave relationship as one concrete cross-scale link that can be investigated further with this framework.

      (1.5) Claims that Allen connectivity produces higher phase gradient directionality (PGD) than local connectivity appear limited to delta oscillations at very specific coupling strengths and applied currents. Few parameter combinations show significantly higher PGD for Allen connectivity, and these are generally low PGD values overall.

      We agree that the original comparison did not sufficiently establish whether the Allen-connectivity advantage held beyond a small number of parameter choices.

      In the revised manuscript, we expanded the analysis to a 10 × 12 grid of excitatory coupling strengths and Poisson stimulus magnitudes for Allen, local, and uniform connectivity. Across this full (g<sub>AMPA</sub>,V<sub>stim</sub>) grid, Allen connectivity shows higher per-band maximum PGD than local or uniform connectivity, especially in the theta, alpha, and beta bands (Fig. 5b). The delta-band difference is small, consistent with the reviewer’s observation that the original delta-band result did not clearly separate Allen from local connectivity.

      In the revised manuscript, Fig. 5c and Fig. 5d show how PGD varies with stimulus magnitude and coupling strength individually. The full per-band PGD heatmaps for all three connectivities, together with the per-band Allen-minus-opponent gap bars, are provided in Figure 4-figure supplement 6. These additions show that the Allen-connectivity trend is not limited to a single representative operating point.

      (1.6) Broadly, it’s unclear how this computational framework can study memory, learning, sleep, sensory processing, or disease states, given the disconnect between simulated intracellular voltages and the local field potentials or other electrophysiological measurements typically used to study cortical traveling waves. While computationally impressive, the practical research applications remain vague.

      We thank the reviewer for this important point. To bridge the gap between our simulations and experimentally measured data, such as local field potentials (LFP), we used a post-hoc LFP estimation pipeline and added a dedicated Methods subsection describing it (LFP estimation from intracellular voltage). The key idea is that LFP primarily reflects the net transmembrane synaptic current in a local population, which we can reconstruct directly from the voltage traces and the connectivity used in the simulation. Please see Methods section LFP estimation from intracellular voltage.

      This pipeline allows us to test whether the macroscopic traveling-wave structure identified in the voltage traces is also present in an LFP-like signal. In the revised manuscript, we added a new Results section and show, in Fig. 6, side-by-side voltage- and LFP-based snapshots, the voltage-LFP PGD scatter (Pearson r = 0.78-0.89 per band), and per-band PGD comparisons across Allen, local, and uniform connectivity on the LFP signal. These results indicate that our conclusions are not restricted to intracellular voltage and provide a closer bridge to LFP-based experimental measurements.

      (1.7) The paper needs a clearer explanation for why medium coupling (100%) eliminates waves in Allen connectivity (Figure 6) while stronger coupling (150%) restores them.

      Thank you for requesting this clarification. The revised analysis suggests that the non-monotonic relationship between coupling strength and wave activity reflects an interaction between network synchrony and spatial organization. At weak coupling, the network has enough coordination to support propagating waves. At medium coupling, increased synaptic drive pushes the network into an asynchronous irregular state that disrupts coherent wave fronts. At strong coupling, rhythmic synchronization is re-established and again supports wave propagation.

      To substantiate this explanation quantitatively, we measured the Kuramoto order parameter R(t)=|〈e<sup>iϕ(x,t)</sup>〉<sub>x</sub>| from the generalized phase ϕ(x, t) of the band passed voltage field (see updated Quantitative measurement of neuronal activity in Methods) and reduced it to the maximum over each 1-s recording window. We then swept the same ten g<sub>AMPA</sub> values (0.005-0.050 nS) and twelve stimulus magnitudes used for the PGD sweeps, for Allen, local and uniform connectivity, in all five frequency bands. The new analysis is presented in Fig. 5: panel (e) shows synchrony versus g<sub>AMPA</sub> for the three connectivities, panel (d) shows the matching PGD curve, and panel (f) shows the synchrony-PGD scatter with Pearson r per connectivity.

      The synchrony curve captures the main peak-dip-recovery structure of the PGD curve. For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses to a 75 % suppression in the medium-coupling window g<sub>AMPA</sub> = 0.020-0.035 nS, and recovers near unity at g<sub>AMPA</sub> ≥ 0.040 nS. Uniform connectivity follows the same U-shape with a slightly earlier dip. Per-band versions of the PGD- and synchrony-vs-coupling curves are provided in Figure 5-figure supplement 1, showing that the peak-dip-recovery profile holds across all five canonical bands for Allen and uniform connectivity. Across the 120- point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson 0r ranging from ≈ 0.30 to ≈ 0.92 across band-connectivity combinations; per-band scatters in Figure 5-figure supplement 2). Local connectivity follows a different trajectory: its synchrony is moderate at weak coupling and decays monotonically with gAMPA without recovering at strong coupling. This is consistent with local connectivity not supporting large-scale propagating waves at strong coupling, so the peak-dip-recovery interpretation applies mainly to Allen and uniform connectivity. We have added this synchrony analysis to the revised manuscript:

      “To diagnose the mechanism behind this profile, we measured the Kuramoto order parameter R(t) = |⟨e<sup>iϕ(x,t)</sup>⟩<sub>x</sub>| from the generalized phase field ϕ(x, t) of each frequency band and recorded its maximum over each simulation window (Quantitative measurement of neuronal activity). The resulting synchrony curve (Figure 5e) resembles the trend of the maximum PGD well (Figure 5d). For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses in the medium-coupling window, and recovers at strong coupling. Across the full 120-point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson r = 0.30- 0.92, all p < 10−3; Figure 5f, with per-band scatters in figure Supplement 2). Therefore, the PGD trough at medium coupling may be a synchrony trough: increased synaptic drive pushes the network into an asynchronous irregular state, while strong coupling re-establishes rhythmic synchronization that supports wave propagation. This synchrony-to-wave bottleneck seems more significant to the networks with long-range connectivity (Allen, uniform), partly because long-range connections can augment synchrony in the coupled neuronal network [6]. We also note that although uniform connectivity is able to achieve almost an identical level of synchrony to that of Allen connectivity, the PGD remains much lower due to the loss of spatial organization within.”

      (1.8) Does using a single connectivity parameter (ρ = 300) across all regions miss important regional differences in cortical connectivity density?

      We thank the reviewer for raising this important point. We agree that a single global density parameter misses region-to-region variation in local synaptic density, and we have extended the Discussion (Construction of the Allen connectivity) to state this limitation alongside the cell-type-specific connectivity and synaptic-weight limitations discussed in response to Point 1.3. Specifically, we now write:

      “At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      This paragraph is placed directly after the cell-type and weight-heterogeneity limitations added in response to Point 1.3.

      Reviewer #2 (Public review):

      (2.1) This work presents a spiking network model of traveling waves at the whole-brain scale in the mouse neocortex. The authors use data from the Allen Institute to reconstruct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in the neocortex.

      Overall, the results are interesting and shed new light on the dynamic organization of activity across the neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      We thank Reviewer 2 for the positive assessment and accurate summary of our work.

      (2.2) Description of Algorithm 1: While the Methods section clearly explains the density parameter ρ, the statement on line 358 concerning the “ideal” average number of connections is a little unclear. The authors should explicitly clarify that ρ is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity. The ρ parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in the mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      We thank the reviewer for raising this important point about our connectivity algorithm and simulation. We have clarified in the revised Methods (Use of the voxelized connectivity data, Methods 4.2) that ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity. Specifically, we now write:

      “The density parameter ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity: higher values of ρ yield more connections per neuron and thus higher biological realism, at the cost of greater memory and runtime [7].”

      A careful accounting of the biological scales involved (synapse density per neuron, total cortical population) and incorporating region or cell-type-specific connection density is left as a future direction. We have noted this in the Discussion subsection:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      (2.3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      The reviewer is correct, and we are grateful for the prompt to be more precise. The Results phrasing around Fig. 2 has been softened to avoid any implication of narrowband rhythmicity, and now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Moreover, we have revised the Introduction to describe [11] as demonstrating traveling waves in broadband (5-40Hz) activity, making clear that traveling waves can occur without requiring narrowband oscillations (see also our response to your related point below). In the revised manuscript, we also separate the broadband activity into canonical frequency bands and analyze the wave activity in each band independently.

      (2.4) Line 217: The authors should clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013) or differ from them.

      We thank the reviewer for this suggestion. We agree that [12] is an important experimental benchmark. A direct quantitative comparison is difficult because the original data were not aligned to the Allen Brain Atlas CCF used in our simulations. We therefore revised the Discussion to identify this comparison as a future direction, alongside the visual-cortex bidirectional-wave data of [13]:

      “A more detailed quantitative comparison with experimental cortical-wave studies, such as the cortex-wide voltage-imaging data of [12] or the bidirectional visual-evoked waves reported by [13], is left as a future direction.”

      (2.5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      We thank the reviewer for raising this important point. The reviewer is correct that higher-frequency oscillations tend to be more spatially localized, which can inherently reduce PGD when measured globally. We therefore revised this analysis by separating the broadband signals into canonical frequency bands and comparing PGD within each band.

      In the revised manuscript, we no longer interpret cross-band PGD differences as evidence for a frequency-to-spatial-scale relationship. Instead, we report the level of macroscopic wave activity within each canonical band. This per-band comparison is summarized in Fig. 5b, which reports the maximum PGD in each band (mean ± SEM across the entire (g<sub>AMPA</sub>,V<sub>stim</sub>) grid) for the three connectivities. Allen connectivity shows higher per-band PGD than local or uniform connectivity in the theta, alpha, and beta bands, without requiring an interpretation of PGD differences across frequency bands.

      For completeness, we also computed the per-band mean(Allen) − mean(opponent) PGD gap across the full parameter grid (10 g<sub>AMPA</sub> × 12 V<sub>stim</sub> = 120 points per connectivity). The result is presented in Figure 4-figure supplement 6: panel (a) gives the per-band maxPGD heatmaps over the (gAMPA, Vstim) grid for each connectivity, and panel (b) gives the per-band Allen-minus-opponent mean-PGD gap. The mean gap against Local is −0.004 in delta, +0.027 in theta, +0.036 in alpha, +0.036 in beta and +0.004 in gamma, and against Uniform is +0.015, +0.044, +0.052, +0.040 and +0.011, respectively. The gap is largest in alpha and broadly concentrated in theta-alpha-beta, rather than in delta as we had originally written. We report these per-band gaps descriptively because they summarize one parameter sweep per connectivity rather than independent biological or simulation replications, and we do not interpret the differences across bands as a meaningful frequency dependence. We have added this analysis to the revised manuscript.

      (2.6)Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detection. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      We agree with the reviewer that the relevant anatomical blocks are already present in the data. Our intent was to relate the sampling procedure to stochastic block models for network generation, not to community-detection algorithms. We have revised this passage in the Discussion subsection Construction of the Allen connectivity: ”In essence, our algorithm belongs to the family of stochastic block models [14], where the block structure is given by the voxelization of the Allen Brain Atlas and the inter-block connection probabilities are set by the voxelized projection strengths.”

      (2.7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      Thank you for raising this important point. We agree that conduction delay is directly tied to axonal propagation time and therefore to anatomical path length. Our original wording was intended to note that the relevant path length cannot be approximated reliably by Euclidean distance in the 3-D coordinate space. We have revised this passage in the Discussion sub-section Cortical model and simulation to state that conduction delay is linked to the white-matter path length of the connection, while noting that our model lacks the actual axonal path geometry through the cortical manifold:

      “However, we did not include conduction delay in our study. Conduction delay is thought to have a proportional relationship with the white matter path length of the connection between two neurons [15]. In our model, although we have the 3D positions of neurons, we do not have geometric information about the cortical manifold. Two neurons can be very close in the Cartesian coordinates measured by the Euclidean distance, but very far in terms of the length of the actual connection in the brain. Therefore, incorporating accurate conduction delay in the model is an important future direction.”

      (2.8) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Thank you for this reference. We have added a citation to [16] in the Discussion subsection Quantitative measurements of 3-D traveling waves, acknowledging that methods for 3-D wave analysis do exist while noting that most published algorithms are designed for 2-D data:

      “There are many techniques available for identifying and measuring large-scale neuronal spatiotemporal patterns [17, 18]. While methods for detecting wave dynamics in three-dimensional data do exist [16], most published algorithms are designed to analyze 2-D data, so our simulation data, which is intrinsically 3-D, presents new challenges for measurement.”

      (2.9) Line 28: It is important to note that the Davis et al. (2020) reference is not actually in the beta band, but instead in the broadband (5-40 Hz). This distinction is important because it demonstrates that waves can occur in neural data without requiring narrowband oscillations.

      Thank you for this correction. We have moved the [11] citation out of the beta-band group in the Introduction and reframed it as a broadband (5-40Hz) reference, so that the sentence now reads:

      “These waves are observed at different frequencies during various brain activities, ranging from slow-wave activity [8, 9], sleep spindles [19], to faster oscillations in alpha [20, 21], beta [22, 23], and gamma [21, 13] frequency bands, as well as in broadband (5-40Hz) activity [11].”

      This distinction is important because it makes clear that traveling waves do not require narrowband oscillations. The revised manuscript therefore separates the broadband activity into canonical frequency bands and compares wave activity within each band.

      (2.10) Line 46-49: This sentence could be clearer, for example, by specifying “certain dynamics” in more precise terms.

      We apologize for the imprecision in the original manuscript. We have clarified the sentence in the Introduction to specify that the dynamics of interest are coexistence patterns of local and global activity in the network, as described in the cited reference [24]:

      “It has been shown that in a coupled neuronal network, the coexistence of global wave activity with locally asynchronous states only arises when the number of oscillators is high enough, where local and global activities can coexist [24].”

      (2.11) Figure 1b(i): Small typo in the label for this panel.

      This has been corrected. Thank you for catching it.

      (2.12) Lines 121-122: It may be important to note that spiking neural networks can also generate self-sustained activity (Vogels and Abbott, JNeurosci, 2005; Kumar et al., Neural Computation, 2008). This self-sustained activity is a form of internally generated “frozen” noise that is fundamentally different from externally imposed noise sources (such as Poisson external input) (Destexhe and Contreras, Science, 2006). Waves appear in this self-sustained activity, as well (Davis et al., Nature Communications, 2021), supporting the generality of this phenomenon.

      Thank you for pointing out these important references. We have added citations to [25], [26], [27], and [24] in the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity:

      “Spiking networks of this scale can also generate self-sustained activity in similar regimes [25, 26], which differs fundamentally from externally imposed noise [27], and traveling waves have been reported under such conditions [24].”

      In the revised manuscript, we also changed the stimulation protocol from constant current to Poisson input, which is more similar to the stochastic input that cortical neurons receive in vivo. Traveling waves observed in our model under this Poisson drive therefore complement, rather than depend on, the self-sustained-activity regime emphasized by the cited works.

      (2.13) Lines 139-140: “anterior-posterior macroscopic waves in both directions” and ”in the reverse direction right after each other” could be clearer. In addition, the study from Aggarwal et al. (Nature Communications, 2022) could be relevant to note at this point.

      We thank the reviewer for these wording suggestions and the relevant reference. In the revised manuscript, we reran the simulations and updated Figure 2 accordingly. The new representative simulation shown in Fig. 2 emphasizes a coherent wavefront sweeping along the anterior-to-posterior axis, and the original passages describing consecutive opposite-direction waves have been removed from the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity, which now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Bidirectional propagation can still occur in the model, but we have chosen not to present it as a focal result of this revised manuscript; consequently, the specific phrasings flagged by the reviewer no longer appear in the Results text.

      We agree that [13] is an important reference, and we now discuss it in the Comparing with experimental data subsection as a potential benchmark for future quantitative comparison with our model.

      (2.14) Figure 7: The caption for this figure could be clearer.

      Line 281: typo ”Ermentrou”.

      Line 418: typo ”excitatory”.

      The typos (“Ermentrou” → “Ermentrout”, and the “excitatory” typo at line 418) have been corrected. The original Figure 7 has been removed from the revised manuscript; its content (PGD versus stimulus magnitude and coupling strength, and PGD by dominant frequency bucket) has been reorganized across the new Figs. 4-6.

      References

      (1) Steriade M, Mccormick DA, Sejnowski TJ. Thalamocortical Oscillations in the Sleeping and Aroused Brain. Science. 1993;262(5134):679-85. Available from: <GotoISI>:// WOS:A1993MD95200029.

      (2) Ye Z, Bull MS, Li A, Birman D, Daigle TL, Tasic B, et al. Brain-wide topographic coordination of traveling spiral waves. BioRxiv. 2023:2023-12.

      (3) Rockland KS.What do we know about laminar connectivity? Neuroimage. 2019;197:772-84.

      (4) Harris JA, Mihalas S, Hirokawa KE, Whitesell JD, Choi H, Bernard A, et al. Hierarchical organization of cortical and thalamic connectivity. Nature. 2019;575(7781):195+. Available from: <GotoISI>://WOS:000496159900061https://www.nature.com/articles/ s41586-019-1716-z.pdf.

      (5) Buzs´aki G, Mizuseki K. The log-dynamic brain: how skewed distributions affect network operations. Nature Reviews Neuroscience. 2014;15(4):264-78.

      (6) Bazhenov M, Rulkov NF, Timofeev I. Effect of synaptic connectivity on long-range synchronization of fast cortical oscillations. Journal of neurophysiology. 2008;100(3):156275.

      (7) Morrison A, Aertsen A, Diesmann M. Spike-timing-dependent plasticity in balanced random networks. Neural computation. 2007;19(6):1437-67.

      (8) Massimini M. The Sleep Slow Oscillation as a Traveling Wave. Journal of Neuroscience. 2004;24(31):6862-70. Available from: https://dx.doi.org/10.1523/jneurosci. 1318-04.2004 https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729597/pdf/ 0246862.pdf.

      (9) Liang Y, Song C, Liu M, Gong P, Zhou C, Kno¨pfel T.Cortex-Wide Dynamics ofIntrinsic Electrical Activities: Propagating Waves and Their Interactions. The Journal of Neuroscience. 2021;41(16):3665-78. Available from: https://www.jneurosci.org/ content/jneuro/41/16/3665.full.pdf.

      (10) Aggarwal A, Luo J, Chung H, Contreras D, Kelz MB, Proekt A. Neural assemblies coordinated by cortical waves are associated with waking and hallucinatory brain states. Cell Reports. 2024;43(4):114017. Available from: https://www.sciencedirect.com/ science/article/pii/S2211124724003450.

      (11) Davis ZW, Muller L, Martinez-Trujillo J, Sejnowski T, Reynolds JH. Spontaneous travelling cortical waves gate perception in behaving primates. Nature. 2020;587(7834):4326. Available from: https://doi.org/10.1038/s41586-020-2802-y.

      (12) Mohajerani MH, Chan AW, Mohsenvand M, LeDue J, Liu R, McVea DA, et al. Spontaneous cortical activity alternates between motifs defined by regional axonal projections. Nature Neuroscience. 2013;16(10):1426-35. Available from: https://doi.org/ 10.1038/nn.3499https://www.nature.com/articles/nn.3499.pdf.

      (13) Aggarwal A, Brennan C, Luo J, Chung H, Contreras D, Kelz MB, et al. Visual evoked feedforward–feedback traveling waves organize neural activity across the cortical hierarchy in mice. Nature Communications. 2022;13(1):4754. Available from: https://doi.org/10.1038/s41467-022-32378-xhttps://www.nature. com/articles/s41467-022-32378-x.pdf.

      (14) Holland PW, Laskey KB, Leinhardt S. Stochastic blockmodels: First steps. Social Networks. 1983;5(2):109-37. Available from: https://www.sciencedirect.com/science/ article/pii/0378873383900217.

      (15) Lemar´echal JD, Jedynak M, Trebaul L, Boyer A, Tadel F, Bhattacharjee M, et al. A brain atlas of axonal and synaptic delays based on modelling of cortico-cortical evoked potentials. Brain. 2022;145(5):1653-67.

      (16) Budzinski RC, Nguyen TT, Min´o-Calero J, Davis ZW, Muller LE. Analyzing transientevoked neural activity in three-dimensional cortical recordings. Physical Review Research. 2023;5(1):013012.

      (17) Townsend RG, Gong P. Detection and analysis of spatiotemporal patterns in brain activity. PLOS Computational Biology. 2018;14(12):e1006643. Available from: https: //doi.org/10.1371/journal.pcbi.1006643.

      (18) Gutzen R, De Bonis G, De Luca C, Pastorelli E, Capone C, Allegra Mascaro AL, et al. A modular and adaptable analysis pipeline to compare slow cerebral rhythms across heterogeneous datasets. Cell Reports Methods. 2024;4(1). Available from: https: //doi.org/10.1016/j.crmeth.2023.100681.

      (19) Muller L, Piantoni G, Koller D, Cash SS, Halgren E, Sejnowski TJ. Rotating waves during human sleep spindles organize global patterns of activity that repeat precisely through the night. Elife. 2016;5. Available from: https://www.ncbi.nlm.nih.gov/pubmed/27855061https://www.ncbi.nlm. nih.gov/pmc/articles/PMC5114016/pdf/elife-17267.pdf.

      (20) Zhang H, Watrous AJ, Patel A, Jacobs J. Theta and alpha oscillations are traveling waves in the human neocortex. Neuron. 2018;98(6):1269-81. e4.

      (21) van Kerkoerle T, Self MW, Dagnino B, Gariel-Mathis MA, Poort J, van der Togt C, et al. Alpha and gamma oscillations characterize feedback and feedforward processing in monkey visual cortex. Proc Natl Acad Sci U S A. 2014;111(40):14332-41. Available from: https://www.ncbi.nlm.nih.gov/pubmed/25205811.

      (22) Bhattacharya S, Brincat SL, Lundqvist M, Miller EK. Traveling waves in the prefrontal cortex during working memory. PLoS Comput Biol. 2022;18(1):e1009827.

      (23) Rubino D, Robbins KA, Hatsopoulos NG. Propagating waves mediate information transfer in the motor cortex. Nature Neuroscience. 2006;9(12):1549-57. Available from: https://doi.org/10.1038/nn1802.

      (24) Davis ZW, Benigno GB, Fletterman C, Desbordes T, Steward C, Sejnowski TJ, et al. Spontaneous travelling waves naturally emerge from horizontal fiber time delays and travel through locally asynchronous-irregular states. Nature Communications. 2021;12(1):6057.

      (25) Vogels TP, Abbott LF. Signal propagation and logic gating in networks of integrate-and-fire neurons. Journal of Neuroscience. 2005;25(46):10786-95.

      (26) Kumar A, Schrader S, Aertsen A, Rotter S. The high-conductance state of cortical networks. Neural Computation. 2008;20(1):1-43.

      (27) Destexhe A, Contreras D. Neuronal computations with stochastic network states. Science. 2006;314(5796):85-90.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines Müller glia (MG) reprogramming in the uninjured mouse retina through a combination of Notch signaling inhibition and AAV-induced proliferation. Building on their prior work showing that Cyclin D1 overexpression and p27^Kip1^ knockdown (CCA) promotes MG proliferation with very limited neurogenesis, the authors now demonstrate that Rbpj deletion alone induces a modest degree of MG-to-neuron conversion without proliferation, in agreement with recent work in the field. However, combining Rbpj deletion with CCA-mediated proliferation substantially enhances MG dedifferentiation and the generation of retinal neuron-like cells. Through genetic lineage tracing, histological analyses, and single-cell transcriptomics, the authors provide evidence that MG-derived cells acquire molecular features of bipolar (ON, OFF, and rod bipolar) and amacrine neurons. Most MG-derived cells appear to survive long-term (up to 9 months).

      Strengths:

      Overall, the study is carefully designed and executed, and the manuscript is clearly written with well-presented figures. While the work does not significantly expand the repertoire of neuronal types generated from mammalian MG beyond what has been previously reported in the field, it provides a valuable and improved strategy for inducing robust MG proliferation and neurogenesis in the mammalian retina.

      Weaknesses:

      (1) It would be better to include a negative control AAV when evaluating the effect of CCA AAV in the Rbpj KO background. This could help distinguish the specific contribution of the CCA construct from potential effects of intravitreal AAV injection itself, which can induce mild inflammation, known to influence MG reprogramming.

      To address this concern, in the revised manuscript we included the result from Rbpj KO eyes injected with a negative control AAV (AAV<sub>7m8</sub>-GFAP-GFP) (Fig. S14a). MG reprogramming efficiency, quantified as the proportion of tdT<sup>+</sup>Otx2<sup>+</sup> cells among total tdT<sup>+</sup> cells, was then compared between the AAV-GFP–treated and Rbpj KO–only eyes. At 4 months post-injection, the percentage of tdT<sup>+</sup>Otx2<sup>+</sup> cells in the AAV-GFP–treated eyes was comparable to that of Rbpj KO alone (Fig. S14b–c), and substantially lower than in the CCA-treated eyes. Together, these results indicate that the enhanced MG reprogramming observed in the Rbpj KO+CCA group is driven by transgenes expressed rather than by nonspecific effects of AAV or injection.

      (2) The extent of MG transduction by the CCA AAV is not clear. As quantifications are normalized to total MG (GFP^+^ or TdTomato^+^) or retinal length, it would be useful to clarify whether near-complete transduction is assumed, or if additional information on transduction efficiency can be provided.

      In our previous study (Wu, Liao, et al., 2025, eLife), we have demonstrated that high-dose (4E10vg/injection) AAV7m8 effectively transduced the whole retina, with near-complete MG transduction observed in the vicinity of the injection site, as evidenced by virtually all MG expressing GFP in these regions. In the revised manuscript, we clarified the transduction efficiency in Line 108-110 on Page 5 and Line 625-626 on Page 27.

      (3) In Figure S10, the reduced MG proliferation observed in the CCA + Rbpj deletion group could also potentially reflect decreased GFAP promoter activity in dedifferentiated MG following Rbpj deletion. Alternatively, MG-derived cells may be more fragile under these conditions.

      We thank the reviewer for these excellent insights. We agree that a down-regulation of GFAP promoter activity following Rbpj-mediated dedifferentiation is a highly plausible explanation for the moderate reduction in proliferation, as lower promoter activity would diminish AAV transgene expression. We have included this possibility in the data interpretation (Line 176-178, page 8). Regarding the alternative possibility of increased cell fragility, we agree that cell death cannot be ruled out, but occasional apoptotic cells over a long period of time are difficult to capture experimentally.

      (4) In the CCA + Rbpj deletion condition, do MG undergo single or multiple rounds of cell division?

      We have previously demonstrated that MG typically undergo a single round of cell division in wild type mouse retina following CCA treatment (Wu, Liao, et al., 2025, eLife). Given our observation that Rbpj deletion suppresses CCA-induced MG proliferation (Fig. S11), it is unlikely that the addition of Rbpj deletion would trigger multiple or continuous rounds of cell division beyond the single-round baseline established by CCA alone. While we did not re-evaluate cell division kinetics in the current study, we reason that CCA similarly drives MG to undergo a single round of division in the Rbpj KO context.

      (5) What fraction of neuron-like cells (bipolar- and amacrine-like) arises from proliferation versus direct transdifferentiation? Quantification of MG-derived cells expressing neuronal markers (e.g., Otx2, HuC/D), with and without EdU labeling, would help distinguish these mechanisms.

      The percentages of MG-derived cells expressing neuronal markers with and without EdU labeling, were shown in Fig 3d-e and Fig S19d-e. In the Rbpj KO-only group, neuron-like cells arise exclusively through direct transdifferentiation without cell division, as no EdU incorporation was detected in Rbpj-deficient MG. In this group, a small fraction of MG-derived cells expressed the neuronal marker Otx2 or HuC/D (Fig 3e, Fig S19e). In contrast, the Rbpj KO+CCA group achieved a substantially higher neurogenesis rate, with a significant proportion of Otx2+ or HuC/D+ MG-derived cells also being EdU+ (Fig. 3d, Fig. S19d), indicating that they arose through de novo neurogenesis. By subtracting the contribution of direct transdifferentiation observed in the Rbpj KO-only group, we estimate that majority of MG-derived neuron-like cells in the Rbpj KO+CCA group were generated through proliferation-mediated de novo neurogenesis.

      (6) In Figure S18a, the authors state that "while the neuron-like clusters were best classified as BC-like and AC-like based on their distinct marker gene expression, they also exhibited mixed expression of genes associated with other retinal neuronal types, including RGC markers (e.g., Tubb3, Myt1l, Grin1) and photoreceptor markers (e.g., Crx, Prom1, Epha10, Gucy2e, Scg3) (Fig. S18a), suggesting that the regenerated cells exist in a hybrid state" and "MG derived neuron like cells also expressed genes characteristic of RGCs and photoreceptors, indicating enhanced lineage". However, many of these genes are not specific to RGCs or photoreceptors and are instead broadly expressed in retinal neurons or enriched in bipolar/amacrine populations. Therefore, it is unclear whether these cells exhibit hybrid RGC or photoreceptor identity.

      We thank the reviewer for this insightful comment and for pointing out the need for greater precision in our terminology regarding these markers. While individual markers may lack absolute, 100% cell-type exclusivity, genes such as Tubb3 and Gucy2e serve as widely accepted lineage-associated genes that characterize RGC and photoreceptor programs, respectively (Soto et al., 2008; Sato et al., 2018; Sotani et al., 2024). We have revised the manuscript to replace terms "RGC-specific genes" and "photoreceptor-specific genes" with "RGC signature genes" and "photoreceptor signature genes", respectively. Furthermore, these RGC- and photoreceptor-signature genes are co-expressed across the entire Otx2+ MG population rather than being segregated into distinct, specialized subpopulations (Fig. 4d, Fig. S20). This uniform distribution indicates that these cells possess a hybrid transcriptional program that concurrently incorporates elements of both RGC and photoreceptor identities.

      (7) The authors provide a thorough molecular characterization of MG-derived cells through immunostaining and single-cell sequencing. However, their morphological features, synaptic connectivity (e.g., synaptic marker expression), and electrophysiological properties remain largely uncharacterized. While these experiments may be technically challenging, this limitation should be discussed.

      We agree with the reviewer that characterizing the precise morphological features, synaptic connectivity, and electrophysiological properties of MG-derived cells is a crucial step for any neuronal regeneration study, and we acknowledge that this represents an important limitation of our current study.

      As demonstrated by snRNA-seq data, the MG-derived neuron-like cells exhibit an incompletely mature state, characterized by hybrid transcriptomic signatures. By immunostaining, we did not observe any MG-derived cells with photoreceptor outer segment or typical RGC morphology. Therefore, it is highly likely that these cells have not established functional synaptic connectivity or acquired mature electrophysiological properties. Performing functional or circuitry assessments at this stage would be premature.

      We have added a comprehensive discussion regarding this limitation, along with future directions for long-term functional validation, in the revised manuscript (Line 502-515 on Page 22).

      (8) The conclusion that CCA + Rbpj deletion induces neurogenesis without compromising MG supportive functions or retinal homeostasis appears somewhat oversold. This claim is primarily based on gross retinal morphology and ZO-1 staining. Given the extent of MG dedifferentiation and ectopic cell generation in the ONL and INL, it is likely that retinal function is affected. Functional assessments (e.g., ERG) would be required to support this conclusion. The authors should consider tempering this statement.

      To address the concern raised by the reviewer, we performed electroretinography (ERG) to evaluate both scotopic and photopic retinal function in the Rbpj KO+CCA-treated eyes compared to contralateral untreated controls (Supplementary figure S23e-h). In addition, we conducted optomotor response testing to assess whether visual behavior is affected following treatment (Supplementary Figure S23d). The results demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal function.

      (9) Regarding the mechanism by which CCA-induced proliferation enhances MG reprogramming in the Rbpj knockout background, one plausible explanation is that chromatin states (e.g., histone modifications and DNA methylation) are transiently reset during DNA replication and cell division. While this alone may be insufficient to activate neurogenic programs, it could synergize with Rbpj deletion to allow neurogenic transcription factors (such as Ascl1, Otx2, NeuroD1, and NeuroD2) to access previously inaccessible chromatin regions, thereby promoting MG reprogramming.

      We thank the reviewer for the insightful suggestion on the model, which aligns well with our experimental findings. Our snATAC-seq data demonstrate that CCA-induced proliferation broadly increases chromatin accessibility at key neurogenic loci, including Neurod2, Dll1, and Otx2, in active MG compared to resting MG (Figure 6f–h). This chromatin remodeling alone is insufficient to drive neurogenesis, as CCA-only treated MG largely revert to a quiescent glial state. However, when combined with Rbpj deletion, which derepresses downstream neurogenic transcription factors such as Ascl1 and Neurog2 by relieving Notch-mediated transcriptional repression, these newly accessible chromatin regions can be effectively occupied and activated by the available neurogenic factors. The concept that cell division facilitates epigenetic resetting to enhance reprogramming efficiency is well established in the somatic cell reprogramming field, where proliferation rate is directly proportional to reprogramming success by promoting the erasure of lineage-restrictive epigenetic marks and the re-establishment of new transcriptional circuits. In the revised manuscript, we incorporated this mechanistic discussion to provide a more comprehensive interpretation of how proliferation and Notch inhibition converge to promote MG neurogenesis in Line 462-477 on Page 20-21.

      Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able to proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here, they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly - and through multiple methods including immunostaining, single-nucleus RNA sequencing, and in situ hybridization - that coupling notch inhibition with cell cycle reactivation induces the expression of neuronal markers in mammalian MG. The snRNA-seq data are particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Weaknesses:

      Whether the newly generated neurons are functionally integrated remains unclear, and the effect of the manipulation on the function of the retina was not tested. Imaging data suggests that many of the newly generated neurons persist for months, but often appear mislocalized. It is also not clear if the manipulation of MG affects long-term MG function. Cell death was not evaluated, and although the authors evaluated the long-term effect on tight junctions, this data was not quantified, and further analysis on morphology or function was not done. Control eyes were untreated, not vehicle-injected.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The transgenic line may be Glast-CreERT, not Glast-CreERT2.

      We appreciate the reviewer for bringing this to our attention. The formal allele symbol for this transgenic line is Tg(Slc1a3-cre/ERT)1Nat, while this strain is generically classified as "Cre/ERT2" by the Jackson Laboratory. Some published studies referred to this line as Glast-CreERT and others as Glast-CreERT2. To maintain consistency with the formal allele symbol, we have adopted "Glast-CreERT" throughout the revised manuscript.

      (2) For snATAC data in Figure 6 e,f, and Figure 19b. It is most likely gene activity, not gene expression, since these are snATAC, not snRNA data.

      For this inaccurate terminology, we have corrected all relevant figure labels and associated text in the revised manuscript to clearly state "gene activity" instead of "gene expression."

      (3) Some text in Figure 6 is a bit too small to read.

      We have increased the font size of the text elements in Figure 6 to ensure readability and have also reviewed all other figures for consistency. Revised figures with improved legibility have been included in the updated manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) There are multiple instances where further elaboration of methods or tools in the test would improve readability and comprehension by a broader audience. It would be helpful to (early, often, and clearly) explain precisely which cell types are labeled in your mouse line and how. Someone unfamiliar with the mouse line may struggle to understand what is labeled by the tdT or GFP. Likewise, it would help to consistently define what cell types are labeled by tdT+ vs. Sox9+, tdT+, etc.

      We have added a clear and detailed description of the mouse lines and labeling strategy early in the Results section, specifying which cell types are labeled by tdT and GFP and how the labeling is achieved. We have also ensured that the definitions of cell type identifiers (e.g., tdT<sup>+</sup> for MG-derived cells, Sox9<sup>+</sup>/tdT<sup>+</sup> for MG remaining in a glial state) are consistently stated upon first use and maintained throughout the manuscript to improve readability for a broader audience. In addition, we added headings for the quantification graphs to improve readability in all quantification figures.

      (2) It is unclear what the difference is between Figure 1c and S1c, and these should be quantified as the % of positive cells, as described in the text.

      We have removed Figure S1c and moved Figure 1c to supplementary figure 1. The MG labeled by EdU and Sox9 or Otx2 were quantified as % of the EdU+ MG.

      (3) S2e: Clarify what pixel level means, is this pixel intensity?

      Yes, "pixel level" in Figure S2e refers to pixel intensity. We apologize for the ambiguous wording and replaced "pixel level" with "pixel intensity" in the revised figure legend to ensure clarity.

      (4) Figure 2: In the magnified image of the GFP+, Sox9- cell, the GFP is also very faint. Could these cells be dying? Analysis of the expression profile of these cells (or ruling out apoptosis) would better support a dedifferentiation argument.

      The faint GFP signal observed in GFP<sup>+</sup> Sox9<sup>-</sup> cells is a sign of ongoing dedifferentiation rather than cell death. This is likely due to chromatin remodeling during reprogramming. A similar decrease in reporter signal intensity during MG dedifferentiation has been previously reported by Le et al. 2024, 2025, supporting the interpretation that reduced fluorescence is a characteristic feature of this process. It is possible that a small fraction of GFP<sup>+</sup> Sox9<sup>-</sup> cells may undergo cell death over an extended period, which would be difficult to detect using apoptosis assays. Our long-term survival experiments demonstrate that more than 80% of MG-derived neuron-like cells survive for at least 9 months following treatment (Figure 7), indicating that majority of these cells are viable. The discussion is included in line 112-114 on page 5.

      (5) Figure 3: The Crx labeling appears everywhere except the identified cell. This seems the opposite of the point you are making.

      Crx signal of the MG-derived cell (tdT<sup>+</sup> Crx<sup>+</sup>), which is pointed out by arrowhead, is in a ring-like pattern. This pattern is consistent with the euchromatin region in inverted nucleus of rod. Crx labeling appears in other cells in the image as Crx is highly expressed in native photoreceptors.

      (6) I think it would be nice to address, in the discussion, the apparent disorganization and mislocalization of cells in the long-term images.

      We thank the reviewer for highlighting this critical observation. During retinal development, precise laminar positioning of neurons is guided by a coordinated interplay of cell-intrinsic transcriptional programs and extrinsic cues including cell adhesion molecules, guidance factors, and interactions with neighboring cells. In the adult retina, many of these developmental cues are no longer present or active, which likely contributes to the failure of MG-derived neurons to migrate to their appropriate laminar positions. Interestingly, the vast majority of our divided MG cells remained localized within the outer nuclear layer (ONL). Because the ONL is the physiological location of photoreceptors, this preferential position could serve as an advantageous baseline layout for driving targeted photoreceptor differentiation in future work. To address reviewer’s feedback, we have expanded our discussion section (Line 543-559, page 23-34) to cover the mechanisms underlying this structural disorganization and its downstream implications for functional circuit integration.

      (7) I'm not convinced that ZO1 alone is sufficient to suggest MG function normally or that retinal homeostasis is maintained. I suggest tempering that conclusion in the text.

      For the revision, we have performed additional experiments to address this concern. The optical coherence tomography (OCT) images revealed that retinal layer organization and ONL thickness were comparable among the uninjected eyes, GFP AAV-injected control eyes, and CCA-treated eyes, demonstrating that overall retinal architecture was well-preserved (Fig. S23a–c). Optomotor response testing revealed no significant differences in visual acuity across groups, suggesting that visual function remained intact (Fig. S23d). Furthermore, electroretinography (ERG) demonstrated that scotopic and photopic a- and b-wave amplitudes were unaffected by the treatment, confirming that light responses from photoreceptor and inner retinal neuron were preserved (Fig. S23e–h). Taken together, these findings demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal structure and functional visual circuitry.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Thank you for the helpful comments and criticisms. We provide exciting additional data, in particular a CRISPR actin-binding motif mutant and FRAP analysis of an exon15e-GFP transgene, both further supporting the importance of the IDR in thin filament stability. We believe that these additional experiments provide compelling evidence supporting our conclusion and substantially advance the current limited body of knowledge surrounding the role of IDRs in structural proteins.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      However, some of the phenotypic analysis, in particular the bending of the sarcomere, likely upon mechanical challenge by muscle contractions, needs more detailed investigations to be fully convincing.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

      Weaknesses:

      (1) The western blot quantifications of Zasp isoform expression are weak. No error bars are indicated in the quantifications; the quantifications appear to be more qualitative than quantitative. According to band intensities, the long Zasp isoforms seem to be less present compared to the shorter ones, even in the flight muscles.

      We have now included quantifications with error bars for the Western blots in our resubmission. It is important to keep in mind that the main point in figure 1B is that there are plenty of exon15e-containing isoforms in IFM, in contrast to other tissues with very limited exon15e-containing isoforms. This is confirmed by the analysis of RNA-seq data in figure 1C, and of course, by the flightless phenotype of the exon15e mutant.

      (2) The phenotypic analysis of the sarcomere appears somewhat superficial throughout the paper. Only Zasp52 and phalloidin are shown; no other Z-disc or thick filament proteins. At least myosin stainings and overview images are important to better judge the phenotypic variations. Are the variants between individuals or regional in the same muscle?

      Our images are representative of the observed phenotypes. Phenotypes are consistently present across all individuals, as reflected in our replicates. Interestingly, they appear not to be randomly interspersed among the sarcomeres but concentrated in certain regions of muscle more than others. Full images are available in the online repository FigShare.

      (3) EM images would benefit from better quantification.

      We do not believe that EM images can be meaningfully quantified, because of the many selection steps preceding image acquisition.

      (4) Other proteins were not analysed with the FRAP-based turnover assay for comparison in wild type and mutant. All Z-proteins might turn over faster in the mutant with the defective Z-disc.

      This is the point we are trying to make. The Zasp52 IDR appears to stabilize the Z-disc and is likely involved in fastening a variety of proteins to it.

      Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

      Weaknesses:

      (1) It is not clear in Figure S1A where exon 15e fits within the Zasp52 locus schematic. This is important as a premise of this paper describes this region to be key, and proof from multiple prediction programs would lend more weight to the prediction of the exon being largely disordered. Inclusion of the discussed short linear motifs, comparison with Canoe or LBD3 for similarities and/or an Alphafold structure would help make the authors' point (colorized with known domains).

      We added a bar below figure S2A to show the region corresponding to exon 15e. We used three disorder prediction programs and one structure (order) prediction program. The majority of exon15e is completely disordered and of very low confidence score, and thus uninformative to display as an AlphaFold structure. Likewise, IDR’s are very difficult to classify, therefore we cannot say much more than that LDB3, Zasp52, and Canoe contain IDRs, with Zasp52 and Canoe both having a putative actin-binding domain within the IDR. We now provide data on the function of the ABD in this resubmission.

      (2) Interesting that immobilization rescues the deterioration phenotypes. The authors should explain in more detail how this was done to avoid dehydration/starvation of the flies.

      We provided more details in materials and methods.

      (3) There is a lot of discussion about the potential function of the IDR region, specifically a putative actin binding motif or other 'ordered' regions that may contain short linear motifs. It would strengthen the findings to show which of these may be essential for Zasp52 function in the IFM. The ability to bind actin could be tested biochemically, and/or smaller deletions could be made to unequivocally test the role of the ABD vs other predicted motifs using genetics. If some of these regions are more ordered, where do they lie within, and do they form a predicted fold or structure that gives insight into function?

      We now provide data on the function of the ABD showing that deleting it has almost no phenotypic defects. That means the IDR is largely/entirely responsible for the observed phenotypes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Western blot in Figure 1B needs proper quantification. A ratio between long and short isoforms in the same muscle type might be informative. Is it known which epitope the antibody recognises? Can a GFP insertion that also labels all isoforms be used as verification? Quantifications are also needed in Figure 2A.

      We have added quantifications of all Western blots (Fig. 1B, 2A, and 2A’). The ratio between exon 15e-containing and total Zasp52 in the same muscle type is included in the lowest bar graph in Fig. 1B. The full-length antibody is polyclonal and was raised against Zasp52-PR which contains all ordered domains; the anti-LIM antibody was raised against the last three LIM domains (both are described or referenced in the materials and methods section). Such a GFP insertion cannot exist due to the complex splicing patterns of Zasp52.

      (2) The name of the deletion allele could be specifically indicated in Figure 1A below the red bar.

      Done.

      (3) It would be useful to indicate the order group names in Figure S2B since species names are hard to read.

      For the version of record we provided high-resolution images, where species names can be read. Drosophilids, Ephemeroptera and Odonata are indicated.

      (4) The inverted spelling of the numbers for the control in Figures 2C and 5H is strange.

      Changed to normal spelling.

      (5) The bending of the myofibril at the Z-disc is a really interesting phenotype. However, it seems it is not always visible; at least it is visible in many myofibrils shown in Figure 3B, but in none in Figure 3E, same genotype, just different staining. Hence, I wonder if this bending could be force-induced by the cutting of the thorax during tissue preparation. It would be useful to display some overview images to allow the reader to judge the quality of the tissue preparation, indicating from where the high magnification view shown was taken. The same is true for Figure 5.

      Overview images are available on FigShare. Note that you can see some “H-zone actin” sarcomeres in Fig. 3B, as well as some mildly bent ones in Fig. 3E. We generally selected images that best demonstrated the phenotype described. Furthermore, neither phenotype is fully penetrant so we cannot expect to see it everywhere. Lastly, it is always possible that phenotypes are affected by preparation, since it is impossible to know what the myofibrils look like in situ. However, all samples were prepared using the same protocol with replicates, and since we see a phenotype in our mutants and not in the control, this indicates that something is different between the two.

      (6) The same applies to the visualisation of the "hyper-contracted" phenotype; again, it seems to be an all-or-nothing phenotype in the zoom shown. An overview image should be shown. The zoom in Figure 4E would benefit from displaying phalloidin in a separate channel. Are actin filaments pulled out of the Z-disc? The latter is often seen in non-perfect cuts in wild-type, but the accumulation at the M is curious. It would be informative to locate the ends of the thick filaments in these cases or quantify thick filament lengths; do these invade the Z-discs? This can easily be done by a myosin staining.

      Is this a regional effect or does it depend on the individual or on the preparation? I am surprised to also see the "hyper-contraction" in 10% of wild-type 5-day adults.

      See previous response where we include overview images. Single-channel images are available; it is visible that actin filaments are not pulled out of the Z-disc. Phenotypes are consistent across individuals as evidenced in our replicates but do tend to be concentrated in certain regions of muscle.

      (7) The EM images would benefit from more overview images. At the moment, we only see a single sarcomere from wild type and mutant, with no quantification of the phenotype. Can the authors see the invading thin filaments into the M-band? The disrupted Z-disc phenotypes are impressive. What is the age of the animal shown in Figure 4?

      We have a panel displaying several mutant sarcomeres. Due to the selectivity and challenges of the EM preparation process, we do not believe we can perform meaningful statistics on them. It sometimes looks like myosin heads are visible in the H zone which may support the presence of thin filaments in the H zone (Fig. 4B and C). However, the quality of these particular EM images is not high enough to identify thin filaments. All phenotypes shown are from 3-week-old animals.

      (8) Is UH-3 GAL4 expressed at the adult stage?

      Yes, from 36 h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (9) Figure 6 would strongly benefit from a myosin staining. Do thin and thick filament lengths scale? It seems that overlap is reduced in the double hets. How can this be envisioned with Z-disc stability? Is myofibril diameter reduced?

      We searched for non-additive differences in myofibril diameter but were unable to detect any.

      (10) What is the FRAP turnover rate of a long Zasp-GFP compared to a short one in wild type? A difference would indicate that it is really the IDR domain that keeps Zasp52 longer at the Z-disc, instead of an indirect effect caused by Z-disc morphology

      We have newly added FRAP data of a GFP-tagged exon 15e construct which displays much lower turnover. This indicates that the IDR does indeed retain Zasp52 at the Z-disc.

      Reviewer #2 (Recommendations for the authors):

      (1) The total protein stain should also be included if it is used for quantitation in Figures 1B and 2A-A'.

      These are available on FigShare.

      (2) It is a bit confusing that the Alphafold plot is inversely correlated with the other 3 prediction programs, although this is explained in the legend. Maybe an Alphafold structure would help make the authors' point (colorized with known domains).

      The AlphaFold structure is almost entirely low-confidence disordered region except for the structured domains so we do not believe it would be helpful to include.

      (3) The title of Figure 8 says 'Certain ex15e defects are rescued by immobilization.' What other defects are not rescued? If true, these should be shown.

      There was a full rescue. We deleted the word “certain”

      (4) Please include a brief explanation of the spatiotemporal expression of UH3-Gal4.

      From 36h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (5) Statistics should be added to Figure 8E.

      Figure 8E (now 9E) has statistics.

      (6) The dark blue color used for integrin staining in Figure S3 is difficult to see. Changing this color may help visualize differences. Also, pointing them out with arrows, etc., will help clarify abnormalities.

      We have described these differences in the figure caption. Single-channel images are available for viewing in any color in FigShare.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive comments on our manuscript. We are pleased that they considered the core behavioral findings important and robust, especially the results showing that the magnitude and variability of others’ donations affected the magnitude and variability of participants' donations, respectively. We also appreciate their acknowledgement of the strengths of the experimental design, large sample sizes, the incentive-compatible and across-domain measures included in Experiment 4, and combined behavioral and computational approaches.

      We agree that the manuscript would benefit from greater clarification in several areas, further analyses, and more cautious interpretations. In the revised manuscript, we plan to clarify the rationale for sequentially presenting social information, the role of prediction responses, the theoretical motivation of examining the variability of others’ donation, the use of the between-subjects design, and the motivation for the RL framework. We also agree that the lack of a non-social repeated-donation control condition limits the interpretation of the phase effects. Our design permits strong inferences about differences in donation changes across different conditions, but it cannot establish that the phase-related changes are exclusively attributable to social information exposure. We will revise the wording accordingly, moderate the causal language, and explicitly discuss the limitations of our design.

      To strengthen the behavioral analyses, we plan to supplement the current mixed-effects linear models with models that reflect the repeated-measure structure of the task, including random slopes for the phase. We will also add statistics in the generalization results section and test the asymmetry between generous vs. stingy social influence. Moreover, the reviewers raised an important concern regarding the associations between psychopathy and the donation change. Because psychopathy is negatively correlated with initial donations in several experiments, absolute donation changes may partly reflect the distance between initial donation and the observed donation mean. We therefore plan to reanalyze the psychopathy effects by using signed donation changes and trial-level discrepancies between participants’ initial donations and the observed social information. These additional analyses will enable a more direct and precise assessment of whether psychopathy is associated with greater susceptibility to social influence.

      We further agree that the comparison and validation of the computational models should be strengthened. In the revised manuscript, we plan to clarify that the learning models are intended to describe the updating beliefs about a group-level donation norm from sequential social information, rather than learning about a single donor. We will expand the candidate model set to include non-learning models, such as models based on the actual social mean, a running average. We will also model the prediction phase and the second donation phase separately. This will help identify the models that provide explanatory values for both predictions of others’ donations and individual donation behaviors. In addition, because the models were not estimated via a Bayesian framework, we agree that the term “posterior predictive checks” is inappropriate. We will rename these analyses. We will also rerun the parameter and model recovery analyses using empirically informed noise levels separately for the prediction and donation phases. In addition, we will implement model-evaluation processes, such as cross-validation, that better reflect prediction for new participants.

      In Experiment 4, we plan to directly report the association between social-information-use measures in the perceptual and the donation task to strengthen the domain-generality effect. We will additionally examine whether social susceptibility in the perceptual task is associated with psychopathy by including all trials, including those in which participants moved away from or beyond the social value.

      Finally, we will correct the reporting and presentation issues identified by the reviewers, including the social information use equation in the perceptual task, the pseudo-SD of individual donations formula, supplementary figure captions, and task duration. We will also provide fuller experimental materials and make the preregistration links more prominent. In addition, we intend to make the analysis code, model-fitting scripts, and data available during the revision process.

      We greatly appreciate the editors’ and reviewers’ thoughtful suggestions, which will help us substantially strengthen the manuscript. We are grateful for the opportunity to address these important points and believe that the planned revisions will enhance the manuscript’s clarity, robustness, and its contribution to the understanding of social influence in donation behaviors.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablast/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      We agree that CRISPR/Cas9 editing of bulk D7 cultures complicates interpretation because this population contains both activated B cells and PB/prePCs. We will therefore revise the text to distinguish phenotypic effects measured across the bulk D7 culture from the downstream single-cell analysis focused on cells along the prePC-to-PC trajectory. In particular, our interpretation of IRF4 and BLIMP1 function in prePCs is based primarily on the D9 scRNA-seq analysis, in which cells arrested in the activated B cell compartment are not used to define the perturbed PC-trajectory states. We will clarify this analytic design in a future revision and temper language implying that all effects arise exclusively within prePCs.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      We agree that the divergence between IRF4 and PRDM1 perturbations could reflect differences in protein turnover, or hierarchical dominance, in addition to developmental timing. We will revise the Discussion to state that our data support a temporally ordered model in which IRF4 acts early to license the prePC-to-PC transition and BLIMP1 consolidates the terminal state, but that the current experiments do not exclude alternative explanations related to hierarchical dominance or degradation kinetics. We will also note in the revised Discussion that degron-based perturbations, rescue experiments, and gain-of-function analyses would be needed to resolve the functional ordering of IRF4 and BLIMP1 with higher temporal precision.

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. (1) Direct testing of selected regulatory elements would make the causal claim stronger. (2) Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fudlly demonstrated mechanism of gene regulation.

      We agree that the current data support the motif lexicon primarily as a predictive model for differential TF occupancy rather than as a fully causal mechanism of gene regulation. We will therefore revise the relevant text in the Results and Discussion. The EMSA data directly test nucleotide-dependent binding preferences, and the CUT&RUN/multiome analyses show that these motif variants are differentially associated with IRF4- or BLIMP1-bound DEG-linked OCRs. However, direct causal testing of endogenous regulatory elements, for example by base editing of selected ISRE/EICE variants, will be required to determine whether these variants are sufficient to predictably alter gene activity in differentiating plasma cells.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations.

      Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study.

      Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20<sup>+</sup> cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation

      We agree that CD138 alone does not fully define terminal PC maturation. In the revised manuscript, we will clarify that CD138 was interpreted in the context of a broader maturation profile, including CD20 downregulation, ICAM2 upregulation, IRF8 loss, IRF4/BLIMP1 expression, Ki-67 loss, and antibody secretion. The D7 PB population was proliferative and CD138<sup>-</sup>, whereas D21 CD20<sup>-</sup> cells were largely Ki-67<sup>-</sup> and included CD138<sup>+</sup> cells, supporting their progressive maturation. We will include Ki67 analysis in CD138<sup>-</sup> and CD138<sup>+</sup> cells at D21 in the revision.

      We agree that the motivation and culture conditions for this experiment required a clearer explanation. The purpose of the D7 sort-and-reculture experiments was to identify the developmental window enriched for cells competent to generate PCs, thereby defining the stage at which IRF4 and PRDM1 should be perturbed. Sorted D7 CD20<sup>+</sup> actB cells and CD20<sup>-</sup>CD38<sup>+</sup>CD27<sup>+</sup> PBs were both placed into the same D7-D14 differentiation conditions, allowing a direct comparison of their PC-generating competence under identical culture conditions. We will clarify this design in the Results and Methods. We have not tested whether returning D7 CD20<sup>+</sup> cells to D0 conditions involving CD40L stimulation restores PC differentiation, and we will now acknowledge in the revision that the CD20<sup>+</sup> fraction may contain cells with distinct intrinsic differentiation potential, including cells differentiating into non-PC states.

      What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20<sup>+</sup> fraction would be enriched in these cells.

      We agree that the CD20<sup>+</sup> D7 cells may contain anergic or memory B cell precursors. However, this does not alter the interpretation that the CD20<sup>-</sup> (CD38<sup>+</sup>/CD27<sup>+</sup>) PBs contain a PC precursor population. Even if memory B cells are generated in the CD20<sup>+</sup> fraction by D7, based on the sorting experiments, their presence would have little-to-no effect on developing PC precursor populations.

      Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138<sup>+</sup>CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      We agree that the heterogeneous nature of CD20<sup>-</sup> cells complicates the interpretation. However, we would like to emphasize that all D7 CD20<sup>-</sup> cells are Ki67<sup>+</sup> whereas all D21 CD20<sup>-</sup> cells are nearly all Ki67<sup>-</sup>. Though we concede these could include recently proliferated PCs at D21, it is consistent with this population becoming quiescent. To better address this question in a future revised version, we will include direct measurements of Ki67 levels in CD138<sup>+</sup> and CD138<sup>-</sup> cells at D21.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      We agree that the current data do not by themselves prove that IRF4 acts earlier than BLIMP1 in developmental time. We will revise the text to state that the data are consistent with a temporally ordered model, rather than demonstrating strict sequential action. The latter interpretation is based on the distinct IRF4 KO stunted PC state observed by D9 scRNA-seq (Fig. 3C), together with the stronger early phenotypic effect of IRF4 loss (Fig. 3A). However, because both factors are mutually reinforcing and because perturbations were not performed across multiple time points, alternative explanations remain possible, including differences in editing efficiency, protein stability, and feedback-dependent TF decay. We will modify the text to better explain the rationale behind this interpretation while also acknowledging alternative interpretations that do not involve sequential IRF4-BLIMP1 functions (see response to Reviewer 1).

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c?

      We thank the reviewer for identifying these points of confusion. We will revise Fig. 3B to display outlier events and frequencies more clearly and revise the figure legend to clarify the donor origin of the displayed UMAPs and corresponding supplemental analyses. We will also clarify that Fig. 3B and Fig. 3C represent distinct readouts: flow cytometry measures IRF4 and BLIMP1 protein abundance, whereas scRNA-seq resolves transcriptional states after perturbation. Thus, the “stunted PC” state in IRF4 KO cells is defined at a transcriptional level as a population positioned between prePCs and PCs, not as a population retaining normal IRF4 or BLIMP1 protein expression. We further clarify that the apparent reduction in proliferative plasmablast-like cells at D9 likely reflects both the later timepoint relative to D7 and differences between transcriptional cell-cycle gene scores and Ki-67 protein persistence.

      What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells?

      The signature genes defining the scores in Fig. S3D are listed in Table S2. We will clarify the source of these genes in the figure legend of a future revised version.

      Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      We thank the reviewer for this suggestion to highlight IRF4, BLIMP1 and exemplar target genes in the various populations. We will update Fig. S3 to show transcript levels of IRF4, PRDM1 and an example of one of each of their target genes in unperturbed cells to demarcate their normal expression pattern.

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

      We agree that the data do not exclude the possibility that BLIMP1 is required before the perturbation window but is less continuously required once cells have progressed beyond a defined prePC stage. Our statement that BLIMP1 is not required to initiate the prePC-to-PC transition is based on the observation that PRDM1 KO cells did not accumulate in the prePC or stunted PC intermediate states observed after IRF4 loss. We will revise the text to make this interpretation more precise by stating that BLIMP1 appears less important than IRF4 for progression into a PC-like transcriptional state but is required for efficient consolidation of the mature PC program. We will also acknowledge that differences in editing efficiency, protein persistence, and timing of BLIMP1 action could contribute to the observed differences in phenotypes.

    1. Author response:

      The authors thank the reviewers for their thorough and fair assessment of our manuscript. We are currently working to edit the manuscript based on the critiques and guidance offered by the reviewers. This will consist of fixing grammar and typos, expanding the material and methods section to include more information on the behavioral assays, modifying graphs for clarity between visuals and interpretations, and correcting our mistakes in neuroanatomical labeling.

      Public Reviews:

      Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      We agree with the reviewer that our light testing data suggests complexity in the response to indirect white light within certain cave populations. In addition, an expanded sample size would likely provide clarity on whether individuals fall within two groups, behavior that looks like direct light or infra-red light, that could relate to genetic variation within cavefish populations. We are currently working to reassess the current data and determining the best course of action for continued studies related to light exposure.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      We agree with both reviewers that the methodological section on behavior needs to be expanded. We are currently editing our methods section to include more detail on water exchanges, odor preparation, timing, biological replicates, and binning.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      We agree that our annotation was incorrect or more accurately mislabeled in our write-up of the preprint and submitted manuscript. Therefore, we have now re-assessed the regions using the tissue cleared and light sheet collected zebrafish atlas, Adult Zebrafish Brain Atlas (AZBA; doi: 10.7554/eLife.69988). We are now editing the resubmission in the following manner: our initial labeling of the ventromedial thalamus will be changed to the lateral olfactory tract (nLOT) of the pallium and the preoptic region to the ventromedial thalamus (VM). We will provide a comparable z-slice of the AZBA segmentation file to illustrate the similarity in position. This would also support a known continuous circuit of olfactory integration, with information flowing from the lateral olfactory tract-to the piriform cortex-to the thalamus. We also agree that a more accurate assessment in Astyanax would require HCR in situ hybridization of markers for those specific brain regions or a neurocomputational brain atlas for adult Astyanax populations. Finally, we assert that this small dataset is preliminary at best and only provides regions of shared activity that could explain anything from perception related processes to relay of odor signaling unrelated to perception. Further work with larger sample sizes and additional populations are currently underway for a follow-up study on the neurobiological basis of olfactory processing and perception in adult cavefish.

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      We agree with the reviewer that the brain mapping section lacks a sufficient sample size and no control group for comparison (surface fish). However, we found the variation in pERK intensity (notably the putative nLOT) to be informative and decided to include the dataset in the manuscript. We are currently working to fill in these data gaps by sampling all populations and increasing the Pachon cavefish sample size. This will be a follow-up study as mentioned above in the last rebuttal paragraph.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      We agree with the reviewer that the hybrid results setup a promising follow up project to map these traits genetically. We are continuing to test odor perception in other hybrid populations and have started Quantitative Trait Locus mapping experiments.

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

      Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      We agree with the reviewer and direct their attention to the same critique by reviewer 1. We believe this is predominantly preliminary data that was included due to the conspicuous increase in pERK signal from the putative lateral olfactory tract (nLOT) for food and decay exposed cavefish. We are continuing to work on odor stimulated brain mapping and look forward to publishing a comprehensive dataset across wildtype and hybrid populations.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      We agree with the reviewer that the ethograms are challenging to read, especially due to our use of different colors for odor categories. We are currently preparing alternative graphs for displaying ethograms that reduce confusion and make following bout transitions for individual traces and bout probabilities for populations easier on the eyes.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    1. Author response:

      We thank the editors and reviewers for their thoughtful evaluation and constructive feedback. We are pleased that all three reviewers recognize the importance of mapping Cas9-induced sister chromatid exchanges (SCE) as a previously invisible repair outcome, and that the RDCP analysis is a notable feature of the study.

      We note that since our manuscript was posted, two companion studies in Science have provided direct biochemical evidence for the TRAIP-dependent pathway we discussed:

      (1) Fujisawa & Labib (Science, 2026; DOI: 10.1126/science.aeh2300) showed that TTF2 bridges CDK1-phosphorylated TRAIP to DNA Polymerase epsilon in the replisome, triggering mitotic CMG helicase disassembly, fork cleavage, and repair via SCE. Loss of this pathway reduced replication stress-induced SCE approximately two-fold in mouse ES cells.

      (2) Can et al. (Science, 2026; DOI: 10.1126/science.aeh1834) independently identified the same CDK1-TTF2-TRAIP axis in Xenopus egg extracts and validated it in HCT116 cells, showing that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions.

      We will incorporate these references in the revised discussion while still framing our RDCP observations as consistent with, rather than definitive proof of, this pathway.

      Below we briefly address the main points raised in the public reviews.

      Reviewer #1:

      We agree that the mechanistic interpretation of the RDCP signature should be presented more cautiously. We will reframe the URR/TRAIP discussion as a model, replacing language such as "direct genetic evidence" with "consistent with." We will add a summary table of RDCP data. We will also expand the description of rescued SCE calls. We will clarify what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) - we think that there is a notable difference at the bulk vs. at the single-cell level depending on the nature of the assay. Fig.1 fonts, labels, and pileup plot descriptions will be improved.

      We agree that the high-SCE subpopulation is particularly interesting and we cannot currently distinguish higher RNP uptake, a permissive cell-cycle state, altered expression, or stochastic variation. This may be better explored by future co-assays with sci-L3-Strand-seq.

      Reviewer #2:

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We will add a brief discussion on this limitation in extrapolating the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements.

      We agree that "a single Cas9 cut" should be revised to "Cas9 targeting of a single genomic locus" to accurately reflect the experimental design. We will clarify the possibility of multiple rounds of cutting at the same sites. We will also clarify "non-local" by modifying Fig.1 - we used this term specifically to include the possibility of inter-sister NHEJ.

      Reviewer #3:

      We will improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

    1. Author response:

      We thank the editors and reviewers for their careful evaluation and constructive suggestions. During the review process, we identified and corrected several presentation, terminology, citation, and figure-legend errors, and these corrections have been incorporated into the current version of preprint. Following the reviewers’ comment, we have also revised the title of the manuscript. We are now preparing a substantive revision that will distinguish more clearly between experimentally supported conclusions and hypothetical regulatory relationships. We plan to examine gene expression at an earlier time point after Br-c knockdown, further investigate the relationships among Kr-h1, chinmo, and E93 using additional RNAi experiments, characterize the cuticular phenotypes of precocious prepupae at higher magnification, and determine whether severe E93-knockdown individuals exhibit evidence of a repeated pupal developmental program. After completing these experiments, we will cautiously revise the proposed regulatory model and moderate conclusions that are not directly supported by the current evidence.

  2. Aug 2026
    1. Author response:

      The following is the authors’ response to the original reviews

      eLife assessment

      This study investigates the role of the bile acid receptor TGR5 in adult hematopoiesis of the mouse model. The findings are potentially useful because the loss of TGR5 leads to dysregulation of bone marrow adipose tissue (BMAT) that has emerging regulatory functions. However, the study is still incomplete because the mechanism of TGR5 is not clear, the stromal cells expressing TGR5 have not been well defined, and there is not strong evidence for the role of TGR5 in recovery from transplant stress.

      We thank the eLife editorial team for handling our manuscript. In our revised version, we took into consideration the suggestions of the reviewers, which we believe have significantly improved the quality of the current study. In summary, our new data provide further evidence that TGR5 is expressed in both hematopoietic cell lineage and stromal cells of the bone marrow (BM), including subpopulation analyses for both. While steady-state hematopoiesis remains intact in TGR5 knockout mice, we demonstrate that loss of TGR5 significantly impacts progenitor reconstitution under stress conditions, alters the BM adipose tissue (BMAT) homeostasis in a sex-specific manner and influences the balance of stromal cell differentiation. In particular, TGR5 deficiency resulted in reduced regulated BMAT and an accumulation of adipocyte progenitors, correlating with improved hematopoietic recovery following BM transplantation. The BMAT decrease is further observed in physiologically and pathophysiologically relevant contexts such as aging and obesity, where TGR5 deficiency is associated with a decrease in the myeloid bias. Collectively, our findings support a previously unrecognized role for TGR5 in maintaining BM niche integrity and highlight its potential as a modulator of hematopoietic support, although the precise molecular mechanisms still need to be elucidated.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      Alonso-Calleja and colleagues explore the role of TGR5 in adult hematopoiesis at both steady state and post-transplantation. The authors utilize two different mouse models including a TGR5-GFP reporter mouse to analyze the expression of TGR5 in various hematopoietic cell subsets. Using germline Tgr5-/- mice it's reported that loss of Tgr5 has no significant impact on steady-state hematopoiesis, with a small decrease in trabecular bone fraction, associated with a reduction in proximal tibia adipose tissue, and an increase in marrow phenotypic adipocytic precursors. The authors further explored the role of stroma TGR5 expression in the hematopoietic recovery upon bone marrow transplantation of wild-type cells, although the studies supporting this claim are weak. Overall, while most of the hematopoietic phenotypes have negative results or small effects, the role of TGR5 in adipose tissue regulation is interesting to the field.

      Strengths:

      This is the first time the role of TGR5 has been examined in the bone marrow.

      This paper supports further exploration of the role of bile acids in bone marrow transplantation and possible therapeutic strategies.

      We thank the reviewer for pinpointing the strengths of our study.

      Weaknesses:

      (1) The authors fail to describe whether niche stroma cells or adipocyte progenitor cells (APCs) express TGR5.

      Using the TGR5:GFP reporter model, we identified GFP<sup>+</sup> cells in the stroma-enriched CD45-Ter119-CD31- population that contains the adipogenic progenitor cells (APC).

      These data, along with the corresponding gating strategy are outline in Figure 6A and B.

      We found the subpopulation analyses within the stroma gate challenging at the individual mouse level given the limited cell numbers for these progenitor populations. We attempted to circumvent this by concatenating the individual files per genotype (WT and TGR5:GFP) for two independent experiments. We were surprised to find that there is little to no GFP expression in the APC population. Nevertheless, we found that the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup>, multi-potent stem cell-like population (Ambrosi et al., 2017), reproducibly showed a TGR5:GFP positivity comparable to that of Ly6<sup>lo</sup> monocytes. We believe this would be compatible with our results showing an increase in CFU-F in Tgr5<sup>-/-</sup> mice, but we remain cautious about its interpretation. Results are shown in Author response image 1 for the concatenated flow plots and numerically in Error! Reference source not found. considering all individual mice analyzed (i.e., non pooled) for completeness. We kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 1.

      Flow cytometry gating strategy used to identify stroma subpopulations in the stroma enriched CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup> gate and their GFP signal for BM cells in TGR5:GFP mice. Subpopulations were immunophenotypically defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (adipogenic progenitor cells (APC)) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations) (Ambrosi et al., 2017). Results are shown as the concatenated data of all mice in each phenotype for two independent experiments, as indicated.

      Author response table 1.

      Frequencies in the total live cell gate for adipogenic progenitor cells (APC) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations, MPSC-like) (gating as in Ambrosi et al., 2017) in the experiments presented in Author Response Figure 1, expressed as average +/- 95% confidence interval for the two independent experiments. The number of mice per genotype and per experiment is indicated in parenthesis. **Paired, two-tailed Student’s t-test statistical analysis for the combined experiments (n=8 per group) indicates p < 0.01 for GFP expressing cells in APC versus MPSC-like populations.

      (2) Although the authors note a significant reduction in bone marrow adipose tissue in Tgr5-/- mice, they do not address whether this is white or brown adipose tissue especially since BA-TGR5 signaling has been shown to play a role in beiging.

      The nature of BMAT and how it relates to brown, white or brown/beige adipose tissue has been a persisting question in the field. Our understanding is that BMAT is currently considered as a distinct adipose depot that is neither white nor brown/beige (Sebo et al., 2019; Suchacki et al., 2020). BMAT does not express UCP1 to an appreciable extent, with reports showing that detectable expression possibly stems from contamination by tissues surrounding bone (Craft et al., 2019). Beyond this consideration, as the regulated BMAT in Tgr5<sup>-/-</sup> mice is almost absent, determination of the brown/beige vs white nature of the little regulated BMAT that remains would be technically extremely challenging.

      (3) In Figure 1, the authors explore different progenitor subsets but stop short of describing whether TGR5 is expressed in hematopoietic stem cells (HSCs).

      We have added these data to the manuscript as part of Figure 1C.

      We have further expanded our data in Figure 1D and Figure 1–figure supplement 1A with the expression of TGR5:GFP in megakaryocyte progenitors (Lin<sup>-</sup>cKit<sup>+</sup>Sca1CD150<sup>+</sup>CD41<sup>+</sup>) as shown in Author response image 2.

      Author response image 2.

      A, representative flow cytometry gating strategy used to identify megakaryocyte progenitors (MkProg) and GFP positivity in TGR5:GFP mice and their wild-type controls. B, frequencies of GFP<sup>+</sup> cells in MkProg population in the BM of 8-12-week-old male TGR5:GFP mice and their controls (n=3 for wild-type control mice, n=4 for TGR5:GFP mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Two-tailed Student’s t-test (B) was used for statistical analysis. p-values (exact value) are indicated.

      Finally, we have completed the characterization of BM progenitor populations to include the erythroid lineage, thus covering the main hematopoietic populations. We have added these data to Figure 1E and Figure 1–figure supplement 1B.

      (4) Are there more CD45+ cells in the BM because hematopoietic cells are proliferating more due to a direct effect of the loss of Tgr5 or is it because there is just more space due to less trabecular bone?

      We observe an average 20% increase in CD45<sup>+</sup> cell counts in baseline Tgr5<sup>-/-</sup> mice. The absolute volume of bone and BMAT lost in these animals does not account for 20% of the medullary cavity volume, so we speculate that the increase in CD45<sup>+</sup> counts is not solely due to increased available volume for these cells.

      (5) In Figure 4 no absolute cell counts are provided to support the increase in immunophenotypic APCs (CD45-Ter119-CD31-Sca1+CD24-) in the stroma of Tgr5-/- mice. Accordingly, the absolute number of total stromal cells and other stroma niche cells such as MSCs, ECs are missing.

      These data are now included in the manuscript and in Author response image 3. Although we detect an increase in the relative frequency of APCs (Figure 7A), on a per-leg basis we did not detect an increase in APC numbers per leg (Author response image 3). We did however observe a decrease in total CD45<sup>-</sup>Ter119sup>-</sup>CD31sup>-</sup> stromal-enriched cells in Tgr5<sup>-/-</sup> mice. Nevertheless, given that only 2–5% of stromal cells are recovered in single-cell suspensions compared with native tissue (Coutu et al., 2017; Gomariz et al., 2018), we consider results for absolute stromal cell numbers prone to over interpretation. Our conclusion, therefore, remains one of relative enrichment of immunophenotypic APCs, supported by in vitro findings of increased adipogenesis and CFU-F formation after plating equal cell numbers.

      Endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) were also quantified, with no differences observed between groups.

      Author response image 3.

      Absolute number of adipocyte progenitor cells (APC), stroma cells (CD45-Ter119CD31<sup>-</sup>) and endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) (n=5 for both Tgr5<sup>+/+</sup> and Tgr5<sup>-/-</sup> mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Unpaired, two-tailed Student’s t-test was used for statistical analysis. p values (exact values) are shown.

      (6) There are issues with the reciprocal transplantation design in Fig 4. Why did the authors choose such a low dose (250 000) of BM cells to transplant? If the effect is true and relevant, the early recovery would be observed independently of the setup and a more robust engraftment dataset would be observed without having lethality post-transplant. On the same note, it's surprising that the authors report ~70% lethality post-transplant from wild-type control mice (Fig 4E), according to the literature 200 000 BM cells should ensure the survival of the recipient post-TBI. Overall, the results even in such a stringent setup still show minimal differences and the study lacks further in-depth analyses to support the main claim.

      We thank the reviewer for this comment. On the one hand, we respectfully disagree on the relevance of the effect size, as Tgr5<sup>-/-</sup> mice recover from low platelet and leukocyte counts significantly faster than wild-type controls. Tgr5<sup>-/-</sup> recipients recovered their platelet levels faster than the Tgr5<sup>+/+</sup> recipients, showing values consistently above 200.000/µL one week sooner; platelet levels below this threshold are considered a risk factor for bleeding events (Morowski et al., 2013; Vannini et al., 2019). In addition, on day 15 post-irradiation, Tgr5<sup>-/-</sup> recipients had higher neutrophil levels, at over the 500 cells/µL-threshold value for infection risk (Morowski et al., 2013; Vannini et al., 2019), but this difference was no longer statistically significant upon stringent multiple-comparison correction (uncorrected p-value 0.028, multiplecomparison corrected p-value 0.108). Underlining the relevance, in a clinical setting, G-CSF is routinely administered to patients daily to enhance myeloid recovery even if the acceleration of recovery is by 1-2 days (Trivedi et al., 2009).

      From the perspective of mortality, we agree that it is higher than expected and constitutes a limitation of our work. Note that we discovered a mistake in data plotting during the preparation of the revised version of the manuscript which renders the mortality curve no longer statistically significant, with a p value that changed from 0.0254 to 0.0676. We have duly changed our interpretation in the text to that of a trend towards higher survival.

      Regarding the mortality in our experiments being higher than expected, we have unfortunately suffered from cases of “swollen muzzles syndrome” in our facilities that have greatly hampered our ability to perform myeloablation experiments (Garrett et al., 2018), as even sublethal doses have resulted in the appearance of side effects that were reasons for euthanasia under Swiss legislation. For example, a strong reduction in mobility requires immediate euthanasia. All experiments were performed blinded to genotype allocation, so we can reasonably exclude experimenter bias. Finally, it could be argued that mice with more marked symptomatology leading to euthanasia are more likely to have hematopoietic deficits, which in our case was mostly seen for WT animals. We have therefore chosen to report mortality alongside the longitudinal assessment of peripheral blood counts to ensure clarity of the coherent effect and full transparency. We strongly believe that this disclosure, when examined as a limitation in the discussion section, aligns with the open science efforts advocated by eLife. Unfortunately, it is beyond the scope of this manuscript and the authors' timeline to rederive the Tgr5<sup>-/-</sup> colony in our new facility and perform new bone marrow transplantation studies with titrating doses of BM.

      Lastly, the choice of 250,000 BM cells serves as the control dose per recommended standards for competitive repopulation assays (Purton and Scadden, 2007). Quoting from this methodological reference paper: “In our experiments, each recipient mouse receives cell doses … together with 2x10<sup>5</sup> competing congenic bone marrow. We have found these cell doses sufficient to detect both reductions (Purton et al., 2006) and increases (Janzen et al., 2006; Walkley et al., 2005) in HSC numbers in different mutant mice… Furthermore, caution should be used when designing competitive repopulation assays, as it has been shown that the reliability of this assay is critically dependent on the numbers of HSCs present in the populations being assessed: when too few or too many HSCs (recipients of <1 1x10<sup>5</sup> or >2x10<sup>7</sup> bone marrow cells each from donor and competing sources) are present, the data may not be meaningful (Harrison et al., 1993).”

      Moreover, our lethal radiation rescue experiments with 2.5x105 cells are designed to deliver a minimal hematopoietic source known to rescue the great majority of animals in standard conditions, and thus to best mimic the situations when enhancement of hematopoietic recovery would be clinically meaningful (as used by our group members in (Naveiras et al., 2009; Tratwal et al., 2020; Vannini et al., 2019)). In our hands and in the absence of “swollen muzzles syndrome”, this approach leads to 80-95% overall survival, which mirrors the clinical setting of autologous transplantation that our transplants are meant to mimic. In clinical practice strategies, enhancing hematopoietic recovery would be most useful to improve morbimortality in either alternative hematopoietic progenitor allotransplants, known associated with delayed engraftment, or in the case of febrile neutropenia associated to bacteriemia, which affects 11-15% of both hematopoietic autologous or allogeneic transplant patients (Gil et al., 2013; Gooley et al., 2010). In summary, septic neutropenia was unfortunately modelled by the rescue experiments presented in Figure 7 C-J with swollen muzzles syndrome concomitant to the reconstitution, but we strongly believe that this complication enhances the clinical relevance of our data.  

      (7) Mechanistically, how does the loss of Tgr5 impact hematopoietic regeneration following sublethal irradiation?

      As delineated in the previous point, we have been seriously conditioned by cases of “swollen muzzles syndrome” (Garrett et al., 2018), which has stopped us from proceeding with more irradiation experiments for this particular study. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis is now a focus for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      (8) Only male mice were used throughout this study. It would be beneficial to know whether female mice show similar results.

      We thank Reviewer #1 for this question, as it led us to perform the characterization of steady-state hematopoiesis and morphological bone and BMAT evaluation of young female mice, yielding new findings that strengthen our manuscript. In summary, we have found that the decrease in BMAT is sexually dimorphic, with young females not showing reduced BMAT levels. Conversely, we did find a trend towards lower trabecular bone content in females. We present these data as part of Figure 3.

      Reviewer #2 (Public Review):

      Summary:

      In this manuscript, the authors examined the role of the bile acid receptor TGR5 in the bone marrow under steady-state and stress hematopoiesis. They initially showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. They further demonstrated that TGR5 knockout significantly decreases BMAT, increases the APC population, and accelerates the recovery upon bone marrow transplantation.

      Strengths:

      The manuscript is well-structured and well-written.

      We thank Reviewer #2 for this comment.

      Weaknesses:

      The mechanism is not clear, and additional studies need to be performed to support the authors' conclusion.

      We agree with Reviewer #2 that more studies are needed to understand the role of TGR5 in the hematopoietic system. We have been hampered in our studies of stress hematopoiesis because of frequent cases of swollen muzzles syndrome (Garrett et al., 2018), which prevented us from conducting additional experiments involving myelosuppression. Furthermore, the identification of the mechanism that links changes in the adipocyte differentiation axis and hematopoietic support, which is more complex than initially thought, has become a top priority for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      Recommendations For The Authors:

      Reviewer #2 (Recommendations For The Authors):

      (1) Figure 1: the authors showed the presence of TGR5 in hematopoietic cells using a GFP report in mice. What's the expression pattern of TGR5 in the nonhematopoietic cells? For example, adipocytes, stromal cells, endothelial cells, etc. In addition to analyze the percentage of TGR5-GFP+ cells, the authors should also quantify the expression levels of TGR5 in various hematopoietic and niche components.

      We thank Reviewer #2 for this question, which we have addressed as points 1 and 3 from Reviewer #1. The expression of TGR5:GFP signal in stromal cells, HSCs, and the various hematopoietic compartments is now presented respectively as new panels in Figure 6A-B, Figure 1C-D, Figure 1–Figure Supplement 1A-B, and Figure 1figure supplement 2A, as well as Author response image 1 for stromal progenitors. Please note that specifically for the stromal compartment, we kindly requested Reviewer #1’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section, as we are concerned about over interpretation for these rare APC and multi-potent stem cell-like subpopulations.

      Consistent with the typical low-abundance expression of G protein-coupled receptors, where ligand-mediated activation is more relevant than absolute receptor abundance, TGR5 is also lowly expressed in most cell types. Although technical limitations prevent us from directly quantifying TGR5 expression in adipocytes (due to the difficulty of isolating these populations from bone marrow), previous studies have reported TGR5 expression in human BMSC-derived adipocytes (Velazquez-Villegas et al., 2018) . For similar reasons, we could not quantify TGR5 expression levels by RT-qPCR in the bone marrow niche, as isolating enough cells from each compartment to reliably detect a low-abundance receptor is technically challenging.

      Regarding endothelial cells, previous studies have shown TGR5 expression in vascular endothelial cells (Kida et al., 2013) as well. In our dataset the number of endothelial cells recovered was limited, largely due to the lack of a dedicated endothelial isolation protocol for the BM. Even so, the data we collected are shown as Author response image 4 and Author response table 2 and suggest that TGR5:GFP level in endothelial cells is lower than in the stromal and hematopoietic compartments. As for Author response image 1 and Author response table 1, we remain cautious about the interpretation of GFP expression in these low frequency stromal and endothelial populations and we kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 4.

      Representative flow cytometry gating strategy used to identify endothelial cells in BM, defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> and their GFP positivity in TGR5:GFP mice.

      Author response table 2.

      Frequency of GFP-expressing cells in the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> gate for endothelial cells presented in Author Response Figure 4, expressed as average (standard deviation). The number of mice per genotype and per experiment is indicated in parenthesis.

      (2) Figure 2: Regarding the competitive transplantation, the authors should bleed the mice every 4 weeks and show the dynamics of donor chimerism up to 16 or 20 weeks. The difference at the 3-week time point is very tiny and is this change significant? There is no statistical significance shown in Supplemental Figure 2 Panel C at a 3-week time point.

      The full data for repetitive monthly bleedings in primary, secondary and tertiary transplants is now shown in Figure 2J and Figure 2-figure supplement 1I-K. The statistically significant difference on the first month, which represents a 26% loss in short-time progenitor phenotype, is small but potentially clinically relevant as it is these progenitors that sustain early hematopoietic recovery from severe leucopenia and thrombocytopenia. Indeed, the hematopoietic phenotype described in Figure 7 C-J for accelerated rescue of Tgr5<sup>-/-</sup> recipients with wild-type bone marrow would be coherently associated to this short-term progenitor phenotype.  

      (3) Figure 3: in addition to the irradiation/transplantation model, the authors should also use alternative models for stress hematopoietic, for example, 5-FU, etc.

      This is an excellent suggestion, but unfortunately beyond the scope of this manuscript due to the move of one of the co-senior authors to another institution. Follow-up studies are however planned with ablative chemotherapy models relevant to the hematopoietic transplant setting (non 5-FU).

      (4) To get a better understanding of the role of TGR5 in the bone marrow niche, the authors should get and/or generate the TGR5 floxed mice and this will allow the authors to delete it from specific cell populations using cell-type specific Cre. In addition, it would be very interesting to investigate further the molecular mechanisms downstream of TGR5.

      We much appreciate this comment and the interest it reflects in TGR5 BM biology. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis are now a focus for one of the involved laboratories and should produce concrete follow-up manuscripts soon.  

      References

      Ambrosi, T.H., Scialdone, A., Graja, A., Gohlke, S., Jank, A.-M., Bocian, C., Woelk, L., Fan, H., Logan, D.W., Schürmann, A., Saraiva, L.R., Schulz, T.J., 2017. Adipocyte Accumulation in the Bone Marrow during Obesity and Aging Impairs Stem Cell-Based Hematopoietic and Bone Regeneration. Cell Stem Cell 20, 771-784.e6. https://doi.org/10.1016/j.stem.2017.02.009

      Coutu, D.L., Kokkaliaris, K.D., Kunz, L., Schroeder, T., 2017. Three-dimensional map of nonhematopoietic bone and bone-marrow cells and molecules. Nat Biotechnol 35, 1202–1210. https://doi.org/10.1038/nbt.4006

      Craft, C.S., Robles, H., Lorenz, M.R., Hilker, E.D., Magee, K.L., Andersen, T.L., Cawthorn, W.P., MacDougald, O.A., Harris, C.A., Scheller, E.L., 2019. Bone marrow adipose tissue does not express UCP1 during development or adrenergic-induced remodeling. Sci Rep 9, 17427. https://doi.org/10.1038/s41598-019-54036-x

      Garrett, J., Sampson, C.H., Plett, P.A., Crisler, R., Parker, J., Venezia, R., Chua, H.L., Hickman, D.L., Booth, C., MacVittie, T., Orschella, C.M., Dynlachta, J.R., 2018. Characterization and Etiology of Swollen Muzzles in Irradiated Mice. Radiation Research 191, 31. https://doi.org/10.1667/RR14724.1

      Gil, L., Poplawski, D., Mol, A., Nowicki, A., Schneider, A., Komarnicki, M., 2013. Neutropenic enterocolitis after high-dose chemotherapy and autologous stem cell transplantation: incidence, risk factors, and outcome. Transpl Infect Dis 15, 1–7. https://doi.org/10.1111/j.1399-3062.2012.00777.x

      Gomariz, A., Helbling, P.M., Isringhausen, S., Suessbier, U., Becker, A., Boss, A., Nagasawa, T., Paul, G., Goksel, O., Székely, G., Stoma, S., Nørrelykke, S.F., Manz, M.G., Nombela-Arrieta, C., 2018. Quantitative spatial analysis of haematopoiesis-regulating stromal cells in the bone marrow microenvironment by 3D microscopy. Nat Commun 9, 2532. https://doi.org/10.1038/s41467-018-04770-z

      Gooley, T.A., Chien, J.W., Pergam, S.A., Hingorani, S., Sorror, M.L., Boeckh, M., Martin, P.J., Sandmaier, B.M., Marr, K.A., Appelbaum, F.R., Storb, R., McDonald, G.B., 2010. Reduced mortality after allogeneic hematopoietic-cell transplantation. N Engl J Med 363, 2091–2101. https://doi.org/10.1056/NEJMoa1004383

      Harrison, D.E., Jordan, C.T., Zhong, R.K., Astle, C.M., 1993. Primitive hemopoietic stem cells: direct assay of most productive populations by competitive repopulation with simple binomial, correlation and covariance calculations. Exp Hematol 21, 206–219.

      Janzen, V., Forkert, R., Fleming, H.E., Saito, Y., Waring, M.T., Dombkowski, D.M., Cheng, T., DePinho, R.A., Sharpless, N.E., Scadden, D.T., 2006. Stem-cell ageing modified by the cyclin-dependent kinase inhibitor p16INK4a. Nature 443, 421–426. https://doi.org/10.1038/nature05159

      Kida, T., Tsubosaka, Y., Hori, M., Ozaki, H., Murata, T., 2013. Bile acid receptor TGR5 agonism induces NO production and reduces monocyte adhesion in vascular endothelial cells. Arterioscler Thromb Vasc Biol 33, 1663–1669. https://doi.org/10.1161/ATVBAHA.113.301565

      Morowski, M., Vögtle, T., Kraft, P., Kleinschnitz, C., Stoll, G., Nieswandt, B., 2013. Only severe thrombocytopenia results in bleeding and defective thrombus formation in mice. Blood 121, 4938–4947. https://doi.org/10.1182/blood-2012-10-461459

      Naveiras, O., Nardi, V., Wenzel, P.L., Hauschka, P.V., Fahey, F., Daley, G.Q., 2009. Bonemarrow adipocytes as negative regulators of the haematopoietic microenvironment. Nature 460, 259–263. https://doi.org/10.1038/nature08099

      Purton, L.E., Dworkin, S., Olsen, G.H., Walkley, C.R., Fabb, S.A., Collins, S.J., Chambon, P., 2006. RARgamma is critical for maintaining a balance between hematopoietic stem cell self-renewal and differentiation. J Exp Med 203, 1283–1293. https://doi.org/10.1084/jem.20052105

      Purton, L.E., Scadden, D.T., 2007. Limiting factors in murine hematopoietic stem cell assays. Cell Stem Cell 1, 263–270. https://doi.org/10.1016/j.stem.2007.08.016

      Sebo, Z.L., Rendina-Ruedy, E., Ables, G.P., Lindskog, D.M., Rodeheffer, M.S., Fazeli, P.K., Horowitz, M.C., 2019. Bone Marrow Adiposity: Basic and Clinical Implications. Endocr Rev 40, 1187–1206. https://doi.org/10.1210/er.2018-00138

      Suchacki, K.J., Tavares, A.A.S., Mattiucci, D., Scheller, E.L., Papanastasiou, G., Gray, C., Sinton, M.C., Ramage, L.E., McDougald, W.A., Lovdel, A., Sulston, R.J., Thomas, B.J., Nicholson, B.M., Drake, A.J., Alcaide-Corral, C.J., Said, D., Poloni, A., Cinti, S., Macpherson, G.J., Dweck, M.R., Andrews, J.P.M., Williams, M.C., Wallace, R.J., Van Beek, E.J.R., MacDougald, O.A., Morton, N.M., Stimson, R.H., Cawthorn, W.P., 2020. Bone marrow adipose tissue is a unique adipose subtype with distinct roles in glucose homeostasis. Nat Commun 11, 3097. https://doi.org/10.1038/s41467-020-16878-2

      Tratwal, J., Bekri, D., Boussema, C., Sarkis, R., Kunz, N., Koliqi, T., Rojas-Sutterlin, S., Schyrr, F., Tavakol, D.N., Campos, V., Scheller, E.L., Sarro, R., Bárcena, C., Bisig, B., Nardi, V., de Leval, L., Burri, O., Naveiras, O., 2020. MarrowQuant Across Aging and Aplasia: A Digital Pathology Workflow for Quantification of Bone Marrow Compartments in Histological Sections. Front Endocrinol (Lausanne) 11, 480. https://doi.org/10.3389/fendo.2020.00480

      Trivedi, M., Martinez, S., Corringham, S., Medley, K., Ball, E.D., 2009. Optimal use of G-CSF administration after hematopoietic SCT. Bone Marrow Transplant 43, 895–908. https://doi.org/10.1038/bmt.2009.75

      Vannini, N., Campos, V., Girotra, M., Trachsel, V., Rojas-Sutterlin, S., Tratwal, J., Ragusa, S., Stefanidis, E., Ryu, D., Rainer, P.Y., Nikitin, G., Giger, S., Li, T.Y., Semilietof, A., Oggier, A., Yersin, Y., Tauzin, L., Pirinen, E., Cheng, W.-C., Ratajczak, J., Canto, C., Ehrbar, M., Sizzano, F., Petrova, T.V., Vanhecke, D., Zhang, L., Romero, P., Nahimana, A., Cherix, S., Duchosal, M.A., Ho, P.-C., Deplancke, B., Coukos, G., Auwerx, J., Lutolf, M.P., Naveiras, O., 2019. The NAD-Booster Nicotinamide Riboside Potently Stimulates Hematopoiesis through Increased Mitochondrial Clearance. Cell Stem Cell 24, 405-418.e7. https://doi.org/10.1016/j.stem.2019.02.012

      Velazquez-Villegas, L.A., Perino, A., Lemos, V., Zietak, M., Nomura, M., Pols, T.W.H., Schoonjans, K., 2018. TGR5 signalling promotes mitochondrial fission and beige remodelling of white adipose tissue. Nat Commun 9, 245. https://doi.org/10.1038/s41467-017-02068-0

      Walkley, C.R., Fero, M.L., Chien, W.-M., Purton, L.E., McArthur, G.A., 2005. Negative cellcycle regulators cooperatively control self-renewal and differentiation of haematopoietic stem cells. Nat Cell Biol 7, 172–178. https://doi.org/10.1038/ncb1214

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public reviews):

      Weaknesses:

      Reliance on self-reports

      A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.

      We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”

      Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.

      Statistical analysis and interpretation of the dynamics matrix

      Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.

      We thank the reviewer for raising this important methodological point. We address it in three parts.

      (i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.

      (ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T<sup>2</sup>, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.

      (iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”

      Terminology (controllability, stability, sensitivity, Gramian)

      Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.

      This is an excellent and important point. We have made the following changes throughout the manuscript.

      (i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.

      (ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”

      (iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.

      (iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.

      (v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.

      Reviewer #2 (Public reviews):

      Online recruitment, selection effects, and remuneration

      Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).

      We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”

      Intervention ongoing during the second block

      Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.

      This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”

      Observation noise

      As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.

      We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.

      Reliance on formal model comparison

      Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.

      We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.

      We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.

      Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.

      Statistical limitations; Bayesian inference

      The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.

      We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.

      Reviewer #3 (Public reviews):

      Dual meanings of controllability

      An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”

      We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.

      Strategy reminder and possible reappraisal

      As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.

      This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.

      Mechanism of distancing effects (eye movement and oculomotor avoidance)

      Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.

      We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”

      Recommendations for the Authors:

      Reviewer #1 (Recommendations for the Authors):

      (1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?

      Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.

      (2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?

      We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.

      (3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.

      References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.

      (4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.

      The typo has been corrected.

      (5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.

      We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.

      (6) Could the authors develop the rationale behind the bias in Equation 1?

      We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.

      (7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.

      We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.

      (8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).

      We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.

      (9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.

      The sentence has been rephrased: ”Cumulative model weights (w<sub>j</sub>) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”

      (10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.

      We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.

      (11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.

      We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.

      (12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.

      We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.

      <(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.

      Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.

      We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      (14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?

      The regression equation has been corrected to DV = β<sub>0</sub> + β<sub>1</sub>IV + β<sub>2</sub>G + β<sub>3</sub>(IV × G) + ϵ, making the interaction term explicit.

      (15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?

      The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.

      Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.

      Beyond the visual trajectory overlays in Figure 4C, we have now added R<sup>2</sup> and peak cross-correlation as a quantitative measure of individual fit quality. Mean R<sup>2</sup> across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R<sup>2</sup> values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.

      (16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?

      Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.

      (17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.

      Interaction effects are now included in the relevant figure caption (Figure 5).

      (18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?

      Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?

      This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.

      (19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?

      This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.

      (20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?

      Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.

      Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.

      (21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?

      The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.

      (22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).

      Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.

      (23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.

      We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.

      (24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.

      The supplementary tables have been restructured to a hierarchical format.

      (25) Table I9 is not referenced in the manuscript.

      A reference to this table has been added in the appropriate Results section.

      (26) It’s a shame that neither data nor analysis code has been made available.

      Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).

      Reviewer #2 (Recommendations for the Authors):

      (1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.

      The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.

      (2) p 5: ’on [not in] the recruiting platform’.

      Corrected.

      (3) p 8: Clarify notation of x<sub>t</sub>t = 1<sup>T</sup> (and same for u). What does this mean?

      The notation has been corrected.

      (4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.

      The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.

      (5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.

      References have been added at each key definition.

      (6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.

      This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.

      (7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?

      A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.

      (8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?

      Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.

      (9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?

      We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.

      (10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.

      Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.

      (11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].

      This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      Reviewer #3 (Recommendations for the Authors):

      (1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.

      More information about the power analysis have been added to the Participants section of the main text.

      (2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).

      We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.

      (3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”

      Thank you for this comment. Please see our response to the public review comments above, which we hope address this.

      (4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?

      See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.

      (5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.

      We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination. Other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain.

      We sincerely thank the editors and reviewers for the thorough review of our manuscript. We appreciate the insightful comments, which have significantly contributed to improving our work. In this revision, we addressed all the concerns posed by the reviewers, including conducting additional analyses and providing the data generated and codes used in this study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics, and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      We thank the reviewer for the comments and suggestions regarding our manuscript. They have been invaluable and instrumental in enhancing our work. Please see R1-1 to R1-8 below for our corresponding responses and revisions.

      Comments:

      (R1-1) Figure 1A, on the top of the figure, ChR2 may be replaced by ChR2-mCherry, as only mCherry is fluorescent. And also, it’s somewhat surprising that in AON and Pir regions (where only axon fibers should be labelled as red), most fluorescence appeared dot-like and looked more similar to cell body instead of typical fiber. The authors may want to double-check this.

      (a) Thank you for pointing out the missing fluorescent marker for ChR2 expression in the label of Figure 1A. We distinguished the transfection of ChR2 in olfactory bulb (OB) neurons from axonal terminals in AON and Pir by identifying the expression patterns in these three regions (see Figure 1). In OB, the cellular nucleus (i.e., labelled in blue by DAPI) was surrounded evenly by ChR2-mCherry expression (red fluorescent, indicated by green arrows in Figure 1), which clearly traced the neuronal cell body shape. Meanwhile, the expression of mCherry in AON and Pir consisted of concentrated red fluorescent dots that were sparsely assembled near the cell nucleus, indicating the presence of synaptic boutons at the axonal terminals. Further, the patterns of expression here did not trace the shape of the neuronal cell body.

      (b) In this revision, we amended the label of Fig. 1A to “ChR2-mCherry”.

      In this revision, changes were made to the confocal images of AON and Pir, based on Figure 1, for clarity and the figure caption was also edited accordingly.

      (R1-2) The authors primarily presented 1 Hz stimulation results. What is the most biologically relevant frequency (e.g., perhaps firing frequency under natural odor stimulation) among all frequencies that were used?

      (a) There are various firing rate and oscillatory frequency of neural activities within the olfactory system (i.e., OB, Pir) in rodents during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      The piriform cortex (Pir) also exhibits natural oscillatory activity across various frequency bands, including slow, theta, beta, and gamma oscillations [9-11]. As in OB mitral and tufted cells, slow oscillations (< 1.5 Hz) in Pir are also correlated with the respiratory rhythm, indicating the intrinsic respiration-related oscillatory properties within the olfactory system [12]. However, while it has been shown that the primary burst of firing in the anterior Pir is locked to respiration, its coding strategies differed from OB during olfaction [13].

      Despite not being able to entirely mimic the evoked neural activities under natural odor stimulation due to the synchronizations induced by our optogenetic stimulations, our frequencies were chosen to be within the range of firing rates of neurons in the olfactory system under natural circumstances. Our selection of 1 Hz frequency for optogenetic stimulation was based on the balance of experimental simplicity of inducing neural excitation and the biological significance of low-frequency neural oscillations in OB, Pir and AON. The robustness of evoked brain-wide long-range activations is one of our criteria for the selection of 1 Hz stimulation frequency at OB, considering that the goal of this study is to examine long-range olfactory networks. Furthermore, we also examined other frequencies of stimulation covering theta, beta and gamma frequency bands, which provided insights into the response characteristics within the olfactory networks.

      (b) In this revision, we added a brief statement in the Results section that 1 Hz was the most biologically relevant frequency based on discussions in R1-2a above.

      (R1-3) In Figure 2, the statistical thresholding is confusing: in the figure legend, it was stated that “t > 3.1 corresponding to P < 0.001” but later “further corrected for multiple comparisons with thresholdfree cluster enhancement with family-wise error rate (TFCE-FWE) at P < 0.05”? Regardless of the statistical thresholding, such BOLD activation seemed to be widespread (almost whole-brain activation). Does such activation remain specific to the optogenetic stimulation, or something more general (e.g., arousal level change)? Furthermore, how those results (I assume they are group-level results) were obtained was not described very clearly. Is it just a simple average of individual-level results, or (more conventionally) second-level analysis?

      (a) We thank the reviewer for drawing our attention to clarity issues regarding the generation of group-level BOLD activation maps and statistical thresholding methods used.

      In brief, we first generate the BOLD activation map of individual animal through conventional general linear model (GLM) analysis. These individual activation maps then underwent a two-step statistical thresholding method to generate the group-level results. Firstly, uncorrected one-sample t-tests were conducted with a threshold of P < 0.001. Secondly, these thresholded activation maps were further corrected using nonparametric inference with threshold-free cluster enhancement multiple comparison correction of family-wise error rate (TFCE-FWE, P < 0.05) before they were averaged. Hence, the BOLD activation maps presented in our manuscript were the result of group-level analysis instead of individual-level analysis.

      (b) We note the reviewer’s concern about whether the observed widespread BOLD activations were caused by the animal’s general brain state (e.g., arousal levels) rather than the specificity of optogenetic stimulation. First, such widespread activations were only specific to 1 Hz stimulation of OB neurons and OB afferents at AON (Figs. 2A, B and 3A, B) whereas activations were localized to regions in the primary olfactory network with increasing stimulation frequencies. Furthermore, varied neural activity adaptation properties were observed (Fig. 3A, B) following repeated 1 Hz optogenetic excitation of OB neurons and OB afferents at AON indicating that such widespread propagation of evoked neural activity was not driven by the animal’s general brain state, which would be relatively random across animals as we interleaved the presentation of each stimulation frequencies. While we cannot discount that certain characteristics of brain states differ under anaesthetized and awake conditions, we showed that light (1.0% isoflurane) anaesthesia minimally affects the BOLD fMRI activations and/or the propagation of optogenetically-evoked neural activity across thalamo-cortical, hippocampal-cortical and vestibulo-cortical networks [14-20].

      Second, the neural representations in OB have been shown to be brain state-independent to ensure the high fidelity of olfactory inputs to higher olfactory cortices, despite differences in spontaneous baseline activity across states [21-24]. In fact, OB neural activity (i.e., slow to delta oscillations) remains highly coupled to respiration rhythms under low anaesthesia as in awake conditions [4]. Further, low-dose isoflurane (i.e., 1% as in the present study) does not affect cortical gamma oscillations (40 Hz) [25-28]. Although odorant-evoked responses in olfactory cortices (e.g., Pir and olfactory tubercle, Tu) can be modulated by overall brain state, various interactions between such cortical circuits to process odor inputs persist under anaesthesia [29-31].

      (c) In this revision, we clarified the two-step statistical thresholding method that was applied to generate the group-level BOLD activation maps, as discussed in R1-3a above, in the captions of Fig. 2 and the Methods section of the manuscript.

      In this revision, we also included a statement in the Discussion section that the widespread BOLD fMRI activations were specific to the 1 Hz optogenetic stimulation and not the animal’s general brain state, as discussed in R1-3b above.

      (R1-4) In Figure 2, why use AUC to quantify the activation, not the more conventional beta value in the GLM analysis?

      (a) We thank the reviewer for raising this concern. We are aware of the more conventional beta value, b, in GLM analysis and have used them for quantification and comparison of the amplitudes of fMRI activations in our previous rodent fMRI studies [18,32,33]. However, in this study, we chose to utilize area under curve (AUC) as it offers a more comprehensive measure of BOLD signal change over time, including shape, duration, and magnitude, thereby capturing the bulk of neural activities and their dynamics throughout the stimulation period. b primarily represents the peak amplitude of BOLD responses (i.e., the % BOLD signal change) [34] and can be constrained by the assumptions and limitations of the GLM analysis, such as the shape of the canonical hemodynamic response function (HRF). Therefore, AUC provides greater accuracy in capturing different aspects of neural responses across various brain regions, such as transient peaks and/or sustained responses.

      (b) In this revision, we have included the justifications for using AUC over the more conventional b value to quantify BOLD fMRI activations, as in R1-4a above, in the Methods section.

      (R1-5) For Figure 2D, the way that it was quantified can be better described as “relative” activation within one condition, and I don’t know how to interpret the comparison among the relative fraction of activated regions. Perhaps comparison using percentage change (i.e., beta values) is more straightforward.

      (a) We thank the reviewer’s comment and suggestion here. “BOLD activation strength” is the AUCs normalized to the respective sum of BOLD activations of their respective stimulation target. We are of the opinion that “strength” can accurately reflect our purpose for conducting this comparison, which is to distinguish the brain networks that were predominantly recruited by AON- or Pir-driven neural activities compared to OB stimulation. 

      (b) We conducted a comparison using relative fraction for fair comparison considering the differences in absolute BOLD signal amplitudes upon different stimulation targets. As we found the overall amplitudes of activations evoked by Pir stimulation were appreciably weaker, our normalization step ensures that the contribution of Pir is not underrepresented. Without this normalization process, the Pir stimulation seems not to activate most of the downstream targets based on the relatively weak activations. As discussed earlier in R1-4a above, comparison using percentage change/b value would be skewed towards the absolute amplitude of BOLD responses and would be erroneous for subsequent interpretation of brain networks that are primarily recruited by OB vs. AON vs. Pir.

      (b) In this revision, we chose not to include additional statements as we have already described in the Results section the reasons for normalizing AUC to reflect and subsequently compare the BOLD activation strength across various networks that were recruited by optogenetic stimulation of three distinct olfactory regions (i.e., OB, AON and Pir). 

      (R1-6) For Figure 3, it may be more convenient for readers to include the results of 1st activation for direct comparison. The current layout makes it difficult to make direct, visual comparisons among all 3 activations. Again, I think using beta values (instead of AUC) may be more conventional.

      (a) We agree that it will be more convenient for readers if we include the 1st activation maps in Figure 3 for direct comparison. Please refer to R1-4 for the usage of AUC rather than beta values.

      (b) In this revision, we modified Fig. 3 to include activation maps from the 1st fMRI session for easy comparison with the 2nd and 3rd sessions.

      (R1-7) Can the DCM results (at least part of it) be verified using the current electrophysiological data? For example, the long-range inhibitory effective connectivity of AON is rather intriguing. If that can be verified using the electrophysiology data, it would be really great. In the current form, the DCM and electrophysiology results seem to be totally unrelated.

      (a) We thank the reviewer for raising this concern, and it’s a great suggestion to causally link the outcomes from dynamic causal modeling (DCM) analysis and electrophysiology findings.

      (b) In principle, the recorded local field potentials (LFPs) can be treated as the neural dynamic component in DCM analysis under specific conditions. They are as follows:

      (1) LFPs are recorded at the identical brain state as the optogenetic fMRI experiments with identical stimulation paradigms.

      (2) The electrode locations of LFP recordings must match the nodes of the a priori matrix defined for the existing DCM analysis (Fig. 4A).

      However, in the present study, we did not have LFP recordings at the entorhinal cortex (Ent), which will likely influence the modeling of effective connectivity strength as one of the nodes defined is dropped from the analysis. Previous studies have showed that changes in the a priori matrix will result in different connectivity estimations [35-38]. We chose to model the connectivity involving Ent due to its documented interactions with the primary olfactory cortices and hippocampal regions [39,40]. Hence, it would be incomplete without modelling the contributions from Ent.

      (c) In this revision, we decided not to redefine the a priori matrix by removing Ent as a node in the DCM analysis to estimate new effective connectivities. Instead, we acknowledge the absence of Ent recordings.

      (R1-8) In Figure 6, it would be great if the adaptation of BOLD and electrophysiology signals can be correlated at the brain region level. The current figure only demonstrated there is adaptation in the electrophysiology recording, but did not show if such adaptation is related to the BOLD adaptation.

      (a) We thank the reviewer for pointing out that we should associate the electrophysiology and fMRI data to validate the olfactory adaptation. As such, we conducted additional correlation analysis between LFP power and BOLD signal profiles. The method for computing the correlation coefficient is described below:

      (1) For each 1 Hz optogenetic stimulation target (i.e., OB, AON, and Pir), regions of interest (ROIs) that have both LFP recordings and BOLD activations were selected. These ROIs were AON, Pir, ventral caudate putamen (vCPu), visual cortex (V1), and ventral hippocampus (vHP) for OB stimulation; OB, Pir, vCPu, V1, amygdala (Amg), and vHP for AON stimulation; and OB, AON, vCPu, V1, Amg, and vHP for Pir stimulation. Note that we excluded the respective stimulated region as a ROI because of BOLD fMRI signal dropout caused by the implanted optical fibre.

      (2) We first calculate the individual animal AUC differences of 2nd stimulation session vs. 1st session and 3rd session vs. 1st session in LFP power and BOLD signal profiles, respectively.

      Subsequently, the mean of the differences for each ROI across animals was then calculated.

      (3) Correlation coefficients and P values of AUC differences between LFP power and BOLD signal profiles were computed using Spearman nonparametric correlation.

      (b) The relationship between the differences of LFP power and BOLD signal profiles showing neural adaptation was described by the correlation coefficient, r, and the corresponding P value. The r and P values upon the three distinct stimulations were OB: r = 0.77 (P = 0.01), AON: r = 0.46 (P = 0.13), and Pir: r = - 0.06 (P = 0.85), respectively. This indicates that the decrease of LFP power is highly correlated with the decrease of BOLD activations upon OB and AON stimulation compared to those upon Pir stimulation.

      (c) In this revision, we added the description of the correlation analysis conducted, as in R1-8a above, in the Methods section.

      In this revision, we also added several statements in the Results section to indicate that neural adaptation observed is significantly related to the decreased BOLD activations when stimulating OB excitatory neurons or OB afferents at AON, as discussed in R1-8b above. Additionally, we also included the outcomes of the correlation analysis as a new Supplementary Fig. 7.

      Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics, and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) The neurons in the Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.

      (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, and clear and provide much information with the most appropriate plots and graphics. The study’s amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      We are grateful for the reviewer’s appreciation of our study’s breadth and depth, and the quality and quantity of our data. The recognition of the strengths in our methods and the inclusion of single-animal-level data in the supplementary information is highly encouraging. We are also appreciative for the reviewer’s constructive comments below, which has help us to address specific weaknesses of the study and thereby improving the manuscript text and flow. We have read through the comments carefully and have made the necessary corrections. We hope that our responses to R2-1 to R2-11 below has sufficiently addressed all the reviewer’s concerns.

      Weaknesses:

      (R2-1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an ‘enormous’ field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (a) As per the reviewer’s suggestion, we displayed the representative EPI images after preprocessing at a matrix size of 128 x 128 with a pixel size of 0.25 x 0.25 mm in Supplementary Figure 2. Up-sampling was not applied out-of-plane. Note that the atlas-based region-of-interest (ROIs) used to extract BOLD signal profiles (i.e., indicated by colored overlays in Supplementary Figure 2) were drawn based on the up-sampled EPI images (from the acquired 64 x 64 to 128 x 128). Notably, the piriform cortex (Pir) and amygdala (Amg) are visible with no appreciable signal dropout at these regions. To clarify, the Pir in our study represented the anterior Pir.

      Signal dropout is pronounced in the posterior ventral regions of the brain (e.g., Bregma- 4 mm, below the Ent) and as expected in a localized region where the implanted optical fiber region that targeted the AON.

      (b) In this revision, we included as a new Supplementary Figure 8 and added a statement in the Results section to describe the absence of appreciable signal dropouts at OB, Pir and Amg, which could affect subsequent quantitative analyses.

      (R2-2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (a) We thank the reviewer for drawing our attention to the missing information regarding the timeline of each optogenetic fMRI experiment. We conducted fMRI experiments immediately after the surgical procedure for implanting the optical fiber cannula. Such an implantation procedure typically last for 45mins under 1.2-2.0% isoflurane. The animal is then moved from the surgical table to the magnet to begin the optogenetic fMRI experiment. Intervals between two fMRI scans were within one minute. As the stimulation frequencies (i.e., 1 Hz, 5 Hz, 10 Hz, 20 Hz and 40 Hz) were pseudorandomized, the interval between any two identical stimulation frequency was on average 30 minutes. In this case, we are measuring fast adaptation (on the order of tens of minutes) rather than long-lasting adaptation. Further, as the fMRI experiment was conducted immediately following fiber implantation, the risk of gliosis around the fiber tips is expected to be low.

      (b) The main representation of olfactory adaptation is the attenuation of neural responses upon continuous or repeated stimuli [41,42]. Several fMRI studies in rodents have reported that olfactory adaptation was detected in primary olfactory cortices (i.e., OB, AON, and Pir) upon repeated odor stimulation [43,44]. Specifically, these studies showed robust decreased responses in primary olfactory cortices under repeated odor stimulation with an odor interval of 30 minutes, which were comparable to those observed in our study upon repeated 1 Hz optogenetic stimulation. As such, the pronounced neural adaptation that we demonstrated when stimulating OB excitatory neurons or OB afferents at AON is unlikely to be caused by less efficient light activation.

      (c) In this revision, we added statements in the Methods section, as discussed in R2-2a above, to clarify the timeline of our optogenetic fMRI experiment.

      (R2-3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/day in a row, or am I mistaken?

      (a) We appreciate the reviewer’s comment here and understand the concerns regarding the absence of an age-matched control group with saline administration instead of D-galactose. We acknowledge that the inclusion of a control group would have provided a more direct comparison to evaluate the differences between healthy and aged animals. However, we were unable to include this group in the present study due to logistical constraints.

      To clarify, our experiments were conducted on healthy rats at the median age of 14 weeks to ensure the animals had reached adulthood. The median age of the D-galactose injected animal was 18 weeks. In terms of the lifespan of rats, they reach sexual maturity at approximately 6 weeks of age [45], at which point we conducted the optogenetic viral vector injection. The age of rats at 14 and 18 weeks can both be regarded as teenage adult [45,46]. The primary purpose of the experiments on the aged animal model and corresponding analyses was to provide some potential clues into the dysfunction of olfactory networks at the system level as animals aged. Despite the age difference between the two groups, we believe that the comparison between healthy and aged rats still offers valuable insights into the ageing-related dysfunctions within the olfactory system. 

      (b) Note that D-galactose was injected into each rat 56 times in total over the whole experiment (i.e., one injection per day over the course of 8 weeks).

      (c) In this revision, we acknowledge that the comparisons made were not with age-matched healthy controls in the Discussion section.

      In this revision, we also clarified that D-galactose was administered once daily for 8 weeks (i.e., a total of 56 injections) in the Methods section.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Minor:

      (R2-4) The look of the activation maps will greatly improve if the displayed brain slices are coronal (as in the supplementary figures) with no angle.

      We thank the reviewer for this suggestion. The purpose of such a display was to maintain the consistency of displaying the overlay of BOLD fMRI activation maps on the 3D renders of rat brain anatomical images. In addition, such 3D display can better impress readers that the optogenetically evoked BOLD activations were indeed long-range and brain-wide. 

      As such, in this revision, we decided to maintain the existing displays of BOLD activation maps in the main figures.

      (R2-5) The choice of testing different stimulation frequencies deserves more justification. Why was it performed in the OB, which is naturally activated on the breathing rhythm?

      (a) We thank the reviewer’s comment here to seek further clarification. OB is widely recognized as the first stage of olfactory information processing in the brain. By activating the OB excitatory neurons, we can ensure that our optogenetically-evoked activations are olfactory-related.

      The choice of various frequencies ranging from low (i.e., 1 Hz) to high (e.g., 40 Hz) was made to cover a range of firing rate and oscillatory frequency of neural activities neural oscillations that exist within the rodent olfactory system during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      (b) In this revision, we added a statement in the Results section to further describe our justifications for testing various stimulation frequencies.

      (R2-6) The absence of natural (odor) stimulation should be acknowledged as a limitation for the findings of this study in the Discussion.

      We thank the reviewer for the suggestion. In this revision, we made the acknowledgement in the Discussion section. 

      (R2-7) Line 46: overstatement, the role of the Piriform cortex at the system level is well known and was not discovered by this study. (see Gordon Shepherd’s book: The Synaptic Organization of the Brain, Figure 10.6)

      As per the reviewer’s suggestion, we decided to use “distinguishes” rather than “uncovers” for an accurate description of our study’s findings, which demonstrated the differences between the piriform cortex and anterior olfactory nucleus in driving the downstream activations of olfactory and non-olfactory targets.

      (R2-8) Line 75: please rephrase for an easier understanding.

      As per the reviewer’s suggestion, we edited this statement to “Further, the need to present a multitude of odor combinations also makes it challenging to efficiently and reliably interrogate long-range olfactory networks and their properties”.

      (R2-9) Line 313: the intermediary region between the Piriform and Entorhinal cortices is the Perirhinal Cortex. No doubt about that.

      As per the reviewer’s comment, we modified the corresponding statement by adding “such as the perirhinal cortex” and cited the appropriate references [47,48].

      (R2-10) Line 391: there is no evidence from this study to exclude that the downstream targets would be less activated because of a decreased OB activation in favor of broad inhibition.

      As per the reviewer’s suggestion, we modified the statement by replacing “rather than” with “in addition to” for a more precise statement.

      (R2-11) Line 585: which was/were the regressor/s used in GLM? Any convolution with common HRFs?

      We regarded the block-designed optogenetic stimulation as a task. The regressors are then obtained by convolving the expected task-evoked neural activity (i.e., boxcar function for block design) with the canonical hemodynamic response function, HRF (SPM12, Wellcome Department of Imaging Neuroscience, University College London, UK).

      In this revision, we added a statement in the Methods section to provide clarification on the regressors used in GLM and the convolution step with canonical HRF.

      References

      (1) Kay, L. M. & Stopfer, M. Information processing in the olfactory systems of insects and vertebrates. Semin Cell Dev Biol 17, 433-442 (2006). https://doi.org/10.1016/j.semcdb.2006.04.012

      (2) Cang, J. & Isaacson, J. S. In vivo whole-cell recording of odor-evoked synaptic transmission in the rat olfactory bulb. J Neurosci 23, 4108-4116 (2003). 

      (3) Margrie, T. W. & Schaefer, A. T. Theta oscillation coupled spike latencies yield computational vigour in a mammalian sensory system. J Physiol 546, 363-374 (2003). https://doi.org/10.1113/jphysiol.2002.031245

      (4) Fontanini, A. & Bower, J. M. Variable coupling between olfactory system activity and respiration in ketamine/xylazine anesthetized rats. Journal of Neurophysiology 93, 3573-3581 (2005). https://doi.org/10.1152/jn.01320.2004

      (5) Kay, L. M. et al. Olfactory oscillations: the what, how and what for. Trends in Neurosciences 32, 207-214 (2009). https://doi.org/10.1016/j.tins.2008.11.008

      (6) Gervais, R., Buonviso, N., Martin, C. & Ravel, N. What do electrophysiological studies tell us about processing at the olfactory bulb level? Journal of physiology, Paris 101, 40-45 (2007). https://doi.org/10.1016/j.jphysparis.2007.10.006

      (7) Beshel, J., Kopell, N. & Kay, L. M. Olfactory bulb gamma oscillations are enhanced with task demands. J Neurosci 27, 8358-8365 (2007). https://doi.org/10.1523/JNEUROSCI.119907.2007

      (8) Martin, C., Beshel, J. & Kay, L. M. An olfacto-hippocampal network is dynamically involved in odor-discrimination learning. J Neurophysiol 98, 2196-2205 (2007). https://doi.org/10.1152/jn.00524.2007

      (9) Kay, L. M. Theta oscillations and sensorimotor performance. Proc Natl Acad Sci U S A 102, 3863-3868 (2005). https://doi.org/10.1073/pnas.0407920102

      (10) Lowry, C. A. & Kay, L. M. Chemical factors determine olfactory system beta oscillations in waking rats. J Neurophysiol 98, 394-404 (2007). https://doi.org/10.1152/jn.00124.2007

      (11) Vanderwolf, C. H. & Zibrowski, E. M. Pyriform cortex beta-waves: odor-specific sensitization following repeated olfactory stimulation. Brain Res 892, 301-308 (2001). https://doi.org/10.1016/s0006-8993(00)03263-7

      (12) Fontanini, A., Spano, P. & Bower, J. M. Ketamine-xylazine-induced slow (< 1.5 Hz) oscillations in the rat piriform (olfactory) cortex are functionally correlated with respiration. J Neurosci 23, 7993-8001 (2003). https://doi.org/10.1523/JNEUROSCI.23-22-07993.2003

      (13) Miura, K., Mainen, Z. F. & Uchida, N. Odor representations in olfactory cortex: distributed rate coding and decorrelated population activity. Neuron 74, 1087-1098 (2012). https://doi.org/10.1016/j.neuron.2012.04.021

      (14) Leong, A. T. et al. Long-range projections coordinate distributed brain-wide neural activity with a specific spatiotemporal profile. Proc Natl Acad Sci U S A 113, E8306-E8315 (2016). https://doi.org/10.1073/pnas.1616361113

      (15) Chan, R. W. et al. Low-frequency hippocampal–cortical activity drives brain-wide resting-state functional MRI connectivity. Proc Natl Acad Sci U S A 114, E6972-E6981 (2017). https://doi.org/10.1073/pnas.1703309114

      (16) Leong, A. T. L. et al. Optogenetic fMRI interrogation of brain-wide central vestibular pathways. Proceedings of the National Academy of Sciences 116, 10122-10129 (2019). https://doi.org/10.1073/pnas.1812453116

      (17) Leong, A. T. L., Wang, X., Wong, E. C., Dong, C. M. & Wu, E. X. Neural activity temporal pattern dictates long-range propagation targets. Neuroimage 235, 118032 (2021). https://doi.org/10.1016/j.neuroimage.2021.118032

      (18) Leong, A. T. L., Wong, E. C., Wang, X. & Wu, E. X. Hippocampus Modulates Vocalizations Responses at Early Auditory Centers. Neuroimage 270, 119943 (2023). https://doi.org/10.1016/j.neuroimage.2023.119943

      (19) Wang, X. et al. Functional MRI reveals brain-wide actions of thalamically-initiated oscillatory activities on associative memory consolidation. Nat Commun 14, 2195 (2023). https://doi.org/10.1038/s41467-023-37682-8

      (20) Xie, L. et al. Brain-wide resting-state fMRI network dynamics elicited by activation of single thalamic input. Nat Commun 16, 11247 (2025). https://doi.org/10.1038/s41467-025-66104-0

      (21) Li, A., Gong, L. & Xu, F. Brain-state-independent neural representation of peripheral stimulation in rat olfactory bulb. Proc Natl Acad Sci U S A 108, 5087-5092 (2011). https://doi.org/10.1073/pnas.1013814108

      (22) Lang, J. et al. Odor representation in the olfactory bulb under different brain states revealed by intrinsic optical signals imaging. Neuroscience 243, 54-63 (2013). https://doi.org/10.1016/j.neuroscience.2013.03.057

      (23) Chery, R., Gurden, H. & Martin, C. Anesthetic regimes modulate the temporal dynamics of local field potential in the mouse olfactory bulb. J Neurophysiol 111, 908-917 (2014). https://doi.org/10.1152/jn.00261.2013

      (24) Wachowiak, M. et al. Optical dissection of odor information processing in vivo using GCaMPs expressed in specified cell types of the olfactory bulb. J Neurosci 33, 5285-5300 (2013). https://doi.org/10.1523/JNEUROSCI.4824-12.2013

      (25) Hudetz, A. G., Vizuete, J. A. & Pillay, S. Differential effects of isoflurane on high-frequency and low-frequency gamma oscillations in the cerebral cortex and hippocampus in freely moving rats. Anesthesiology 114, 588-595 (2011). https://doi.org/10.1097/ALN.0b013e31820ad3f9

      (26) Joliot, M., Ribary, U. & Llinas, R. Human oscillatory brain activity near 40 Hz coexists with cognitive temporal binding. Proc Natl Acad Sci U S A 91, 11748-11751 (1994). https://doi.org/10.1073/pnas.91.24.11748

      (27) Murthy, V. N. & Fetz, E. E. Oscillatory activity in sensorimotor cortex of awake monkeys: synchronization of local field potentials and relation to behavior. J Neurophysiol 76, 3949-3967 (1996). https://doi.org/10.1152/jn.1996.76.6.3949

      (28) Tallon-Baudry, C., Bertrand, O., Delpuech, C. & Pernier, J. Stimulus specificity of phaselocked and non-phase-locked 40 Hz visual responses in human. J Neurosci 16, 4240-4249 (1996). https://doi.org/10.1523/JNEUROSCI.16-13-04240.1996

      (29) Murakami, M., Kashiwadani, H., Kirino, Y. & Mori, K. State-dependent sensory gating in olfactory cortex. Neuron 46, 285-296 (2005). https://doi.org/10.1016/j.neuron.2005.02.025

      (30) Wilson, D. A. & Yan, X. Sleep-like states modulate functional connectivity in the rat olfactory system. J Neurophysiol 104, 3231-3239 (2010). https://doi.org/10.1152/jn.00711.2010

      (31) Schreck, M. R. et al. State-dependent olfactory processing in freely behaving mice. Cell Rep 38, 110450 (2022). https://doi.org/10.1016/j.celrep.2022.110450

      (32) Gao, P. P., Zhang, J. W., Chan, R. W., Leong, A. T. L. & Wu, E. X. BOLD fMRI study of ultrahigh frequency encoding in the inferior colliculus. Neuroimage 114, 427-437 (2015). https://doi.org/10.1016/j.neuroimage.2015.04.007

      (33) Gao, P. P., Zhang, J. W., Fan, S. J., Sanes, D. H. & Wu, E. X. Auditory midbrain processing is differentially modulated by auditory and visual cortices: An auditory fMRI study. Neuroimage 123, 22-32 (2015). https://doi.org/10.1016/j.neuroimage.2015.08.040

      (34) Goddard, E. & Mullen, K. T. fMRI representational similarity analysis reveals graded preferences for chromatic and achromatic stimulus contrast across human visual cortex. Neuroimage 215, 116780 (2020). https://doi.org/10.1016/j.neuroimage.2020.116780

      (35) Friston, K. J., Harrison, L. & Penny, W. Dynamic causal modelling. Neuroimage 19, 12731302 (2003). https://doi.org/10.1016/s1053-8119(03)00202-7

      (36) Friston, K. J., Kahan, J., Biswal, B. & Razi, A. A DCM for resting state fMRI. Neuroimage 94, 396-407 (2014). https://doi.org/10.1016/j.neuroimage.2013.12.009

      (37) Bernal-Casas, D., Lee, H. J., Weitz, A. J. & Lee, J. H. Studying Brain Circuit Function with Dynamic Causal Modeling for Optogenetic fMRI. Neuron 93, 522-532 e525 (2017). https://doi.org/10.1016/j.neuron.2016.12.035

      (38) Zeidman, P. et al. A guide to group effective connectivity analysis, part 1: First level analysis with DCM for fMRI. Neuroimage 200, 174-190 (2019). https://doi.org/10.1016/j.neuroimage.2019.06.031

      (39) Salimi, M. et al. Disrupted connectivity in the olfactory bulb-entorhinal cortex-dorsal hippocampus circuit is associated with recognition memory deficit in Alzheimer's disease model. Sci Rep 12, 4394 (2022). https://doi.org/10.1038/s41598-022-08528-y

      (40) Chen, Y. N., Kostka, J. K., Bitzenhofer, S. H. & Hanganu-Opatz, I. L. Olfactory bulb activity shapes the development of entorhinal-hippocampal coupling and associated cognitive abilities. Curr Biol 33, 4353-4366 e4355 (2023). https://doi.org/10.1016/j.cub.2023.08.072

      (41) Pellegrino, R., Sinding, C., de Wijk, R. A. & Hummel, T. Habituation and adaptation to odors in humans. Physiol Behav 177, 13-19 (2017). https://doi.org/10.1016/j.physbeh.2017.04.006

      (42) Sinding, C. et al. New determinants of olfactory habituation. Sci Rep 7, 41047 (2017). https://doi.org/10.1038/srep41047

      (43) Zhao, F. et al. fMRI study of olfaction in the olfactory bulb and high olfactory structures of rats: Insight into their roles in habituation. Neuroimage 127, 445-455 (2016). https://doi.org/10.1016/j.neuroimage.2015.10.080

      (44) Zhao, F. et al. fMRI study of the role of glutamate NMDA receptor in the olfactory adaptation in rats: Insights into cellular and molecular mechanisms of olfactory adaptation. Neuroimage 149, 348-360 (2017). https://doi.org/10.1016/j.neuroimage.2017.01.068

      (45) Sengupta, P. The Laboratory Rat: Relating Its Age With Human's. Int J Prev Med 4, 624-630 (2013). 

      (46) Quinn, R. Comparing rat’s to human’s age: How old is my rat in people years? Nutrition 21, 775-777 (2005). https://doi.org/10.1016/j.nut.2005.04.002

      (47) Burwell, R. D. & Amaral, D. G. Perirhinal and postrhinal cortices of the rat: interconnectivity and connections with the entorhinal cortex. J Comp Neurol 391, 293-321 (1998).

      (48) Kajiwara, R., Takashima, I., Mimura, Y., Witter, M. P. & Iijima, T. Amygdala input promotes spread of excitatory neural activity from perirhinal cortex to the entorhinal-hippocampal circuit. J Neurophysiol 89, 2176-2184 (2003). https://doi.org/10.1152/jn.01033.2002

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPES.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also present a clear limitation and challenge for interpretation. As a reader, even those who are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain an understanding of how they all fit together.

      Indeed, it seems like there are many threads running together in the paper, which makes it challenging to find the through-line of the key findings, or to understand how they might relate to some pre-existing hypotheses, rather than merely interesting patterns detected in the data. In the Introduction and Discussion, it seems as if the key question is to understand the pathways by which MPEs impact cognition, but this is a rather broad topic, so it is not clear exactly what the authors are aiming at with this question and study design.

      As an example, authors operationalize frontal theta power as an index of cognitive control demand, and one of the pathways by which MPEs impact cognition. But this point becomes somewhat circular, since it is not clear how or why the Mismatch x Strength interaction in frontal theta reflects that demand. It would have been better to set this pattern up in the Introduction as a theoretically driven hypothesis, since it currently appears more like a post-hoc interpretation. This is mirrored by how the issue is first brought up in the Introduction, where it states somewhat vaguely: "whether MPEs are followed by an increase in frontal theta... warrants closer examination".

      Again, we appreciate the reviewer’s thoughtful feedback on where the manuscript can be clearer, especially given the rich set of results it reports. Following the reviewer’s guidance, we restructured and revised the Introduction to further motivate the hypotheses that (a) MPEs increase both attention/arousal (grounded in studies of Event Segmentation Theory) and cognitive control (given findings on reward prediction errors and other types of prediction errors), and (b) there are greater increases in these processes triggered by strong compared to weak MPEs. To better link these hypotheses to resulting statistical tests, we note that hypothesis (a) was tested in our trial-level regression models in Fig. 1 by examining main effects of Strength, whereas hypothesis (b) was tested in our models in the Mismatch x Strength interactions. On point (b) and potentially circularity, we note that previous work indicates that frontal theta scales with negative reward prediction errors; as such, we hypothesized that stronger MPEs would elicit more frontal theta, as evidenced by a robust Mismatch x Strength interaction during the probe period.

      Later in the results, there are findings relating frontal theta to pupil dilation, posterior alpha suppression and then subsequent memory. It was hard to understand how all the findings might be linked together functionally or conceptually. Are the authors potentially postulating a mediating or mechanistic pathway, in which the MPE leads to increased cognitive control (frontal theta), which then leads to enhanced subsequent memory of those events? If this is the case, then maybe a formal path analysis would be the best way to test or state this hypothesis. It would also be useful to specify more clearly how the pupil components and alpha suppression factor into this mediating path, since it was not clear.

      Relatedly, the authors suggest that internal attention and arousal also play relevant roles in this pathway, but these are also not clear. In some cases, it is stated as if this is a distinct pathway from the cognitive control one, since there is a focus in the results on the independence of frontal theta and posterior alpha, but elsewhere they seem to be treated as two aspects, or distinct steps, within a single pathway. Again, these different threads of the findings were quite challenging for the reader to follow. Pathway analyses, such as with multiple mediation or moderated mediation, could be a useful way to address this question. For example, it seems as if readiness-to-remember is another behavioral outcome (like subsequent memory) that could be used in the search for mediators.

      We thank the reviewer for highlighting these ambiguities in the original submission and for the thoughtful encouragement to leverage mediation models to more formally test the hypothesized relationships between MPE magnitude and constructs of control, attention, and arousal. In the revision, we now more clearly hypothesize that the effects of strong MPE-driven increases in attention and arousal might be explained, in part, by cognitive control (as indexed by frontal theta) upregulating attention and arousal. To more explicitly test this model of the relationships between our measures, as recommended by Reviewer #1, we now include multivariate mediation analyses to assess whether, at a trial-level, changes in posterior alpha and the immediate pupil effect PC3 are explained in part by increases in frontal theta. Because changes in posterior alpha following MPEs and PC3 scores did not predict subsequent memory, mediation analyses addressing the hypothesis that our attention/arousal measures mediate the effect of frontal theta on subsequent memory were not conducted. Following insights from a cross-correlation analysis, as recommended by Reviewer #2, we tested an additional model to examine whether the effects of MPE magnitude on frontal theta were explained in part by changes in posterior alpha. Examination of the posterior distributions of the indirect effects did not favor our hypothesized model, nor the alternative model that attention upregulates control. Altogether, these outcomes suggest that increases in control, attention, and arousal following strong MPEs may be elicited independently. Yet, we also note that the current set of experiments may be underpowered for these mediation analyses; future work can further investigate the directionality between these effects. Finally, with respect to the readiness-to-remember findings, they were removed in the interest of space, as recommended by Reviewer #2.

      At the minimum, it would be quite helpful to have diagrammatic figures that specify the hypothesized and observed relationships between independent variables (Strength, Mismatch), physiological indices (pupil dilation components, frontal theta, posterior alpha) and key outcome measures (accuracy, RT, next-trial retrieval success, subsequent memory), so that the reader can refer back to them as each component of the analyses is conducted.

      To further increase conceptual clarity, we also followed this helpful suggestion, adding diagrammatic figures to illustrate our hypotheses regarding interactions between the effects (Fig. 3a) and to summarize the observed relationships between MPEs, control, attention, and arousal (Fig. 6).

      Minor Points:

      Many figures had x-axes showing a pupil component or EEG power metric broken down by quartile or quintile. Yet nowhere is it ever explained why this graphical (or analytic?) approach is used and what it reflects, or how it is decided which break down to use (quartile/quintile). If the data are analyzed as a correlation, why is a scatterplot not shown instead?

      In the linear mixed effects models, continuous values were used to assess relationships between variables. In the figures, the continuous variables were binned into quartiles or quintiles for ease of visualization. We opted to visualize the data using this approach, rather than with a scatterplot, given the large number of trials. We updated the figure captions to clarify the approach.

      It was surprising that, unlike readiness-to-remember, which was analyzed via logistic regression and odds-ratio, subsequent memory was not analyzed in the same fashion (i.e., as a binary outcome variable predicted by frontal theta), rather than in a reverse chronological one (subsequent memory predicting frontal theta). Historically, it was the case that subsequent memory was analyzed in this manner, but that was before the era in which trial-level linear mixed-effect models were in wide usage, as they are implemented in this study. Thus, the choice seems like a wasted opportunity or a step backwards analytically.

      We thank the reviewer for this encouragement and agree with the point. In the revision, we note that the readiness-to-remember results were removed in the interest of space and clarity, as was recommended by Reviewer #2. With respect to the subsequent memory analyses, they are now analyzed via logistic regression.

      Reviewer #2 (Public review):

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The methods are rigorous, and the claims are mostly supported by the data, but there are a few weaknesses or places that could be improved:

      (1) The authors conduct PCA analysis to identify different components of the pupillary response to MPE and relate them to behavior. Specifically, the authors identify components PC3 and PC4, which they interpret as related to MPE. However, some parts of the interpretation could be clearer or better justified:

      (a) The authors refer to PC4 as "post-decision cognitive processing". But, given that RT was between .5-.7s, and PC3 peaked after more than 1s, wouldn't it be cautious to interpret PC3 as postdecision as well?

      Thank you for raising this point. Given that pupil is a relatively sluggish response, it is possible that both components reflect post-decision cognitive processing, even if PC3 peaks before PC4. Following the reviewer’s guidance to adopt more cautious language, we replaced “post-decision” with “post-MPE”.

      (b) MPEs overall elicit longer RTs in this study, suggesting that long RT is a behavioral marker of MPE. Nonetheless, the authors argue on p. 12: "Altogether, these findings indicate that when stronger mnemonic predictions (as indexed by shorter RTs) were violated." And, PC3 is correlated with shorter RTs for mismatches, meaning that behaviorally, these trials were more similar to matches. Thus, how do the authors interpret shorter versus longer RTs for MPEs, and what processes do these RT reflect?

      We thank the reviewer for stressing the need for greater clarity regarding the relationships between RT and the constructs of interest. With respect to RT, we interpret the condition-level difference in RTs between mismatch and match trials as a behavioral marker of an MPE. However, when comparing mismatch trials within a given strength condition to each other (i.e., an analysis at the trial-level), shorter RTs may reflect a stronger prediction, greater certainty that the probe is a mismatch, and therefore the experience of a stronger MPE. Note that while larger PC3 scores were associated with shorter RTs for mismatches (Fig. 2b, right), the mismatch RTs in the largest PC3 quartile were still longer than those on match trials in the corresponding quartile (in other words, there was still a condition-level difference in RTs in the largest PC3 quartile, suggesting that the mismatch trials in this bin are behaviorally still likely to be different from match trials in the corresponding bin).

      To clarify these relationships and our interpretation, we modified the referred to text: “(as indexed by shorter strong mismatch RTs).” Moreover, we added text to the Discussion, further delineating our reasoning here and the implications of our findings for understanding the mechanisms giving rise to and triggering by MPEs of varying strengths. This includes adding an explicit summary of the logic and findings that notes that, at the trial-level, shorter RTs may reflect stronger predictions; at the match/mismatch condition-level, longer RTs may reflect the experience of a MPE. For PC3, shorter RTs (trials with a stronger prediction) in the Strong Mismatch condition were associated with a larger pupillary response. The added text notes that “while longer mean RTs for mismatches compared to matches are a behavioral marker of a MPE, within-condition differences in RTs (i.e., between mismatch trials) may reflect more subtle differences in MPE magnitude, with shorter RTs reflecting stronger predictions and thus stronger MPEs. Strong mismatch RTs were used in the mediation models as a proxy measure of MPE magnitude; this estimate may be noisy because RTs in this experiment are likely sensitive to factors independent of mnemonic prediction strength (e.g., preparatory attention (Supplementary Fig. 6) or memory strength of the mismatch probe). This limitation may have additionally reduced sensitivity to detecting indirect effects.”

      (2) The brain to pupil relationship (p. 13-14): If I understand correctly, this was done on a trial-by-trial basis, but the high temporal resolution allows doing the analysis in a time-resolved manner - does brain activity at a certain time point preceding/following the pupil response correlate with the pupil response? It might be that cognitive control influences attention mechanisms or vice versa (because there is some overlap in the response). Although not testing causality, this temporally resolved correlation would be an interesting way to start probing how signals might influence each other.

      Thank you for this suggestion. We now report a cross-correlation analysis (Fig. 3d) that suggests that cognitive control increases precede attention decreases at retrieval, whereas in response to a strong MPE, cognitive control increases follow attention increases. There were no significant clusters for the temporal relationships between frontal theta and pupil, nor for posterior alpha and pupil.

      (3) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. However, are they specific to this task? I think it would benefit the paper to show that these relationships are potentially specific to making match/mismatch memory decisions, versus, e.g., any stimulus processing. For example, the authors could run the same analyses locked to stimuli in the study phase, anticipating a different pattern, if indeed these findings are specific to the associative memory task.

      Many of the associations between our measures indeed did not show an interaction with Mismatch. Due to jitter in the ISI in the study phase, some of the analyses in the retrieval phase cannot be performed in the exact same way for the study phase. We will leave these questions to be addressed in future research. We agree that an important question for future research is to address whether these responses depend on making match/mismatch decisions and now include consideration of this point in the Discussion: “Finally, the magnitude of observed increases in control, attention, and arousal following strong MPEs may be influenced by the decision-making process engaged when making match/mismatch judgments. Not all MPEs necessitate behavioral responses. Whether similar magnitudes in neurocognitive responses and consequences for learning are observed upon detection of an MPE, but in the absence of a decision remains unclear.”

      (4) During memory retrieval (i.e., before the probe), the authors find that frontal theta, a marker of cognitive control, was associated on a trial-by-trial basis with more posterior alpha (i.e., less alpha suppression, potentially reflecting less attention), and that this association was stronger for weaker predictions. The authors interpreted this as weaker predictions necessitating more cognitive control, and that more cognitive control was recruited specifically in trials where retrieval included less content (memory reinstatement) to attend to. Generally, cognitive control is recruited to facilitate memory retrieval. If so, one possible interpretation is that this correlation reflects cognitive control effort that has failed to produce enough memory reinstatement. The other possibility is that this correlation reflects more specific retrieval of the correct probe, without retrieval of interfering items (i.e., overall less content). I believe that the former explanation predicts that this correlation would be associated with longer RTs (more difficult decisions), while the latter predicts shorter RTs (easier decisions due to successful retrieval), at least for matches.

      Thank you for these insightful comments. Because this analysis is not key to the main questions about MPEs and given both reviewers’ concerns that the manuscript can be overwhelming for the reader given the sheer number of findings reported, we opted to move this point from the main text to the Supplement. However, following the reviewer’s guidance here, we conducted the proposed analyses and tested these alternative accounting by modeling RTs. The results favour the former interpretation:

      “Greater control being associated with less attention could reflect failure in controlled retrieval efforts to reinstate sufficient memory evidence of the probe and thus fewer retrieval products to which attention is allocated. Alternatively, greater cognitive control could increase the likelihood of retrieval success, eliciting selective retrieval of the correct probe and inhibition of interfering items. To address these alternatives, we examined how the association between frontal theta and posterior alpha related to the difficulty of a trial, as assayed by RTs. The former failure-of-control account would predict that a stronger positive frontal theta- posterior alpha association would relate to longer RTs, whereas the greater-retrieval-specificity account would predict that a stronger positive association would relate to shorter RTs. In a model predicting RTs as a function of frontal theta, posterior alpha, Strength, Mismatch, and their interactions, we found a two-way frontal theta × posterior alpha interaction (β=0.016, CI=[0.001, 0.031], p=0.033), such that a stronger positive association between frontal theta and posterior alpha predicted longer RTs. This relationship did not differ as a function of Strength (no frontal theta × posterior alpha × Strength interaction: β=-0.015, CI=[-0.033, 0.004], p=0.124), Mismatch (no frontal theta × posterior alpha × Mismatch interaction: β=-0.016, CI=[-0.050, 0.018], p=0.342), or interact with Strength and Mismatch (no frontal theta × posterior alpha × Strength × Mismatch interaction: β=0.024, CI=[-0.014, 0.062], p=0.218). Together, these outcomes support the idea that during memory retrieval, positive coupling between frontal theta and posterior alpha may reflect failure or inefficiency of cognitive control efforts to rapidly accumulate mnemonic evidence to which to attend in support of a memory decision.”

      (5) In section 3, the authors found a positive relationship between alpha during memory retrieval and PC3 during MPE. If I understood correctly, this means that less attention during retrieval (less suppression) is correlated with a stronger PC3 response. How do the authors interpret this? Maybe along the same lines as in (5), specifically retrieving the correct information (i.e., less retrieved content to attend to) means a stronger prediction, leading to a stronger MPE, and a stronger MPE response, as reflected by PC3?

      We appreciate this comment, as it highlights a need for greater clarity here. The observed relationship was actually negative (Fig. 4f), meaning that more attention during retrieval was associated with a stronger PC3 response. We suspect that the lack of clarity here may be due to the original statement that “there was a positive relationship between posterior alpha suppression during memory retrieval [and PC3 scores]”. To increase clarity, we have modified this statement to “there was a negative relationship between posterior alpha during memory retrieval [and PC3 scores]”. We interpret the negative relationship with posterior alpha (i.e., positive relationship with posterior alpha suppression) to indicate “that greater attentional allocation during memory retrieval, which occurs when memories are stronger and more retrieval products can be reinstated and attended to (Fig. 1e; Fig. 4c), predicts the magnitude of immediate pupil responses to MPEs.”

      (6) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it's not a fair comparison, and could change participants' overall criterion for old/new decisions. However, one possibility would have been to test only the weak prediction; this could have given some specificity to the neural subsequent memory findings.

      We were indeed concerned about the change in decision criterion and did not include match items for this reason. Nonetheless, this is a useful suggestion that future work could include the weak match items for a better comparison of subsequent recognition memory. We now comment on this in the Discussion: “Whether control-associated enhancements in learning are specific to learning from MPEs can be further tested in future work by including a test of subsequent memory for weak match probes as an additional control condition for comparison.”

      (7) The authors nicely characterized the different PC of pupillary MPE response. But, with respect to subsequent memory, they only present pupil size. Unless there is some methodological reason that prevents testing subsequent memory on the PC, I think this will be very informative about the potential mechanisms underlying memory of MPE.

      The pupil PCs were not associated with subsequent memory, though there are some interesting trends in the 48-delay condition which could be explored in future work. These findings are now reported in the Supplement (Supplementary Fig. 16).

      (8) This paper includes many interesting findings, and I am not sure how they all come together into a cohesive mechanistic understanding of MPE response and subsequent memory. I think the paper would benefit from either a conceptual mechanism figure or, in the Discussion, have a summary of a proposed mechanism integrating the findings together.

      We thank the reviewer for stressing this point, which also was raised by Reviewer #1. To better emphasize the novel contributions of the work and to assist the reader’s understanding of the key findings, we: (a) revised the Introduction to more explicitly describe our hypotheses about the relationships among our measures as they relate to MPE responses and subsequent memory; (b) now report mediation models to more directly test these relationships and include diagrams of the hypothesized relationships; and (c) include a schematic figure at the end of the Results to highlight the key mechanistic relationships supported by the data.

      (9) Relatedly, the section "Immediate, strength-sensitive neurocognitive impacts of MPEs" does not link the arguments to specific data points, so it's hard to follow which data specifically the authors are interpreting.

      The discussion in this paragraph rests on the Strength × Mismatch interactions observed in Figures 1d-f, summarized in the first sentence of the section. To increase clarity, we changed the title of this section to “Neurocognitive impacts of MPEs are strength-sensitive”.

      (10) If I understand correctly, the authors did not find improved memory for strong compared to weak MPE. First, I think this behavioral result should be incorporated in the main paper and in the interpretation of the results. Second, given that the neural effects the authors tested either correlated with memory for strong MPE or did not show a relationship with memory, what neural/pupil response could explain memory for weak MPE?

      Thank you for raising these points. The behavioral result and the frontal theta trial-level regression model for subsequent memory are now described in the main results and included in the Discussion. As noted in the Discussion, the trial-level regression model indicates pre-probe and post-MPE frontal theta effects may explain memory for weak (and strong) mismatch probes. While the magnitude of probe-period MPE frontal theta is additionally predictive of memory for strong mismatch probes (as indicated by the Subsequent Memory × Strength interaction in the trial-level regression model in Fig. 5c and the logistic regression in Fig. 5b), frontal theta during this period does not additionally enhance memory for strong mismatch probes above that of weak mismatch probes (Fig. 5a). We added more discussion on why memory for strong and weak mismatch probes in the current experiment did not differ. We also discuss potential directions for future research to further probe mechanisms underlying MPE-driven learning.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is recommended that the authors determine whether formal path analyses, testing for mediation and moderation, would provide a useful approach from which to better integrate the disparate set of findings and make clear their causal/functional implications.

      At a minimum, adding a diagrammatic figure is recommended to visually depict the key components of the study (independent variables, physiological indices, outcome measures) and how they relate to each other both conceptually (ideally in a theoretically hypothesized manner) and in terms of the observed findings. Such a figure will help the reader keep track of the many types of findings and results threads, and with the goal of better organizing the results into a clearer narrative through-line.

      Thank you for the suggestions to add mediation analyses and diagrammatic figures. To more formally address our hypothesis that increases in attention/arousal following MPEs are explained in part by increases in cognitive control, we tested two mediation models: one with posterior alpha as the measure of attention and the other with pupil PC3 scores. We now diagram this hypothesis in Fig. 3a. Given the outcomes of the cross-correlation analysis suggested by Reviewer #2, we also tested an additional model where posterior alpha might explain the impact of strong MPEs on frontal theta (now diagrammed in Fig. 3e). We did not find credible evidence for an indirect effect in any of the models, suggesting that strong MPE-driven increases in control, attention, and arousal may be elicited independently. Finally, in a newly added summary diagram (Fig. 6), we highlight the observed effects of strong vs. weak MPEs on RTs, control, attention and arousal; differences in control, attention, and arousal for strong vs. weak predictions preceding the MPE; trial-level mismatch-specific or strong-specific effects on subsequent memory; and mismatch-specific interactions between processes during retrieval and responses to MPEs. We hope these revisions address Reviewer #1’s concerns and that these figures help the reader keep track of the key hypothesized relationships and main findings.

      Reviewer #2 (Recommendations for the authors):

      (1) The relationship between event segmentation and prediction errors has been reviewed recently in two papers (Nolden et al., 2024, Neuroscience & Biobehavioral Reviews; Rouhani et al., 2024, JOCN for a potentially relevant computational model). I wonder if insights from these papers can inform the Introduction/Discussion of the current manuscript.

      Thank you for these suggestions. These papers are now incorporated into the Introduction and the Discussion.

      (2) Brod et al. (2022, Psych. Bull. Rev.) have previously reported increased pupil dilations for MPE correlating with subsequent memory, specifically for strong prediction errors. I think it's worth including this paper in the Introduction as the finding is highly relevant. The Brod paper might provide more direct evidence of "MPE-related increases in pupil size" than the evidence the authors provide (p. 3).

      Thank you for this suggestion. This paper is now incorporated into the Introduction and Discussion.

      (3) I'm confused about the temporal analysis: "Trial-level regression analyses were conducted on the frontal theta, posterior alpha, and pupil time series from the associative retrieval test to identify temporal clusters that were sensitive to the factors of Mismatch (i.e., mismatch vs. match probes) and/or Strength (i.e., strong vs. weak associative pairs). For each participant and each time point, a linear regression model testing main effects of Mismatch and Strength, and a Mismatch × Strength interaction was run using R to compute beta weights for each regressor." What regression exactly was run? A separate model for each participant and time point? Across trials, then? Later, the authors mention that beta weights were averaged across participants and t-tests and permutation tests were conducted. However, if the data were averaged, what t-test was conducted? And how was the permutation test conducted? It's also unclear what the authors mean by "the sign of each participant's beta weights" - what sign?

      Thank you for raising this point. A regression was run separately for each participant and each time point, across trials: neurocognitive measure ~ Mismatch + Strength + Mismatch: Strength. Beta weights were averaged only for visualization; t-tests were conducted on each set of beta weights (across participants), separately for each time point. We modified the text of the Methods to describe our procedure more clearly.

      (4) The associative memory and recognition accuracy data are presented as d'. In addition, the authors should provide hits and false alarms to facilitate a better interpretation of the results.

      We now report these outcomes in Table 1 and refer to them in the main text.

      (5) The authors argue regarding the frontal theta that "Qualitatively, the main effect of Strength emerged later than the main effect of Mismatch, suggesting that the increase in frontal theta evoked by the probe was more sustained for weak compared to strong trials (or, as a corollary, that the greater control elicited by strong MPEs enabled more rapid resolution of conflict and ultimate choice selection)." It was unclear to me how that stems from the data.

      This statement has been removed altogether.

      (6) Especially in Figure 1, I think clarity can be improved if the authors would indicate the specific subsection they are referring to in the text, because even within, e.g., 1d, there are different graphs, so mentioning which graph is relevant for which statement would be helpful to the reader.

      Thank you for this suggestion for improving the clarity of the manuscript. Fig. 1d-f includes multiple parts because we wanted to show the raw data (the mean time series on the left) as well as the model coefficients (time series on the right). We are hopeful that, given that each subplot is titled and that the main text refers to both the subplots on the left and on the right, this will be clear to the reader as is. We welcome further guidance if this remains a concern.

      (7) This seems highly speculative to me: "Elevated PC4 scores on weak hits, where no error was made, may instead reflect retrieval practice and thus internally oriented attention. On these trials, when cueelicited retrieval may have been weaker, the probe may have provided additional support for pattern completion of the learning episode for that association, and this engagement in memory retrieval may have elicited pupil dilation (c.f., Strength effect in Figure 1f). Overall, PC4 may therefore reflect an attentional orienting response that, depending on the relative success of memory retrieval and probe identity, may direct attention internally or externally. " (p.13). In my opinion, the authors make a lot of assumptions about underlying processes. I'd consider removing.

      We removed this text.

      (8) In Figure 3c, should the x-axis be PC3 quintile? And (c) is not in the figure caption.

      Thank you for this note. We addressed these issues.

      (9) The carryover effects the authors report are interesting, but they seem detached, and it is unclear how they fit with the additional findings. In a paper that already includes many findings, I'd recommend either integrating better or removing.

      Following this guidance, we removed the readiness-to-remember findings in the interest of space and clarity.

    1. Author response:

      We thank the reviewers for their time and insightful comments. We are also grateful for their appreciation of the unusual circumstances that led to the publication of the manuscript in its current form.

      In response to reviewer #2’s question about replication, we provide additional details here. The error in our retracted original publication affected only the imaging results. Despite this, we reproduced key behavioural experiments by generating additional datasets (rather than simply rechecking records and authenticating results) and therefore have full confidence in our behavioural findings. We are very happy to share some of these replication experiments below:

      Author response image 1.

      Data showing replication of key experiments. From L-R these data replicate those shown in Figure 2e, Figure 4h, Figure 4c, Figure 4d.

      The conceptual framework of the original study remains valid. It was actually formulated based on the behavioural data and before any physiological recordings were made. We believe that it still represents the most parsimonious explanation for the observed behavioural results. Multiple behavioural findings support a model in which multisensory training leads to the recruitment of visual γd Kenyon cells into an otherwise olfactory memory trace. These include: (1) the requirement for γd KC output during olfactory retrieval following multisensory training; and (2) the sequential learning experiments, which were originally designed to test this model and provide independent evidence for its predictions. We nevertheless agree that the loss of the physiological data reduces the amount of evidence supporting the proposed mechanism. The nature of the physiological changes following multisensory learning remains an important question that we intend to address in future work.

      We also thank the reviewers for identifying the unfortunate typo in the Abstract, which we believe contributed to the confusion over the dopamine receptors tested in this circuit. We selected these receptors based on our in-house single-cell transcriptomic expression data, together with published evidence indicating their specific expression in the neurons of interest. We have also responded to reviewer comments about the clarity of the figures and added additional labels to figure 2, to clarify the experimental paradigm in each case.

    1. Author response:

      We greatly thank the Reviewing Editor, Senior Editor and the reviewers for their constructive and thoughtful feedback, as well as for recognizing the significance and strengths of our study. We are very encouraged by overall positive assessment and appreciate these insights to strengthen the manuscript. Below, we outline our plans to address the key points raised in reviewer#1’s public review. We also note that reviewer#2 did not identify any weaknesses in the study and are thankful for this positive evaluation.

      Point 1: Relating in vivo and in vitro findings and discussing the relevance of myelin clearance in AD. As suggested, we will elaborate our discussion to cover the relationship between our in vivo and in vitro findings and present a more unified picture of GPR34 function. We will also highlight the relevance of myelin clearance in Alzheimer’s disease independent of amyloid plaque burden.

      Point 2: in vivo assessments and phenotypes. While we recognize the value of expanding our in vitro and cellular findings to in vivo and cognitive measures, we believe this additional assessment is beyond the scope of this current study. At least, we will expand our discussion to relate our findings of GPR34-medilated myelin pathology in the contexts of neurodegeneration and cognitive deficiency in AD.

      Point 3: Myelin engulfment versus lysosomal degradation. We agree that the clear distinction between impaired myelin engulfment and altered lysosomal degradation is an important point. Although additional experiments to address these scenarios are beyond the scope of the study, we will perform targeted pathway analysis of existing transcriptomic datasets focusing on the phagosome and lysosomal degradation pathways along with the in vivo findings related to CD68, which could offer an interpretation.

      Point 4: Comparison of transcriptional signatures. As suggested, we will compare the transcriptional signatures induced by myelin exposure in WT iMGLs with our 5xFAD RNA-seq datasets, as well as publicly available datasets from human Alzheimer’s disease microglia, to unravel the potential converged and distinct features.

      Point 5: Reconciling our findings with previous GPR34 studies. We will expand our discussion to compare and clarify our current findings with the previous studies and cover potential reasons for those different observations. We will also cover the current limitations of existing GPR34 pharmacological tools compounds, including their limited blood-brain permeability, which precludes the effective use of those tools for proposed in vivo studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers, visual learning ability has been modest and difficult to study. To put a number on it, in this paper, animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild-type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others, i.e., some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's a description of more variation within the set of 64 for each strain, two different parent sets per strain, different sexes, conditioned and unconditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day), or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e., the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which a stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding, and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here, and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm, allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

      Following our provisional response to the reviewers, we have implemented these additions to the manuscript:

      (1) We have added 7 additional tables (Table 1-6 and table 8) to the supplementary that detail how many individual flies were used in which of the seven separate experiments, how the individuals were distributed across genotypes, replicates and sexes, and how many were filtered out before the final analysis. At the bottom of each table, we added a short description of the type of experiment and a brief explanation of the filtering. The four smaller experiments where we tested the two mutant lines and one wild-type DGRP line were used primarily to test and validate the experimental platform and the behavioural paradigm used for the main experiment. In these four experiments we tested green place learning, blue place learning, green colour learning, blue colour learning in four separate batches of flies. In each of these experiments we used 192 individuals (64 individuals x 3 genotypes x 4 experiments = 768 individuals in total). In the multiday experiment we used 64 flies per genotype per each of the four groups of sequences of learning paradigms, in total 512 individual flies. Here, each of these 512 individuals were retested in different learning paradigms over 4 days (Table 5). The main experiment was the green place learning (Table 6) where we tested all 88 DGRP lines and again the two mutant lines was used to obtain the majority of the main results and conclusions (64 individuals x 90 genotypes = 5760 individuals). Lastly, additional 896 individuals were measured in the experiment using blue place learning paradigm to test the consistency of learning behaviour as opposed to colour bias within genotype (Table 8). In summary, in all experiments, we have always measured behaviour in 64 flies per genotype (full loading of the behavioural platform), and they were distributed almost entirely evenly across replicates, sexes, and conditions (control vs conditioned). No individual was reused across experiments. In most cases, after filtering the data, 60 or fewer individuals were used in final analyses. For the very few deviations from this experimental design (which occurred due to unforeseen events such as dropped/sick vials, flies flying away or accidentally squished during setup, skewed number of males and females etc.) we added a short explanation in the text below the tables. In total, across all reported experiments in this study, we measured behaviour in 7936 individuals.

      (2) We have added a schematic visual representation of classical measurement of individuality (variance of the distribution of behaviour within genotype where genetically identical individuals are raised in the same environment), entropy-based measurement of individuality (residual individuality) and the change in residual individuality, as we use them in this study (Figure 2D). We also provide a list of different DH<sub>resid</sub> measures and what distributions are being compared across the DH<sub>resid</sub> in the same figure. We hope this will serve as a more intuitive explanation of individuality and help readers interpret and follow more easily the results that we report after this figure.

      (3) In the same vein, we added another schematic visual representation to Figure 3 (Figure 3F) where we depict how distributions of individual behaviour may change with every decision and how this change translates to (or can be read out from) the change in residual individuality. We have also renamed the X axis of Figure 3E to “DH<sub>resid</sub> Start”, so that it is clearer what is measured here and matches the explanation in Figure 2D.

      (4) We have added two additional supplementary figures where the reader can inspect in more detail how the distributions of individual behaviour change longitudinally across time for each genotype in control and conditioned (Figure 3 – Figure supplement 2 and  Figure 3 – Figure supplement 3). From these figures one can glance how variance as well as the shapes of the distributions change as the flies learn in the conditioned setting, and how they remain largely the same in the control where flies behave spontaneously. We have added a sentence in the main text to introduce these figures: “We found that the distributions of individual behaviour were broader and their shapes changed substantially over the course of the experiment for the conditioned flies, and not for the control flies Figure 3 - figure supplement 2, Figure 3 - figure supplement 3).”

      (5) As noted in the first provisional response to reviewers, we changed the sentence “In every individual, behaviour is shaped by deterministic, genetic factors and by environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.” to “In every individual, behaviour is shaped by fixed genetic factors and by variable environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.”

      (6) Some sentences were edited in the results so that we can correctly refer to the newly added tables and figures. Context, meaning or interpretation of the results in these sentences was not altered.

      (7) While adding the new table references to the text, we noticed a typo that propagated in the previous version where the number of flies used in the main experiment was stated to be N= 5238, when in fact it should have been N=5239. This is now fixed.

      We once again thank the reviewers for their comments and suggestions – we believe their suggestions helped us improve the presentation and interpretability of our study and we hope the reviewers and readers will agree with this as well.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed olo system. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having fewer cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The capture of a rich dataset, including cone imaging and eye motion, is valuable. Benchmarking with the prior literature, suggesting that good visual acuity can be maintained despite a 50% loss in cone density, is impressive. However, it is known that cone density varies dramatically from the peak cone density location in the foveal center to even a location a few degrees outside of the fovea. In addition, there is a high degree of subject-to-subject variation in peak cone density. Given that the C stimulus is hollow in the middle, the stimulus does not actually hit the location of the peak cone density but must land slightly outside of it. Therefore, considering the actual cone density of where the stimulus lands will be important to discuss and/or analyze.

      The reviewer is correct that the cone density will vary dramatically with distance from the foveal center. However, importantly, in our experiment the Landolt C stimulus is fixed in the world and the eye is free to move across it. Therefore, the subject can direct their gaze to any part of the letter, rather than it being fixed to the hollow center of the letter. In the worst case, if the subject kept their gaze fixed at the center of the letter and were viewing the largest letter corresponding to the worst acuity measured in our experiments (20/100), the cone density on average would be 13% lower at the edge of the letter than at the center. However, it is unlikely that a subject would have fixated in this manner, and the vast majority of letters shown during the experiments were much smaller than 20/100.

      For completeness, we have calculated the peak cone densities for each of our subjects using the cone density centroid method described by Reiniger et al (2021). We have added these numbers and a description of the method to the Subjects section in Methods and Materials on lines 339-344.

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to addressing this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e., that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that, in a patient's eye, with cone dropout, there will be gaps in the mosaic. Not considering any other non-photoreceptor-related reasons for visual acuity loss, which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery, which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      Cone loss does manifest differently in different retinal degenerative diseases, and in this work we implement dropout on a cone-by-cone level. To address the reviewer’s points, we have added a description of the range of spatial manifestations of cone loss across a range of diseases to the Discussion section on lines 320-328, and emphasize that we focus on one particular manifestation in this paper.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      We thank the reviewer for their helpful comments and constructive feedback.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The patient recruitment limitation seems to be a bit artificial here. This is indeed a limitation, but perhaps not the primary limitation or motivating factor. Consider removing/rephrasing this motivation.

      We agree with the reviewer’s comment, and have removed it from the abstract and removed its framing as a limitation in the Introduction section in two instances on lines 31 and 36-37.

      Was the peak cone density quantified, and the location of the peak cone density determined? Reporting the range of eccentricities over which the C stimulus lands relative to the peak cone density location, as well as the actual cone density that is being used to sample the C stimulus on the retina, seems to be important for contextualizing this study. It is a bit too simple to only consider the percentage of cones that are reduced.

      We have added the peak cone density for each subject to the Subjects section in Methods and Materials. This experiment did not require fixation; rather, the Landolt C stimulus was fixed in space and the subject could move their eye freely across it, meaning that different parts of the fovea may have sampled the letter on different trials. At the highest dropout percentage, where acuity was the worst, the letter size was 20/100 at threshold, or 25 arcmin. If the subject were to fixate with their peak cone density at the center of the letter, we have computed that the average decrease in cone density at the edge of the letter (12.5 arcmin away) would be 13%. The majority of trials in the experiment showed letters that were much smaller than this, and would have been subject to even less variation in cone density.

      Do the authors have any idea about the approximate size of the cones in healthy subjects compared to diseased eyes? Importantly, if the cones in patients are larger due to the dropout of their neighbors, then the retinal coverage area would be larger due to their larger size, and the amount of light that can be coupled into larger cones may also be larger. Can this be modeled or discussed?

      In our implementation, we did not emulate a change in cone size, and rather modeled the loss as discrete holes in an otherwise intact retina. We have added text to the Discussion section on lines 320-328 to make the distinction between this form of cone loss and other forms where cones appear to fill in for their neighbors resulting in a contiguous mosaic of lower density overall.

      Acknowledging some of the shortcomings of this approach for simulating the patient condition could be improved. It may be worthwhile to tone down the premise of this paper if these cannot be adequately explained.

      In order to tone down the premise of the paper, we have made the following changes to the text.

      We now emphasize on lines 65-69 in the Introduction section that we focus specifically on the impact of cone loss on acuity without modeling downstream factors.

      In addition to the description of other diseases that we added in response to a previous comment, we have also added the following text on lines 313-318 of the Discussion section:

      “... factors beyond the photoreceptors also play a role in shaping vision under retinal degeneration. In this work, we did not model any downstream factors such as shorter outer segments (Foote 2018), retinal rewiring (Jones 2016, Lee 2021), or ganglion cell hyperactivity (Kramer 2023). Instead, we sought to characterize vision in the presence of cone loss at the lowest possible level, considering only the decrease in sampling power at the retinal input.”

      In the methods, it is not completely clear the rationale for determining the appropriate size of the C stimulus. How is visual acuity determined if the C stimulus size is not changed?

      A more careful explanation of how the C stimulus size is set is warranted.

      In the experiments measuring visual acuity, the C stimulus size was selected by a QUEST staircase on each trial. For each dropout condition, we ran 4 interleaved QUEST staircases with 20 trials each. This is described in the “Acuity Threshold Experiment” section in the main text (lines 99-100) and in Materials and Methods (line 407). To clarify further, we have updated the following sentence on line 423:

      “For each condition, we ran 4 interleaved QUEST staircase procedures (Watson and Pelli, 1983) with 20 trials per staircase, which varied the size of the Landolt C on each trial.”

      What is the clinical visual acuity of the subjects being tested? It seems important to report this if the authors want to use their C stimulus as a proxy for clinical visual acuity.

      The subjects being tested have excellent acuity. In Figure 1, we can see that their adaptive-optics-corrected acuity for the baseline 0% dropout condition ranges from approximately 20/10 to 20/12.5 across the 4 subjects. We have added the following statement to the “Subjects” section in Materials and Methods (line 339):

      “All subjects self-reported to have normal vision.”

      Given that the title of the paper emphasizes the role of eye motion, it seems that a more careful analysis of the magnitude and type(s) of eye motion could be added. There are eye motion data provided in the supplemental figure, but it is not completely clear how this eye motion data is actually being used to derive meaningful information about visual acuity.

      We performed analyses to determine whether there seemed to be a significant difference in eye motion patterns between the cone and pixel dropout conditions, which was described in the section “Analysis of Eye Motion Data” and in Supplementary Figure S1. In that figure, we show that for all 4 subjects there is no significant difference in the iso-density contour area containing 68% of their eye motion data. We suggest in the paper that due to the pseudorandom presentation of trials and the limited duration of those trials, subjects were unlikely to adapt or adjust their eye movement, and that instead their natural eye motion served as a data collector that improved acuity.

      What is the accuracy of the eye motion and cone dropout stimulation delivery in the fovea? Given the small size of the cones, it seems that this is one of the most challenging locations of the eye to test with this new olo technology.

      Eye tracking and targeted light delivery are crucial in the AOSLO system and the reviewer is correct to point out that this is most difficult to achieve at the foveal center. To address this concern, we have done some simple modeling and have added the following text to the Cone-by-Cone Stimulation section in the Methods and Materials.

      “This latency, combined with other factors such as diffraction and residual aberrations, limit the ability to restrict the light to only the targeted cone. Considering a 543-nm focus through a 7.2 mm pupil, a random tracking error with a full-width-at-half maximum (FWHM) of 0.5 arcminutes (Harmening et al. (2014)), a 0.0125 diopter residual defocus error (maximum error given the step sizes of 0.025 diopters in the AOSLO defocus controller), an average cone spacing of 0.5 arcminutes (Wang et al. (2019)), and a Gaussian cone acceptance aperture with a FWHM that is 0.5 times the inner segment diameter (Macleod et al. (1992)), we estimate that each targeted cone receives 5.41 times more light than its nearest neighbor. This means that the ’dead’ cones cannot be fully excluded from the visual processing. Furthermore, the light leakage reported in Fong et al. (2025) further adds to the signal of non-targeted cones.

      Nevertheless, it is important to point out that the information about the stimulus (Landolt C in our case) is sampled at the targeted cone’s location and so, although nearby stimulated cones might detect light, they do not contribute to any increases in the sampling process. This is analogous to adding defocus blur to letters in the pixel dropout condition as neither situation will improve the spatial information.”

      References

      Reiniger, J.L., Domdei, N., Holz, F.G., Harmening, W.M.: Human gaze is systematically offset from the center of cone topography. Current Biology 31(18), 4188–4193 (2021)

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work. Below, we provide point-by-point responses to the comments from the two peer reviewers, and we hope these revisions are satisfied by you and the reviewers.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Thank you for this suggestion. We have acknowledged this limitation in the Discussion by adding the paragraph below (see the Third point):

      “Despite of these, several questions deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if the documented facilitation of yak nutrition by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition? Third, our study site is located at approximately 3,200 m elevation, relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 272-286 in the Discussion section.

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015).

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

      Thank you for this suggestion. We have added more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we have included the sentence of “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.”

      See these revisions in Line 333-335 in the Methods section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Thank you for the suggestion. We have acknowledged this limitation in the Discussion, by adding a paragraph as: “Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 284-286 in the Discussion section.

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Thank you for your suggestion. We have acknowledged this limitation in the Discussion, by adding the paragraph as “Second, if the documented facilitation of yaks by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition (Yang et al., 2026)?”

      See these revisions in Line 276-279 in the Discussion section.

      Reviewer #1 (Recommendations for the authors):

      Although no doubt a bit sensitive, it would have been better to reveal a bit more about how pikas were removed.

      We have provided more details about how pikas were removed in the no-pika treatment, by adding “For the no-pika treatment, pikas were trapped once every two weeks using 30 live traps (25 cm high × 25 cm wide × 40 cm long) within each plot and relocated elsewhere in the study site.” in the Methods section. We didn’t recorded how many pikas were removed from the corresponding plots, so no data were available for this point.

      See these revisions in Line 411-413 in the Methods section.

      The authors also missed a few relevant papers worth citing, including Badingqiuying, R. B. Harris, and A. T. Smith. 2018. Summer habitat use of plateau pikas (Ochotona curzoniae) in response to winter livestock grazing in the alpine steppe Qinghai-Tibetan Plateau. Arctic, Antarctic, and Alpine Research 50 (1): e1447190

      We have cited this key paper in Line 75 in the Introduction section.

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the Stress Gradient Hypothesis (SGH). Thus, we deleted the description on SGH. However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate densities of pikas has the best beneficial effects on yak growth. We added a separate paragraph in discussion to have a clear discussion.

      We have made the major revisions below to address these concerns.

      (1) We have revised the title into “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands ”.

      (2) We have deleted all the statements about facilitation and competition predicted by the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We added a paragraph in discussion (Line 248-259) to have a clear discussion on the humped relation between yak weight gains and pika burrow densities as “Because of the natural variations in pika density in the pika-present treatment, we were able to obtain a hump-shaped relationship between yak weight gains and pika burrow densities in these plots. Compared with the absence of pikas, the facilitative effect reached its maximum at approximately 200 burrows/ha but became competitive at densities exceeding 400 burrows/ha (Figure 3C). This result reveals that pika density modulates the net outcome for yak weight gain, with a facilitation peak at ~200 burrows/ha and a competition onset above 400 burrows/ha. Our findings offer empirical evidence for the non-monotonicity theory, under which the competition-facilitation balance varies with population density: facilitation dominates at low densities, competition at high densities, and these density-dependent shifts may underpin community stability and productivity (Zhang, 2003; Zhang et al., 2015). The theory further holds that the facilitation threshold, not the competition-facilitation transition, is the critical factor governing the stability of interacting species or communities (Zhang et al., 2015).”

      Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two places where we discussed non-significant interactions: Line 175–177 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Figure 3F, I and Appendix 1—figure 1, table 5,8).” and Line 190–192 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Appendix 1—figure 2, table 10).”.

      We have cross-checked both the Results section and the Figures sections mentioned above, and confirmed that they are consistent now.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Thank you for the suggestion. Now we have added the Appendix 1—table 3 and Appendix 1—table 7 for the model results of generalized additive models (GAMs) for Figure 2C and Figure 3C that plotted with smoothed lines in Appendix 1.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      We have provided the Appendix 1—table 13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Appendix 1.

      We have also added a sentence of “A summary of all statistical models used in the study is available in Appendix 1 table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We have revised this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Figure 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction

      Line 53 - I wouldn't describe small mammals like rodents as keystone species. They can have strong impacts on ecosystems, but not disproportionate relative to population size, which is a key part of that definition. It may be true when you talk specifically about pikas later on, but not small mammals as a general category.

      We have replaced this term with “consumers” here, see Line 63.

      Lines 58-61 - Good hook

      Thank you for this positive comment.

      Lines 69-74 - I don't think the stress gradient hypothesis is the right pitch for this work. The SGH posits that facilitation increases with abiotic stress, but no abiotic stressors were measured in this study. Population density interacts with abiotic factors in the SGH, but population density in and of itself is not an abiotic stressor. The papers you cite here all look at the interaction between population density and abiotic stressors (e.g., water availability). So these lines set me up to expect a stress gradient in your experiment that didn't exist, then left me confused later on. It would be better to highlight the strength of your work (lines 70-76 pose interesting questions and predictions) rather than trying to make it fit within the SGH.

      Thank you for pointing out this problem. We have deleted the statements about SGH here.

      See these revisions in Line 88-91 in the Introduction section.

      (2) Materials and Methods

      Lines 498-509 - Did you verify beforehand that no plants within the enclosures had been grazed on? How did you know that the consumption or clipping was specifically from those pikas?

      We have clarified here by adding “Before cage installation, we carefully checked the plants within each plot and removed those that had been previously grazed or damaged by herbivores.” in the Methods section. In this case, we made it sure that the consumption or clipping was specifically from those pikas with the cages.

      See the revisions in Line 349-350 in the Methods section.

      Lines 509-512 - I would be careful calling this preference. It's really just a record of what they consumed along paths they were walking, which could be about accessibility and convenience as much as preference.

      We have replaced “diet preferences” with “diet composition” here, see Line 359 in the Methods section.

      Line 514 - Clarify the specific question or hypothesis you're testing with this field survey, beyond just generally testing associations.

      We have modified the sentences here as “In July 2021, we investigated the potential facilitation of pikas on yaks mediated by the poisonous Stellera forbs under unmanipulated field conditions in the study site.”.

      See Line 365-366 in the Methods section.

      Lines 532-534 - The intro for this paper sets it up to be about density-dependent movement from competition to facilitation, but the experiment here is set up to compare pika presence/absence. The framing of the paper needs to be adjusted to better align with this experimental design.

      The same issue as mentioned above. We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the balance between competition and facilitation as predicted by the Stress Gradient Hypothesis (SGH). However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate density of pika has the best benefical effect on yak. We added an separate paragraph in discussion to have a clear discussion about this point in the Discussion section.

      We have made the major revisions below to address this concern.

      (1) We have revised the title as “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands”.

      (2) We have deleted the statements about facilitation and competition and the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We kept the discussion on the humped relation between yak weight gains and pika burrow densities (Figure 3C). We added an separate paragraph in discussion to have a clear discussion about this point. For details, see Line 248-259 in the Discussion section.

      Lines 566-569 - Why did you need to simulate pika clipping when you already had pika presence/absence treatments?

      We conducted Stellera removal treatment by simulating poisonous plant clipping behaviors of pikas because we want to confirm that the shifts in abundance of this dominant poisonous plant species is the key mechanism in driving pika-yak facilitation in our system. If we simply looked the differences in yak weight gain in the pika presence/absence treatments, it should be difficult to secure the underlying mechanism. In addition to the reduction in abundance of the poisonous Stellera, pikas may cause a variety of shifts vegetation properties including plant productivity and diversity, and soil disturbances that can exert direct and indirect effects on yak foraging activities, and thus their weight gains.

      Lines 566-569 - Is there another citation you can give showing that Stellera forbs taller than 20cm are both preferred by pikas and exert greater impacts on plant/animal communities? Those are big assumptions that need to be better supported or explained. If there was a logistical reason that you didn't remove all of the smaller forbs, that needs to be laid out as well.

      We have added one new citation here to support this method here. We have revised this section as “To simulate the clipping behavior of pikas, we clipped only those Stellera forbs exceeding 20 cm in height. This threshold was chosen based on previous observations that pikas preferentially target large forbs of this size (Liu et al., 2009).”

      Liu W, Zhang Y, Wang X, Zhao JZ, Xu QM, Zhou L. 2009. The relationship of the harvesting behavior of plateau pikas with the plant community (In Chinese). Acta Theriologica Sinica 29:40-49.

      Also, we have deleted the sentence of “and can exert significant impacts on the plant community and on yak grazing behaviors (Z.Z., field observations)” mentioned above, because these patterns were observed only by the authors in the field and lack supporting data.

      See these revisions in Line 419-422 in the Methods section.

      Line 587 - sample size per treatment? Was it consistently one species that you measured, and if so, which one? If you collected different species or a mix of species across treatments, then you can't really compare the nutrient values because you haven't accounted for interspecific variation.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g).

      We have revised this section as below:

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Lines 603-613 - Please provide the dependent variables in each model, as well as any interaction terms, in addition to the random effects. Please also state explicitly what you were trying to test with each of these models.

      There are two models that use the tweedie family (forb and sedge bite rate). Indeed, we need to include the tweedie power parameter to help understand the mixture of the three families. We have included the p-value (power) in the Appendix 1—table 13 in the Appendix 1. We chose to use tweedie because a normal gaussian family model fitted the results poorly.

      We have added the Appendix 1—table 13 in the Appendix 1, which provides the summary of all statistical models used in the study, including response variables, model type, distribution family (with Tweedie power parameter where applicable), interaction terms, random effects structure, and sample sizes.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Lines 613-614 - Tweedie is a category of distributions that includes quite a few different options, including Gaussian, Poisson, and Gamma distributions, some of which are normal and some of which are not. So, justifying the use of Tweedie distributions in your model structure doesn't really make sense, and it doesn't really tell me which distribution each model pulled from. Please clarify specifically which distributions you used for each model and why.

      We have now added the reason why we used Tweedie distributions by adding the sentence of “There were two models that used the tweedie family (forb and sedge bite rate). We chose to use tweedie because a normal gaussian family model fitted the results poorly” in Line 468-470 in the Statistical analyses section.

      We have also included the p-value (power) for forb and sedge bite rate in the Appendix 1—table 13 in the Appendix 1.

      Lines 615-618 - Provide citations for R packages described in the text.

      We have provided all the related citations for all R packages described in the text, as listed below.

      glmmTMB: Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Maechler, M., & Bolker, B. M. (2017). glmmTMB Balances Speed and Flexibility Among Packages for Zero-inflated Generalized Linear Mixed Modeling. The R Journal, 9(2), 378–400. https://doi.org/10.32614/RJ-2017-066

      mgcv: Wood, S.N. (2017). Generalized Additive Models: An Introduction with R (2nd edition). CRC Press. AND Wood, S.N. (2011). Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(1), 3–36. https://doi.org/10.1111/j.1467-9868.2010.00749.x

      DHARMa: Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi-Level / Mixed) Regression Models. R package version 0.4.6. https://CRAN.R-project.org/package=DHARMa

      tidyverse: Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L.D., François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M., Pedersen, T.L., Miller, E., Bache, S.M., Müller, K., Ooms, J., Robinson, D., Seidel, D.P., Spinu, V., Takahashi, K., Vaughan, D., Wilke, C., Woo, K., & Yutani, H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686

      See these revisions in Line 472-475 in the Statistical Analyses section.

      (3) Results

      Line 127 - Somewhere in the intro or methods, describe what clipping is and why the pikas do it.

      We have provided more details about the clipping behaviors of pikas by adding “Notably, pikas often clip (although do not eat) the wolf poison S. chamaejasmehas because its relatively large size hinders predator detection (Fan et al., 1998).” in Line 330-332 in the Methods section.

      Line 128 - Change from "preferred" to "consumed greater proportions of"

      See the correction in Line 146 in the Results section.

      Line 136 - You measured yak weight once a month, so give the result in monthly weight gain rather than daily.

      Here we preferred to keep the unit of daily weight gains (converted from monthly ones), as this is the standard presentation for livestock growth performance, also see Fig. 1 in Odadi et al., 2011 Science’s paper.

      W. O. Odadi, M. K. Karachi, S. A. Abdulrazak, T. P. Young, Science 333, 1753–1755 (2011).

      Lines 138-140 - The linear models, as you described them in the methods (lines 603-620), don't test for hump-shaped relationships. Please update the methods to explain how you tested the density relationship and how you got this interpretation.

      The description for Figure 3C here showed estimates from a GAM (family: gaussian), and we did not use a linear model in this figure.

      We have now added a new model summary Appendix 1—table 7 for this Figure 3C in the Appendix 1.

      Line 145 - I don't think you can claim that the total available forage was more nutritious for yaks, as you picked plants that yaks happened to be chewing along a grazing path. You would need to take samples from a consistent set of plant species at random locations to make this claim.

      Sorry for this confusion. We modified “the total available forage” as “the major available forage” here (see Line 172) because we collected the same forage plant species and analyzed their nutrients.

      The same as mentioned above, to clarify the sampling methods, we have also revised this point in the Methods section to provide more detail on how forage samples were collected and their quality were analyzed. See these revisions in Line 439-447 in the Methods section.

      (4) Discussion

      Line 166 - Need to address inconsistency in how you talk about pika density. Here you talk about the impacts of pikas at moderate densities, which I think is a fair claim. Elsewhere, you talk about density-dependence, which I don't think you really measured, given that all your sampling was either in the absence of pikas or within a narrow window of densities that can all be categorized as moderate.

      Done! As mentioned above, we have removed the term of “density-dependence” in the whole manuscript, but keep the term of “moderate density” in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph was removed here), and the References section.

      Line 176-178 - This is a really cool finding.

      Thank you for this positive comment.

      Lines 179-181 - You didn't test competition between plant species, and you didn't measure light, soil moisture, or soil nutrients. So you can suggest competition as a potential mechanism, but you can't say definitively that's what is happening.

      We now have lowered our tone here as “We speculate that these improvements in food availability and nutrition for yaks may be due to the release of grasses and sedges from competition with the forbs for limiting above- and below-ground resources”.

      See the revision in Line 208-211 in the Discussion section.

      Lines 184-186 - You did a good job of it here, suggesting a likely potential mechanism at play without claiming it is for sure happening when it hasn't been measured.

      Thank you for this positive comment!

      Lines 188-199 - Strong paragraph. The impacts of large herbivores on smaller animals are well-studied, but reciprocal impacts are often overlooked.

      Thank you for this positive comment!

      Lines 201-205 - You didn't test the stress gradient hypothesis because there was no abiotic gradient. You also did not take any measurements during outbreaks, so you cannot claim to have compared low-moderate to outbreak pika densities. I think the paper would be much stronger if you removed the stress gradient hypothesis and instead focused more on the literature around facilitation between mammals of different body sizes, as you do in lines 205-209.

      We agreed that our work didn’t specifically design to test the stress gradient hypothesis (SGH) between pikas and yaks, so we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Line 209-215 - Again, you didn't test a competition-facilitation balance because you never tested or demonstrated competition. One of the main strengths of this paper is demonstrating facilitation, so build on that strength rather than referencing things you didn't measure.

      The same as mentioned above, we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Lines 218-222 - Not an accurate description of the relationship between herbivore diet and body size. Larger herbivores typically tolerate lower-quality plants in order to consume sufficient calories, but plenty of them do this via mixed feeding. Grazing in large herbivores and livestock is usually due to specifics of the digestive tract (e.g., hindgut fermentation) rather than specifically about body size.

      We have deleted this description of the relationship between herbivore diet and body size here. Instead, we have modified this statement as “The coexistence of a diverse of herbivore species with different diet selections and size classes can lead to an “compensatory effect” on grass and forb biomass that helps to maintain a balance and diverse plant community” in the Discussion section.

      See these revisions in Line 234-237 in the Discussion section.

      Lines 245-251 - Paragraph addresses an important point. Lines 247-249, though, overstate what you measured. There's no measurement of livestock production or biodiversity in the study.

      We have replaced the term of “livestock production and biodiversity” with “livestock growth performance” here. See Line 288-298 in the Discussion section.

      (5) Figures

      Figure 2C-D - This applies to all figures, but you need to plot the best-fit line generated from your model instead of using geom_smooth, which is what these lines look like. You can do this using functions like predict or ggpredict. Using these smoothed lines implied non-linear relationships that you didn't actually test for.

      We have redrawn Figures 2C and 2D to use model estimates directly and have included the related model summaries as Appendix 1—table 3 and table 4 in the Appendix 1.

      Figure 3B - This figure doesn't show yak weight gain in the presence of pikas. Instead, it shows weight loss when pikas are absent. It's a subtle difference, but very important for interpretation. Yak can maintain weight just fine without pikas as long as Stellera are absent, too. Your results consistently show no Pika x Stellera interactions, but that doesn't match your significance values here. Need to double-check and explain that.

      We have revised the descriptions for Figure 3 and 4 in the Results section, by emphasizing that the absence of pikas REDUCED weight gains of yaks, INCREASED toxic plant abundance, and REDUCED the quantity and quality of palatable grasses and sedges. We have revised these sections as below:

      Abstract (see Line 51-53)

      “Compared to the pika-present treatment, pika removal dramatically increased cover of the poisonous Stellera forbs by two-fold, reducing the abundance and protein content of palatable grasses and sedges, yak foraging efficiency, and yak weight gain by up to 42%.”

      Results (see Line 153-167, Line 169-175, Line 184-192)

      Also, we did find significant Pika x Stellera interactions for yak weight gains, we have provided these details in the Appendix 1—table 5 and table 6 in the Appendix 1.

      (6) Recommendations

      Lines 245-257 - Would recommend combining the last two paragraphs into one.

      We have combined the last two paragraphs, see Line 288-298 in the Discussion section.

      Line 530 - Can you replace large with a more precise measure of area?

      We are unable to provide a more precise measure of area here, so we have deleted the description of “in a large area”, but we have also added the note of “in the study site” by the end of the sentence to better describe the location of the plots.

      See the revision in Line 413 in the Methods section.

      Line 541 - Does this mean +/- 7.8 standard deviations? If so, how big is that range in kg?

      Here should be “115±7.8 kg”, we have done this correction in Line 392 in the Methods section.

      Figure 3 - Would be helpful to use "Stellera/pika present" and "Stellera/pika absent" rather than saying "No Stellera/pika" since you also use the No. abbreviation for numbers a lot in this figure.

      We have revised all the related Figures for this issue in Figure 3, 4, and Appendix 1—figure 1,2.

      Figure 3F and I - Show letters for significance on these two plots as well, even if it is just a row of a's.

      We have redrawn Figure 3F,I to address this issue.

      Figure S1 - Applies to all boxplots. Be consistent about showing significance, even if it is a row of a's indicating no difference between treatments.

      We have re-drawn Appendix 1—figure 1,2 to address this issue.

      Table S1 - Something happened with the line numbers, so they are in the table instead of on the left side.

      We have fixed this problem for Appendix 1—table 1 in the Appendix 1.

      Table S1 - Applies to all tables. Include the type of model that you ran (including distribution if not Gaussian) in the table legend.

      We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Table S9 - This legend has a good description, including the type of model you used and what you were testing. Apply this more detailed legend to the rest of the tables.

      Again! We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides new insight into the regulation of cell organization and division in Trypanosoma brucei through the control of a kinesin motor protein by a polo-like kinase. The authors present solid evidence from rigorous biochemical and imaging analyses showing that phosphorylation modulates kinesin function and cellular organization. However, direct in vivo evidence that PLK phosphorylates kinesin-G is lacking.

      We performed experiments to investigate the effect of PLK inhibition on the phosphorylation of KIN-G in vivo in trypanosome cells by immunoprecipitation and mass spectrometry. The new results showed that treatment of trypanosome cells with GW843682X, a potent TbPLK inhibitor validated previously in procyclic trypanosomes, reduced the phosphorylation levels on Thr301 and Ser569 of KIN-G by ~27% and 100%, respectively. The partial reduction in Thr301 phosphorylation after GW843682X treatment could be attributed to slower dephosphorylation of phosphorylated Thr301 after GW843682X was added to the cell culture. Nonetheless, these new results demonstrated that KIN-G is an in vivo substrate of PLK.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript identifies the orphan kinesin KIN-G as a substrate of Polo-like kinase (TbPLK) in Trypanosoma brucei and demonstrates that phosphorylation of Thr301 inhibits KIN-G microtubule binding and disrupts its cellular function. Using a combination of in vitro kinase assays, phosphosite mapping, microtubule binding and gliding assays, and in vivo complementation with phosphomimetic and phosphodeficient mutants, the authors link TbPLK-mediated regulation of KIN-G to defects in centrin arm integrity, FAZ elongation, Golgi organization, flagellum positioning, and division plane placement. The study provides a mechanistic advance in understanding how TbPLK regulates centrin arm biogenesis and integrates KIN-G into the growing regulatory network controlling hook complex and FAZ assembly. Overall, the work is technically strong, internally consistent, and builds logically on previous studies from this group and others.

      Strengths:

      A major strength of the manuscript is the clear mechanistic link between phosphoryltion of Thr301 and loss of microtubule binding activity. The use of phosphomimetic (T301D) and phosphodeficient (T301A) mutants in an RNAi-rescue framework provides a clean and convincing demonstration of functional relevance in vivo. The integration of biochemical assays with detailed cell biological phenotyping (centrin arm length, FAZ elongation, basal body segregation, and cytokinesis markers) is particularly effective and makes the central conclusion robust. The observed phenotypic cascade from centrin arm defects to FAZ and division plane abnormalities is also well aligned with existing models of trypanosome morphogenesis.

      Weaknesses:

      My (more or less main) concern relates to the interpretation of the Golgi phenotype. The conclusion that phosphorylation of KIN-G "impairs Golgi biogenesis" is currently based on fluorescence microscopy using TbGRASP and Sec13 markers and on quantification of the number and distribution of Golgi/ERES puncta in binucleated cells. While these data convincingly demonstrate altered Golgi/ERES number and spatial organization, they do not distinguish between true defects in Golgi biogenesis or duplication and alternative possibilities such as fragmentation, vesiculation, or mislocalization of Golgi membranes. Given the central role of Golgi-centrin arm organization in the proposed model, ultrastructural analysis (for example, by EM or electron tomography) would greatly strengthen this aspect of the study by providing direct evidence for structural alterations of the Golgi and its association with the centrin arm and ERES. Such data would elevate this part of the manuscript from a descriptive fluorescence phenotype to a true structural cell biological insight. I appreciate that this experiment goes beyond the current dataset, but it would substantially enhance the mechanistic depth of the Golgi-related conclusions and strengthen the causal chain linking centrin arm defects to Golgi abnormalities. However, I have to confess, the inclusion of such data would make this reviewer particularly enthusiastic about the work. If this is not feasible, I would recommend tempering the wording of "Golgi biogenesis" to a more conservative description, such as altered Golgi organization or duplication, and explicitly acknowledging the limitations of fluorescence-based analysis for this conclusion.

      Thanks for these very constructive comments, which are very well taken. We totally agree with this reviewer on these points. Since it is not feasible for us to perform EM or electron tomography, we have revised the manuscript to describe the effect of KIN-G phosphorylation on the Golgi as “altered Golgi duplication” rather than “Golgi biogenesis”. We also explicitly acknowledge the limitations of fluorescence-based analysis of the Golgi for this conclusion and suggest that further characterization with EM or electron tomography would allow one to reveal the potential structural alterations of the Golgi and its association with the centrin arm.

      An additional conceptual point concerns the dual role of TbPLK in centrin arm regulation. TbPLK is known to promote centrin arm biogenesis through phosphorylation of TbCentrin2, yet in this study, TbPLK phosphorylation of KIN-G negatively regulates centrin arm assembly. This dual positive and negative regulatory role is intriguing but could be discussed more explicitly. The manuscript would benefit from a clearer conceptual framework addressing how phosphorylation of KIN-G might serve as a temporal or spatial switch to restrain KIN-G activity at specific stages of centrin arm assembly.

      This is a great point. However, we are not sure whether the previous work on TbPLK phosphorylation of TbCentrin2 could lead to the conclusion that TbPLK promotes centrin arm biogenesis through phosphorylation of TbCentrin2. In the published work (de Graffenried et al., MBoC, 2013), trypanosome cells expressing the phospho-deficient mutant TbCentrin2-S54A only showed minor growth defects, exhibiting growth defects after 5 days (de Graffenried et al., MBoC, 2013). Cells expressing the phosphomimic mutant TbCentrin2S54D, however, showed very strong growth defects (de Graffenried et al., MBoC, 2013). The effects of TbPLK phosphorylation on TbCentrin2 appear to be quite similar to that of TbPLK phosphorylation on KIN-G, although the KIN-G-T301A mutant does not have growth defects (up to 5 days in our experiments). It appears that the primary role of TbPLK in regulating TbCentrin2 and KIN-G is negative regulation. Nonetheless, we have added more discussion on these regulatory roles of TbPLK in the revised manuscript.

      Finally, a schematic model summarizing the proposed regulatory pathway from TbPLK phosphorylation of KIN-G to centrin arm assembly, FAZ elongation, division plane placement, and Golgi organization would aid the reader.

      Thanks for this suggestion. We made a schematic model to summarize the roles of KIN-G and its regulation by TbPLK. This is included in Figure 8.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as a misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding, and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei supports the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do not formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      This is a great point that is very well taken. We also had been puzzled by the observation of no growth defects of T301A mutant. This comment enlightened us. From the new experiments we performed to compare the phosphorylated peptides of KIN-G in cells treated and non-treated with the PLK inhibitor GW843682X, we calculated the percentage of phosphorylated Thr301 in non-GW843682X-treated cells. We found that the percentage of peptides containing the phosphorylated Thr301 is ~14% of the total Thr301-containing peptides (Fig. 1H). This result indicates that T301-P is indeed a small minority of the population in the asynchronous trypanosome cells. It is possible that T301 phosphorylation may occur at a specific cell cycle stage such as early S-phase, during which PLK and KIN-G co-localize at the centrin arm.

      (b) Published work links PLK to cell division, FAZ elongation, etc.. The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc..

      Yes, previous work discovered essential roles of TbPLK in basal body segregation, centrin arm biogenesis, FAZ elongation, and cytokinesis. These functions of TbPLK correlate with TbPLK’s localization to multiple subcellular structures, the basal body, the centrin arm, and the new FAZ tip, and are attributed to the regulation of its substrates at these structures. At the basal body, TbPLK phosphorylates SPBB1, which is required for basal body segregation. Defects in basal body segregation can lead to defective flagellum positioning and FAZ elongation. At the new FAZ tip, TbPLK regulates the cytokinesis regulator CIF1, which is required for cytokinesis. At the centrin arm, TbPLK phosphorylates TbCentrin2 at S54 and KIN-G at T301 (and S569, which was newly identified as an in vivo TbPLK site and has not yet been characterized). However, cells expressing TbCentrin2-S54A have very weak growth defects, and cells expressing KIN-G-T301A have no detectable growth defects. In contrast, cells expressing TbCentrin2-S54D and cells expressing KIN-G-T301D have strong growth defects. Therefore, the essential role of TbPLK in centrin arm biogenesis apparently is not attributed to the phosphorylation of TbCentrin2 and KIN-G. It is possible that phosphorylation of other centrin arm-localized protein(s) by TbPLK may be essential for centrin arm biogenesis, but this possibility remains to be explored.

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      We performed experiments and presented the data in Fig. 1H. We also included commentary in the revised manuscript on the points about the potential role of T301 phosphorylation. Thanks for these great comments that significantly improved the manuscript.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      We treated trypanosome cells with a potent PLK inhibitor GW843682X, which was previously demonstrated to inhibit TbPLK activity in vitro and mimic TbPLK knockdown in vivo in trypanosomes, and immunoprecipitated KIN-G for mass spectrometry. We compared the KIN-G peptides identified by mass spectrometry from trypanosome cells treated with or without GW843682X, and found that two phosphosites (T301 and S569) were reduced by ~27% and 100%, respectively, after GW843682X treatment (Fig. 1G). These results provided evidence to support that TbPLK phosphorylates KIN-G in vivo.

      Reviewer #3 (Public review):

      Summary:

      Here, the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during the early S-phase of the cell cycle. Centrin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescence to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      Some of the broader conclusions are not directly supported by the data. For example, the title states "Polo-like kinase phosphorylation of the orphan kinesin KIN-G negatively regulates centrin arm biogenesis in Trypanosoma brucei," but the data do not directly address the specific role of TbPLK in phosphorylating KIN-G in cells. Moreover, some of the more specific conclusions in the paper, for example, that "phosphorylation of KIN-G" causes various cellular defects, are a bit of an overstatement. The supporting data rely on the expression of a phospho-mimetic mutant of KIN-G. Presumably, phosphorylation in cells is a normal part of KIN-G regulation, and it is not just phosphorylation, but rather hyperphosphorylation that is being mimicked by the mutant. Some rewording of the specific conclusions is warranted, and the broader conclusion would be better supported with additional experimental evidence.

      This is a great point that is very well taken. We performed new experiments to address the in vivo phosphorylation of KIN-G by TbPLK. We treated cells with a potent PLK inhibitor, GW843682X, and then immunoprecipitated KIN-G for mass spectrometry to identify changes in phosphorylation. We found that the phosphorylation levels of T301 and S569 were reduced by ~27% and 100%, respectively, confirming that these two sites are in vivo TbPLK phosphosites.

      We also calculated the ratio of phospho-T301 versus non-phospho-T301 in non-treated cells and found that phospho-T301 accounts for ~14% of the total KIN-G protein. This new result suggests that it is the hyperphosphorylation that causes growth defects. We have revised the manuscript accordingly to reflect this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Several statements use rather strong causal language (for example, "thereby impairing Golgi biogenesis, FAZ elongation, and division plane placement"). While the phenotypic correlations are convincing, direct causality is largely inferred from prior literature. Slightly tempering this wording could improve precision. It would also be helpful to state explicitly in figure legends the number of cells analyzed per condition and the number of independent experiments for each quantified phenotype.

      Thanks for these constructive comments, which we agree and appreciate greatly. We have revised the manuscript accordingly to improve precision. The total number of cells analyzed, and the number of independent experiments were included in the figure legends.

      Reviewer #2 (Recommendations for the authors):

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      Thanks for this comment. We have deleted the discussion about TbPLK activities that were published previously.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Figure 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Figure 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces the motility of KIN-G. These statements are contradictory.

      We meant to say that there was a slight but insignificant effect. We agree that such statements are somewhat contradictory and, hence, have been deleted. Thanks.

      (3) p. 8, and Figure 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      We have deleted the wording “ventral side”, as it is not necessary. Thanks.

      (4) Figure 7B. Please explain the labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      Thanks very much. We included two sentences in the revised manuscript to explain this point.

      The sentences read as follows: “The nascent posterior is formed near the mid-portion of the NFD cell through microtubule bundling and cytoskeleton remodeling during late stages of the cell cycle (Wheeler et al., 2013). Consequently, the NFD cell inherits the old, existing cell posterior, whereas the OFD cell inherits the newly formed or nascent cell posterior.”

      (5) Figure 4, 5, and 7: "% Cells" is reported. Please indicate the total number of cells that were examined.

      The total number of cells were included in the figure legends.

      Reviewer #3 (Recommendations for the authors):

      (1) The manuscript should be carefully edited for minor grammatical errors.

      Thanks. We have carefully proofread the manuscript and corrected the grammatical errors.

      (2) A general conclusion is that TbPLK phosphorylation of KIN-G in cells is critical for regulating its motor activity. However, this relies on the expression of phospho-mimetic mutants, which bypass TbPLK. Thus, there really is no direct evidence provided to support the specific role of TbPLK other than the in vitro phosphorylation data. Some additional experiments to assess the specific role of TbPLK in phosphorylating KIN-G in cells would lend support for the general conclusion. Is it possible to deplete or inhibit TbPLK and show that this impacts the phosphorylation of KIN-G in cells?

      This is a great point that is very well taken. We performed a new experiment by inhibiting TbPLK with a potent PLK inhibitor, GW843682X, and then immunoprecipitating KIN-G for mass spectrometry to identify the phosphorylation levels before and after GW843682X treatment. We were able to confirm that T301 and S569 phosphorylation was reduced after treatment. This confirms that TbPLK phosphorylates T301 and S569 of KIN-G in vivo in trypanosome cells.

      Specific:

      (1) Figure 2: For microtubule gliding assays, representative videos should be included as supplementary data. Also, when the data do not show a significant difference between KIN-G and KIN-G + TbPLD-K70R, then the authors should not state there is a slight difference, as this is not supported by the data.

      We have deleted the statement. We have included the representative videos for all the experiments presented in Figures 2 and 3. Thanks.

      (2) Figure 3: As stated above, for microtubule gliding assays, representative videos should be included. Moreover, for KIN-G-T301A, the authors say that the gliding activity was "moderately, but insignificantly, reduced." Again, if the difference is not significant, it cannot be concluded that there is a difference compared with the wild-type protein.

      We have deleted the statement. Thanks.

      (3) Some of the headings in the Results section are not accurate. Regarding the in vivo results, the heading "Phosphorylation of Thr301 in KIN-G by TbPLK causes defective cell proliferation" seems to be an overstatement. It may be more accurate to state that "Expression of a phospho-mimetic mutant of KIN-G causes defective cell proliferation." The same comment applies to the other headings that follow this one. Presumably, there is a population of phosphorylated KIN-G in cells, and phosphorylation/dephosphorylation is a normal part of its regulation. In the text, the authors might more accurately conclude that hyperphosphorylation causes the defects they are seeing.

      This is a great point. We agree and as we responded above, we have revised the manuscript, from the title to the main text, to reflect the point that hyperphosphorylation of T301 by TbPLK causes the defects. Thanks very much for these great comments!

    1. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback. We are grateful that the experimental design and the evidence for genetically canalized nutrient resorption efficiency (NuRE) were recognized as compelling and important, with clear implications for predicting wetland nutrient cycling under salinization. We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them, with particular attention to the weaknesses noted.

      Reviewer #1 raised two important concerns regarding the scope and generality of our conclusions. First, the salinity treatment spanned only one growing season, so the genetic canalization we document specifically refers to the absence of a plastic response to an acute salt shock; whether long-term, chronic or multigenerational salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question. Second, the test of nutrient limitation control relies on resorbed N:P and N: K ratios as proxies rather than direct nutrient manipulation, and the metabolomic analysis is used primarily to validate stress effectiveness. We accept these criticisms and will address them as follows.

      Regarding the temporal scope of the treatment, we will explicitly state in the Discussion that our conclusion of canalization pertains to short-term acclimation to an acute salt shock, and we will discuss the scenarios under which chronic, more severe or multigenerational exposure could trigger plastic, acclimatory or transgenerational responses. Given our team’s prior work on parental and transgenerational effects in clonal plants, we will frame these as testable hypotheses for future research. We will also acknowledge that the physiological mechanisms underlying the lack of a plastic increase in NuRE (for example, phloem loading or senescence-associated gene expression) are not directly resolved in the present study and will propose targeted molecular investigations as a natural next step.

      Regarding the nutrient limitation tests, we will clearly acknowledge in the Discussion that the resorbed N:P and N:K ratios provide an established but indirect proxy for nutrient limitation, and we will discuss how direct nutrient addition experiments could provide stronger causal evidence, while deepening the integration of the metabolomic profiles with genotype-level NuRE variation where feasible. We will also quantitatively assess the potential collinearity between ecotype and phylogeographic lineage among the Chinese populations.

      Reviewer #2 raised three substantive issues regarding the interpretation and presentation of our results. First, although latitude emerged as a significant predictor in the variation partitioning, the relatively low R² values indicate that much variance remains unexplained, and our interpretation should be more cautious; in addition, the functional significance of the “inverted” nutrient limitation (slopes greater than 1 for resorbed versus green N:P) needs further explanation. Second, the single moderate salinity level (10 ppt) does not allow assessment of whether more extreme stress might trigger a plastic response. Third, the ecotype analysis applies only to the Chinese populations, as ecotype classification was not available for non-Chinese populations, and this should be stated explicitly in the Results. We fully agree with these points and will revise accordingly.

      To address the concern about latitude, we will temper the language in the Results and Discussion, explicitly noting the limited proportion of variance explained by latitude and acknowledging the role of unmeasured factors, while retaining the study’s central message that genetic and phylogeographic origin outweigh short-term plasticity in shaping NuRE. We will also expand the Discussion to explain the functional significance of the “inverted” nutrient limitation, that is, why P and K are resorbed more completely relative to N, whether as a strategy to maintain optimal N:P:K ratios or a reflection of the higher costs and lower availability of N.

      To address the salinity gradient concern, we will acknowledge in the Discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response and will justify future dose-response experiments to identify the threshold at which canalization of NuRE might be overcome.

      To address the ecotype limitation, we will explicitly state in the Results that the ecotype analysis in Figure 4 is based only on the Chinese populations, for which ecotype classification was available.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study uses technically compelling long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The authors further propose the provocative model that post-translational mechanisms, rather than the transcriptional-translational processes, may contribute to circadian regulation of neuronal excitability.

      We are pleased that our study is recognized as valuable and that our technically very challenging long-term in vivo recordings and computational modeling are appreciated. We agree that we are proposing a provocative model that opposes the current hypothesis in chronobiology, which suggests that all observed biological circadian rhythms are outputs of a transcriptional-translational feedback loop (TTFL) clock. Instead, we suggest that a cell comprises, in addition to the TTFL clock, other posttranslational feedback loop (PTFL) clocks without the need of daily transcription and daily degradation of its core elements. While the circadian TTFL clock is entrained to the daily light-dark cycle, the circadian PTFL clocks are suggested to be entrained to other daily cues such as to the availability of pheromone, or the daily changing levels of hormones and second messenger levels. Our novel hypothesis proposes that the TTFL and PTFL clocks are coupled and linked, constituting an adaptive, plastic network that can tune and phase-lock to different external and internal Zeitgeber signals. However, we certainly do not claim that the TTFL circadian clock is not at all involved in the circadian control of the ORN’s circadian membrane potential rhythms. We clarified our manuscript accordingly.

      However, the evidence for circadian firing in these neurons […] remains incomplete.

      As requested by the reviewers during the previous round of review, we had provided the results of RAIN analysis (Thaben and Westermark, 2014) of individual animals in our first revision (Fig. 4A), which clearly shows that two-thirds of the population express circadian rhythms in key attributes that we used to characterize the spontaneous spiking activity. In the initial review it was assumed by the reviewers that phase alignment of dispersed rhythms would bias interpretations of rhythmicity. After having established with RAIN that individuals show circadian rhythms, albeit dispersed across the population due to the lack of a zeitgeber in DD conditions, phase-alignment of recordings from DD animals is a valid next step to prepare the data for statistical analysis across the population. Phase-alignment of desynchronized rhythms is a generally accepted and necessary method employed in chronobiology (e.g., for insect ORNs: Gosh et al., 2024). It is proven as prerequisite to find and statistically analyze rhythmicity in complex, desynchronized data.

      As we explained in the previous rebuttal and in our first revision, rhythms in electrical activity of insect ORNs cannot be easily synchronized by the light-dark cycle alone, but appear to require daily cycles of pheromone, as shown in other moth species (Gosh et al., 2024). Please be aware that the animals that we used here have never been exposed to pheromone, as we state in the Methods. Therefore, this lack of pheromone exposure can explain why about one third of our experimental population is not expressing any daily or circadian rhythmicity in spiking attributes. This is an important result of our manuscript, reported for the first time for Manduca sexta, providing evidence for our hypothesis that it is not the LD-entrained circadian TTFL clock that governs electrical activity rhythms in ORNs. As we explained in the first revision, and clarified further here in the second revision, we cannot phase-synchronize our animals with cycles of pheromone application in our experimental paradigm because we are researching circadian rhythms in spontaneous spiking activity and not pheromone responses. We failed to obtain phase-alignment with a single pheromone pulse the night before the experiments started. These data were added as supplementary Figure to Fig. 3 in the first revision. Here, we further revised our manuscript to clarify this important finding.

      Thus, as requested by the reviewers in the initial review, we could successfully confirm our previous results of circadian firing in ORNs and the disruption of these circadian rhythms with Orco antagonist OLC15 with RAIN. In the current review, the reviewers raise no further specific critical points or comments that would doubt our careful rhythm analysis of our long-term recordings. Thus, we conclude that we provided clear evidence for our central, exciting new finding. For the first time we demonstrated an unexpected new task for Orco: Orco controls the circadian firing pattern in the spontaneous activity, and thus, of the ORN’s membrane potential, via its property as leak/pacemaker channel.

      However, the evidence […] for post-translational modification of Orco as the underlying mechanism remains incomplete.

      We agree with the reviewers that there are many more experiments and combined efforts of biochemists, structural biologists, and electrophysiologists required to provide complete evidence for post-translational modification of Orco and to reveal the underlying mechanism of its circadian control. It is beyond the scope of the current manuscript to provide all details of post-translational control of Orco.

      The reviewers asked previously for additional evidence that Orco transcription is not controlled via the TTFL clock. As requested, we provided extended qPCR–based evidence that Orco, in contrast to timeless, is not controlled by the TTFL clock on the transcriptional level (Fig. 6 in Revision 1). Furthermore, we added a new result in Revision 1 to demonstrate cAMP-dependent post-translational modulation of Orco open-time probability (Fig 9 in Revision 1) at a ZT at which antennal cAMP levels are low (Schendzielorz et al., 2015). We already showed in Flecke et al., 2010, that the addition of cAMP at different ZTs increases the spontaneous spiking activity only at specific ZTs. Here, we show that the effect of cAMP depends on Orco. Since Orco’s circadian role is not mediated via TTFL control, it can be concluded that post-translational mechanisms provide daily/circadian temporal control. In this additional Figure we provide statistically significant proof that, in agreement with our model-prediction, Orco’s circadian control of the ORN spontaneous activity could be mediated via the second messenger cAMP. ZT-dependent input for Orco would be provided via daily changes in cAMP levels (Schendzielorz et al., 2015).

      In contrast, the study does provide strong evidence that the application of cyclic nucleotides can modulate Orco-dependent activity at a single time point, and reports that the temporal pattern of Orco transcript abundance is not circadian.

      We appreciate that the reviewer confirms that our revised manuscript with additional experiments now provides strong evidence that cAMP modulates Orco-dependent spontaneous activity of M. sexta ORNs. Since we already published that cAMP levels expresses daily rhythms in M. sexta antennae (Schendzielorz et al., 2015), and in vivo cAMP infusion increases spontaneous activity and sensitizes pheromone detection (Flecke et al., 2010), and our computational model here proves that circadian modulation of open time probability of Orco is sufficient to explain our experimental data, it is sufficient for our conclusions to test just the one specific zeitgeber time when endogenous cAMP levels are low and pharmacological cAMP increase has the strongest impact. To further reveal complete ZT-dependence of cAMP modulation of Orco´s control of spontaneous activity is beyond the scope of the current manuscript and not part of the current research question. The structure of Orco is extraordinarily conserved during evolution, thus, the cited experimental results from other laboratories and other species showing that Orco is a hub for posttranslational modification are very likely generalizable to different insect species. We clarified the manuscript accordingly. In Drosophila, Orco has at least 5 phosphorylation sites for protein kinase C (PKC), is cAMP-dependently sensitized, and has a Ca<sup>2+</sup>/calmodulin binding site that orchestrates the localization of the OR-Orco heteromer to the cilia. However, in fruit fly and other insects, so far, it can only be speculated how circadian control is provided for Orco, since there are no other publications that examine the circadian regulation of Orco in detail. We clarified our manuscript accordingly.

      To summarize, the logical conclusion based on our newly provided data is that the current hierarchical hypothesis in chronobiology based solely on a circadian TTFL clock that controls Orco transcription does not explain our findings in hawkmoth ORNs. Therefore, we suggest a new systemic hypothesis based upon coupled TTFL and PTFL circadian clocks that can also reconcile otherwise inconsistent data published for insect and mammalian circadian clocks (please see reviewed data in: Stengl and Schneider, 2024). We clarified our manuscript in the second revision and added a new Figure 10 to illustrate our novel hypothesis.

      However, the findings are incomplete to exclude a role for transcriptional-translational mechanisms and their associated multi-layered controls in circadian regulation.

      We certainly do not imply excluding a role for the TTFL clock in (indirectly) affecting circadian control of the membrane potential of ORNs. The new qPCR experiments added in the first revision clearly show that the circadian control mediated via Orco is not an output of the TTFL clock via transcriptional control of Orco. Instead, we predict links between a PTFL membrane clock comprising Orco as hub to integrate posttranslational control and the TTFL nuclear clock. We clarified our manuscript accordingly, adding a new Figure 10 to further illustrate and visualize our hypothesis. The predictions of this systemic hypothesis will be challenged in further experiments that, however, are beyond the scope of the current manuscript.

      Joint Public Review:

      This manuscript puts forward the provocative idea that a posttranslational feedback loop regulates daily and ultradian rhythms in neuronal excitability. The authors used in vivo long-term tip recordings of the long trichoid sensilla of male hawkmoths to analyze spontaneous spiking activity indicative of the ORNs' endogenous membrane potential oscillations. This firing pattern was disrupted by pharmacological blockade of the Orco receptor. They then use these recordings together with computational modeling to predict that Orco receptor neuron (ORN) activity is required for circadian, not ultradian, firing patterns. Orco did not show a circadian expression pattern in a qPCR experiment, and its conductance was proposed to be regulated by cyclic nucleotide levels. This evidence led the authors to conclude that a post-translational feedback loop (PTFL) clockwork, associated with the ORN plasma membrane, allows for temporal control of pheromone detection via the generation of multi-scale endogenous membrane potential oscillations. The findings will interest researchers in neurophysiology, circadian rhythms, and sensory biology. However, the manuscript has limited experimental evidence to support its central hypothesis and is undermined by several assumptions that underlie their data analysis and model builds, as well as insufficient biological data including critical controls to validate and/or fully justify the model the authors are proposing.

      We want to remind our reviewers that we used “ORN” as abbreviation for olfactory receptor neuron (= sensory receptor neurons, a.k.a. olfactory sensory neuron (OSN)) and not for Orco receptor neuron, although we focus on the function of Orco. Accordingly, our central finding is that Orco as ion channel is required for the daily/circadian modulation of spontaneous action potential activity generated by the olfactory receptor neurons in the absence of pheromone stimulation.

      We do not understand the specific basis for the conclusions of the reviewers. Therefore, we ask to please specify what experimental evidence is missing to support our central hypothesis that Orco is not directly TTFL- but PTFL clock-controlled, and to name specifically what the “several assumptions” are that undermine our careful data analysis and model builds. Which specific argument in our previous rebuttal was wrong, was not conclusive? Furthermore, please specify your claim that “critical controls are missing”. Which controls are missing for which experiments? We did add a new figure panel in the first revision to demonstrate that neither the addition of DMSO (the OLC15 solvent) nor the repeated attachment of the recording electrode, which was necessary to obtain paired datasets, altered the spontaneous spiking activity (Fig 1B in Revision 1). Furthermore, we expanded the time series of qPCR data and added tim as positive control to Orco (Fig 7), strengthening our argument that Orco expression is not under TTFL control. Dose-response curves of various Orco agonists and antagonists have been published before (see our references in Revision 1) and are therefore not repeated here.

      Our newly added data confirm what our modeling predicted: cAMP increases spontaneous ORN activity dependent on Orco. Previous publications provide evidence for daily rhythms in cAMP concentrations in hawkmoth antennae (Schendzielorz et al., 2012).

      As is true for any other hypothesis, a hypothesis can only be falsified but not validated and needs to be tested by many experiments from many laboratories over a long time until it will be replaced by the next hypothesis that better explains accumulating contradicting evidence. We are very much looking forward to experimental challenges of our provocative new hypothesis by colleagues in the field of olfaction and of chronobiology. We are convinced that our manuscript will greatly stimulate the field, possibly provoking a paradigm switch in chronobiology and in olfactory research.

      Strengths:

      The authors raise several intriguing model-based hypotheses regarding the mechanisms that underlie the generation of olfactory rhythms. The electrophysiological approach and the long-term recording paradigm are elegant and technically impressive. In the revised version, the authors have added additional qPCR data supporting the lack of rhythmic Orco transcript expression and included a new figure suggesting that cAMP can modulate Orco conductance.

      We thank the reviewers for their acknowledgement of our careful work and hope that our further revisions and clarifications help to argue our case.

      Major weaknesses:

      (1) The cAMP experiment was only conducted at one time-point, which is insufficient to support the central claim that "AMP and cGMP may have ZT-dependent effects on Orco conductivity".

      We agree with the reviewers and revised our discussion accordingly to clarify that in this manuscript it is not our central claim that cAMP and cGMP may have ZT-dependent effects on Orco conductivity. Instead, our data show for the first time that Orco controls circadian rhythms of spontaneous activity of ORNs and that the circadian rhythmicity of spontaneous activity is lost when Orco is blocked. Therefore, we provide novel experimental evidence that Orco is a prerequisite to the circadian rhythmicity of spontaneous activity and thus, to circadian rhythms in the membrane potential of ORNs. Furthermore, as requested by the reviewers we provided clear evidence in the first revision that Orco is not controlled at the transcriptional level by the TTFL clock, in contrast to the TTFL clock protein TIMELESS. Thus, it follows logically that Orco is under post-transcriptional control. Since cAMP levels show circadian oscillations and Orco is gated by cAMP (Fig 9 in Revision 1), we used our computational model to show that a cAMP-dependent increase in Orco conductance alone, via daily oscillating concentrations of cAMP, is sufficient to explain our findings. Therefore, we propose here that daily/circadian oscillations of cAMP modify spontaneous spiking activity via Orco on a posttranscriptional level. But it certainly does not provide all evidence for respective mechanisms of how this cAMP modulation of Orco is obtained, since this is beyond the scope of the current manuscript.

      Since we realized that it is difficult for our readers to visualize a circadian PTFL membrane clock we added a new hypothesis-Figure (Figure 10) and considerably focused and clarified especially the discussion of our manuscript. We pointed out that a membrane-associated signalosome that comprises delayed negative feedback mechanisms, and, thus, constitutes an oscillator, a “membrane clock” that generates oscillations. Based upon our data we propose a membrane-associated signalosome constituting a PTFL circadian clock with Orco as central element. This PTFL membrane clock generates superimposed ultradian and circadian rhythms in its outputs: rhythms in the membrane potential, Ca<sup>2+</sup>, and cAMP levels. The PTFL clock comprises positive feedforward elements that upregulate its outputs, resulting in more depolarization, higher Ca<sup>2+</sup>- and higher cAMP levels. Via the clock’s delayed negative feedback mechanisms these outputs are downregulated, again, resulting in hyperpolarization, decreasing Ca<sup>2+</sup>- and cAMP levels. This signalosome comprises the pacemaker channel Orco as a central hub that is controlled via changes in voltage, Ca<sup>2+</sup>, and cAMP levels. Nevertheless, we predict coupling between the multiscale PTFL membrane clock and the TTFL circadian clock in the nucleus to obtain stable circadian rhythms. As likely mechanism of coupling we predict that Ca<sup>2+</sup>- and cAMP-dependent kinases interlink both types of clocks, thereby obtaining robust and at the same time flexible interlinked cellular rhythms.

      We hope to now successfully clarify and to visualize our central hypothesis of our manuscript that Orco is not directly controlled by a TTFL circadian clock but is a central element of a membrane-associated posttranslational feedback loop clock (PTFL) clock that is linked to but not forced by the TTFL clock which is predicted to control intracellular Ca<sup>2+</sup> homeostasis in a circadian rhythm.

      (2) The revised manuscript continues to rely heavily on prior publications or defers key mechanistic questions (or important manipulations) to future studies. In its current form, the evidence presented remains insufficient to support the central claim that a PTFL constitutes the primary underlying circadian clock mechanism. The proposed model is intriguing, but the data provided do not yet directly demonstrate the novel mechanism.

      We do not understand why the reviewers considers it to be problematic that we “continue to rely heavily on prior publications”. Certainly, we built upon previous publications of our lab as well as on manifold experimental data published by other laboratories in the field of insect olfaction. Our ample citations demonstrate that we have an overview both of the current state of literature and relevant previous literature, dating back to the very first experiments that pioneered pheromone transduction in insects. Based upon our extensive knowledge and experimental data collected in insect olfaction and based on very careful, critical, rigorous analysis of our data and data published by others, we were able to come up with a novel interpretation of the current literature about insect olfaction that differs considerably from the current main views. We consider this to be our strength and judge it as good scientific practice and not a flaw of our work. However, since we do not focus on OR-Orco heteromers and their function in pheromone/odor transduction in the cilia in the current manuscript, we considerably shortened this part of the discussion, avoiding pointing out that highly sensitive moth pheromone transduction greatly differs from less sensitive general odor transduction in Drosophila. Furthermore, since here we focus on cAMP, but not on cGMP-dependent modulation of Orco, we also deleted/considerably shortened this part of our discussion.

      We certainly agree with the reviewers that, while the data provided in the current manuscript are a logical basis for developing our novel hypothesis, they are not a direct and sufficient demonstration of proof and we are not able yet to directly demonstrate and explain the novel mechanism predicted. We would like to point out that if we provided this final proof, it would not be any more a novel hypothesis, but only a novel finding.

      We agree with the reviewers that our provocative hypothesis requires rigorous testing by many further experiments, hopefully not only by our laboratory, but hopefully stimulating new experimental challenges by other laboratories employing different species. But certainly, these experiments with proof-of-principle will take many years and are beyond the scope of our current research paper.

      As per eLife’s assessment system we would like to ask the reviewers to provide detailed feedback as to which experiments/results within the scope of this manuscript would complete this work, or how they think this study should be framed in the light of the results that we obtained. Nevertheless, we hope that with our current careful review the reviewers will be more convinced by our arguments and experiments as valid basis for our provocative new hypothesis.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We appreciate that the reviewers provided an overall positive assessment of our manuscript and offered constructive suggestions for improvement. All three reviewers noted that a key strength of our study is the implementation of a gut microbiome model for the characterization of interbacterial antagonism pathways such as the type VI secretion system (T6SS) that approaches natural complexity. They note our work represents a significant advance in microbiome research, and generates resources that will be of use to many researchers in the field. Two of the reviewers point out that the complexity of our model limits the nature of measurements we can make, and suggest we temper the strength of the some of the conclusions we draw. As noted in more detail below, in our revised manuscript, we have used more precise wording to characterize our findings, and we are more explicit about the connection between the measurements we made and what we can conclude about the physiological role of the T6SS in the gut microbiome.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain-replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have a significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally coevolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets, including isolate genomes, metagenomes, and ICE distribution maps, will be a valuable community resource, particularly for researchers interested in strainresolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage, therefore, requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      We thank the reviewer for this thoughtful summary of our study. We were glad to read they conclude our work will have a significant impact on the microbiome field and that the resources we have developed will be of value to the community.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      We appreciate this detailed and nuanced characterization of the strengths of our study.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and the mechanistic role of the T6SS.

      We acknowledge that in some places, our manuscript could benefit from greater precision in the language we use when linking the outcomes we observe in our study to their potential underlying causes. Specific revisions we made to address this concern are described below.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple nonexclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      We agree that multiple mechanisms could explain why populations of certain species carrying a T6SS decline over time, and why for others, the T6SS contributes to long-term persistence. Our use of the term “fitness cost” to describe the phenomenon of decline observed for P. vulgatus carrying the T6SS was not meant to imply any particular underlying mechanism, but was rather our attempt to characterize the phenotypic outcome we observed in simplified terms. We note that ecological context is an important determinant of the fitness cost or benefit of any given trait, and our study sheds light on the importance of the presence of the WildR community and the mouse intestinal environment to the fitness contribution of the T6SS to B. acidifaciens and P. vulgatus. Nonetheless, to avoid implying an overly simplistic interpretation of our results, we have modified our language in the manuscript in several places when describing the role of the T6SS in species persistence in mice colonized with the WildR community.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      We acknowledge that our study does not fully resolve the physiological processes under selection that mediate role of the T6SS in maintaining B. acidifaciens populations in WildR-colonized mice. Indeed, several of the outcomes of T6SS activity the reviewer lists, such as target cell killing and nutrient release, are inextricably linked and thus inherently difficult to disentangle. We note that we did attempt to measure higher-order community effects of T6SS activity with metagenomic sequencing, but acknowledge that this approach may not have been sufficiently sensitive to detect small community shifts mediated by a relatively low-abundance species. To address the concern that our current framing implies more of a mechanistic understanding that our study achieves, we have substituted “ecological” for “physiological” where appropriate throughout the manuscript.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      We are not entirely clear what the reviewer means by “differences between native and recipient hosts”, but we agree that additional quantitative studies will be needed to address the generalizability of our findings. Future studies are also needed to address the mechanistic basis for the difference in the benefit conferred by the T6SS that we observed between P. vulgatus and B. acidifaciens.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      We thank the reviewer for specifically highlighting the potential pertinence of this literature to our study. Indeed, we did not cite studies indicating a link between T6SS activity and the uptake of DNA and other resources released by targeted cells. As we note above, the release of intracellular contents from target cells is an inevitable consequence of the delivery of lytic effectors. Thus, distinguishing between fitness benefits conferred from the elimination of competitor species and those arising from scavenging the nutrients released during this process is not straightforward. Measuring the benefits deriving from the uptake of certain released molecules, such as DNA, was not immediately feasible in the system employed in this study and instead we focused on the direct lytic consequences of the effectors delivered via the T6SS. We revised the Discussion to include reference to these possible downstream benefits of T6SS activity (Lines 476-479).

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state that the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing the lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable, practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      The reviewer points out that studying the T6SS activity in M. schadleri would potentially expand the generality of our claims. We agree that having an isolate of this species along with genetic tools for its manipulation would allow us to probe the importance of the T6SS in the gut microbiome more broadly. At the suggestion of the reviewer, we have added explicit mention of the potential benefit of studying the T6SS in this organism to the Discussion (lines 538-539), an endeavor that lies outside of the scope of the current study.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      We agree with the reviewer that invoking fitness costs, selective advantages or physiological burdens should be done cautiously, and have made revisions to our manuscript where we acknowledge that more precise language was needed (line 43, 416, line 417). However, we would also argue invoking fitness costs and benefits when describe strain persistence dynamics in mice has substantial precedent in the literature (Feng et al. 2020, Brown et al. 2021, Park et al. 2022, Segura Munoz et al. 2022), to list a handful of representative examples published by different groups). It is unclear to us what additional in vivo growth measurements could be taken to substantiate our claim that the T6SS provides a fitness benefit to B. acidifaciens during prolonged gut colonization, or that carrying the ICE imposes a fitness cost on P. vulgatus during longterm colonization. Our in vitro experiments evaluating the competitiveness conferred by T6SS activity provide a measure of insight into its fitness benefits, but as our in vivo strain persistence data and the work of many others show, in vitro measurements do not necessarily capture in vivo parameters.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence. This insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      We thank the reviewer for the positive feedback and are glad they agree our study provides new insight into the role of interbacterial antagonism in natural communities.

      Weaknesses:

      Overall, there is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      The reviewer makes a good point that the complexity of the experimental system we employ precludes some lines of experimentation that would yield more mechanistic information. As the reviewer notes, we were aware of the tradeoff between mechanistic resolution and ecological realism when selecting our experimental system. Our deliberate choice to favor biological complexity over mechanistic clarity in this study stemmed from our perception that a major gap in understanding of the T6SS and other antagonism pathways lies in defining their ecological function in complex microbial communities.

      Reviewer #3 (Public review):

      Summary:

      Shen et al. investigate the contribution of the type VI secretion system of Bacteroidales in the gut microbiome assembly and targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing B. acidifaciens to persist in the gut.

      Strengths:

      Using a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and directly examining targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      We thank the reviewer for their enthusiasm toward our study.

      Weaknesses:

      Some conclusions are based on only four mice per condition. The author should consider increasing the sample size.

      We agree that in some experiments it would be beneficial to increase the sample size from four mice. However, the experiments we performed for this study are time and resource-intensive. Additionally, the experiments on which we base our primary conclusions were all independently replicated with similar results. Given these factors, we determined that the extra confidence that might be afforded by increasing our sample size did not merit the delay in publication and investment in resources that would be required.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to disturbances in the microbiome, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

      We agree that investigating the role of the T6SS during perturbations of the microbiome is a key next step for this work and thank the reviewer for highlighting this important future direction.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Beth A. Shen et al. present a comprehensive and carefully executed study investigating the ecological role of the type VI secretion system (T6SS) in maintaining bacterial strains within a native, complex gut microbiome derived from wild mice. By integrating genetic manipulation, metagenomic and sequencing-based tracking approaches, in vitro competition assays, and gnotobiotic colonization experiments, the authors provide compelling evidence that the T6SS functions primarily as a persistence factor rather than a determinant of initial colonization.

      The study is conceptually strong and addresses an important gap in our understanding of how interbacterial antagonistic systems operate in complex, native microbial communities. The manuscript is generally well organized, the data are clearly presented, and the main conclusions are supported by robust experimental evidence. In particular, the demonstration that T6SS-encoding ICEs confer context- and hostdependent fitness effects, including transient benefits and potential long-term costs, adds important nuance to prevailing models of microbial competition.

      That said, several aspects of the study would benefit from clarification and deeper mechanistic discussion. Addressing the points below would further strengthen the rigor and interpretability of the work. Overall, this is a strong and interesting manuscript, requiring some revisions.

      We greatly appreciate the positive summary of our work by the reviewer, which highlights the multi-faceted approach we took to address gaps in our understanding of interbacterial antagonism in the microbiome. It is our hope that the reviewer agrees that our revisions of the manuscript, based on their feedback, clarify our methods and strengthen the interpretability of our work.

      Major comments

      (1) The competition assays in Figure 2B suggest that T6SS-dependent fitness effects are most pronounced among members of the order Bacteroidales. However, these experiments primarily measure population-level competitive outcomes rather than direct T6SS-mediated targeting events. In addition, the limited number of non-Bacteroidales strains included in the assay makes it difficult to conclude that T6SS activity is strictly restricted to closely related taxa.

      The authors should either temper their conclusions regarding target specificity or clarify that these data reflect competitive outcomes rather than direct evidence of targeting. Expanding the discussion to acknowledge these limitations would improve interpretative accuracy.

      With regards to the measurement we employed for assessing T6SS-mediated targeting, we acknowledge that this is, to a degree, an indirect way of determining T6SS targeting. However, there is extensive precedent in the literature for the use of similar assays in assessing targeting by many contact-dependent antagonism systems including the T6SS in Bacteroidales (Russell et al. 2014, Chatzidaki-Livanis et al. 2016, Wexler et al. 2016) and many Proteobacteria (e.g. (Hood et al. 2010)), the T4SS in Xanthomonas citri (Souza et al. 2015), the CDI system in Escherichia coli (Aoki et al. 2005), and the Esx system in Streptococcus intermedius (Whitney et al. 2017). In these studies, targeting was demonstrated by specific depletion of the competitor strain in the presence of a strain encoding an active antagonism system. We acknowledge that the competitive index we report Figure 2B reflects the relative population levels of both species in the assay, and thus does not directly show target species depletion. We opted to use this metric to display the data in the manuscript as a way of efficiently encapsulating and comparing many strain combinations in a single figure, and because the competitive index differences we observed in these derive from differences in target species growth yields (see Author response image 1, indicating growth yields from a representative strain pairing).

      Author response image 1.

      The T6SS of B. acidifaciens targets a WildR-derived P. vulgatus strain. CFUs indicate populations of the indicated strains after co-culture of wild-type or T6SS-inactivated B. acidifaciens with P. vulgatus. Data represent means and standard errors (n=3, *P<0.01, t-test with log ><0.01, t- test with log transformed data)

      We additionally acknowledge that more extensive testing is needed to fully understand the target range of the Bacteroidales T6SS. In our study, we assessed targeting of every WildR species that was readily culturable, which to the best of our knowledge, represents the broadest panel of targets for the Bacteroidales T6SS to be tested to date. We limited our testing to these strains, as the goal of these experiments was to gain insight into which co-residents of the WildR could be targeted by B. acidifaciens. We agree that testing of a broader cross-section of potential targets has merits, but this would require targeted cultivation strategies to obtain these organisms, and lies outside the scope of the current study. We have revised the manuscript to clarify that the target range testing encompassed the diversity of isolates available (p. 10, lines 231-236).

      (2) The bae1 gene encoded in Bacteroides caecimuris F12 contains a frameshift mutation. It would be valuable for the authors to comment on whether such frameshift mutations are a common genomic feature among gut-associated Bacteroides species in murine models. In addition, comparative analysis of human gut metagenomic datasets could reveal whether homologous effector proteins are present in commensal Bacteroides populations, and whether these homologs exhibit similar disruptive mutations.

      More broadly, the manuscript would benefit from a discussion of whether expression of a fully functional bae1 effector might impose a fitness cost on Bacteroidales members, for example, through metabolic burden or altered resource allocation. This is particularly relevant in light of recent studies demonstrating that T6SS effectors can drive physiological trade-offs by modulating metabolic dynamics (PMID: 40592326). Integrating this perspective would strengthen the evolutionary interpretation of effector mutagenesis.

      We agree with the reviewer that the functional and evolutionary significance of the point mutation in bae1 merits further investigation. Following the reviewer's suggestion, we looked in our own datasets and available public datasets from mouse and human microbiomes for evidence of bae1 inactivation. Unfortunately, the gene is present at a low enough frequency that these analyses were inconclusive. In our own metagenomic data from WildR mice, we did not obtain sufficient sequencing depth to assess the frequency at which bae1 is inactivated across genomes. We found a single complete copy of bae1 identical to that of B. acidifaciens in one published mouse microbiome-derived MAG, and detected fragments of the gene in a number of publicly available isolate and MAG genomes, but these were too low of quality to assess whether or not the gene was intact.

      As to whether or not bae1 expression imposes a fitness cost in the producing organism, we think this is unlikely to be significant, given that the impacts of Bae1 will be neutralized by the accompanying immunity protein. We speculate that the point mutation in the B. caecimuris gene is more likely to have arisen through genetic drift than as a result of selection.

      (3) Quantification and tracking of ICE transfer in vivo. In Figure 4D, the authors assess the abundance of resident P. vulgatus populations in germ-free mice co-gavaged with wild-type strains and derivatives carrying either the intact ICE or ICE ΔtssC. Because both ICE variants are capable of horizontal transfer, it is essential to clearly describe how the authors distinguish between (i) the original wild-type strain, (ii) engineered donor strains, and (iii) recipient strains that have newly acquired the ICE or ICE ΔtssC.

      Clarification of the specific molecular or sequencing-based strategies used to discriminate these populations is necessary to ensure accurate interpretation of the colonization dynamics.

      In this experiment, the P. vulgatus strains we introduced which carried the ICE (either the wild-type version or ICE DtssC) also contained an erythromycin resistance cassette (ermG) inserted distal to the ICE insertion site. Populations of the ICE-containing strain were quantified by either qPCR targeting the ermG gene (Fig. 4D, Supplemental Fig. 4D) or by plating on erythromycin-containing media (Fig. 4F). Endogenous P. vulgatus populations were quantified by qPCR targeting the ermG insertion site, which is disrupted in the marked strain. These methodological details have been added to the figure legend for clarity. We acknowledge that transfer of the ICE between introduced and endogenous populations is possible, and would not be detected by these metrics. To assess whether this occurs, we performed ICE-seq analyses on samples collected from mice colonized by the WildR and P. vulgatus ICE at early (7 days) and late (56 days) time points. These analyses revealed that overall, ICE distribution in this experiment was similar to that observed in mice colonized with the WildR alone (Figure 4A and Author response image 2). They additionally provided corroborating evidence that the population of ICE-containing P. vulgatus declined over the course of the experiment. Importantly, the only ICE insertion site we detected in P. vulgatus in these samples was that found in the introduced P. vulgatus strain. Previous studies show that GA1-containing ICE can insert at numerous locations in Bacteroides sp. genomes, a finding supported by our mapping of the ICE insertion sites from in vitro transfer experiments (Supplemental Fig. 4C) (Garcia-Bayona et al. 2021). Thus, our ICE-seq detection of a sole P. vulgatus ICE insertion site indicates that transfer of the element between P. vulgatus populations is likely not occurring in our experiments.

      Author response image 2.

      ICE-seq analysis indicates that introduction of P. vulgatus ICE into WildR-colonized mice has little impact on ICE distribution among endogenous strains. Graphs show frequency of mapped ICE junctions deriving from the indicated species as determined by 5¢ or 3¢ ICE-Seq analysis of DNA extracted from fecal samples collected either 7 or 56 days post-gavage of the WildR and P. vulgatus ICE into germ-free mice.

      (4) The analysis of fitness trade-offs associated with ICE acquisition in P. vulgatus convincingly demonstrates that the benefits of ICE transfer are transient and contextdependent. However, the mechanistic basis of these trade-offs remains underexplored. While the study primarily attributes both benefits and costs to T6SS-mediated antagonism, the ICE likely encodes additional genes that could influence metabolism, regulation, or stress responses.

      We agree with the reviewer that there are many mechanistic questions remaining regarding the benefits and costs associated with ICE acquisition, and acknowledge that we have not investigated the fitness contributions of ICE-encoded genes other than the T6SS. Indeed, as we noted in our discussion of the results from introducing P. vulgatus carrying the ICE into WildR-carrying mice, our data suggest that ICE genes outside the T6SS may be beneficial (lines 463-465). At the reviewer’s suggestion, we have reiterated the importance of considering the fitness contribution of genes beyond the T6SS in determining ICE distribution in the WildR community (line 527).

      Minor comments:

      (1) In lines 319 and 333, the manuscript refers to "Supplemental Figure 3F" and "Supplemental Figure 3G," respectively. However, the provided Supplemental Figure 3 appears to end at panel E. Please clarify or correct these references.

      We have modified the text to reference the correct figure panels.

      (2) Line 1043: The notation for "OD600" should be corrected for consistency and accuracy.

      The notation for OD600 has been updated to be consistent throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      Minor comments:

      (1) Line 144. I would be careful of using "strong correlation, in this sentence. Although it shows a higher correlation than lab mice. Also, the labels in Figure 1A for mouse WildRF7 are confusing and not well explained in the figure legend.

      We modified line 147 (new line in edited manuscript) to say “positive correlation” rather than “strong correlation” to better represent the result. We also revised the legend for Figure 1A to better explain the samples of WildR F7 that were analyzed.

      (2) Line 155. It's unclear which strains were isolated from the WildR community, and the reason for isolating only 15 strains. Also, Supplemental Figure 1 shows 17 isolates, not 15.

      We apologize for the confusion here. We isolated 17 strains, which is the number of distinct strains we were able to readily culture from this community. We obtained genome sequences for 15 of these, and were able to assemble a genome for one more of the strains from metagenomic data.

      (3) Line 236. Is it known what makes B. uniformis resistant to B. acidifaciens carrying a T6SSS? Does it have an orphan immunity protein?

      We do not know why B. uniformis is not targeted by B. acidifaciens under the conditions of our experiments. It does not encode homologs of the immunity genes bai1 or bai2, and does not appear to be intrinsically resistant to targeting by this T6SS given that it is effectively targeted by P. vulgatus carrying the ICE (Fig.4b).

      (4) Line 295. There is a consistent decline in C. acid abundance after 27 days in Figures 3B and 3C. How do you explain this? Is the endogenous B. acid expanding to outcompete C. acid exo since the total C. acid exo remains constant when gavaging 100x B. acid exo?

      We believe that the eventual decline in the introduced population of B. acidifaciens is likely due to a fitness cost imposed by the erm resistance marker we employed. We noted this phenomenon when describing the results depicted in Fig. 3F-H, but neglected to include this explanation earlier. This oversight has been corrected (lines 315-317).

      References

      Aoki, S. K., R. Pamma, A. D. Hernday, J. E. Bickham, B. A. Braaten and D. A. Low (2005). "Contact-dependent inhibition of growth in Escherichia coli." Science 309(5738): 1245–1248.

      Brown, E. M., H. Arellano-Santoyo, E. R. Temple, Z. A. Costliow, M. Pichaud, A. B. Hall, K. Liu, M. A. Durney, X. Gu, D. R. Plichta, C. A. Clish, J. A. Porter, H. Vlamakis and R. J. Xavier (2021). "Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors." Cell Host Microbe 29(9): 1351-1365 e1311.

      Chatzidaki-Livanis, M., N. Geva-Zatorsky and L. E. Comstock (2016). "Bacteroides fragilis type VI secretion systems use novel effector and immunity proteins to antagonize human gut Bacteroidales species." Proc Natl Acad Sci U S A 113(13): 3627– 3632.

      Feng, L., A. S. Raman, M. C. Hibberd, J. Cheng, N. W. Griffin, Y. Peng, S. A. Leyn, D. A. Rodionov, A. L. Osterman and J. I. Gordon (2020). "Identifying determinants of bacterial fitness in a model of human gut microbial succession." Proc Natl Acad Sci U S A 117(5): 2622-2633.

      Garcia-Bayona, L., M. J. Coyne and L. E. Comstock (2021). "Mobile Type VI secretion system loci of the gut Bacteroidales display extensive intra-ecosystem transfer, multispecies spread and geographical clustering." PLoS Genet 17(4): e1009541.

      Hood, R. D., P. Singh, F. Hsu, T. Guvener, M. A. Carl, R. R. Trinidad, J. M. Silverman, B. B. Ohlson, K. G. Hicks, R. L. Plemel, M. Li, S. Schwarz, W. Y. Wang, A. J. Merz, D. R. Goodlett and J. D. Mougous (2010). "A type VI secretion system of Pseudomonas aeruginosa targets a toxin to bacteria." Cell Host Microbe 7(1): 25–37.

      Park, S. Y., C. Rao, K. Z. Coyte, G. A. Kuziel, Y. Zhang, W. Huang, E. A. Franzosa, J. K. Weng, C. Huttenhower and S. Rakoff-Nahoum (2022). "Strain-level fitness in the gut microbiome is an emergent property of glycans and a single metabolite." Cell 185(3): 513-529 e521.

      Russell, A. B., A. G. Wexler, B. N. Harding, J. C. Whitney, A. J. Bohn, Y. A. Goo, B. Q. Tran, N. A. Barry, H. Zheng, S. B. Peterson, S. Chou, T. Gonen, D. R. Goodlett, A. L. Goodman and J. D. Mougous (2014). "A type VI secretion-related pathway in Bacteroidetes mediates interbacterial antagonism." Cell Host Microbe 16(2): 227–236.

      Segura Munoz, R. R., S. Mantz, I. Martinez, F. Li, R. J. Schmaltz, N. A. Pudlo, K. Urs, E. C. Martens, J. Walter and A. E. Ramer-Tait (2022). "Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomes." ISME J 16(6): 1594-1604.

      Souza, D. P., G. U. Oka, C. E. Alvarez-Martinez, A. W. Bisson-Filho, G. Dunger, L. Hobeika, N. S. Cavalcante, M. C. Alegria, L. R. Barbosa, R. K. Salinas, C. R. Guzzo and C. S. Farah (2015). "Bacterial killing via a type IV secretion system." Nat Commun 6: 6453.

      Wexler, A. G., Y. Bao, J. C. Whitney, L. M. Bobay, J. B. Xavier, W. B. Schofield, N. A. Barry, A. B. Russell, B. Q. Tran, Y. A. Goo, D. R. Goodlett, H. Ochman, J. D. Mougous and A. L. Goodman (2016). "Human symbionts inject and neutralize antibacterial toxins to persist in the gut." Proc Natl Acad Sci 113(13): 3639–3644.

      Whitney, J. C., S. B. Peterson, J. Kim, M. Pazos, A. J. Verster, M. C. Radey, H. D. Kulasekara, M. Q. Ching, N. P. Bullen, D. Bryant, Y. A. Goo, M. G. Surette, E. Borenstein, W. Vollmer and J. D. Mougous (2017). "A broadly distributed toxin family mediates contact-dependent antagonism between gram-positive bacteria." Elife 6(Jul 11): e26938.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnk1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      Thank you for the excellent suggestion. We have added mouse in situ hybridization data curated from Allen Brain Atlas and found mRNA encoding WNK1 and WNK2 highly expressed in the hippocampus compared to WNK3 and WNK4. We also have included WNK1 knockdown data from cell lines (Figure S7A-F).

      Whole body WNK1 knockout is embryonically lethal, and we do not have access to brain tissue specific WNK knockout animal models. In addition, knockout of WNKs from brain tissue samples from animals is not very efficient from our experience and therefore, we included data from cell lines. In other non-neuronal cell lines, WNK1 knockdown replicates the effect of WNK463 (Figure S7A-D). However, in SHSY5Y cells, WNK1 knockdown did not replicate the effects of WNK463 on pAKT levels (Figure S7EF). This suggests tissue-specific effects of WNKs, and it also supports our suggestion that cooperativity among WNK family members is required in neuronal cells. This further supports our conclusion that WNK463 is an ideal tool to test our hypothesis in this study as it targets all 4 WNKs (WNK1-4) and furthermore, WNK463 was reported in the literature to inhibit only the four WNKs out of more than 400 kinases tested, indicating more selectivity than many small molecules used to target other enzymes.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      Thank you for the comment. As suggested, we tried to measure insulin in mouse hippocampal tissue lysate, and the levels fell way below the detectable range (78 - 5000 pg/mL) of the mouse insulin detection kit (Abcam: AB285341) used.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      Thank you for the suggestion. The conclusion has been rephrased as suggested.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      The background information related to figure 5 has been simplified as suggested. Response regarding the controls used is provided in the response to critique 5 as below.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      GAPDH has been used as a loading control wherever applicable for Western blots on lysates (see Figures: 5D, 3F, 4E, 4C) In other cases, such as in Figure 5E, the analysis of pAS160 is normalized to total AS160 as this is more appropriate compared to GAPDH. For IP experiments such as Figure 5G, GAPDH is not an applicable control as we are using purified protein fragments in this case. For IP experiments (Figure 5K, 5L, 5C), IP proteins have been normalized to the input protein levels serving as a loading control for the IP because GAPDH is an intracellular protein which necessarily is not pulled down along with the proteins being IP’ed. Therefore, in this case, GAPDH is not a valid loading control. The choice of our loading controls used are very well supported by previous publications from our lab and other labs working on WNK pathways.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      Thank you for the suggestion. Figure 1K is already at the end of the figure 1 and it shows the conclusion of that figure. Figure 4A have been placed at the end of the figures as suggested. Other schemes are only added at the end of the figures as suggested.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

      Thank you for your comment. The suggested images have been replaced with higher resolution ones.

      Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on a overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

      Thank you for the positive response and acknowledging that we have satisfactorily addressed all of your critiques.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      (1) The principal weakness of the manuscript lies in the interpretation of biological robustness. The authors identify network topologies that sustain oscillatory behaviour despite perturbations to the system or parameters. However, in many cases, this persistence is due to the presence of partially redundant oscillatory motifs within the network. While this observation is interesting and of clear value for circuit design, framing it as evidence of evolutionary robustness may be misleading. The “mutant” systems frequently exhibit altered oscillatory properties, such as changes in frequency or amplitude. From a functional cellular perspective, mere oscillation is insufficient — preservation of specific oscillation characteristics is often essential. This is particularly true in systems like circadian clocks, where misalignment with environmental cycles can have deleterious effects. Robustness, from an evolutionary standpoint, should therefore be framed as the capacity to maintain the functional phenotype, not merely the qualitative behaviour. 

      We agree with the reviewers that our framing conflated qualitative robustness (continued oscillation) with functional robustness (oscillation with preservation of properties such as frequency), and that the latter is a more meaningful definition of robustness in an evolutionary context. We have edited the manuscript to remove any suggestions that our results can explain the multiple-oscillator architecture of circadian clocks. We have also added a paragraph in the discussion that highlights the distinction between qualitative and functional robustness, as well as the other ways in which our training environment differs from an evolutionary context.

      Locations of changes:

      Abstract, second-to-last sentence

      Results, final section, first paragraph

      Discussion, first paragraph

      Discussion, new paragraph

      (2) A secondary limitation is that, despite the methodological advances, the scale of the systems explored remains modest. While moving from 3- to 5-node systems is non-trivial, five elements still represent a relatively small network. It is somewhat surprising that the algorithm does not scale further, particularly when considering the performance of MCTS in other domains — for instance, modern chess engines routinely explore far larger decision trees. A discussion on current performance bottlenecks and potential avenues for improving scalability would be valuable.

      We thank the reviewers for raising this important point. We have edited the manuscript to specify that we faced two distinct bottlenecks in scaling our experiments. The first is the runtime and scaling of the underlying Gillespie simulations, which become much more expensive as circuit size increases. The second is our use of the original (“vanilla”) MCTS algorithm, without the deep-learning value and policy networks that have driven the dramatic gains in domains such as Go and chess. We have added a new Discussion paragraph that identifies these two bottlenecks, followed by a paragraph on future methodological enhancements that goes into more detail about potential improvements to the algorithm. Using a power-law extrapolation of the current scaling, we estimate that without further methodological improvements the largest tractable network is approximately 7 nodes, while AlphaZero-style scaling could plausibly extend the approach to roughly 19-node circuits. We have also added an order-of-magnitude estimate of the 5-node search space (≈2×10<sup>9</sup> topologies) to give the reader a more concrete picture of the current scale.

      Locations of changes:

      Results, final section, end of paragraph 1 (Gillespie runtime as the practical bottleneck)

      Discussion, new paragraph 4 (bottlenecks and possible improvements)

      Discussion, new paragraph 5 (projected scaling under deep-learning-based extensions)

      Introduction, paragraph 4, and Discussion, paragraph 1 (search-space size)

      Methods, new section “Estimation of search space size”

      (3) It is worth noting that the emergence of oscillations in a model often depends not only on the topology but also critically on parameter choices and the nature of the nonlinearities. The use of Hill functions and high Hill coefficients is a common strategy to induce oscillatory dynamics. Thus, the reported results should be interpreted within the context of the modelling assumptions and parameter regimes employed in the simulations.

      We agree that the modeling assumptions substantially impact the interpretation of the results, and we have expanded the description of our modeling framework to make these assumptions explicit. To clarify, our model does not use Hill equations directly. Instead, cooperative binding is represented as a sequential, mass-action binding process. In addition to a new Methods section explaining our model in more detail, we have added a Methods section to show analytically that the effective Hill coefficient in our system is always ≤2. It also mentions an important practical benefit of using sequential binding rather than explicit TF dimerization, which is that it improves the size and scaling behavior of the reaction system. 

      Locations of changes:

      Results, section 2, paragraph 1 (clarification that no Hill function is imposed; effective Hill coefficient ≤2)

      Methods, section 1, new subsection “Sequential binding yields Hill coefficients ≤2”

      Recommendations for the authors:

      (4) It would be helpful to include the explicit reaction equations and corresponding reaction rates used in the simulations, to facilitate reproducibility and better understanding of the modelling assumptions.

      We have added the explicit reaction equations, the corresponding rate parameters, and the bounds used during random parameter sampling. The Results section describing our model now reports the rate parameters used in the simulations shown in the paper. Additionally the new subsection at the beginning of the Methods presents the full stochastic model. Finally, the bounds used for parameter sampling are now included in-line in the corresponding Methods subsection. The original tables of parameter values and sampling bounds (Tables 1 and 2) have been retained.

      Locations of changes:

      Results, section 2, paragraph 1 (rate parameters in main text; Table 1 retained)

      Methods, section 1, new subsection “Stochastic model of a transcription factor network”

      Methods, section “Random sampling” (explicit parameter bounds; Table 2 retained)

      (5) Sustained oscillations are notoriously difficult to observe in Gillespie simulations due to stochastic noise. Could the authors comment on whether simulation times were sufficiently long to distinguish sustained oscillations from transient dynamics?

      We agree this is an important methodological point. We have clarified that simulations were run for 11.1 hours. Because nearly all the oscillators we found had periods below 100 minutes and most oscillators had periods below 12 minutes, almost all were observed for at least 6.6 cycles, which we believe is sufficient to distinguish sustained oscillations from transient dynamics.

      Locations of changes:

      Results, section 2, paragraph 2

      (6) While the manuscript alludes to broader applications of the proposed method, it would be beneficial to elaborate on these possibilities. Clarifying how this tool could be extended to other types of dynamical behaviours or biological questions would strengthen the impact.

      We have rewritten the final paragraph of the Discussion to give concrete illustrative examples of how the framework can be extended beyond oscillator design, including the discovery of more complex design principles and the design of synthetic multicellular circuits such as multi-cell type cancer therapies and morphogenetic patterning circuits.

      Locations of changes:

      Discussion, final paragraph

      (7) It remains somewhat unclear whether the contribution lies primarily in the novel application of reinforcement learning or whether there are methodological innovations within the algorithm itself. Are there existing tools with similar objectives, and how does this work improve upon them in terms of performance or capabilities?

      We have edited the Discussion to state more directly that the principal contribution of this work is the novel application of reinforcement learning to the problem of network topology design problem, rather than a fundamentally new RL algorithm. We also contextualize CircuiTree relative to other topology-search approaches and articulate where the use of RL provides a concrete advantage in navigating large combinatorial search spaces.

      Locations of changes:

      Discussion, end of paragraph 3

      (8) Further details on the training of the reinforcement learning algorithm would be appreciated. Was it trained a priori, and if so, how much data was required?

      CircuiTree is not pre-trained. Like AlphaZero, it learns exclusively from simulations performed during the search itself, with no externally curated dataset or a priori domain knowledge. We have clarified this in two places in our manuscript.

      Locations of changes:

      Introduction, beginning of first paragraph

      Results, section 1, paragraph 4

      (9) Some discussion on how computational time scales with increasing network size would be valuable.

      As mentioned above, we have added (i) an order-of-magnitude estimate of the size of the 5-node search space (≈2×10⁹ topologies) to make the current scale concrete; (ii) a power-law extrapolation of the current algorithm’s scaling that suggests ≈7 nodes is the largest tractable network without further improvements; and (iii) a Discussion paragraph projecting that AlphaZero-style scaling, combined with faster simulators, could plausibly extend the approach to ≈19-node circuits. The methodology for estimating the search-space size is described in a new Methods section.

      Locations of changes:

      Introduction, paragraph 4 (search-space size)

      Discussion, paragraph 1 (search-space size)

      Discussion, new paragraph 4 (extrapolated scaling and 7-node bound)

      Discussion, new paragraph 5 (projected scaling under deep-learning extensions)

      Methods, new section “Estimation of search space size”

      (10) The discussion of knockouts that dampen or restore oscillations is interesting. Was this analysis performed systematically, or were the examples selected through observation? Clarifying the methodology here would add rigour to the interpretation.

      We have clarified that the topologies highlighted in this analysis were selected by observing the fault-tolerant properties of a few topologies that exhibited high mutational robustness in our screen, rather than via a systematic analysis of the screen results. The text now states this explicitly so the reader can interpret the examples accordingly.

      Locations of changes:

      Results, final section, paragraph 3

      Other changes made by the authors

      In addition to the changes above, we have made several minor edits to improve clarity and presentation. We corrected spelling and formatting errors throughout, and improved the clarity of language in the Discussion section. No changes were made to the Figures, Tables, Algorithms, or Supplementary Information.

      We are grateful to the reviewers for their input and believe the manuscript has been substantially improved as a result. We hope the revised version meets with the editors’ and reviewers’ approval.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produces NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells and genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build a mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels, and constructed a mathematical model. The main conclusions are supported by layers of evidence from different assays.

      Weaknesses:

      Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

      Thank you for this assessment of our work. We have now added additional panels with statistical tests to the figures and included explanatory text in the figure legends.

      Regarding the cell-type-specific effects, we would like to offer a more nuanced view of this. Some of the genes we studied (NIT-GC1 and NIT-GC2) are only expressed in 4 cells in an organism of ~10,000 cells (as we demonstrated by HCR, immunostainings and the analysis of single-cell RNAseq data) and we knocked-down these genes (validated by antibody staining) with two independent morpholinos (to be able to rule out off-target effects). This is as cell-type specific as it gets. NOS is also expressed in a very limited number of cells. In the larval stages, we detected NOS expression only in the four INNOS cells and the pigmented eyes. We could previously show that UV avoidance is only mediated by the cPRC circuit (including INNOS) whereas phototaxis is exclusively mediated by the pigmented eyes Verasztó et al. (2018). Due to this behavioural specificity, we are therefore confident that the effects of the NOS mutations on UV avoidance are due to the lack of NOS from the INNOS cells and not the pigmented eyes. Besides UV avoidance, we have characterised the speed of ciliary swimming without a light stimulus, as well as phototaxis []. In addition, we also tested phototaxis and detected a reduced phototactic reaction in NOS mutants, but not after chemically inhibition of NO production (Figure 3B and 3F). The additional effects of NO are thus well documented in the paper and independent of the function of NO in the UV reaction. In addition, we have now did further quantifications and added new data (Figure 3 – figure supplement 4) to show that NOS mutants show normal lunar periodicity of sexual maturation, similar to wild-type animals. NOS mutants are viable and fertile, are feeding, building tubes and are mating as wild-type animals (these behaviours were not quantified here, we only show the data for lunar periodicity).

      Reviewer #2 (Public Review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      While I appreciate the intent of bringing together a large set of measurements from connectomics and calcium imaging in the framework of a model, the model seemed rather poorly constrained. How many parameters are in the model shown in Figure 6A? How many of them are well constrained by experimental measurements? The authors also don’t perform sensitivity analysis on the parameters of the model. And ultimately, the conclusion over the model in Figure 7 is somewhat trivial within the unitless construction: larger amplitude and longer duration stimuli lead to increased activation of the downstream neuron thought to lead to the downward swim behavior. I could imagine that a large family of models would arrive at this same result, and without units, there is no way to really test it with new behavioral experiments.

      We thank the reviewer for these comments. We have now thoroughly revised the model based on new experiments and carried out a sensitivity analysis. We are also more positive about the usefulness of the model though, for the following reasons.

      General usefulness of the model: With the model we can now reproduce all the qualitative dynamics of the circuit. The modelling also completely changed the way we were thinking about the system. For example, we needed to include a time-limited step in the cPRC transduction cascade leading to NOS activation to capture the time-invariance of the peak. In the future, we can also use this model to generate prediction e.g. about what the response to repeated stimulations can be. In the revised version, we included a new series of measurements of calcium dynamics in the cPRCs under varying duration and intensity of UV light. These important new results gave us further insights and necessitated a revision of the model.

      Unitless model: The model is indeed unitless, since already our input data from calcium imaging represent normalised data and the model was fit to these data. We would need a lot more information to build up a proper ground-up biophysical model (e.g. including capacitance, ionic concentrations etc.). This does not mean that the model is not useful.

      Sensitivity analysis: We have created a pipeline to carry out local and global sensitivity analysis and carried out a variance-based sensitivity and identifiability analysis. These data and code are included in the revised version. In the model, most parameters are identifiable. If we fix only two of the parameters all other parameters can be identified.

      Constrains and lots of possible models: In terms of the constrains, since we don’t have dimensioned quantitites, our constraints are bounds on the parameters. However, the structural identifibiality analysis tells us where we can identify parameters and can guard against sloppiness and overfitting. We only fitted the model on a subset of the data and can reproduce dynamics on a larger set of data. For important cellular interactions, the signs of the interactions are constrained. The model thus also allows us to rule out a large number of models - e.g. we could rule out a very simple model of progression from step to step as it was not possible to fit the data to such a model.

      We have updated the text to reflect these changes, e.g.: “Many of the couplings in our model are constrained (e.g. UV leads to INNOS activation) and e.g. reversing the sign of some of these couplings would not arrive at the same result. Several earlier variants of the model could not be fit to the data, The model is thus well constrained by our physiological experiments and the 106 circuit map.”

      Reviewer #3 (Public Review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      No doubt, the study has been conducted extensively. However, I have a few comments, please see below.

      Page 4: “In contrast, both two- and three-day-old homozygous NOS-mutant larvae showed a strongly diminished UV avoidance response (Figure 3A, B and Figure 3-figure supplement 1B, C).” Instead of using subjective terms like “strongly,” it would be more relevant to provide statistical values. However, I could not locate any means of statistical analysis on larval behaviour. Can the authors indicate the statistical values for all behaviour studies?

      We thank the reviewer for these comments. We have changed the wording and also added the results of statistical analyses to the figures and explanations to the figure legends.

      Page 5: “(D) Vertical displacement in 30 sec bins of wild type and mutant (NOSΔ11/Δ11 and NOSΔ23/Δ23) three-day-old larvae stimulated with 395 nm light from the side, 488 nm light from the top and 395 nm light from the top.” The error bars for WT are too long at the end of the experiment. It is not clear how the authors decided to use this time frame. Did the authors try carrying this out for an extended time period? How did the authors decide on 120 seconds as the time frame for exposure? Authors should provide data on larval behaviour for an extended time.

      The 120 seconds time frame of exposure takes into account the reaction time and swimming speed of the larvae as well as the size of the assay chamber (160 mm water height).

      By the end of a 120 stimulation many larvae tend to accumulate at the bottom of the chamber due to downward swimming and cannot be further tracked. This effect leads to higher variability in the data towards the end of the experiment in the wild-type batches. We have showed both continuous vertical data as well as the binned data. The 30 sec bin was chosen for convenience and for better comparison with our previous paper on UV avoidance behaviour (Verasztó et al. 2018)

      Page 13: “During the UV response, prototroch cilia beat slower than trunk cilia, resulting in a head down stable state (‘rear-wheel drive’). In contrast, during the pressure response prototroch cilia beat faster than trunk cilia, leading to a head-up orientation (‘front-wheel drive’). Testing this hypothesis will require biophysical experiments and mathematical modelling.” Authors should carry out ciliary beating analysis under UV light in the current study with NOS mutant larvae. Since the pressure and UV detection systems are closely related, comparing the difference in ciliary beating is important to 155 demonstrate this hypothesis. Further, did the authors check the Ca sensor GCaMP6s under pressure conditions?

      We thank the reviewer for this suggestion. We have carried out further experiments and analysed the ciliary beat frequency (CBF) of larvae exposed to UV stimulation. We added these data to Figure 3—figure supplement 3. The results (increase of CBF under UV in wild-type but not NOS mutant larvae) were quite surprising to us and falsified our initial hypothesis. We have rewritten the discussion to reflect this important new finding.

      The response of cPRC cilia to changes in hydrostatic pressure has been extensively documented in our recent paper on the mechanism of barotaxis in the Platynereis larva (see Bezares-Calderón et al., https://doi.org/10.1101/2023.02.28.530398).

      Page 18: “strips. One strip contained UV (395 nm) LEDs (SMB1W-395, Roithner Lasertechnik) and the other infrared (810 nm) LEDs (SMB1W-810NR-I, Roithner Lasertechnik).” Authors should test larval swimming behaviour at different wavelengths. Even though they are performed in previous work, the experiment with different wavelengths is necessary to be conducted in NOS mutant larvae in parallel with a control. This will confirm that NOS is principally associated with UV. Further, to demonstrate that this mechanism is associated with ciliary movement, authors need to provide this evidence.

      The diving reaction by non-directional light can only be induced by UV/cyan light, as we have shown previously, and it is mediated by a single UV-opsin photopigment (c-opsin1). The avoidance experiments can thus only be done with UV/cyan light. We also measured swimming 174 behaviour with 480 nm directional light to test phototaxis (Figure 3D and Figure 3—figure 175 supplement 1F).

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) The current introduction focuses on the nitric oxide signaling. It would be helpful for readers to have an introductory section about Platynereis larvae (a total number of neurons, etc) and rationales to use its nervous system as a model to study mechanisms of NO signaling and gating of visual response.

      We added an extra paragraph to give more detail on the number of cells in the larva and why Platynereis larvae are interesting to study to understand how synaptic and volume signalling interact. “The 3-day-old larva has over 9,000 cells classified into 202 neuronal and 92 non neuronal cell types (Verasztó et al., 2025). The synapticly connected subset of the cells in the body form a connectome of over 2,000 cells. Besides synapses, neurons in the larva also signal by volume transmission mediated by a rich repertoire of neuropeptides (Williams et al., 2017) and other modulators (Bauknecht and Jékely, 2017). The transparent and experimentally accessible Platynereis larvae could therefore inform how synaptic and volume signalling interact to mediate behaviour (Jékely and Yuste, 2024).”

      (2) The readers would want to know why these larvae swim downward in response to UV and upward to 480nm light. Is that for maintaining the certain depth from the surface of the water?

      It has been suggested that the ciliary photoreceptor circuit, which senses UV light, and the rhabdomeric photoreceptor circuit, which senses blue light, exchange messages with each other and that the two work together to form a depth gauge. By allowing larvae to swim at their preferred depth, the depth gauge influences where they end up when they become adults.

      (3) Are there splicing isoforms of NOS in the Platynereis dumerilii? If so, do antibodies and probes for in situ distinguish them? In Drosophila, truncated isoforms can inhibit the function of full length isoform, and therefore it was important to use methods to distinguish spicing isoforms.

      We did not identify any alternatively spliced forms of Platynereis NOS in our published transcriptome resources.

      (4) “INRGWs” acronym appears without explanation in the first paragraph of the results.

      Corrected: “the INRGWa neurons (cholinergic interneurons that express an RGW neuropeptide)”

      (5) In Figure 1C, it is difficult to see the projection patterns of individual cell types. Figure supplement 3 can be combined with the current Figure 1.

      We have moved one panel from Figure supplement 3 to the main Figure 1 (panel D) to show the INNOS projections more clearly.

      (6) Add more explanation about NOSp::palmi-3xHA reporter. Is it membrane-targeted reporter with the upstream promotor sequence of NOS? Or is it endogenous NOS that was tagged with palmi-3xHA?

      It is a membrane-targeted reporter driven by the upstream promoter sequence of the NOS gene. The construct is delivered by plasmid injection. We have clarified the description in the text and the figure legend (“The four apical organ cells, but not the eyes, were also labelled with a transiently expressed NOS-reporter transgene. This transgene contains the upstream promoter sequence of the NOS gene that drives a membrane-targeted palmitoylated tdTomato reporter (Figure 1F).” and “Expression of a membrane-targeted reporter driven by the NOS regulatory region (NOSp::palmi-3xHA-Tomato; magenta”). More details are in the Methods section under Transient transgenesis.

      (7) Show lack of anti-NOS immunostaining in NOS mutant to warrant specificity of the antibody. The subcellular localization of NOS in the dendritic arbors of INNOS is essential for the proposed model. The immunostaining image in Figure 4 -figure supplement 3D can be in the main figure. Is NOS also in the axons of INNOS?

      We further optimised the immunostaining with the NOS antibodies in WT and NOS mutant and added the staining data to the main Figure 1G and Figure 3—supplement 1. NOS was observed to be clearly localised in a region in the neuropil corresponding to the dendritic site of the INNOS cells. The staining was completely absent in larvae of both NOS knockout alleles. We have also added these explanations to the text.

      (8) “that was defective in NOS mutant (Figure 5I)” should be corrected as “Figure 5H”.

      We have restructured this part of the text and figure, the data from the Ser-h1 cells are now in Figure 5 - figure supplement 2.

      (9) Related to Figure 7, how do the larvae respond to a sequence of UV light (e.g. ten times 0.5s ON and 0.5 OFF) or ramping up/down UV light? Can the model make any predictions?

      We carried out new experiments and generated new model predictions. In Figure 6 – figure supplement 7 we show how changing the amplitude and duration of a single UV stimulation influences the response and the model output. In Figure 6 – figure supplement 8 we show how the model behaves when we provide repeated stimulations.

      (10) What are the knock down efficiency of NOS and NIT-GC morpholnio?

      We have now quantified the fluorescence intensity by immunostaining in wild-type and morphant larvae. We show the data for NIT-GC1 and NIT-GC2 in Figure 4—supplement 3F. The knock downs are very efficient. For NOS, we did not do morpholino experiments since we have two null alleles (Figure 3 – figure supplement 1).

      Reviewer #2 (Recommendations For The Authors):

      Why is the behavior of the control animals so different in panels B and C of Figure 3? One group reaches only 15 mm vertical position and the other reaches ~60 mm. Is this just batch variation? Am I missing an experimental variable here? If this is indeed batch variation, then some additional text and statistical analyses might help the reader interpret the behavioral data.

      Thank you for pointing this out. This was a mistake in the original Figure 3C of the unit on the Y axis. We have now corrected this. Additionally we have added statistical tests.

      Really, that’s the only, relatively minor issue I could find. This was a pleasure to read, and I learned a lot. Congratulations on an excellent study.

      Thanks a lot for these comments.

      Reviewer #3 (Recommendations For The Authors):

      Page 3, Figure 1 A: The authors reconstructed the cPRC circuit in 3-day-old larvae and detected NOS gene expression in 2-day-old larvae in Figure 1D&E. Can the authors provide a cPRC circuit reconstruction for 2-day-old larvae?

      We do not have a full connectome of the 2-day-old larva. We also show NOS gene expression in three-day-old larvae (Figure 1 – figure supplement 2).

      Page 4, Figure 2: NO produced by UV/violet stimulation to cPRCs: Including a diagram of the larvae would enhance reader understanding.

      We have added a schematic diagram to Figure 2.

      Page 4: Two Platynereis NOS knockout lines (NOSΔ11/Δ11 and Δ23/Δ23) using the CRISPR/Cas9: Could you direct me to the knockout conformation results for NOS knockout lines NOSΔ11/Δ11 and Δ23/Δ23?

      This is shown in Figure 3 – figure supplement 1 (genetic deletion and loss of antibody signal).

      Page 4: “Three day-old but not two-day-old NOS-mutant larvae also showed reduced phototactic behaviour, suggesting a function for NOS in the visual eyes that mediate three-day-old phototaxis” This sentence is unclear. Why is it only three days old but not two days old?

      2-day-old and 3-day-old larvae have very different type of phototaxis. 2-day-old larvae use their eyspots and show non-visual helical phototaxis. 3-day-old larvae use their visual (‘adult’) eyes for visual phototaxis. We have clarified this sentence: “Given that phototaxis in 1 and 2-day-old trochophore larvae is mediated by their non-visual eyespots (Jékely et al., 2008) and in 3-day-old nectochaete larvae by the visual eyes (Randel et al., 2014), these data suggest a function for NOS in the visual eyes (Figure 3D and Figure 3—figure supplement 1E).”

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Detailed experiment setup, if possible, schematic or real setup images would help to replicate the experiments.

      We have added a schematic diagram to Figure 3. 

      Page 5: “All trajectories start at 0 x and y position and time 0 corresponding to 10 sec after the onset of 395 nm stimulation from the side.” How did the authors determine the 10-second duration? Are there any reasons for this choice?

      In our experimental setup we had a limit of tracking individual larvae of approximately 40 seconds. We therefore restricted our analysis to a 10 sec pre and 30 sec post-stimulus interval.

      “(A) Swimming trajectories of wild type (WT, n=32) and NOS mutant (NOSΔ11/Δ11, n=26 and NOSΔ23/ Δ23, n=47) three-day-old larvae.” Authors keep shifting between 2-day and 3-day-old larval data in Figures 1, 2, and 3, causing inconsistency.

      Due to their elongated shape and active muscular contractions it is difficult to carry out calcium imaging experiments with 3-day-old larvae. These experiments were therefore done with 2-day-old larvae. However, we did our behavioural experiments with both 2- and 3-day-old larvae. We observed similar patterns of UV-avoidance behaviour, NOS gene expression, ciliary activity, and mutant phenotype. We are therefore confident that the cPRC responses and circuit activity are similar across these two stages.

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Why didn’t the authors present any statistics on the plots? A statistical test is required to prove that the difference is significant.

      We have added statistical tests to the data in Figure 3.

      Page 5: “Analysis of sGCs in Platynereis indicated that these genes are not expressed in any of the cells of the cPRC circuit (not shown and (Verasztó et al., 2017)).” Hence the data is relevant; please provide this data in supplementary.

      We re-analysed previous single-cell data (Achim et al., 2018) and found no sGC homologues detected in cPRC and INNOS. However, we detected an sGCβ subunit in the INRGWa cells. We added these data to the source data and amended the figure and the text. “Analysis of sGCs in Platynereis showed that these genes were not expressed in cPRC or INNOS cells, we only detected expression of an sGCβ subunit in the INRGWa cells (Figure 6B).”

      Page 10: “Diagram of the mathematical model with the components, interactions, parameters and equations used to model Ca dynamics.” It would be helpful to include a legend explaining the meaning of each term (e.g. what is K, Co, Cp etc) and arrow color. Additionally, the main text should provide detailed descriptions to aid understanding.

      We have added a new diagram of the model and its parameters to Figure 6. In addition, in the Methods section we have an extensive description of the model and all its parameters.

      Page 12: “This activated state is maintained for several tens of second.” Can authors point out the data?

      We point to the relevant figure now. “The high-Ca2+-state is then maintained for several tens of second (Figure 4A, B).”

      Page 13: “In the Platynereis circuit, our mathematical model indicates that the magnitude of the NO-dependent signal depends on the intensity and duration of the UV/violet stimulus.” Did the authors conduct experiments of various intensity and duration apart from modelling? If not, such experiments should be carried out to validate the mathematical model.

      We have carried out new calcium imaging experiment with varying duration and intensity of the stimulation. These experiments were key to the revision of the model because they clearly pointed to the time-invariance of the NO-dependent peak in the cPRCs. These data and their model fits are summarised in Figure 6 – figure supplement 7.

      References

      Bauknecht P, Jékely G. 2017. Ancient coexistence of norepinephrine, tyramine, and octopamine signaling in bilaterians. BMC Biology 15. doi:10.1186/s12915-016-0341-7

      Jékely G, Colombelli J, Hausen H, Guy K, Stelzer E, Nédélec F, Arendt D. 2008. Mechanism of phototaxis in marine zooplankton. Nature 456:395–399. doi:10.1038/nature07590

      Jékely G, Yuste R. 2024. Nonsynaptic encoding of behavior by neuropeptides. Current Opinion in Behavioral Sciences 60:101456. doi:10.1016/j.cobeha.2024.101456

      Randel N, Asadulina A, Bezares-Calderón LA, Verasztó C, Williams EA, Conzelmann M, Shahidi R, Jékely G. 2014. Neuronal connectome of a sensory-motor circuit for visual navigation. eLife 3. doi:10.7554/elife.02730

      Verasztó C, Gühmann M, Jia H, Rajan VBV, Bezares-Calderón LA, Piñeiro-Lopez C, Randel N, Shahidi R, Michiels NK, Yokoyama S, Tessmar-Raible K, Jékely G. 2018. Ciliary and rhabdomeric photoreceptor-cell circuits form a spectral depth gauge in marine zooplankton. eLife 7. doi:10.7554/elife.36440

      Verasztó C, Jasek S, Gühmann M, Bezares-Calderón LA, Williams EA, Shahidi R, Jékely G. 2025. Whole-body connectome of a segmented annelid larva. eLife 13. doi:10.7554/elife.97964.3

      Williams EA, Verasztó C, Jasek S, Conzelmann M, Shahidi R, Bauknecht P, Mirabeau O, Jékely G. 2017. Synaptic and peptidergic connectome of a neurosecretory center in the annelid brain. eLife 6. doi:10.7554/elife.26349

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank both editors and the three reviewers for their positive feedback on the revised manuscript. Following suggestions from Reviewer #1, we cite additional literature to clarify the use of ’fitness potential’ and we have revised the caption of Figure 3 to better explain how we generate the panel of mutants. We have also added a sentence to emphasize that the selection coefficient should match the time-scale of bottleneck effects in the evolution environment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community

      Many thanks to the reviewers and editorial team. This is an insightful and engaging set of reviews with a range of productive suggestions. We’re pleased that the strengths of this approach show through and that this can be a positive contribution to the community.

      We have implemented the large majority of suggestions and believe that the paper is greatly improved with them in place.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing': changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly, attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weaknesses:

      Some secondary interpretations (what is "unique" to age vs global anatomy) may go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think, fairly quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically, within-subject normalization). Relative scaling sharpens the spectral specificity of the spatial maps, while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g., SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age, and the authors spend some time trying to back out confounds/covariates. This section is handled transparently (in general, I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV), but it would be nice to find a measure that is independent of this: controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      Thank you for this concise summary of the work. We’re glad that the strengths of the approach come through clearly.

      This is a nice paper, and I have only a few concrete suggestions:

      (1) High-gamma 

      There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data, but still..) above 30-40 Hz, and these effects are the weakest anyway. Could you add an analysis (e.g., ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step? Or just note this point more clearly?

      Thanks for this suggestion. We agree that there is concern about eye movements for these gamma band analyses. The ICA preprocessing we conducted was relatively thorough, and we were able to remove EOG-related components from the majority of datasets, even though the experimental protocol involved eyes closed at rest. It is possible that some components were missed, but they would require more advanced labelling tools or manual intervention to identify.

      There are alternative automated ICA labelling tools that could be used, but these are either not suitable for MEG (ICLabel; https://labeling.ucsd.edu/tutorial, https://mne.tools/mne-icalabel/dev/api/iclabel.html) or optimised for data from CTF/4D systems (MEGNet; https://doi.org/10.1016/j.neuroimage.2021.118402). It is beyond our capacity to modify one of these tools for CamCAN for the current analyses.

      It was much more straightforward to rerun the analysis without the ICA step to see the impact of reintroducing all ocular artefacts removed from the v1 analysis. These results are shown in Supplemental section XXX and replicated below.

      We have added the following Figure 11 and text to the paper.

      Main Text

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged and low-gamma effects are reduced (see supplemental section A.3). “

      Supplemental materials

      “The age results may be contaminated in some way by residual cardiac or ocular artefacts that are not removed during preprocessing. Though ICA denoising was applied, it is possible that some artefactual components were not identified and removed from the dataset. The results at low frequencies and in the gamma range are most likely to be directly impacted by this contamination.

      To explore the impact this has on our analysis, we reran the core GLM effect of age on the data with no ICA artefact rejection at all, allowing all eye movements and heart rate components to remain in the data (Figure 10).

      This no-ICA analysis has three differences to the original in the publication. The low-frequency decrease with age is stronger and more widespread in the analysis that removes ocular artefacts with ICA. A large negative effect is visible in both analyses, though without ICA several frontal and temporal sensors no longer show significant effects. Similarly, the effect size of the low-frequency effect of age is substantially larger when ICA denoising is applied.

      The low-alpha effect was strongly reduced in the analyses that do not remove artefacts with ICA. A large central-occipital group of sensors shows an effect between 7 and 8.5 Hz in the ICA analysis, but this is reduced to a single sensor at 8 Hz when ICA is not computed.

      At high frequencies, the age effect in frontal sensors is larger with ICA, and the age effect in posterior sensors is larger without ICA, though the position and frequencies of significant effects are largely unchanged. The remaining effects in the high alpha and beta ranges are unchanged by application of ICA.

      Overall, ICA either improves the estimation of age effects (low-frequency, low-alpha, high-gamma) or has negligible effects (high-alpha, beta). Only the posterior low-gamma effect is reduced by ICA. Together, we take this as evidence that our core results are robust to interference by eye movements and that the ICA denoising is working effectively to reduce noise in the analysis.”

      (2) GGMV confound control 

      Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation, you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      This is an interesting area with some tricky interpretation. Thanks for the nudge to help us clarify further.

      We have added the following text and Figures 17 & 18 to the paper to clarify these points.

      Main Text Section 2.8

      “It is important to note that the correlation between age and GGMV does impact the interpretation of the GLM results, but does not prevent the model fit. We explore the model validation and diagnostics in detail in Supplemental Section A.6. In brief, the model is able to separate the unique contributions of age and GGMV. However, the correlation between factors leads to an inflation in the standard error of the estimates. A hypothesis test on these estimates is valid. However, the inflated variance reduces our ability to detect statistically significant partial effects.”

      Supplemental Section A.6

      “There is a strong correlation between age and Global Grey Matter Volume (Pearson’s r=-0.75). The shared variance arising from this collinearity adds nuance to the interpretation of the results, which we explore in more detail in this section.

      Firstly, the sum-square residuals for the group-level model fit including both Age and GGMV are shown as a function of frequency in Figure 17. We see that the residuals broadly follow the overall distribution of variance in the data, peaking at low frequencies and in the alpha range. These frequency ranges are where the strongest signal is visible, but also the highest variability between participants. We would expect that the group model would not perform so well in the points of greatest variability. Importantly, though the residuals are relatively high in the alpha, this is still in the context of a very well-performing model with R2 values of around 80%. “

      “Secondly, the correlation between age and GGMV is not inherently problematic for the GLM, but it does add complexity and nuance to the interpretation of the results. Some additional model validation statistics are shown in Figure 18. “

      “The singular value spectrum of the design matrix indicates whether a design is low-rank, the smallest singular value in this case in 0.37 which indicates that there isn’t a rank deficiency which would prevent us from estimating the model. Though we can estimate the model, correlated regressors can reduce its efficiency. The variance inflation factors for the joint AGE-GGMV model are above 1 for both parametric regressors, indicating that the standard errors of their estimates are inflated. The VIF of 2.35 indicates that the standard errors of this joint model are around sqrt(2.35) = 1.533 times greater than they would be in a separate or uncorrelated model. Though there is no hard rule for this, the literature generally suggests that a VIF above 5 (or sometimes 10) indicates severe multicollinearity.”

      Including additional regressors in the model can change the age estimate by ‘partialling’ out the variance that can be attributed to the other variables and by inflating standard errors. In this specific case of GGMV, the partialled estimates are reduced heterogeneously across space and frequency, and the amount of inflation is at a tolerable level. Overall, the regression is able to separate the unique effects of age and GGMV, at the cost of this inflation in the associated standard errors.”

      Reviewer #2 (Public review):

      This paper describes the application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is a discussion of the limitations of this approach, in comparison to other methods.

      Thank you for this summary and the thoughtful review. The emphasis for this paper is exactly on the points you highlighted, and we’re glad that this has come across well. We agree that a broader comparison to other methods would be a useful addition and have included a series of additional discussion points to address this.

      We will take the opportunity to reclarify that our intention for this method is not to replace other, more complex or targeted approaches, but to establish a more generalisable foundation for their development. We’re fully supportive of other approaches and are working on their application ourselves. On reflection, this was not clear enough in the first submission, and we have added the following text to the introduction to clarify.

      “Reporting of whole-head and full-frequency spectra of effect estimates would make it straightforward to aggregate across studies and eventually enable identification of sub-threshold effects that may be missed in single analyses but are consistent across studies. We argue that this approach provides a generalisable foundation that can support more complex analyses with a frequency component (such as aperiodic slopes, burst detection, and dynamic functional networks) that require more researcher degrees of freedom.”

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels.

      In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands/oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, which explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e, models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      This is an important point, and we are in complete agreement about the importance of giving a balanced description of how this approach fits within the broader literature. We have added the following paragraph to the discussion on limitations to lay this out more clearly.

      “Our approach promotes an exploratory and unbiased approach to quantifying the age effect on neuronal power spectra, which is intended to complement more focused analyses. This has the benefit of reducing researchers’ degrees of freedom and of being broadly generalisable. These come at the cost of a loss in specificity and in sensitivity. Our approach does not specifically quantify features derived from the power spectrum, such as alpha-peak frequency or the aperiodic component of the spectrum. These features are mixed into our full-spectrum estimates but not directly quantified. Thus, they can be challenging to interpret from our approach. Secondly, the mass-univariate approach suffers from a potential loss in sensitivity compared to results that aggregate across spatial or spectral regions that contain consistent results. Where a region or frequency band of interest can be supported from the literature, an approach focusing on a single region has the benefit of reduced noise by averaging estimates from a larger range of observations. Finally, models that consider the whole shape of the spectrum [Donoghue et al. 2021] would also be able to combine information across a range of frequencies rather than depending on a single frequency bin for each estimate. These models have the additional benefit that their parameters are often directly interpretable as features of interest, such as spectral slope or peak frequency.”

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Figure 2A, but their approach cannot test this directly (e.g., in terms of plotting effects of age on frequency, magnitude, and width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      We agree that this point might be misleading within its own figure and have moved the result to the supplemental material with a more lightly phrased wording in the main text. Figure 3 on effect sizes has moved to Figure 2, and a new Figure 3 shows the quadratic effect of age as suggested later in the review.

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged, and low-gamma effects are reduced (see supplemental section A.3).“

      This supplemental section contains some additional content relevant to the next comment.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      As above, we agree that this can be misleading and that our intention got muddled in the heading and writing of this subsection. We are not intending to argue that there are definitely two distinct and different effects. We clarify this in the main text of our first version, but not clearly enough:

      “The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak”.

      We choose to present the first paragraph of this section as a discussion of two effects because this is what is already reported in the literature on changes in alpha power with age. Whilst the decrease in alpha frequency with age is well replicated, the decrease in alpha power is less consistently reported, and the literature shows a highly variable picture of the spatial pattern of this effect. We do not claim a strong dissociation based on these results, but both are reported within the literature.

      Our intention is to provide some clarity to this literature by taking a step back and looking at the ‘lay of the land’ in an unbiased way. With this approach, we see different effects of age on magnitude within a canonical alpha range that are completely separated in frequency. With this perspective, it is feasible that different publications taking different regions of interest, frequency band definitions, and processing options could lead to mixed reports of increases and/or decreases in alpha power with age.

      When single papers that take focused but inconsistent approaches report inconsistent results, we would argue that our approach leads to a substantial interpretational gain on the level of collections of papers in the literature.

      We have added the following content to supplemental section A.1 to clarify and Table 3 illustrates the issue. We have retained the paragraph discussing the possibility that these two effects combine to represent a shift in frequency of a single peak.

      Main text section 3.1, replacing the section on ‘Two dissociable effects…’

      “Reconciling conflicting reports of the ageing effect on alpha power.

      The literature exploring how ageing changes alpha power is heterogeneous. Papers that report results in resting alpha power, either from a canonical band or from an individual peak frequency, include reports of a variety of contrasting age effects. This includes positive correlations with age [Rempe et al., 2023, Stier et al., 2023], negative correlations [Thuwal et al., 2021, Lodder and van Putten, 2011, Medrano et al., 2025, Park et al., 2024], both positive and negative effects separated by space [Hoshi and Shigihara, 2020, Pathak et al., 2022], or null results when correcting for individual frequencies and aperiodic slopes [Scally et al., 2018, Merkin et al., 2023] (see supplemental section A.1 more detailed summary). There is broad variability in methodological approaches, which could account for the variety of results. Stier et al. [2023] suggest that analyses in sensor or source space may lead to different effects. Critically for our work, this variability also prevents formal aggregation of results and meta-analyses that could clarify the picture.

      We have proposed that, by taking a step back and tolerating a reduction in sensitivity, we can map out the whole-head whole-frequency structure of the age effect with minimal researcher degrees of freedom and bias. Investigating age effects as a complete spectrum shows that two contrasting age effects on alpha magnitude coexist in close proximity in space and frequency: an increase with a small effect size in central occipital sensors around 7-8.5 Hz and a decrease with a large effect size across a broad set of occipital, temporal and frontal sensors between 9.5-12.5 Hz. Different data samples and different data analysis choices, particularly the selection of regions of interest, source reconstruction, sensor normalisation, or correction of aperiodic components, might emphasise one effect or the other in each analysis. Whilst we have not conclusively explored all possible variants of these analyses, we have provided a framework that would allow future studies to perform formal comparisons and meta-analyses to resolve this bottleneck.”

      “Relationship between effects on alpha power and alpha individual frequency.

      The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak (Seen qualitatively in Figure 1A). This change in alpha peak frequency is highly replicable [Cesnaite et al., 2023, Dustman et al., 1993, Sahoo et al., 2020, Scally et al., 2018, Pathak et al., 2022, Zibrandtsen and Kjaer, 2021] and is a highly predictive spectral marker of ageing [Stier et al., 2024]. Decreases in alpha peak frequency have been linked to a decline in cognitive performance in healthy ageing [Cesnaite et al., 2023, Finley et al., 2024] and MCI [Garc´es et al., 2013, L´opez-Sanz et al., 2016, Puttaert et al., 2021].

      This compelling possibility that the age effect on alpha is a shift in a single peak is complicated by strong evidence for presence of multiple alpha peaks within individuals [Lodder and van Putten, 2011, Chiang et al., 2011, 2008, Klimesch, 1999], with distinct generators and functional relevance [Sokoliuk et al., 2019]. A complex pattern of changes in power, frequency, and spatial distribution likely underlies age-related change in alpha oscillations. Future work will need to explore all three features at the individual level to clearly illuminate the change.”

      Supplemental section A.1

      “Part of our motivation for this method is that variability in the methodological choices in different publications makes it difficult to aggregate varying results across the literature. For example, though a decrease in alpha peak frequency with increasing age is reported highly consistently, the effect of age on alpha power is much more variable. Table 3 shows a representative sample of publications over the last 20 years that report a change in alpha power with age (note that this is intended to be a representative rather than an exhaustive list). Over half of publications (9/15) report a decrease in alpha power, whilst the remaining publications report an increase (2/15), both increases and decreases (1/15), a decrease but only without correcting for aperiodic slope (1/15), a quadratic effect (1/15), and no effect (1/15). Stier et al. [2023] suggest that the choice of analysis space (sensor space or source reconstruction) is likely a source of discrepancies between studies.

      Critically, it is difficult to reconcile these findings with the information reported in the publications. For example, even within the 9 publications that report a decrease, there is little correspondence in the spatial location of the effect. We argue that focused approaches cannot resolve this issue alone, as targeted analyses are more specific to each dataset and less generalisable.

      In the specific case of the mixed literature on change in alpha power with age, our results show that both effects are present and separated in frequency. It is feasible that the different methodological choices and datasets used by each study in our survey means that one or other of these two effects were emphasised. As a result, the literature may not be mixed in scientific terms, but that a rich pattern of results is obscured by methodological variability.”

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (eg of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      This is an important point, and we have added the following text to clarify, with one exception. Firstly, the normalisation we applied scales the whole spectrum linearly and would change the 1/f intercept but not the 1/f^alpha exponent.

      Main text section 3.3

      “Both absolute and relative power measures are used throughout the literature, but there is little consensus about their interpretation [Sandre and Troller-Renfree, 2026]. We focus on relative power for most of our results. By normalising each participant’s power spectrum by the sum across frequencies, we found that the results gained specificity in frequency band and reduced concern about wide intersubject differences in overall variance. Though this improved some analyses, relative power has important drawbacks. In particular, it can reduce sensitivity to effects that are broadly distributed across the spectrum and uses arbitrary scaling rather than meaningful physical units. We support calls in the literature to report both relative and absolute power measures [Rempe et al., 2023, Sandre and Troller-Renfree, 2026].”

      The authors should give more information on how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for the effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Artefactual components were estimated using standard tools in MNE python. Specifically:

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_ecg

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_eog

      We have clarified the text in the methods to make the overall process clearer.

      We believe that this process has been broadly effective, and we have rejected an average of 2.25 ECG components within each dataset. There remains a strong possibility that residual ECG artefact is present in the data.

      We have added the following text to methods section 4.2

      “Artefactual components relating to eye movements or the heart rate were automatically identified by correlation with the simultaneous EOG and ECG channels. ECG artefacts were identified using cross-trial phase statistics [Dammers et al., 2008] and an automatic threshold based on the sample rate of the data, as implemented in the mne.preprocessing.ICA.find_bads_ecg function in MNE Python. Between 0 and 3 EOG components were rejected in each dataset, with an average of 0.99 (standard deviation: 0.79) across all datasets. EOG artefacts were identified by correlation with the HEOG and VEOG channels, with a threshold set to r = 0.35, as implemented in the mne.preprocessing.ICA.find_bads_eog function in MNE Python. Between 0 and 5 ECG components were rejected in each dataset, with an average of 2.25 (standard deviation: 0.84) across all datasets. The continuous sensor data were then reconstructed without the influence of the components labelled as artefacts.”

      Similar to the response to Reviewer 1, we have not been able to rerun a more advanced ICA algorithm on the data but have repeated the analysis without any ICA to see if including all ocular and cardiac artefacts influences the results. This change does not introduce any new signal components to the results but does attenuate the low-frequency and low-alpha effects. With the assumption that our initial ICA analysis captured the majority of the largest ECG components, we are confident that our core findings are not compromised by cardiac artefacts.

      Please see the response to comments to Reviewer 1 for additional text in the manuscript on this point.

      The authors should clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      We have used the maxfilter files as provided by the CamCAN team and an equivalent implementation defined in OSL-ephys for the Oxford and Cambridge MEGUK data. We have clarified the text in section 4.2

      “All MEG data pre-processing was carried out using MNE-Python [Gramfort, 2013] and OSL-ephys [Quinn et al., 2022, van Es et al., 2025] using the OSL batch pre-processing tools. For the CamCAN data, we proceeded with analysis on post-maxfilter processed data provided by the CamCAN team. Briefly, the data were processed using AA [Cusack et al. 2015] with automatic bad channel detection (limited to 7 channels) with the origin set to the centre of a sphere fitted to the individual’s Polhemus head shape points. Maxfilter signal-space separation was performed with the temporal extension enabled (temporal window of 10 second and correlation threshold of r = 0.98). Head position was continuously estimated and compensated for during periods where the HPI coils were on. After maxfilter processing, head positions were translated into a default head position defined as a point relative to each individual’s origin in a head coordinate frame. Data with the full maxfilter processing and with the head position translation or movement compensation were extracted from the CamCAN database. An equivalent pipeline was implemented in OSL-Ephys and applied to the data from the MEG-UK datasets from Oxford and Cambridge. Files from the MEG-UK Nottingham dataset were processed with third-order gradiometry applied.”

      We found that head position in CamCAN does change with age. We have included the following in Supplemental section A.5

      “The head position of participants within the MEG sensor dewar is an important consideration that has the potential to change the signal-to-noise level of each individual data recording. There are significant differences in head position as a function of age in the CamCAN dataset in Y (front-back) direction indicating that older participants are seated further forward in the dewar than younger participants.”

      As a point of reference, this pattern is consistent with the largest change with age estimated using the un-normalised raw power spectra in Figure 5. A 1-95Hz region shows a change with age that is consistent with older participants sitting further forward in the dewar. This effect is absent in the relative power contrasts. 

      We have not fully explored this final point so have not included it in the main paper, but believe it is a useful addition to the discussion of the reviews.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, both reviewers are enthusiastic about your paper, indicating that it provides compelling empirical support for its key claims and that it represents a valuable theoretical advance for the field. They provide a number of comments that you may want to consider prior to finalizing the manuscript for publication. All the best, Redmond O'Connell

      Reviewer #1 (Recommendations for the authors):

      (1) It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple controls used, and the permutation count. I kept wanting to see this as I was reading to compare back and forth.

      This is a helpful suggestion, we have included the table in a new supplemental section which is referenced from the main text at the start of the results

      Main text section

      “A summary of all GLM analyses carried out in this work is available in supplemental section A.7.”

      With the following table included in supplemental section A.7

      (2) I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      We’re very glad that this section is working well and agree that the practical planning content should have been more constructive! We have added the following text as a guide

      Main text section 3.2

      “We propose the following steps as a practical guide for researchers looking to plan a data sample with a reasonable chance of correctly identifying a particular effect.

      (1) Define research question and identify previous results: Your research question must be defined in advance and well specified. There should be relevant data or literature that can be used to guide your decision.

      (2) Define the smallest effect size of interest and the decision criterion for the sample decision: It is critical to define the parameters of how you will make your sample size decision ahead of time. We recommend considering what the ’smallest effect size of interest’ [Anvari and Lakens, 2021] would be for your question. This is specific to your question and is about more than statistical significance. What effect size would indicate that there is a practically meaningful effect for the future literature to consider?

      (3) Estimate effect sizes to inform your decision: Either by aggregating reported statistics from the literature, or by dedicated processing of previous data, compute an estimate of the effect size. Effect sizes are only estimates, so it is important to compute confidence intervals around your estimate to get a measure of variability.

      (4) Compute power/precision curves assuming these results: Using the estimated effect sizes, compute power curves [Baker et al., 2021] that visualise the relationship between effect size, statistical power, and sample size.

      (6) Select the sample size that meets your pre-defined criteria: Using your definitions from step 2, and when considering the whole power curve, make a decision about what sample size would give you a reasonable chance of replicating the effect of interest.

      (7) Make note of any differences or deviations from this plan during your data collection: There are many practical reasons why your planned sample might not match the data acquired in practice. Such deviations from a plan are ok but should be acknowledged, and any mitigating steps explained [Lakens, 2024].

      (8) Document your process for inclusion in a preregistration or publication: Include details on how a sample size decision was made in your research outputs including preregistrations, preprints, and publications. These details will be useful for future researchers to understand your process and to implement their own.”

      Reviewer #2 (Recommendations for the authors):

      (1) Though they mention in the Discussion, the authors could have noted earlier in Section 2.1 that one does not need to assume a linear effect of age - one could use a polynomial expansion or even local splines within the same GLM framework. Indeed, it seems unlikely a priori that effects of age across the wide range of ages in the CamCAN data are linear for all frequencies.

      This is an important point, and one that was straightforward to implement in our model. Based on feedback on other sections, we have removed the previous Figure 2, moved Figure 3 on effect sizes to the second position, and added a new Figure 3 containing a model of the quadratic age effects. We believe that this is a substantial improvement in the paper.

      Main text section 2.3

      “(2.3) Quadratic effects of age

      The linear effect of age is a convenient and simple regression model. However, it makes a strong assumption that change with age is uniform across the whole age range. A second group-level GLM was computed with an additional regressor to quantify quadratic effects of age, which have been reported in the ageing literature [G´omez et al., 2013, Rempe et al., 2023, Stier et al., 2023]. The spectrum of t-values for the quadratic age predictor (Figure 3) shows significant effects for a U-shaped change with increasing age in the low frequency (1-5 Hz) range in central sensors. Significant inverted-U shaped effects are present in two frequency ranges in the beta band. A low-frequency beta effect is present in occipital and temporal sensors between 16 Hz and 20 Hz, whilst a second high-beta effect is present in central sensors between 24 Hz and 30.5 Hz. The low-frequency and low-beta effects overlap in space and frequency with linear effects, suggesting that the low-frequency change with age has both a linear increase with age and a U-shaped component, whilst the low-beta effect has a decrease with age and an inverted-U shaped component. The high-beta effect does not overlap in frequency with any of the reported linear effects. The effect sizes for quadratic effects range between Cohen’s F 2 values of 0.033 for low frequency to 0.076 for high beta and are generally lower than the effect sizes for linear change.”

      Main text section 2.4

      “(2.4) Sample size planning for effects of age on the neuronal power spectrum

      We use the 95% confidence intervals around the effect sizes to make recommendations for future sample sizes for future samples that plan to replicate these results. The observed power calculations have no bearing on the interpretation of the present results. Instead, they should be used as a general guideline for planning future studies. Table 1 gives a full summary of the peak statistics and future sample size range for the six age effects identified in Figure 1B. These results have implications for future sample planning for resting-state electrophysiology studies of ageing. Using the upper bound of the sample size estimates as a conservative estimate, the linear changes with age that have relatively large effect sizes would have well-powered replications with sample sizes of around 50-60 participants. However, the smaller linear effects and all the quadratic effects would require samples of 200 or more participants to have the same probability of detecting the effect if it is indeed present (Table 1). This indicates that study samples should be planned with the smallest effect of interest in mind and that ageing effects in different frequency bands may not all be well powered within the same sample.”

      Discussion section 3.1

      “Linear increase and inverted-U effects in the beta band. The literature consistently reports an increase in low-beta power with age [Gomez et al., 2013, Heinrichs-Graham and Wilson, 2016, Heinrichs-Graham et al., 2018, Hubner et al., 2018, Koyama et al., 1997, Rempe et al., 2023, Stier et al., 2023, Veldhuizen et al., 1993, Xifra-Porxas et al., 2019]. We observed significant inverted-U-shaped quadratic effects of age in two frequency ranges in the beta band, a posterior low-beta component (centred around 18 Hz) and an anterior high-beta component (centred around 25 Hz). This is broadly consistent with reports of quadratic ageing effects in the beta band in the literature [Rempe et al., 2023], though, to our knowledge, our results are the first to suggest a separation of effects into different parts of the beta range. These spectrum changes may be associated with age-related changes in underlying bursting dynamics [Brady et al., 2020, Power et al., 2023].”

      “Methods section 4.7

      Age plus quadratic age models: To move beyond linear change and explore U and inverted-U shaped effects of age, we fitted a group model with three regressors. One constant term, one z-transformed age regressor, and one z-transformed quadratic age regressor (age − mean(age))<sup>2</sup>.”

      (2) The authors could point out an obvious extension of their approach to mixed-effects models, which could properly model within- and between-participant effects (repeated measures), e.g., for longitudinal effects of ageing, GAMMs, etc, including hierarchical linear models that combine trials and participants in the same model, and potentially model trial-specific/stimulus effects, etc.

      This is an important point; we have added the following paragraph to the discussion section.

      Discussion section 3.5

      “Future extensions

      A clear future extension for this work is to formally incorporate estimates of within-subject variability into a mixed-effect model. These powerful models would enable modelling of both fixed effects and random effects, allowing researchers to account for variation within individuals over time and between individuals. LMMs also provide improved approaches for handling missing observations and unbalanced designs, making them especially useful for longitudinal and hierarchical data. A second expansion of this work could use Generalised Additive Mixed Models (GAMMs) to model non-linear relationships using smooth functions while also accounting for random effects. This may allow for more realistic representations of complex patterns of change across age. Linear Mixed Modelling comes with a substantial increase in researcher degrees of freedom and can be challenging to implement and report accurately [Meteyard and Davies, 2020]. We have used fixed-effects modelling in this work in line with our objectives of maintaining generalisability and minimising researcher degrees of freedom.”

      (3) Do the authors want to comment on why effect sizes in Figure 4Ci are so much higher for the Oxford sample? I know the authors' main point is that smaller samples lead to more variable estimates of the true effect size, but there are also other interpretations for these sample differences, e.g., recruitment differences, scanner differences, etc. Have the authors considered implementing empirically Bayesian approaches (like COMBAT, a type of mixed-effects model) to adjust for site differences in both offset and scaling, which I think should be a fairly simple extension of GLM-spectrum?

      We agree that our explanation of here is somewhat lacking. To be transparent, we have thought long and hard about this difference and cannot identify a clear reason why this difference is so striking. There is no apparent difference in the other covariates, such as head position, age, or gender, which could explain why the Oxford dataset has larger effect sizes in this specific frequency range (the results in the low-frequency and alpha ranges are consistent).

      We have added the following to the main text to highlight dedicated data harmonisation strategies that could be applied in this situation.

      Main text section 2.5

      “This may arise from relatively poor estimates of the population level variability from smaller data samples. It is possible that more systematic differences in the participant sampling, recruitment, and data acquisition process could contribute to between-site differences. At present, our analyses can show that the core effects of ageing are replicable across several datasets, though we have not formally combined these datasets into a single analysis. Formal methods for data harmonisation, such as ComBat [Johnson et al., 2006], could correct for additive and multiplicative differences in data scaling across sites to improve site comparisons.

      (4) There are a number of formatting problems - at least in the PDF provided to reviewers - for some symbols, e.g., "below ¡7 Hz" (I presume "j" was originally "~" or something?). Also, sometimes a figure is cited by a single, hyperlinked number, without the word "Figure".

      (5) Line 402: "influence" = "influenced"

      (6) Line 431: date for Gelman & Loken?

      Thank you for highlighting these, we have done a thorough proofread and fixed a large number of spelling and formatting issues.

    1. Author response:

      We thank the editors and the two reviewers for their careful evaluation of our study, as well as for their positive assessment of the research question, model reproducibility, and the results concerning temporal-order representations. We also agree with the core issues raised in the reviews: the current manuscript has not yet sufficiently distinguished descriptive changes in recurrent dynamics from the functional mechanisms underlying the behavioral advantage; the behavioral effect itself is small and close to the performance ceiling; the phase analyses require more stringent controls; and the specific role of short-term synaptic plasticity in the rhythmic advantage has not yet been directly tested. We plan to add the corresponding analyses in the revised manuscript and to temper several mechanistic claims.

      First, we would like to clarify three aspects of the study design and analysis pipeline. First, the rhythmic and arrhythmic conditions were not performed by two separately trained groups of networks. Each network was jointly trained using balanced batches containing rhythmic-match, rhythmic-non-match, arrhythmic-match, and arrhythmic-non-match trials. Therefore, the behavioral differences were compared within the same independently initialized network. We will revise the relevant descriptions in the Abstract, Results, and Methods to make this joint-training procedure more explicit. Second, Figures 2 and 3–4 used different dimensionality-reduction approaches because they addressed different analytical questions. In Figure 2, dPCA was applied to time-aligned population activity to separate task-related variance into temporal, ordinal-position, and stimulus-related components, allowing us to examine how rhythmicity affected each representational component. In contrast, Figures 3 and 4 used standard PCA to construct a low-dimensional population signal for spectral and phase analyses without explicitly demixing task variables. We will clarify this distinction and the rationale for the two analysis pipelines in the revised Methods. Third, STSP was not introduced as an auxiliary module. Previous theoretical and computational studies have suggested that STSP can contribute to working-memory maintenance (Mongillo et al., 2008; Masse et al., 2019). Based on previous studies, we incorporated STSP alongside persistent neural activity to examine how synaptic and consistent neuronal activities jointly support sequential working memory. Under the current architecture and training settings, our ablation experiments showed that RNNs without STSP had difficulty successfully learning the task. We will emphasize this result and further distinguish the roles of STSP in task learning, delay-period maintenance, and the rhythmic advantage. We will also explore whether vanilla RNNs can successfully learn the same task under alternative hyperparameter settings.

      Our core hypothesis is that temporal regularity improves the encoding of sequential information by organizing recurrent population dynamics and phase structure during the encoding period, and that this organization subsequently influences information maintenance during the delay period and behavioral performance. The current results establish a relationship between temporal regularity and encoding-period dynamics, but direct validation of how this organization influences subsequent working-memory maintenance remains insufficient. In the revision, we will focus on strengthening the link between the encoding and delay periods and test whether population dynamics and phase organization during encoding are associated with subsequent information maintenance and behavioral performance.

      To provide a more direct view of the network dynamics, we plan to add intuitive visualizations of network activity and phase structure, allowing readers to evaluate the reported temporal organization with less dependence on dimensionality reduction, filtering, and complex statistical processing.

      The current behavioral advantage of approximately 0.4 percentage points is small, and network performance in both conditions is close to ceiling. Following the reviewers’ suggestions, we plan to evaluate the stability of this effect across learning, random initializations, and representative hyperparameter settings, and to examine whether the rhythmic advantage becomes more pronounced under higher memory load when ceiling effects are reduced. We will also avoid equating statistical significance directly with functional importance.

      Although the manuscript already includes decoding and perturbation analyses of neural activity and synaptic efficacy during the delay period, we will perform additional delay-period analyses to more directly examine whether the temporal organization established during encoding is associated with subsequent working-memory maintenance. These analyses will help distinguish the contributions of rhythmic input to stimulus encoding and memory maintenance.

      We will also test the role of STSP more directly. The current delay-period shuffle results show that synaptic efficacy makes a functional contribution to delay-period maintenance, but this does not yet demonstrate that STSP specifically supports the rhythmic advantage. We will further illustrate the roles of STSP during encoding and delay-period maintenance. In the revision, we plan to increase the number of independent networks in the perturbation analysis and test the interaction between temporal regularity and STSP disruption.

      Finally, we will further tighten the conceptual framing of the manuscript by more clearly distinguishing stimulus-driven phase organization from self-sustained oscillations, and by clarifying that the analyzed signals are model population signals derived from firing-rate population activity rather than LFP or EEG field potentials. Where the current evidence is insufficient to support interpretations in terms of entrainment or self-sustained oscillations, we will adopt more cautious terminology and revise the corresponding conclusions accordingly. We will also clarify the relationship between our phase-organization results and the findings of Liebe et al. (2025) in the revised Results.

      Overall, the revised manuscript will more precisely frame the study around how temporal regularity improves sequential working memory and how this behavioral advantage is associated with the organization of recurrent population dynamics and synaptic-state representations during encoding. After completing the additional training and analyses, we will also make the corresponding model weights, training configurations, and necessary analysis files publicly available to improve reproducibility.

      We again thank the editors and the two reviewers for their detailed and constructive comments. We will revise the manuscript carefully on the basis of these suggestions.

      References

      Mongillo, G., Barak, O., & Tsodyks, M. (2008). Synaptic theory of working memory. Science, 319(5869), 1543–1546. https://doi.org/10.1126/science.1150769

      Masse, N. Y., Yang, G. R., Song, H. F., Wang, X.-J., & Freedman, D. J. (2019). Circuit mechanisms for the maintenance and manipulation of information in working memory. Nature Neuroscience, 22(7), 1159–1167. https://doi.org/10.1038/s41593-019-0414-3

      Liebe, S., Niediek, J., Pals, M., Reber, T. P., Faber, J., Boström, J., Elger, C. E., Macke, J. H., & Mormann, F. (2025). Phase of firing does not reflect temporal order in sequence memory of humans and recurrent neural networks. Nature Neuroscience, 28, 873–882. https://doi.org/10.1038/s41593-025-01893-7

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

      We agree with the reviewer that neurotransmitter trafficking studies would be interesting, but they would presumably require the use of acetylation mimic mutants (Lys-to-Gln mutations), which we would want to validate by comparison to our semi-synthetic proteins with authentic AcK. Such experiments are planned for a follow-up manuscript, and we will investigate the reviewer’s suggested experiment at that time. Thus, we have not modified the manuscript to address this issue.

      Subsequently, they measure the aggregation speed of the variants in seeded aggregation experiments with preformed fibrils (PFFs) from WT aSyn, and conclude that acK at positions 12, 43, and 80 yields slower aggregation. They reach similar conclusions when measuring seeded aggregation in primary cultures. As far as I understand it, the seeding experiments in cells use seeds that are assembled from partially acetylated alpha-synuclein, but that are made of non-acetylated wildtype alpha-synuclein, and the alpha-synuclein that is endogenous in the cells is also non-acetylated (or at least not beyond what happens in these cells at endogenous levels). It is therefore unclear how the cellular seeding experiments relate to the in vitro aggregation assays with (partially) acetylated substrates.

      We understand the reviewer’s concerns and have modified the manuscript to clarify that the method of in vitro seeding really reports on the impact of acetylation on the elongation phase of aggregation. We have also clarified that this is different than the role that acetylation plays in seeding cellular aggregation with pre-acetylated fibrils. We note that having the monomer population acetylated in cells presents technical challenges that might also be addressed with Gln mutant mimics, and we plan to pursue such experiments in the follow-up manuscript described above.

      Anyway, both aggregation experiments ignore that the structures of aSyn filaments in Parkinson's disease (PD) or multiple system atrophy (MSA) are different from those formed in these experiments, and that, therefore, the observed aggregation kinetics are likely irrelevant for the speed with which disease-relevant filaments form in the brain.

      Finally, the authors describe the cryo-EM structure of mixtures of acK80:WT aSyn filaments, which are predominantly made of WT aSyn, with a previously described structure. Filaments made of only acK80 aSyn have a modified arrangement of this structure, where the now neutral side chain of residue 80 packs inside a hydrophobic pocket. The authors discuss differences between the acK80 structures and those of other structures from in vitro assembled aSyn filaments, none of which are the same as those observed from PD or MSA brains, nor are any attempts made to transfer observations from the in vitro experiments to the structures of disease. The relevance of the cryo-EM structures for human disease, therefore, remains unclear.

      The Conclusion on p.20 mentions an interesting and valuable result: the authors used the acetylated recombinant proteins to determine the extent of acetylation within human protein samples by quantitative liquid chromatography MS (SI, Figures S41-S49). Their conclusion is that "The level of acetylation was variable - no clear trend was observed between healthy control and patients - nor between patients of different diseases (SI, Table S4, Supplementary Data 1)" This result implies that acetylation of aS is not directly related to its pathogenicity, which again adds doubts on the disease-relevance of the results described in the rest of the paper.

      We acknowledge the concerns raised in the above paragraphs and believe that they can all be addressed by clarifying our purpose. The different fibril polymorphs adopted in PD and MSA are likely the result of an interplay of many PTMs and non-proteinaceous cofactors. Therefore, we are not necessarily trying to claim that our AcK80 fold is populated in health or disease, but that by driving Lys80 acetylation, one could push fibrils to adopt this conformation, which is less aggregation-prone. A similar argument has been made in investigations of alpha-synuclein glycosylation and phosphorylation. Our results in Figure 9 imply that Lys80 acetylation could be increased with HDAC8 inhibition. We have revised the manuscript to make these ideas clearer, while being sure to acknowledge the limitations noted by Reviewer #1.

      Reviewer #2 (Public review):

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

      We have noted this shortcoming in revisions of our manuscript, but have not performed new experiments as we do not believe that using vesicles instead would change the conclusions of these experiments (that only AcK43 produces an effect, and a modest one at that).

      It would help the reader to have the experimental details (e.g., buffer, protein/lipid concentrations) for the different assays written in the figure legend.

      We have added additional detail to the figure captions.

      The authors use an assay consisting of mixing 10% fibrils + 90% monomer to investigate the effect of lysine acetylation on aS. However, the assay only probes fibril elongation and/or secondary processes. The current wording can be misleading, and the term aggregation could be replaced by seeding capacity for clarity. For example, the authors state that lysine acetylation at sites 12, 43, and 80 each inhibits aggregation, but this statement is not supported by the data. Instead, the data show that the acetylation at these sites slows down the fibril elongation and thus decreases the seeding capacity of aS fibrils. In order to state that lysine acetylation has an effect on aS aggregation, fibril formation, the author should use an assay where the de novo formation of fibrils is assessed, such as in the presence of lipid vesicles or under shaking conditions.

      As noted in our response to Reviewer #1, we have clarified which phase of aggregation we were investigating in our in vitro experiments.

      It is not clear from the EM data that the structures of the different lysine acetylated variants are different, unlike what is stated in the text.

      We feel that it is clear from structures in Figure 8 and the EM density maps in Figure S38 that the AcK80 fold is indeed different. Although the overall polymorphs are somewhat similar to WT, the position of K80 clearly changes upon acetylation, altering the local fold significantly and the global fold more moderately. We have added backbone RMSD calculations to quantify the differences in WT and AcK80 folds in Figure SX. They differ by ~5 Å in the fibril core region.

      Reviewer #3 (Public review):

      Weaknesses:

      The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      We understand the reviewer’s surprise and have edited the manuscript to clarify that the NCL yields were not unusually low, but were comparable to ncAA yields, and since it is significantly easier to scan AcK positions using ncAAs, we felt that ncAAs are the method of choice in this case.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Cryo-EM data processing: particles were pre-processed in cryosparc and ChatGPT-generated scripts. Are priors on tilt and psi angles defined through this procedure? And are segments from individual filaments kept strictly in the same half-sets for gold-standard estimation of resolution? Absence of psi and tilt priors will lead to worse refinements than otherwise possible. Worse, a mixture of segments from different filaments into the two half-sets may lead to overestimated resolution estimates.

      We thank the reviewer for their thoughtful analysis of our cryo-EM data processing approach. In response, we have modified our approach and added additional commentary on processing to the Materials and Methods section.

      “CryoSPARC (.cs) files were then converted to RELION STAR files using the PyEM csparc2star.py script. Following this data conversion, a custom Python script developed with ChatGPT precisely determined the start and end coordinates of each fibril. This script processed the csparc2star.py star file output to create coordinate pairs based on the cryoSPARC Fibril ID, notably without transferring the original tilt and psi angles from CryoSPARC to RELION. The output of this custom script served as the input for the autopick RELION extraction step.

      To handle curved fibrils, a special segmentation strategy was implemented: the script traced the coordinates along the fibril, generating a new start and end coordinate pair for individal segments, with the segment's end coordinate assigned either after spanning 10 particles (around 50 nm in length) or when the end of the fibril was reached. This resulted in shorter, straighter segments for processing. Finally, the createAutopick function from cryolo_boxmanager_tools.py in crYOLO was utilized to generate a STAR file linking the final particle coordinates to their movie files, which was then used to perform particle extraction in RELION. The standard RELION image processing pipeline then followed, including 2D classification, refinement, and 3D classification.

      This approach resolves the issue with curved fibrils, but also results in all picked fibrils being 50 nm or shorter in RELION. Consequently, segments from the same fibril receive unique fibril IDs in RELION and may be split into different half-maps. Although this avoids random splitting of individual particles without regard to their origin from the same fibril, it could still cause an overestimation of resolution.

      To assess any potential overestimation of resolution, another script was created to reassign the fibril ID of each particle in the RELION STAR file to that of the closest particles in the cryoSPRAC data, effectively ensuring that all particles from an individual fibril have the same fibril ID and are not assigned to different half-maps. This revised STAR file was then used as the input images STAR file for 3D auto-refinement in RELION. A comparison of maps before and after reassigning the fibril IDs shows negligible effects on the resulting structures and on their estimated resolution, indicating that the original procedure did not results in any significant overestimation of the resolution.”

      Author response image 1.

      RELION Cryo-EM maps before (gray) and after (yellow) fibril ID reassignment for (A) WT-A (B) WT-B (C) <sup>Ac</sup>K<sub>80</sub>-A (D) <sup>Ac</sup>K<sub>80</sub>-B, showing that the new processing approaches did not significantly change the fibril structures.

      (2) The 25% acK80 structure in S52 is understood to be wildtype-only. The authors mention in the main text that this is because of the strong density for the K80 side chain. An additional argument would be that an acK80 would leave an unshielded negative charge on the neighbouring E46, as K80 and E46 form a salt bridge in this structure.

      We appreciate the reviewer’s idea and have included a comment on the salt bridge impact.

      (3) It would be valuable to include side views of the density for all reported reconstructions to assess to what extent the beta-rungs are separated.

      The requested side views have been included in Figure SX.

      (4) Methods sections should be moved into the main text of the paper.

      This Materials and Methods portion of Supporting Information has been moved to the main text.

      Reviewer #3 (Recommendations for the authors):

      (1) We suggest removing "all" from the manuscript title as this claim might not hold up in the future.

      We understand the reviewer’s concern. Our title was meant to imply “all currently known” rather than “all” forever, but as this wording is awkward, we have deleted “all” as suggested.

      (2) The white font in Figure 1A is sometimes hard to read, especially on the yellow-green background between amino acids 70-80.

      We have changed this to black font.

      Figure 1 panels C-E are not referenced in the manuscript text?

      References to the Figure 1 panels have been added.

      In panel C, a structure is predicted, but based on what data? Why is there both a small and a big structure in panel C?

      Explanations of the Figure 1C images have been added to the caption.

      (3) Typo ε-acetyllysine -> Nε-acetyllysine

      This has been corrected throughout.

      (4) You report solubility issues during NCL, and these are typically alleviated by the use of chaotropes during ligation. Please specify what you mean by "standard NCL conditions" and include parameters such as guanidine concentration, pH, concentrations, volumes, and temperature.

      These details have now been added to the methods section.

      (5) According to Figure 2D, the desulfurization was incomplete. Please explain.

      The small peak observed next to the product peak is an adduct with sinapic acid (+206 Da), the matrix mixed in for MALDI acquisition. We do not see a sign of +32/64Da peak which would correspond to incomplete desulfurization

      Author response image 2.

      (6) Please include sequences of your constructs for ncAA mutagenesis. Especially, which intein was attached at what position. This is important for other groups to fully understand the production of acetylated aSyn variants.

      The full DNA sequence has been added to Supporting Information. This plasmid has also been reported previously in the referenced publications.

      (7) You state ncAA mutagenesis "yielded 0.11-1.5 mg" aSyn. I guess this is per-liter culture expression medium?

      This has been corrected.

      (8) The resolution/DPI of Figures 4, 5, S14-16, and S18 should be increased. They look blurred compared to Figures 6 or S19.

      These figures have been updated.

      (9) You provided only summaries in Figures 4 and 5 because of space restrictions. However, it would be nice to have at least some selected individual experiments (WT, acetylation K12/43/80) right next to it without the need to switch to S14, S15, or S16.

      Select experimental data has been added to main text Figures 4 and 5 as requested.

      (10) Please increase the size of the microscopy images in Figure 6 (in Figure S19, it looks much better).

      This size of the images in Figure 6 has been increased.

      (11) The labelling A1-A3 in Figure S33 was somehow unclear to me.

      These three panels show different sections of the TEM grid, illustrating heterogeneity in the 25% <sup>Ac</sup>K<sub>12</sub> fibrils that was not observed for fibrils of the other acetylation variants. We have clarified this in the figure caption.

      (12) In Figure 9, you present preliminary results regarding site-specific deacetylation by HDAC8. These results are interesting, but considering the exploratory nature of this in vitro experiment, I suggest toning down the highly enthusiastic discussion of potential in vivo effects.

      We understand the reviewer’s concern and have mitigated the claims of potential impact from these in vitro results.

    1. Author response:

      Reviewer #1 (Public review):

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.   

      We agree with the reviewer that tanycytes that have been described in the third ventricle have not been reported to be myelinated. However, our study was conducted on the hippocampal formation that borders the ventral horn of the lateral ventricles. We will include images of the myelin-forming ependymal cells that we refer to as tanycytes  

      Additionally, the methodologies described in the manuscript lack clarity and controls.

      We will expand on our methods section and include controls.

      For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      Most experiments described in this manuscript were carried out on human brain. However, functional studies will have to be carried out on rodent brain. We will thus process rodent tissue as well to test whether our observed findings are consistent between human and rodent.  The animals used for these experiments were raised for bladder research and are wild type regarding neuronal and glial cells, particularly using the Cy3 channel. Utilizing these brains for our experiments has allowed us to test rodent tissue at both light- and electron-microscopic levels without having to sacrifice additional animals. The consistency of our findings between human and rodent brain further supports that the calcium indicator in the vascular system of these mice did not affect neurons or glial cells.

      Regarding the uptake experiment: When we initially discovered that myelin-forming macroglia form waste-internalizing glial canals within neuronal in spider brain it was unclear where the AQP4-immunoreactive cells were located. The somata of the myelin-forming cells lacked AQP4 immunoreactivity. Suspecting a synergistic interaction between the myelin-forming and AQP4 expressing cells whereby the myelinforming cells create the canal structure that sequesters waste from the neuron and the AQP4-expressing cells create a convective flow toward the waste-internalizing structures. However, unable to locate the somata of these cells, we submerged a freshly dissected spider brain with the attached surrounding tissue intact in physiological spider saline and slowly added blue vital dye solution to test which cells would internalize the dye. We then identified the (blue) cells in the lining of the dorsally located tubular system that we routinely detached from our brain preparations explaining why we were unable to locate these cells.  Immunolabeling of this tubular system revealed the cells that reside in the lining of this tubular system (see Author response images 1 and 2) the original (Figure 8) shows their long slender processes. Interestingly, this system is continuous with the stomatogastric system. 

      Author response image 1.

      Shows the proposed canal system in spiders

      Author response image 2.

      AQP4-immunoreactive cells in the spider primitive ventricular system that we localized due to similar uptake experiments we have conducted in mouse brain.

      We have utilized this method in mouse brain to test the validity of our postulation, that ependymal tanycytes internalize the presented fluorochrome from extracellular spaces and test which areas and structures may be involved in this uptake. As demonstrated in this experiment, the alveus, and a fine network of cell processes within the brain parenchyma show fluorescence, indicative that they internalize the fluorochrome from extracellular spaces. We used goat-coupled fluorochrome to further test with a FITCcoupled secondary antibody that the observed fluorescence is indeed due to uptake of the goat-coupled secondary antibody and not due to intrinsic autofluorescence. Control preparations lacked this fluorescence.  To further test our postulation that the uptake is indeed AQP4-mediated we have applied an AQP4-blocker, which showed a significantly reduced fluorochrome uptake compared to the controls without this blocker.

      As we state in the text, we are aware of the limitations of this experimental design, however, like in our spider experiments we consider these findings helpful as they likely show an overview of the cellular network in the hippocampus that governs waste-uptake and may help identify suitable target areas for similar studies on organotypic tissue cultures utilizing two-photon microscopy.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization. In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

      We respectfully disagree with this comment and hope that the inclusion of additional evidence together with the clear visibility of this canal system in the spider brain will encourage the reviewer to investigate this possibility themselves. We cannot ignore large amounts of amyloid beta-immunolabeled receptacles emanating from tanysomes in swell-bodies and declare them fixation artifacts, particularly when we demonstrate the expression of Presenilin 1 and APP in swell bodies. We furthermore encourage the reviewer to revisit myelinated cells in the brain in both depictions in the available literature and actual brain preparations. We have not been able to locate actual electronmicrographs of longitudinal sections through neurons that show myelination past the axon hillock at the EM-level consistent with our current understanding of myelination. The only depiction of this form of myelination we found were schematic drawings. We will include several new images that show such longitudinal sections through neurons that are easily obtained and we have numerous additional images that we are happy to share. In all our preparations (mouse, rat and human) the myelination pattern is consistent with the images we will depict in new figures 1 and 2.  Not to bring attention to this inconsistency would be dishonest scientific conduct.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      As mentioned in our response to reviewer 1 we have now included additional experimental evidence that demonstrates the myelinated ependymal cells and additional gene expression experiments. 

      This becomes particularly important because the manuscript repeatedly interprets Luxolpositive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serialsection analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

      We agree with the reviewer that it is important to correctly investigate and describe cellular structure. The first author of this manuscript is a 30-year veteran of published cellular ultrastructure and the three-dimensional reconstruction of cells and entire cell networks. We have spent the last five years to try and confirm our current understanding of cellular structure, and we are unable to reproduce our current identification of cell structure. One example is mentioned in response to reviewer 1. We are unable to find electonmicrographs in publications that show myelination of neurons consistent with the countless schematic depictions available online and in the literature. This includes publications about myelination. Our findings are all consistent with the new images we will include in our revised manuscript. Longitudinal sections through neurons are easily obtained and neurons can be followed well beyond the axon hillock.

      A second example that is inconsistent with our current understanding of cells and biochemical processes in cells are ‘astrocytes’ and ‘reactive astrocytes’. We will show in the revised manuscript swell-bodies have no defined cytoplasm that every cell requires to fulfil basic cellular functions required for survival. It is very apparent to a structural expert that swell-bodies lack cytoplasm. The ‘consistency’ of this lacking cytoplasm is ‘inconsistent’ with fixation artifacts. In Author response image 3 we demonstrate that immunolabeling for AQP4 shows two types of structures. (1) Immunolabeled structures that are void of immunolabeling in their lumina and show receptacle-like immunoreactivity along the outside (panels I,J,L,M image below) consistent with swellbodies as indicated by the provided amyloid beta-stained and Luxol H&E-stained swellbodies that are associated with immunolabeled receptacle-shaped structures on the outside (panels K,N). A structurally trained eye quickly recognizes that these are not immunolabeled cells but are consistent with swell-bodies and referred to as ‘reactive astrocytes’ in the literature. We show in panel O what an immunolabeled cell looks like, with slender processes and immunolabeled cytoplasm. In this image in panels G1-4 we demonstrate that swell-bodies contain AQP4 mRNA explaining why they are immunoreactive for this protein. It is important that we bring attention to these details and do not randomly describe structure just based on a signal. This is exactly the point we make and I hope that the reviewer recognizes our expertise in cellular neuroscience and in particular in recognizing cellular structure. 

      Again, I urge the scientific community not to dismiss our findings but to actually study the validity of our findings. 

      Author response image 3.

      In the revised manuscript we will include additional experiments that we have carried out, that have also increased the sample number of tested human brain. We have in total so far investigated 13 different human brain samples, six of which are AD-affected. 

      We will revise our discussion to explain our observations in context with the literature better and include so much evidence in support of our hypothesis that it would be unreasonable to dismiss all this compelling and logical evidence.

      This research was not planned; it resulted from our accidental discovery in the spider system when our animals struggled with early onset neurodegeneration that we needed to address. This is when we recognized the waste-internalizing role of myelin in giant spider neurons. As we will discuss in the revised manuscript, such systems are highly conserved throughout evolution, and this is what made us realize that the only images of myelinated neurons we could find were either schematic drawings, single cross sections through myelinated cell profiles or very high magnification insets that also did not show the actual neuron that is myelinated.

      One can argue that there are both myelinated and unmyelinated axons. However, the varicose projections clearly originate in the ependymal lining and double-label for AQP4. This is consistent with our postulation and inconsistent with our current understanding of myelination. 

      I sincerely urge the neuroscience community to re-visit myelination in the brain, we have done this for the past five years with extensive experience in this field and the only hypothesis that is supported by our findings is presented in this manuscript. Please do not dismiss these findings, the spiders have uncovered a waste canal system in the brain and putting this system in place will help us to gain a better understanding regarding neurodegenerative diseases. Lastly, and maybe most importantly, understanding how this system works in spiders and how the myelin is structurally anchored to microtubule that were missing in our degenerating spiders has allowed us to identify the cause for this sudden neurodegeneration and rescue our tropical, cold-blooded spider colony by installing a new heating system and raising the room temperature so that the coldsensitive microtubules no not dissociate anymore.  

      We would like to thank the reviewers to strengthen the content of this manuscript with their critical comments, we hope that our revision will help clarify some doubts.

      (1) Pasquettaz R, Kolotuev I, Rohrbach A, Gouelle C, Pellerin L, Langlet F. Peculiar protrusions along tanycyte processes face diverse neural and nonneural cell types in the hypothalamic parenchyma. Journal of Comparative Neurology. 2021;529(3):553. doi: 10.1002/cne.24965. PubMed PMID: edsgcl.646848213.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors for their positive assessment of our manuscript, and all the referees for their constructive comments, which have substantially improved this work. We welcome the opportunity to address referee #2's points for the public record, as they highlight key theoretical nuances and valuable future research directions.

      (1) We appreciate the reviewer's point that, in a noisy biological system, the increased motion parallax provided by a rich 3D depth structure naturally aids in separating translation from rotation. We fully agree on this point. Our argument aimed at highlighting a fundamental theoretical distinction. Pure algebraic decomposition algorithms are mathematically capable of solving heading on flat planes. The fact that human perception often shows biases in these zero-depth conditions, unless extra-retinal cues are present, suggests that the visual system does not rely on a generalized, global de-rotation algorithm. Instead, it relies on heuristic, depth-dependent structural signals (like motion parallax and retinal curl). We maintain that while depth certainly reduces noise, its strict necessity points toward an ecologically grounded control strategy rather than a noisy global decomposition process.

      (2) We think the reviewer raises a valid point regarding the exact neural locus of retinal curl encoding. It is true that Graziano et al. (1994) utilized centred spiral stimuli rather than the spatially offset curl geometries defined in our task. We view the spiral tuning of MSTd not as a direct, one-to-one mapping of full-field retinal curl, but rather as the foundational computational building block required to extract such a signal. We fully agree with the reviewer that identifying exactly where and how this population response is decoded into a unified, gaze-relative retinal curl signal remains an exciting and open empirical question for future neurophysiological research.

      (3) We acknowledge the reviewer's call for transparency here. The transition from well-documented multiplicative gain fields (which modulate response amplitude based on eye position) to a direct, localized gaze-centered inhibitory drive is indeed a theoretical abstraction in our model. We utilized this localized inhibition as a functional mechanism to demonstrate how sensory evidence and spatial priors might competitively interact within a standard Mexican-hat recurrent architecture. While gain fields clearly establish that parietal networks integrate gaze position, we agree that the exact local-circuit implementation mapping these gain fields to the specific inhibitory dynamics we modeled has yet to be empirically established.

      (4) The reviewer smartly questions whether the 3-5 second behavioral biases delay emerges from the 2.4-second smoothing window used in our computational flow manipulation. It is important to clarify that this 2.4-second window was used solely to stabilize the computed curl signal against high-frequency gait oscillations. This smoothing was restricted strictly to the modeling phase of the controller and neural model and was not applied to the participants' responses. The gradual build-up of their perceptual bias over 3-5 seconds represents their own intrinsic temporal integration of this trajectory, independent of the smoothing parameters used to smooth the curl used in the fitting of the controller and neural modelling.

      (5) We concede the reviewer's point that citing a computational model (Layton & Browning, 2014) does not constitute direct empirical evidence for a relationship between MSTd heading preferences and their receptive field locations. Our intention was to highlight a successful theoretical framework that elegantly organizes known properties of MSTd into a system capable of bypassing global de-rotation. We readily acknowledge that direct, single-cell empirical validation of this specific topographic relationship is currently lacking in the literature, and we appreciate the reviewer ensuring this distinction is clearly noted for the record.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental "nuisance" signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. While the evidence for the role of curl signals is convincing and advances our understanding of vision-based navigation, the work's impact would be strengthened by situating these findings among other cues that contribute to heading estimation, and by clarifying both the time course of these computations and their generalizability across different navigational contexts.

      We thank the editors and reviewers for their insightful feedback and positive assessment of our study. In this revised version, we have made substantial modifications to better situate our findings within the broader landscape of cues contributing to heading estimation, while also clarifying the time course of these computations and their generalizability across different navigational contexts. These points are included in new sections in the revised discussion.

      In addition, and following eLife guidelines, we have moved the methods to the end and make sure that the manuscript reads well without needing to go through methods first.

      Next, we address all the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      We appreciate Reviewer #1’s very positive feedback. Incorporating the perspective of ‘incidental’ sensory signals is a valuable suggestion that aligns perfectly with our findings. We agree that this perspective significantly strengthens the impact of our paper.

      In the revised version we have added a last section in the Discussion (Generalizability and Testable Predictions) to comment on the functional utility of 'incidental' signals and incorporated the suggested references. In addition, in the same heading, we briefly elaborate on the predictions and generalizability of the model and possible manipulations that might affect the integration between sensory evidence (curl signal) and straight-ahead prior.

      Reviewer #1 (Recommendations for the authors):

      It would be great if the authors could discuss the implications and predictions of their model.

      First, from a broader perspective, the study forms an important piece in the emerging recognition that incidental sensory signals are not a nuisance to the sensorimotor system, but contain functionally relevant and effectively used visual signals (Rolfs & Schweitzer, 2022). The authors may want to appreciate their contribution to this perspective in the discussion of the impact of their results. Indeed, a similar shift in perspective has been realized in the recognition that saccade-induced motion signals are not entirely suppressed but play a functional role for gaze correction (Schweitzer & Rolfs, 2021).

      Second, the authors could spell out additional predictions of their proposal: What are experimental manipulations that could shift the balance between relying on a straight-ahead prior and sensory estimation of curl? What would happen in extreme cases of such sensory evidence? When would priors become overwhelmingly influential?

      As commented in the public review we have now included these two aspects in the discussion.

      Minor point: After equation 11, the authors may want to specify that, like position, gaze g is also coded as {x,y}, just like position p.

      While the neural model details have been moved to Appendix 2 (including this equation), we added text before Equation 22 (previously Eq. 11) clarifying that gaze is encoded in image coordinates. We omitted point index i because the equation applies to all image points relative to a given gaze g.

      Reviewer #2 (Public review):

      We appreciate the reviewer’s feedback regarding the formalization of our reference frames. We agree that certain definitions were implicitly assumed rather than explicitly stated. We have revised the manuscript to provide all necessary self-contained information, ensuring that the geometry of the task response and the definition of heading are unambiguous. In the last paragraph of the revised introduction, we make clear the response frame of reference which is also included in the caption of fig 1. Also, we have addressed the gap between the task response (in world coordinates) and the functional role of the controller. This is particularly discussed in the discussion (section: reference frames) in which we also provide (and rule out) potential alternatives to our response biases. We also address all the other points raised by the reviewer.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:

      (a) To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given, and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.

      In our study, participants were instructed to report their “perceived direction of self-motion” by aligning a rotational encoder (steering wheel) with the direction they felt they were moving within the 3D simulated scene. Consequently, participants reported their instantaneous heading in a world-centered reference frame, from which the 3D trajectories were reconstructed. Since the reviewer had to infer this information, we have clarified this point at the end of the introduction, legend of figure 1, methods (now at the end of the ms.) and discussion to ensure it is immediately evident.

      Participants were informed that the initial heading (i.e. θ<sub>0</sub> in our controller nomenclature) was oriented “straight ahead” relative to their body which was aligned longitudinally with the experimental room. We have modified Figure 1B and revised the Methods section to explicitly clarify this initial alignment and the instructions provided to participants.

      In the revised manuscript, we have clarified that while the participant’s report is world-centered, the retinal curl provides a gaze-relative heading signal. Although this was already mentioned, we emphasize this point. In natural navigation toward a fixated target, a world-centered vector is often unnecessary; an error signal indicating heading relative to fixation is sufficient (as the reviewer also notes). However, the initial alignment of the heading within the 3D scene allows the brain to “calibrate” this internal controller, mapping the retinal curl signal onto the 3D world coordinates required for the task. Ad commented above, a new section in the discussion addresses and hopefully clarifies the relation with the controller.

      The reviewer also asks how we can be certain that participants were reporting in world coordinates rather than an alternative frame, such as “heading relative to the fixation target.” We believe our “Cancelled Curl” (and over-cancelled) conditions provide the most compelling evidence to rule out this alternative. In these conditions, the physical position of the fixation target in the scene remained identical to the unaltered flow condition. If participants were simply reporting heading relative to the fixation target’s spatial location, the observed biases should have persisted regardless of the flow manipulation. Instead, the bias vanished when the curl was removed. This causal evidence proves that the bias is driven by the retinal motion signal (curl) rather than the spatial orientation of the eyes or the target’s position in the scene. Furthermore, the temporal evolution of the response supports a world-centered integration (in agreement with Warren et 2001 Nat Neuro.). For simulated straight paths, the perceived heading remains straight for the first few seconds (consistent with the initial world-centred alignment), with biases only emerging after approximately 3 seconds of integration (a point we elaborate on in our response to Reviewer #3). Had participants been responding based on a simple gaze-relative reference frame from the onset, these biases would have manifested significantly earlier. We have incorporated these points into the revised Discussion to better frame our findings alongside other cues, such as the Focus of Expansion (FOE) and egocentric visual direction that contribute to heading estimation.

      Finally, we have rephrased the abstract sentence for clarity. However, we maintain that the original premise remains valid once the world-centered initial heading is aligned with the gaze-centered reference frame.

      (b) It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      The reviewer notes that we must be clear about the relationship between curl and heading (relative to fixation) and the variables that affect curl. We also thank the reviewer for encouraging to add the equations that show the relation of curl with additional variables. We have now included these equations in appendix 1.

      Beyond the discrepancy between heading (θ) and gaze (ψ), curl is geometrically determined by translational self-motion speed (v), eye height (h), and pitch (α). More specifically, curl = (v.sinψcosα)/h. The derivation is now included in appendix 1. Since h = dsinα, where d is the 3D distance to the fixation point, we could express cos α as a function of distance. Certainly, there is not a 1:1 map from curl signal to heading relative to gaze (e.g. θ-ψ). Participant would need to know v and eye height plus extra-retinal information. Frenz et al (2003, Vis Res.) showed that people can estimate self-motion directly from optic flow, across different simulated eye height and gaze angle; extra-retinal information can, in addition, provide knowledge to ψ and α. It is then plausible that the visual system can use and transform the curl signal from a qualitative directional cue (i.e. steering left or right of fixation) into a quantitative steering command. By combining curl with knowledge of gaze orientation and eye height, the visual system can resolve ambiguities in the flow field and utilize curl as a more precise error signal for locomotor control. These aspects are now included in the new version of the discussion.

      (2B) I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors need to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.

      We thank the reviewer for this point. We have addressed the alignment of the reference frames in our response to Issues 1a and 2a. Once the initial orientation (θ<sub>0</sub>) is established in the world frame, the controller model generates steering adjustments that directly translate into heading predictions within that same world reference frame. By treating the perceptual report as an output of the locomotor controller, we resolve the discrepancy between the steering task and the reported heading.

      (2c) In addition, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      We respectfully disagree with the reviewer’s interpretation regarding data smoothing. The thin lines in Figure 2 represent the mean 3D paths derived directly from the response variable (θ<sub>t</sub>) across trials of identical conditions for each participant (as detailed in the ‘Computation of Perceived Path’ section). No smoothing or filtering has been applied to these plotted trajectories other than computing the mean across trials. We also wish to remind the reviewer that the raw data and analysis code remain publicly accessible for further inspection. Having said that, we include now a supplementary figure showing an example of raw data responses as a function of time. This figure will be a supplemental figure of main Figure 2 (now provisionally included in the Suppl Information).

      Regarding the visual representation: in earlier versions of the manuscript, we included shaded 95% Confidence Intervals (CIs) in Figure 2. However, this addition rendered the plot overly cluttered and obscured the individual trajectories. We therefore chose to present individual participant means (thin lines) alongside group averages (thick lines) to emphasize inter-subject variability. For clarity, the 95% CIs are explicitly displayed in Figure 3, where the data density is more conducive to shaded areas.

      (3) “...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      We have updated the Discussion to more specifically align our findings with Matthis et al. (2022). We emphasize that our study provides the perceptual validation for their ecological observation that the FOE is often too unstable for reliable use, whereas foveal curl remains a robust signal for path estimation. Our paper provides the causal link, since we manipulate curl in real-time (the ‘cancelled & over cancelled curl’ condition) providing the critical evidence that perceived heading is affected by this signal. The relation with this previous study is made clear in the revised discussion.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      We thank the reviewer for noting that retinal slip (velocity error) is a more critical metric than positional gaze error. We agree that tracking inaccuracies can introduce translational noise into the flow field. The 3° threshold was established based on the eye tracker’s specifications and the naturalistic setup (1-meter viewing distance without head stabilization). Across all participants, the mean positional error ranged from 1.016° to 1.5° (1 deg is 2.08 cm in our setup). We also calculated retinal slip values, which ranged from 0.12 to 0.27 deg/s (X dimension) and 0.12 to 0.23 deg/s (Y dimension). These values are comparable to natural oculomotor drift (Kowler et al., 1979) and are understandably small given the low velocity of the fixation target. We have added this information about retinal sleep at the beginning of the results section.

      Consequently, it is highly unlikely that retinal slip influenced the results. Furthermore, assuming that tracking error remained consistent across fixation conditions, any present retinal slip cannot explain why the bias followed the retinal curl manipulation as predicted by the controller. We therefore consider retinal slip to be an unlikely confounding factor.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good, given that the model has 30 parameters, and these data are pretty low-dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed, they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      We thank the reviewer for the opportunity to clarify the logic behind our modeling choices. We acknowledge that the “separate fits” are inherently less informative due to the high number of free parameters relative to the data. Our primary scientific goal was not to achieve perfect descriptive accuracy via 30 parameters, but to test a specific functional hypothesis through the “joint fit.”

      The Logic of the Joint Fit:

      We agree with the reviewer that the joint fit misses some paths in some conditions. Of course, the joint fit reflects a significant compromise. The “Gain” (the weighting of the curl signal) is likely not a static constant but is dynamically tuned based on task demands, confidence in the visual signal, simulated speed, and so on. By using a single Gain parameter, we intentionally ignore this contextual variability to see how much of the behavior can be explained by a “minimalist” controller. In this sense, the 2-parameter joint model is a deliberate attempt to test this limit. By forcing a single Gain parameter to account for all conditions across both straight and curved paths within one flow manipulation (e.g. unaltered flow) we are asking if a single, fixed linear relationship between retinal curl and steering effort/gain can explain the results. We view the joint fit not as a “perfect” model, but as a stronger test of the curl-based control theory. The fact that a 2-parameter model can capture the direction and scale of biases across such a diverse set of conditions (straight/curved paths, five fixation eccentricities) suggests that retinal curl is a robust signal. Upon closer analysis, these discrepancies between the joint model and the data are most pronounced in the over-cancelled condition which is the one when sensory evidence becomes more ecologically inconsistent with the extra-retinal information (gaze direction). While the joint fit successfully demonstrates that a single parameter can capture the general functional role of curl, it fails to account for the complex sensory re-weighting that occurs in ecologically inconsistent conditions (like ‘over-cancelled’ flow). We have updated the manuscript to discuss these limitations in the “fitting the controller” section, framing the model as a parsimonious first-order approximation rather than a complete description of human heading perception based on a minimal set of parameters.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Figure S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      We acknowledge that the presentation of the neural model requires more clarity regarding its objectives and its relationship to the behavioral data.

      We first wish to clarify the intended scope of the neural ring-attractor model. Our primary goal was not to provide a comprehensive account of behavioral performance across all conditions (which is the role of the controller model), but rather to demonstrate a biologically plausible mechanism that explains the emergence of the “Opposite-to-Gaze” bias. While the controller demonstrates that the bias follows a specific control law, the neural model shows how such a law can emerge from known primate neurophysiology, specifically, spiral-tuned MSTd neurons, gaze-contingent inhibition, and an egocentric “straight-ahead” prior.

      Why Straight Paths are Sufficient for this Objective. The reviewer asks why only straight paths were simulated. In our study, the straight-path condition with eccentric gaze is the purest test of the bias mechanism. Simulating the straight paths allowed us to isolate the interaction between foveal inhibition and the straight-ahead prior without the confounding variable of path-curvature flow. Given the complexity of the neural network’s parameter space, we focused on these conditions to provide a clear neuro-plausible explanation. We have added text when introducing the model (Neural simulations in the Results section) to make clear why we model straight paths only.

      Units: Pixels vs. Degrees. We acknowledge that the use of “pixels” in the plots of internal neural dynamics may appear awkward. The neural network operates on input stimuli that are defined by the pixel resolution of the videos used in the simulations, we used pixels as the native coordinate system to describe the movement of activity peaks within the network’s internal “map.” We have decided to keep the pixel units in these figures.

      Behavioral Output (Meters): Importantly, the final heading estimates produced by the network are not left in pixels. We use a pinhole camera model to reconstruct the 3D trajectories from the neural activity. These results are expressed in meters, allowing for a direct comparison with the human behavioral data.

      Addressing Wild Oscillations and Smooth Paths. The oscillations observed in the instantaneous heading estimates reflect the stochastic nature of the population peak when tracking high-frequency sensory inputs. In our model, the synaptic time constant (τ) was kept relatively small to ensure a fast, low-latency response to changes in self-motion. While increasing τ would have produced smoother internal dynamics, it would also have introduced delays into the control loop. Instead, we chose to maintain this high sensory responsiveness and applied a temporal moving average later to the network’s decoding to reconstruct the 3D trajectories. This is explicitly stated in the section “Heading Estimation and 3D path reconstruction” in the new appendix 2.

      In addition, the neural activity over time is shown in two ways: the heatmap shows the neuron with preferred heading (one can see more oscillations, specially when the fixation point is closer to the centre (eccentricities -2 and 2), due to larger competition between the sensory evidence and the straight-ahead prior. The other way is the decoded heading. In the ring-attractor model, the decoded heading (φ̂) is not determined by a single neuron but is calculated using a population vector average (equation 19). By summing across the entire population, the decoder effectively integrates sensory evidence from many neurons simultaneously. One can appreciate (see e.g. Fig. 5B) that averaged decoding, leads to a smoother resulting estimate (the white dashed line, whose visibility had been improved in the revised version). Behavioral work by Burr and Santoro (2001) suggests that global motion signals (divergence and rotation in optic flow) are integrated over much longer timescales—roughly 1000ms to 3000ms—compared to local motion units (~200 ms).

      In the previous manuscript, we discussed this aspect in lines 424-426. In the new version, we have added text in the Heading estimation and 3D path reconstruction section (now in appendix 2) stating that we smoothed the decoded signal in agreement with this psychophysical evidence before applying the camera model.

      See also our comment on temporal integration in the responses to reviewer #3

      Reviewer #2 (Recommendations for the authors):

      (7) Line 51: "...a functional role of rotational flow components has been largely neglected in both theoretical and experimental work on heading perception." I feel like this statement is too strong and that the authors try too hard to "sell" their findings by underrepresenting previous work. There are numerous studies (many not cited), both behavioral and electrophysiological, that have examined how heading perception depends on pursuit eye movements, either physical movements or visually simulated ones. These studies directly involve rotational flow components, and several of them have concluded that rotational flow components contribute to estimating heading in the presence of eye movements (just one example is Grigo and Lappe 1999). Because these studies generally involved horizontal pursuit of a target on the horizon, rather than tracking a point in the ground plane (like the authors' work), these studies generally did not involve flow fields with retinal curl around the fixation point. But I consider these older studies just a special case of the more general geometry, and they still involve rotational flow components. Moreover, various previous studies have used stimuli for which there was no FOE present in the visible display (either due to simulated rotation or masking out the FOE), and the authors do not seem to give credit to these works either. In addition, several studies have implicated a role of extraretinal signals in perceiving heading during eye movements, so retinal curl cannot explain everything. Rather than overemphasizing the limitations of previous work, the authors would be better served to explain how their findings extend and generalize from these previous studies.

      We thank the reviewer for pointing out this oversight; it was not our intention to overlook previous work. While our original version cited studies considering rotation-related cue, we have now substantially revised the introduction to include previous work and better acknowledge the informative role of rotation. Our central aim remains to distinguish between models that compensate for rotation to recover a heading vector and our proposal that the visual system exploits retinal curl directly as a primary, functional signal for locomotor control.

      We have now updated the Introduction and Discussion to better situate our work within the context of studies (including Grigo & Lappe, 1999 and some additional ones we have included in the new version) that have investigated the informative role of rotational flow. We now clarify that our study extends these findings by investigating the non-uniform rotational patterns (curl) that emerge during ground-plane fixation, representing a more general and biologically ubiquitous case of locomotor control, while acknowledging previous studies that have also considered the potential role of curl generated by gaze fixation.

      (8) Figure 3: I did not understand why there are two purple and two blue curves in the graphs of the middle column. And the caption does not explain this.

      This a very good observation. This was explained in lines 268-273 (previous version). When the gaze is straight-ahead (same direction as heading), there is no curl. However, we introduced positive or negative curl in the altered conditions. These purple and blue lines refer to these trials and show that the bias re-appears in the expected direction when curl is (unexpectedly) added. We think this adds additional evidence to the curl contributing to heading. Even though this was extensively explained we have added text in the caption of figure 3.

      (9) Line 331: What makes the authors think that retinal curl is computed in area MSTd? They should cite studies to support this idea if it has been shown in physiology.

      Evidence was cited in the introduction (Graziano et al 1994) of sensitivity to spiral motion in addition to neuro-computational models that implement this activity also cited (e.g. work of Leyton et al.)

      (10) The neural network model for computing heading from curl requires a "gaze-centered inhibitory drive" that inhibits activity around where the eyes are looking. This is probably a biologically plausible thing, but is there any evidence to support the idea that this signal exists in the parts of the brain where the authors believe these computations to be happening? They simply posit the existence of this gaze-centered inhibition as though it is common knowledge, but they provide no citations nor discuss any previous evidence for its existence.

      While neurophysiological evidence primarily describes this as gain-field modulation, this process frequently involves localized suppression of neural activity to facilitate coordinate transformations. In parietal areas such as LIP and 7a, eye-position signals do not just enhance responses but can also suppress them, effectively shifting the 'center of gravity' of a population response (Read et al 1997; Born et al 2005, cited in the discussion in the revised section re-evaluating the Focus of Expansion). In the context of our ring-attractor model, this functional modulation is most parsimoniously implemented as a gaze-centered inhibitory drive.

      (11) Lines 482-483: Why should perceptual biases related to retinal curl take seconds to show up?? The curl information itself must be present very quickly, perhaps requiring just a few video frames. So what does this imply about mechanisms? The authors throw out this assertion, but it is left hanging without any further analysis or support.

      The time course reflects the integration requirements of complex motion processing. While local flow is processed rapidly, global patterns like retinal curl require longer temporal windows to reach a stable estimate (Burr et al 2001). In our study, this integration is functionally necessary to filter the higher-frequency 'wobble' induced by gait-cycle oscillations. We now discuss the temporal integration aspects under a new heading in the discussion.

      (12) Line 508: "This suggests that the "bias" observed in our perceived headings may reflect the operation of a control law optimized for action rather than a failure of a perceptual system designed for passive estimation." The authors make this statement to justify why perceptual biases are present with unaltered curl. But I don't fully understand the logic. Are they saying that it is not possible to have a set of computations that can do both things accurately? Is it possible to show this theoretically? Moreover, if it is not possible to rule out other possible sources of the biases, such as those described above (reference frame of judgments, eye movements, etc), then is it necessary to invoke this logic?

      Our logic is that the observed 'bias' is not a representational failure, but a functional byproduct of a control law optimized for active steering. In a closed-loop system, the objective is to null the error signal (retinal curl) to maintain a stable path. When observers are asked to make an open-loop heading report, they likely utilize this same control signal, which manifests as a systematic bias toward the 'null' point of the controller as a result of a sustained fixation in discrepancy with the simulated translation/heading.

      We do not suggest that accurate perception and control are theoretically incompatible; rather, we suggest that perception and action rely in the same underlying information (e.g. work of Brenner & Smeets). While other factors, such as coordinate transformations between reference frames, certainly can contribute to the reporting process, our interpretation provides a parsimonious link between the psychophysical data and the underlying steering mechanism. By framing the bias as a consequence of a 'nulling' strategy, we explain not just the existence of the error, but its specific direction and magnitude relative to the fixated target.

      (13) Line 546: "...MSTd would simultaneously code curvature for trajectory estimation and heading across the neural population, with curvature encoded through the spirality of the most active cell and heading through the visuotopic location of its receptive field center." The latter part of this argument seems to imply a relationship between the heading preferences of MSTd neurons and the locations of their receptive fields. I am not aware of any evidence for such a relationship, so the authors should indicate whether this is based on some experimental data or just a speculation.

      We thank the reviewer for this observation. The proposal that heading is signaled by the visuotopic location of active MSTd populations is a core architectural feature of our model and is supported by several lines of evidence.In the Layton and Browning (2014) framework, MSTd is modeled as a visuotopic map of functional 'hypercolumns'. Each hypercolumn contains neurons tuned to a continuum of spiral patterns, but all neurons in a given hypercolumn share a receptive field center at a specific location in visual space. Consequently, the visuotopic location ($x, y$ coordinates) of the maximally active hypercolumn represents the center of motion (heading), while the spirality (the tuning dimension within that hypercolumn) represents path curvature. We have clarified this in the revised discussion (re-evaluating the FoE) to emphasize that this dual-coding scheme arises from the simultaneous representation of 'where' (population map location) and 'what' (spiral tuning) in MSTd.

      Reviewer #3 (Public review):

      The primary limitation of the paper is that it avoids discussion of some of the inevitable complexities of heading perception. The main issue is what exactly is meant by heading. Different behaviors evolve over different timescales. The geometry of retinal motion defines instantaneous heading, which varies widely through the gait cycle. Time-varying information like this is known to be important in the momentary control of balance. Heading can also be thought of as steering the body toward a distant goal, which evolves over longer timescales. The current manuscript appears to be concerned with heading information integrated over a few seconds and seems to provide evidence that heading is indeed integrated over the gait cycle. The issue of the time scale of the computation is touched on, but it is not related to how it might be used in normal walking or what situations it might apply to. Steering toward a distant goal during walking is not a very difficult problem and may not require evaluation of retinal motion, but control of balance is more challenging and may depend critically on curl. Consequently, the timescale of the computation needs to be considered in order to understand what is meant by heading.

      We thank Reviewer #3 the comments regarding the definition of heading at different time scales, the role of the gait cycle, and the temporal integration of the curl signal. These comments have helped us refine the manuscript’s core arguments.

      We agree that “heading” must be precisely defined within the context of the differing temporal demands of balance and steering. While instantaneous heading provides the high-frequency feedback necessary for momentary postural adjustments and balance, our study is concerned with heading as a gaze-relative signal used for the continuous control of a locomotor trajectory. As such, we have revised the manuscript to specify that the perceived heading measured in our task reflects a signal integrated over the gait cycle to filter out the oscillatory noise induced by head bob and sway (mainly in the Discussion section).

      The reviewer correctly notes that gait-induced head bob and sway produce high-frequency oscillations in the curl signal, yet our behavioral results show smooth, slowly evolving biases. The visual system does not react to “instantaneous” curl, which would lead to jittery, unstable heading estimates. Instead, it integrates flow over a timescale roughly commensurate with a full gait cycle (~500–1000ms). This implies a significant temporal integration process. This temporal integration is consistent with evidence (Burr and Santoro,2001, Vis Res) indicating that optic flow signals (radial and rotational components) are integrated over windows of approximately up to 3 seconds to ensure perceptual stability. Neurally, this likely involves the projection from area MSTd to the Ventral Intraparietal area (VIP), a pathway where fast, eye-centered sensory inputs are transformed into stable, body-centered representations suitable for guiding long-term steering behavior (Chen et al. 2011, JNeurosci.). By grounding our definition of heading in these specific temporal and neural constraints, we tried to clarify how the visual system exploits retinal curl for goal-directed action in natural, dynamic environments and relate our findings to recent studies addressing the role of retinal motion on balance (Powell et al. 2026 Bioarx).

      In our implementation, we explicitly address the high-frequency noise introduced by gait dynamics by smoothing the retinal curl signals computed from the stimulus videos before they are fed into the controller. This temporal filtering allows the fit of the controller’s prediction to the response data while remaining robust to the rapid fluctuations of head bob and sway. In contrast, the neural ring-attractor model would not require an external smoothing step; instead, the integration is an emergent property of the system’s architecture that can be controlled with different parameters, as commented above in a response to Reviewer #2. The dynamics of the synaptic weights and the characteristic “leak” in the population activity naturally implement a leaky integration of sensory evidence, ensuring that the decoded heading reflects a sustained estimate rather than an instantaneous response to visual noise.

      We also agree that we avoided discussing some complexities of the heading perception. In the new version, we also include and integrate the distinction between instant heading and future path in different parts of the ms (introduction) and mainly discussion (temporal integration and steering section) which have been revised substantially.

      Reviewer #3 (Recommendations for the authors):

      There are a number of points that require clarification.

      (1) Head bob and sway were included in the stimulus and need to be addressed in both the analysis of the data and the interpretation. The curl signal in the stimulus varied over time, commensurate with normal gait. However, the results don't reflect the same level of variability that would be produced from curl over the gait cycle. This means that the information must be integrated over some longer timescale. It is not clear from the data analysis what this integration is. Is there an implicit integration with the manipulation of the steering wheel? If subjects indeed appear to be able to use curl to evaluate heading over timescales of seconds, this needs to be explicitly addressed, as it is a novel result. This would require parts of the discussion to be changed/expanded to maintain consistency. For example, line 482 talks about the buildup of biases over time.

      As commented above in the public response, the curl estimated from the optic flow algorithm was smoothed before being input into the controller (path fitting and predictions). The smoothing was only applied to the curl signal, not to the participants responses. Also, as mentioned before, the time course of the bias is consistent with integration times of optic flow reported in the literature. All these aspects are now explicitly included in the new display and conditions section (Flow manipulation conditions).

      (2) Since the experiment included curl variability resulting from the gait cycle, some discussion is needed about the role of retinal motion in the control of balance and posture. There is a large literature about the role of flow in controlling gait and momentary adjustments of the body while walking. Additionally, it should be noted that in the task, head bob and sway from 1 prerecorded individual was shown to all subjects. It is known that gait varies significantly between different individuals, and it should be acknowledged that this may lead to differences at the individual subject level for perceiving heading.

      We have included a point in the discussion addressing the different time scales for different use of optic flow signals (postural control vs locomotion).

      We agree with the reviewer that utilizing a single gait profile for all participants may introduce individual differences in perceived heading, as the simulated head motion might not perfectly match each participant’s unique biological gait signature. However, we prioritized stimulus consistency over idiosyncratic accuracy. By ensuring that every participant viewed the exact same motion profile, we could be certain that the systematic 'opposite-gaze' biases observed across the population were driven by our experimental manipulations of gaze and retinal curl, rather than being confounded by variability in head-motion kinematics. We have added an acknowledgement of this point at the first paragraph of the displays and conditions section in the Methods.

      (3) More information is required about the use of the rotating wheel for the measurement of heading. How easy was it to use? What about time delay, and how does this deal with the bob and sway? Does the wheel impose a de facto integration on the perceptual measurement?

      The rotary encoder provided an intuitive steering-wheel interface that participants found easy to operate. To ensure minimal latency (1–5 ms), the device was interfaced via an Arduino Uno and sampled by a dedicated background Python thread, isolated from the visual rendering loop. We have incorporated these technical details into the Methods (Procedure) section.

      (4) Restructuring the description of the models It remains unclear why the dynamics of the neural network are a necessary inclusion in this paper. It seems interesting, but there is no comparison to actual neural data or other related work. Instead, this appears to be a description of what the network is doing, which is already defined by the equations. This needs to be clarified for its exact interpretation with respect to real neural data, and its importance here for understanding the biases that emerge in heading judgments. The paper would flow better if this section were included as supplementary material or omitted from the paper entirely, as it seems to detract from the other points. If this is a description of why the biases are seen in the controller, then the supplementary material is a good place for it.

      We thank the reviewer for this suggestion. We clarify that the neural model is not intended to simulate specific empirical neural data, but rather to provide a biologically plausible implementation of the controller. This allows us to demonstrate how the observed biases emerge from the dynamics of standard cortical architectures, such as ring attractors. This is now mentioned when introducing the neural simulation results.

      Following the reviewer's suggestion, we have moved the neural model equations to Appendix 2 while retaining the simulation results in the main text (Results). We believe it is essential to present not just the abstract controller, but also its functional implementation, as this provides a mechanistic bridge between retinal signals and locomotor behavior.

      (5) For the modeling approaches, the math would be more appropriate for supplementary materials.

      To ensure a better flow of the paper, we have moved the neural model equations to Appendix 2, while Appendix 1 now details the relationship between measured curl and other variables (speed, yaw, pitch, etc.). We have retained the controller model in the Methods section, consistent with the eLife layout where Methods follows the Discussion.

      (6) How do the models and data analysis deal with the influence of the gait cycle in the input? Do they integrate the information over that timescale? If so, the integration time needs to be specified.

      As commented above, to mitigate gait-cycle fluctuations, we smoothed the computed curl signal before inputting it into the controller and applied a similar smoothing process to the neural model’s readout. Using a LOESS filter, the effective integration window was 2.4 seconds. Close to the integration time reported in Burr et al. 2001. These parameters have now been explicitly specified in the Methods section. For the empirical data, we just utilized trial-averaging.

      (7) What does biologically plausible mean in terms of the neural network model? Especially when control wasn't explicitly a variable in the measured behavior of the subjects.

      By biologically plausible, we mean that our model is constrained by neural architectures documented in the primate brain—specifically ring-attractor dynamics, population coding, and gaze-centered gain-fields. Crucially, the network utilizes recurrent connectivity with a 'Mexican-hat' profile (local excitation combined with lateral inhibition). This is a standard and widely accepted motif in computational neuroscience, representing the consensus on how cortical circuits maintain a stable "bump" of activity to represent spatial variables. Rather than introducing ad-hoc mechanisms, we demonstrate that the observed behavioral biases emerge naturally from these established neural components when they are tasked with maintaining locomotor stability. Since we think this aspect was already emphasized, we haven’t added any additional detail.

      (8) There should be more extensive acknowledgement of the body of literature that has challenged the use of the focus of expansion. That section should also include references to work that has investigated extraretinal signals, as they may also be important.

      We have expanded the Introduction and Discussion to more thoroughly acknowledge research challenging FOE-based models and the critical role of extraretinal signals. These updates, which also align with our response to Reviewer #2, provide a more comprehensive context for our model within the existing body of heading and self-motion literature.

      Minor points:

      (1) Line 69 - In self-generated motion, spiral patterns are almost always centered on the fovea, but many physiological experiments present spirals in the peripheral retina. This is incompatible with the motion generated during self-motion. Therefore, clarify whether the type of spiral motion Graziano investigated was centered on the fovea.

      In the experiments conducted by Graziano et al. (1994), spiral stimuli were centered on the receptive field (RF) of the individual neuron being recorded to accurately characterize its tuning. While the reviewer correctly notes that spiral centers often align with the fovea during active steering (due to fixation on a goal), MSTd neurons possess large RFs that provide a comprehensive 'template' system across the visual field. This population-level representation allows the brain to recover trajectory information even when the focus of motion shifts relative to the fovea—for example, during pursuit eye movements or when fixating on landmarks off the direct path of travel. To keep this part of the text brief, we haven’t add more details concerning this study.

      (2) Line 72 - "Magnitude" instead of "amount".

      This has been changed.

      (3) Line 115 - State explicitly whether the scale of the visual stimulus was matched to the scale of the actual natural images shown in VR.

      We have updated the Methods (Displays and conditions) section to explicitly state that the visual scale was veridical. The virtual camera’s parameters were calibrated such that its field of view (91°) matched the physical dimensions of the projection screen (2.03 m × 1.16 m) at the 1.0 m viewing distance. This ensures that the angular size of the objects and motion gradients in the stimulus were 1:1 with the scale of the simulated natural environment.

      (4) Line 144 - While Farneback is a good dense flow estimation algorithm, it is noisy and may impose biases/variability in the calculation of curl. This should be acknowledged.

      We acknowledge that the Farneback algorithm can introduce variability in curl estimation. To mitigate this, we utilized 10 independent renderings of each experimental trial to compute the flow fields. Although this methodology was reflected in the data previously uploaded to our OSF repository, we have now explicitly added this detail to the manuscript (Flow (curl) manipulation conditions). The computed curl used for the modeling was derived from the aggregate of these different runs, ensuring a robust and stable signal that accounts for potential algorithmic noise.

      (5) Line 206 - "Join fits". Is this a technical term? It sounds awkward. Would "Joint fits" make more sense?

      The referee is right. We have corrected this.

      (6) Line 280 - "Consistent with" (typo).

      This has been corrected.

      (7) Line 286 - Typo in title.

      Also corrected to Fitting the controller.

      (8) Lines 482-493 - There should be more discussion on the time course of integrating the stimulus, and the relationship/generalizability to more natural stimuli.

      This part of the discussion (related to integration time) has been changed considerably to include discussion of postural control in addition to locomotion.

      (9) Lines 516-517 - It is mentioned that retinal flow dynamics override the visual direction cue. This may not be generally true, as the reweighting of the cues might depend on things like task demands or actual stimulus context. In the present experiment, the subject only has access to a large moving textured ground plane, and the body is stationary.

      We agree with the reviewer that cue reweighting is highly context-dependent. However, as noted in the original manuscript (Lines 516-517), we specifically stated that retinal flow dynamics 'can' override the visual direction cue, rather than asserting a universal rule. This phrasing was intentional to acknowledge that while flow is a potent signal—especially in the presence of a large, textured ground plane as used in our paradigm—the relative weighting of these cues remains contingent on the specific sensory and task conditions. We believe this remains a fair and cautious interpretation of our findings.

      (10) Lines 524-256 - It is unclear why Matthis et. al. 2022 is cited for this point. Some of the steering literature, like Wilkie Wann & Allison 2006 or Lappi & Mole 2018 (and some of their other work), would be more relevant for the definition of a control law under these circumstances.

      We agree and this part has been changed substantially.

      (11) Lines 528-532 - Warren et. al. 2001 should be cited in this section because their results were interpreted in terms of focus of expansion, but may result from the curl signal (Powell et. al., 2026. The Role of Retinal Flow in Walking. bioRxiv, 2026-02.).

      We agree with this suggestion and the citation has been added.

      (12) Lines 544-546 - Layton and Browning are focused more on path perception from the implemented spiral tuned cells. Because of this, it wouldn't be appropriate to say it is a shift away from FOE-based heading based on this citation alone. There are more models/psychophysical results that would strengthen this claim.

      We agree with the reviewer that Layton and Browning focus specifically on path perception. Our original intention in citing this work was to emphasize the neurophysiological continuum from radial to circular motion (spiral tuning) rather than discrete expansion/rotation channels. However, we have revised this section to clarify that the reliance on retinal curl represents a mechanism for determining the future path (locomotor trajectory) rather than merely instantaneous heading. This distinction acknowledges that while heading is a momentary vector, the integration of curl signals allows the system to anticipate and control the intended path over time—a framing that better aligns with both the cited literature and our proposed controller model. As commented above this is now extensively discussed in the revised version.

      (13) Figures - There are minor visibility issues for some of the figures. In Figure 2, the thick line is unreadable, and in Figure 1, the x-axis labels are crowded.

      The axis in Fig 1 has been modified to avoid crowdedness. In Fig. 2, the thick (average line) has been modified. We hope they are more visible now.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of Go/No Go tasks, which differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analyze dopamine recordings made with fast scan cyclic voltammetry, and find that dopamine signals vary most consistently to cues that signal a required action (Go cues) vs cues signaling action withholding (No Go cues). Through various analyses, they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively, these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, largely showing consistent results with past studies.

      Weaknesses:

      The paper could heavily benefit from some revision to 1) increase clarity of the figures, the methods, and the analysis. 2) The inclusion of many tasks is a strength, but also somewhat overshadows specific points in the data, which could be improved with some revision/reworking. 3) Some conclusions are not fully justified. As shown, support for the conclusion "dopamine reflects action initiation but not controllability or effort" is lacking without more analyses and additional context. 4) Further, the notion that the dopamine signals reported here reflect spatial information could be justified more strongly.

      We thank the reviewer for their detailed evaluation and constructive feedback. We have made substantial revisions to address each concern raised:

      (1) Clarity of figures, methods and analyses

      We have revised the organization of the panels in Figure 1 for clarity.

      We have revised Figure 2 and its caption: we have labelled all comparisons depicted in the figure, and now included a line, “Subject-wise comparisons of dopamine data were made for all alignments”, to clearly show that all statistical tests shown in Figure 2c-e were performed between subjects.

      In the caption of Figure 3, we have now added, “... and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” to clearly show that statistical tests were performed on the trial-wise level for Figure 3c.

      We have now added a table to the Methods section (Table 1), detailing the sample size in each task variant and the number of trials within each No-go classification.

      We have added more information in the Methods section: we now include Videos to show the classified No-go behaviors and other trial types (Go and Free); and provide a schematic of the DLC workflow in the supplementary materials (Supplementary Figure S9).

      (2) Strengthening specific points in the data

      To improve clarity of our main findings, we have revised the layout of the Results section such as including more descriptive headers:

      “Behavioral performance in Go/No-go (“short task”) was unaffected by controllability of reward pursuit”

      “Behavioral performance in Go/No-go/Free (“long task”) was unaffected by controllability of reward pursuit”

      “Motivation to approach the reward magazine was similar between Go and No-go trials”

      “VMS dopamine release encodes reward-related action initiation”

      “VMS dopamine release does not only reflect reward-related action initiation”

      “Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards”

      “Motivational state reflected by No-go behavioral strategy correlates with dopamine signal size during reward approach”

      (3) (4) Conclusions drawn from data

      We have carefully revised our conclusions to accurately reflect our experimental design and analyses, emphasising our core finding that VMS dopamine was consistently increased in Go versus No-go trials, throughout manipulation of the type of trial start (self- and cue-initiated) and effort manipulation (short and long task variants).

      We thank the reviewer for the feedback regarding our VMS dopamine signals reflecting spatial proximity to reward. We performed additional analyses, which we present in Supplementary Figure S10, S11 and S12, to support our interpretation that VMS dopamine encodes spatial proximity to reward.

      We appreciate the reviewer comment relating to the statement, “Dopamine reflects action initiation but not controllability or effort". We have revised the wording of our conclusion to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). This subheading is now changed in the Discussion to “Reward-related dopamine depends on action initiation irrespective of controllability and effort”. The additional analyses that we have performed are shown in Author response image 1entitled: “Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long).

      Additional details on subjects used in each study, analysis details on trialwise vs subjects-wise data, and other context would be helpful for improving the paper.

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures in the original submission of this manuscript.

      To improve the paper, we now include a table in Methods Section 6: Statistical Analysis (Table 1), detailing the sample size in each task variant (with FSCV recordings) and the number of trials within each No-go classification.

      To give more context, we made Author response image 1 to illustrate the number of subjects in each task variant (and their overlap):

      Author response image 1.

      Number of subjects included in each Go/No-go task variant (total n = 21). Values (n) depict overlap of each subject between task variants. One animal was recorded in self-initiated Go/No-go and Cue-initiated Go/No-go/Free (dotted line with arrowheads).

      Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a Go/No Go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI has elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on Go trials, or a longer nosepoke time for No Go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for Go than No Go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.

      The manuscript is well-written, and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions. I have the following comments.

      Re: The lack of relationship between effort to acquire reward in the current study and the magnitude of dopamine release, 1) can the authors unpack this a bit more? 2) Why the difference between the Walton and Bouret studies? Were the shifts in effort requirements comparable across the behavioral tasks? 3) What else could be different between the methodologies?

      We thank the reviewer for the feedback and have responded to each of the three questions below (see points 1-3).

      Firstly, we tried to improve the clarity of our research aims. Our primary comparison throughout the manuscript is between Go versus No-go within each task variant. We ask whether the Go/No-go difference in dopamine signaling persists across different response demands. Thus, testing effort was not central to this main question, but rather a feature of the task that did not affect the Go-No-go dopamine difference.

      (1) Consistent with this aim, we show that VMS dopamine differs between Go and No-go actions persistently across all task variants despite differences in response requirements (action was always accompanied by greater dopamine release compared to action suppression). Our behavioral-training data suggest that the ability to perform short and long tasks differed: rats were first trained to criterion on either a short (∼2 s) or long (∼3 s) Go/No-go variant, with the longer variant requiring substantially more training sessions (short: 18.5 ± 7.6 sessions vs long: 41.1 ± 7.9 sessions; see Author response image 2), indicating behavioral demands were higher for the long-task.

      Author response image 2.

      (2) Regarding the apparent discrepancy with Walton and Bouret (2019), we acknowledge that our original description was imprecise (We wrote: “Previous studies have shown that dopamine signals are influenced by the effort required to obtain rewards”). Our intent was not to suggest a direct contradiction, but rather to emphasize that our findings are consistent with the paper’s broader conclusion that effort encoding by dopamine is limited and highly context-dependent. We have now adjusted the manuscript to better reflect our intent by changing the sentence to, “Previous studies have shown that dopamine signals may be influenced by the effort required to obtain reward but only for particular task conditions (Cousins et al., 1996; Gan et al., 2010; see for reviews, Salamone and Correa, 2024; Walton and Bouret, 2019).”

      For added clarity, these were the main results highlighted in the Walton and Bouret review: Gan et al. (2010) demonstrated that VMS dopamine sensitivity to low-effort costs is prominent early in training (≤ 2 training sessions) and diminishes after extended experience (> 9 sessions). Similarly, Hollon et al. (2014) reported that cue-evoked VMS dopamine primarily tracks reward magnitude with minimal modulation by effort. In line with this literature, our rats were highly trained (≥ 9 sessions until the first recording), and exhibited no difference of average Go minus No-go dopamine between short and long task variants within controllability type (see Author response image 3), supporting the idea that extended training exhibits minimal effort-related modulation of VMS dopamine.

      Author response image 3.

      Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of controllability and effort showed moderate evidence of absence (controllability: BF<sub>10</sub> = 0.324; effort: BF<sub>10</sub> = 0.309). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.101), and the full model including a controllability × effort interaction was the least supported of all models examined (BF<sub>10</sub> = 0.043). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither controllability, effort, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (3) With respect to task comparability and methodological differences, our behavioral paradigm differs in important ways from those highlighted by Walton and Bouret, where effort was often manipulated by training animals to associate cues with different numbers of lever presses within the same session, and typically involved only action initiation. In contrast, our task required both action initiation and action suppression, and changes in response contingencies occurred across separate recording sessions rather than within-session cue-based manipulations. Although these paradigms are not directly comparable, and only had the same dopamine recording technique in common (FSCV), a key takeaway of our results is that regardless of effort differences, VMS dopamine during action initiation is consistently higher than during action suppression.

      I would argue that the cue- vs self-initiated distinction was pretty minor, given that there was a fixed ITI of 5s. How does this task modification compare to those used previously to show that dopamine release corresponds to behavioral controllability? It would help the reader if the authors could spend more time discussing these disparate findings and looking for points of methodological divergence/commonality.

      We agree that clarifying how our manipulation of controllability compares to prior work improves the manuscript, and we have made the necessary adjustments. However, we would first like to correct an incomplete characterization of the task design.

      While the short-task variant used a fixed 5 s inter-trial interval (ITI), the long-task variant employed a variable ITI ranging from 15–25 s. In the long-task variant, the timing of trial onset was less predictable, and we believe this manipulation reduced animals’ ability to precisely estimate when reward pursuit could begin. Under these conditions, whether trials were Self-initiated or Cue-initiated had a substantial impact on animals’ control over the initiation of reward pursuit. That said, we agree that the Self- versus Cue-initiated distinction overall represents a moderate manipulation of controllability compared to those used in studies that focus on controllability.

      A key source of divergence across studies lies in the definition of controllability. We defined controllability as the animals’ ability to choose the time point of beginning the reward pursuit, rather than whether an action was required, and have now added the following sentence in the:

      - Introduction section: “... controllability of reward seeking, defined as the ability to determine when to initiate reward pursuit (Self- vs Cue-initiated trials)...”;

      - Results section: “We defined controllability as the rats’ ability to choose the time point of reward pursuit. In Cue-initiated trials, the time point at which trials could be started was dictated by a cue light, whereas in Self-initiated trials rats were able to choose intrinsically (control) when to attempt a trial start.”;

      - Discussion section

      Importantly, the action requirements for Go, No-go, and Free trials were identical across these trial-start conditions. We found that the degree to which controllability was manipulated in our task was insufficient to modulate the action-specific VMS dopamine signal (Go vs No-go difference), which remained robust across conditions.

      In contrast, controllability has been defined by others as the presence versus absence of an operant action requirement for reward. For example, Goedhoop et al. (2023) directly contrasted operant (lever press required) and Pavlovian (no action required) conditions, removing action execution as a prerequisite for reward. In that context, cues signaling operant control elicited sustained VMS dopamine release, which was interpreted as reflecting anticipation or preparation for executing a learned action. Similarly, Hamid et al. (2021) demonstrated that dopamine “wave” directionality across striatal regions depends on controllability defined by operant versus Pavlovian conditioning.

      Taken together, these comparisons (results from the present study and in the literature) suggest that dopamine sensitivity to controllability may depend on how it is manipulated. We have clarified these methodological distinctions in the revised Introduction, Results and Discussion, and emphasized that more extreme manipulations (such as removing action requirements entirely or increasing uncertainty over trial timing) may be necessary to reveal controllability-dependent changes in VMS dopamine signaling. Alternatively, the apparent discrepancies across studies may primarily reflect differences in the underlying definitions of controllability rather than conflicting results.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Poh et al. investigated whether dopamine release in the ventral medial striatum integrates information about action selection, controllability of reward pursuit, effort, and reward approach. Rats were implanted with FSCV probes and trained in four Go/No Go task variants:

      (1) trials were self-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (2) trials were cue-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (3) trials were self-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued, and effort was increased,

      (4) trials were cue-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued.

      The authors report that dopamine levels rose during Go trials and slowly rose in No Go trials, but this pattern did not differ across task variants that modified effort and whether trials were cued or initiated. They also report that dopamine levels rose as rats approached the reward location and were greater in rats that bit the noseport while holding during the No Go response.

      Strengths:

      (1) Interesting task and variants within the task paradigm that would allow the authors to isolate specific behavioral metrics.

      (2) The goal of determining precisely what VMS dopamine signals do is highly significant and would be of interest to many researchers.

      Weaknesses:

      (1) This Go/No-Go procedure is different from the traditional tasks, and this leads to several problems with interpreting the results:

      (a) Go/No Go tasks typically require subjects to refrain from doing any action. In this task, a response is still required for the No Go trials (e.g., continue holding the nosepoke). The problem with this modified design is that failure to withhold a response on No Go trials could be because i) rats could not continue holding the response, as holding responses are difficult for rodents, or ii) rats could not suppress the prepotent go response. This makes interpreting the behavior and the dopamine signal in No Go trials very difficult.

      We appreciate the reviewer raising this important methodological consideration. We acknowledge that our Go/No-go task differs from traditional paradigms used in humans and primates (e.g. Raud et al. 2020, 10.1016/j.neuroimage.2020.11658; Eagle, Bari & Robbins 2008, 10.1007/s00213-008-1127-6; Roitman & Loriaux 2013, 10.1152/jn.00350.2013).

      However, our design addresses the specific constraints of studying dynamics in freely moving rodents while maintaining the core feature of Go/No-go tasks: requiring suppression of a prepotent response. Our task accomplishes the primary aim of our study, which is to compare VMS dopamine dynamics during action initiation and action suppression, and below we list the reasons why. Therefore, we do not believe that this difference compromises the validity and interpretation of our results.

      It has been suggested for decades that the two main processes governed by mesolimbic dopamine are reward learning and motivated action, and our study aimed to better understand how VMS dopamine integrates reward-related information and motivated action, rather than studying them in isolation. To do so, we trained rats in a modified Go/No-go task.

      More recent work (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019) demonstrates that VMS dopamine signaling incorporates both action and reward-related information, rather than either of the two alone. Importantly, in freely-moving rodents, examining this relationship requires preventing the approach response that occurs when reward delivery is anticipated (Pavlovian bias, go for rewards). Traditional Go/No-go designs that simply require "doing nothing" would not achieve this control in freely-moving rats, as animals immediately approach the reward magazine as soon as reward is inferred (as seen in our Free trials). Thus, we require a No-go condition, as we and others have defined (Syed et al. 2016), whereby animals have to actively suppress the ‘initiation’ response. Action initiation is defined at the beginning of the Discussion section: “... at two distinct points after trial start: 1) when rats began lever pressing (Go), and 2) when rats walked to the reward magazine, either without action requirement (Free) or after successful trial completion (Go and No-go)”.

      To further strengthen our interpretation that we compare action initiation and suppression, and to facilitate cross-species translation of our results (i.e., rodent to human), we also include Free trials, where reward delivery requires no specific action (which are essentially like “doing nothing” trials in traditional tasks). This addition allowed us to directly compare No-go and Free trials, where animals must actively suppress responding while maintaining task engagement, to a condition where no overt action is required for a reward, respectively. The dramatic difference in VMS dopamine between No-go and Free trials demonstrates that VMS dopamine reflects active action suppression during No-go trials, rather than merely the absence of action requirements. This has now been discussed.

      Finally, we only report correct Go, No-go, and Free trials, which differs from that of human go/no-go studies that focus on the failure of appetitive no-go trials (i.e., inhibiting the pre-potent response). In the present study, the dopamine signals that we interpret are restricted to successful trials only: action initiation (moving the lever press), action suppression (i.e., suppressing the prepotent Go response while maintaining their position in the nose-poke port), or no action (no lever press, not staying in the port). While this design differs from human Go/No-go paradigms, we believe our study of correctly performed Go, No-go, and Free trials are necessary for isolating action-dependent components of dopamine signaling in freely moving rats (action initiation vs action suppression vs action free).

      (b) Most Go/No Go tasks bias or overrepresent Go trials so that the Go response is prepotent, and consequently, successful suppression of the Go response is challenging. 1) I didn't see any information in the manuscript about how often each trial type was presented or 2) how the authors ensured that No Go responses (or lack thereof) were reflecting a suppression of the Go response.

      We appreciate the reviewer's attention to this important methodological consideration. The originally submitted version of the manuscript already addressed both concerns raised.

      Trial type presentation frequencies

      The Methods section describes our trial presentation approach: "On recording days, the trial types were counterbalanced. Within a session, Go left, Go right, and No-go trials were presented with 33% probability each, without replacement. For sessions with Free trials, trials were presented with a 25% chance without replacement."

      This design results in overrepresentation of Go trials overall (66% in the short-task; 50% in the long-task), which establishes the prepotent Go response as intended in standard Go/No-go paradigms. To improve clarity, this detail has now been included in the Methods section.

      Ensuring No-go responses reflect suppression of Go response

      Our paradigm incorporates multiple features that ensure successful No-go performance reflects suppression of the prepotent Go response:

      First, the overrepresentation of Go trials (addressed above) establishes response prepotency. Second, during No-go trials, rats must maintain their snout in the nose-poke port for the duration of the action cue, which creates the requirement to suppress the natural tendency to approach rewards (i.e., Pavlovian bias; Jones et al. 2017, 10.1016/j.bbr.2017.05.044; Guitart-Masip et al. 2014, 10.1007/s00213-013-3313-4; Dayan et al. 2006; 10.1016/j.neunet.2006.03.002). In Go trials, such natural bias does not require suppression as the lever can be approached and pressed during the action-cue period. This conflict between the instrumental No-go requirement and the Pavlovian-instrumental bias toward action makes action suppression particularly challenging (consistent with computational accounts of similar paradigms; Lloyd & Dayan 2023, 10.1371/journal.pcbi.1011569; Jones et al. 2017, Guitart-Masip et al. 2014, Dayan et al. 2006).

      Figure 3 provides behavioral evidence of this challenge: animals frequently left the nose-poke port and developed spontaneous motor strategies (such as biting and digging) to stay in the port, suggesting Pavlovian bias interfering with response suppression for rewards. Importantly, all reported No-go data include only correct trials (i.e., those without lever presses), ensuring that the dopamine signal reflects successful response suppression rather than failed Go attempts.

      (2) The authors observe relatively consistent differences in the DA signal between Go and No Go trials after the action-cue onset. However, the response type was not randomized between trial type, so there is a confound between trial type (Go/No Go) and response (lever/nosepoke). The difference in DA signal may have nothing to do with the cue type, but reflects differences in DA signal elicited by levers vs. nosepokes.

      As stated in the Introduction section and discussed in our rebuttal to point 1a, the focus of our investigation is how VMS dopamine signals differ during action initiation versus action suppression for rewards, as this is a central unanswered question in the dopamine field. More recent work demonstrates that dopamine incorporates not only RPE but also action initiation (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019), and our goal is to further our understanding of action-dependent VMS signals during reward pursuit.

      The reviewer suggests that dopamine differences may reflect differences in lever vs. nosepoke rather than cue type (Go vs No-go). We respectfully suggest this concern reflects a misunderstanding by the reviewer of our experimental question. The cue-action relationship is the experimental manipulation itself. It is not possible to study how dopamine encodes instructed action initiation versus suppression without linking specific cues to specific actions. The suggestion to 'randomize' action type across cue types would eliminate the very phenomenon we are investigating: how dopamine signals differ when cues instruct different action requirements.

      Our experimental design specifically compares reward pursuit with action requirements (Go trials: lever press; No-go trials: sustained hold) to reward pursuit without action requirements (Free trials: direct magazine approach). This design allows us to isolate how action initiation and action suppression influence reward-related dopamine signaling, which can reveal how the timing of action initiation influences RPE-dopamine. And which is the point of the study: to show how actions influence RPE dopamine signaling.

      Supporting this interpretation:

      Firstly, trial types were randomly interleaved, and each auditory cue explicitly instructed a specific behavioral response. Our design directly follows established methods demonstrating that VMS dopamine encodes whether actions are initiated or suppressed following action cues (Syed et al. 2016). That study, like ours, intentionally linked cue identity to a specific action requirement to assess how dopamine reflects instructed behavioral control. Thus, the fact that Go and No-go cues map onto different actions is inherent to the question being addressed, not an unintended confound.

      Second, as discussed in our response to point 1b, the asymmetry between Go and No-go trials is theoretically essential. Go trials align with Pavlovian approach tendencies (action initiation to reward), while No-go trials create conflict with this bias by requiring action suppression despite the cue being associated with a reward. This Pavlovian-instrumental conflict makes suppression particularly challenging (Lloyd & Dayan 2023, PLoS Comput Biol 10.1371/journal.pcbi.1011569) and allows us to examine the role dopamine in overriding prepotent responses.

      Third, the inclusion of Free trials (discussed in point 1a) demonstrates that our findings reflect instructed action control rather than simply motor execution. Free trials require neither lever pressing nor nose poke maintenance, yet show dopamine dynamics distinct from both Go and No-go trials, confirming that dopamine signals encode action requirements beyond motor output per se.

      Finally, we demonstrate that VMS dopamine differs in the same trial type (No-go) and can be classified based on different movement patterns (Biting, Digging, Calm). Importantly, the difference in VMS dopamine only appeared after the action was completed, particularly during reward approach (Figure 3). This data argues against the idea that VMS dopamine is particularly tied to the specific operant manipulanda as suggested by the reviewer, but rather, may reflect an internal motivational state for reward.

      Together, the aim of the present study is not to redefine Go/No-go paradigms for rodents, but to utilize this task structure to investigate action-dependent dopamine signalling for rewards, which cannot be answered without the cue-action mapping that we have used.

      (3) Both Go and No Go trials start with the rat having their nose in the noseport. One cue (Go cue) signals the rat to remove their nose from the noseport and make two lever responses in 5 seconds, whereas the other cue (No Go cue) signals the rat to keep their nose in the noseport for an additional 1.7-1.9 s. The authors state that the time between cue onset and reward delivery was kept the same for all trial types, and Figure 1 suggests this is 2 s, so was reward delivered before rats completed the two lever presses? I would imagine reward was only delivered if rats completed the FR requirement, but again, the descriptions in the text and figures are incongruent.

      The reviewer asks whether reward was delivered before rats completed the two lever presses and notes incongruence between text and figures. We respectfully note that these details were stated in the originally submitted version of the manuscript (see below).

      Reward delivery timing

      The reviewer asks whether reward was delivered before rats completed the two lever presses, which refers to the short-task variant. No - reward was always delivered immediately after the second lever press for all Go trials. This is described in Methods Section 3: Behavioral procedures - Self-initiated task variant. For added clarity, we have now added the term “immediately”: “... food pellet dispensed into the reward-magazine immediately.”

      In the Results section, we report that the average latency to complete two lever presses was 1.8s, which closely matches the 1.7-1.9s nose-poke hold maintenance required for No-go trials in the “short” variant. Thus, the time point of reward delivery was matched between Go and No-go trial types.

      Representation of reward delivery timing in figure and text

      The reviewer's confusion appears to stem from the schematic representation in Figure 1 and the task variant structure. There were two overarching task variants with different trial requirements:

      “Short-task” variants: Go trials required two lever presses (completed on average in 1.8s);

      No-go trials required 1.7-1.9s nosepoke maintenance

      “Long-task” variants: Go trials required a ‘rewarded’ press to occur 2.7-3.2s after cue onset (completed within ~3s); No-go trials required 2.7-3s nosepoke maintenance.

      For simplicity in depicting action-cue onset in Figure 2c, we used grey shading with a speaker icon at approximately 0-2s and 0-3s to represent these two variants. This schematic representation was not intended to indicate the precise reward delivery time, which (as stated in the Methods) occurred only upon successful completion of trial requirements in the short-task variant, or 2s after successful completion of trial requirements in the long-task variant. To improve clarity, we have adjusted the legend of Fig. 2 for more clarity, adding “Shaded gray area depicts approximate duration of action-cue onset for “short” and “long” task variants.”, and included more information under Results: “In the short-task variants, action-cues switched off after trial completion and a reward was delivered immediately” and “n the long-task variants, action-cues switched off after trial completion, or in the case of Free trials after 3s, and reward was delivered 2s later (Figure 1a).”

      (4) The manuscript is difficult to understand because key details are not in the main text or are not mentioned at all. I've outlined several points below:

      (a) The author's description in the manuscript makes it appear as a discrimination task versus a Go/No Go task. I suggest including more details in the main text that clarify what is required at each step in the task. Additionally, providing clarity regarding what task events the voltammetry traces are aligned to would be very useful.

      We respectfully note that the requested details were already present in the originally submitted version of the manuscript (see below). However, we acknowledge that the task design is complex and may benefit from additional clarity in the main text to aid reader comprehension.

      Behavioral task

      The reviewer suggests our task appears more like a discrimination task than a Go/No-go task. We acknowledge that our paradigm differs from traditional Go/No-go tasks used in humans and primates, as discussed in our responses to points 1a and 2. However, we classify this as a Go/No-go task because it shares the defining feature: requiring action initiation (Go) and suppression (No-go). This classification is consistent with established rodent literature examining action initiation versus suppression (Syed et al. 2016).

      Moreover, as discussed in our response to point 1a, we included Free trials specifically to demonstrate that No-go trials require active suppression rather than discrimination alone. The distinct dopamine dynamics across Go, No-go, and Free trials confirm that our task captures action initiation, action suppression, and action-free states, which is an important contrast that we needed to address our research question about action-dependent dopamine signaling.

      The key requirements for each trial type and task variant are described at the beginning of the Results section. A full description of each step required in the task was provided in the Methods Section 3: Behavioral procedures, to avoid repetition in the Results. Specifically:

      Self-initiated Go/No-go task variant (“short”)

      Self-initiated Go/No-go/Free task variant (“long”)

      Cue-initiated Go/No-go and Go/No-go/Free task variant

      To improve clarity, we have added a reference in the Results section to the Methods: Behavioral procedures for additional procedural details.

      Voltammetry trace alignment

      The events to which voltammetry traces are aligned were stated in the legend:

      “... when traces were aligned to action-cue onset…”

      “... aligned to the time when animals departed the nose-poke port…”

      “... we realigned traces to the moment animals arrived at the reward magazine… “

      Figure 2c legend: "c) Traces aligned to action-cue onset"

      Figure 2d-e legend: “d) Traces aligned to nose-poke exit and e) magazine arrival….”

      However, to improve clarity, we have now added “aligned to action-cue onset.. “ to make the trace alignment more immediately apparent when results are first presented, and added: “Dopamine data were aligned to events of interest: action-cue onset, nose-poke exit, and magazine arrival.”

      (b) How many subjects were included in each task variant? The text makes it seem like all rats complete each task variant, but the behavioral data suggest otherwise. Moreover, it appears that some rats did more than one version. Was the order counterbalanced? If not, might this influence the DA signal?

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures, where each task variant description includes the corresponding sample size (“Self-initiated Go/No-go task (‘short’; n =9)”, “Self-initiated Go/No-go/Free task (“long”; n = 5)”, “A total of n = 11 and n = 15 were included in the Cue-initiated Go/No-go and Cue-initiated Go/No-go/Free tasks, respectively.”

      For added clarity, we have also added a table for the separation of No-go trials and the number of subjects that it has come from in the Methods.

      Task variant completion

      Not all animals completed all task variants. As stated in Methods Section 4: Real-time dopamine recordings and analysis, animals had to achieve >60% success rate for each trial type on at least two consecutive training sessions to proceed to recording. Other reasons include electrode degradation before all recordings could be completed (See Author response image 1).

      Training order

      We trained four cohorts of animals. One cohort was trained first in the short-task variant

      (Cue-initiated Go/No-go), and the remaining three cohorts were trained first in Cue-initiated Go/No-go (3s) long-task variant (i.e., without Free trials, behavioral and FSCV data not presented in the manuscript). Training order was not fully counterbalanced due to constraints described above.

      Following additional analyses, our data suggest that training order did not influence our core comparison of Go minus No-go dopamine. To directly address whether training order influenced dopamine signals, we separated animals based on whether they were first trained in the Cue-initiated Go/No-go (2s) or the Cue-initiated Go/No-go/Free (3s). We calculated the average dopamine of each rat during the action-cue period and then calculated the difference between them (Author response image 4). We observed absence of evidence of a difference between the groups.

      Author response image 4.

      Average Go minus No-go dopamine reveals no effect of the initial training variant (2s-first vs 3s-first) or test task version (Go/No-go vs. Go/No-go/Free). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of the initial training variant and test task version both showed moderate evidence of absence (initial training variant: BF<sub>10</sub> = 0.378; test task version: BF<sub>10</sub> = 0.367). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.134), and the full model including an initial training variant × test task version interaction was the least supported of all models examined (BF<sub>10</sub> = 0.069). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither the variant animals were first trained on, the task version administered at test, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (5) There is a major challenge in their design and interpretation of the dopamine signal. Both trial types (Go and No Go) start with the rat having their nose in the noseport. An auditory cue is presented for 2-3 s signaling to the rat to either leave the noseport and make a lever response (Go trial) or to stay in the noseport (No Go trial). The timing of these actions and/or decisions is entirely independent, so it is not clear to me how the authors would ever align these traces to the exact decision point for each trial type. They attempt to do this with the nose-port exit analysis, but exiting the noseport for a Go trial (a rat needs to make 2 lever presses and then get a reward) versus a No Go trial (a rat needs to go retrieve the reward) is very different and not comparable.

      We respectfully disagree with the reviewer’s assertion that our data alignment approach is problematic. Aligning neural activity to specific behavioral epochs that occur at different times and across conditions is a widely used method to investigate the relationship between neural activity and behavior. Just to mention some examples: data collected with fiber photometry (e.g., Tan et al. 2026, doi: 10.1038/s41386-026-02368-4; Hart et al. 2024, doi: 10.1016/j.celrep.2024.113828) and voltammetry (Hamid et al. 2016, doi: 10.1038/nn.4173; Syed et al. 2016, doi:10.1038/nn.4187).

      Alignment method

      We intentionally designed the task so that overall action timing is matched between trial types (as described in our response to point 3), while specific behavioral epochs occur at different times. This allows us to compare dopamine dynamics during comparable behavioral events across Go and No-go trials (e.g., nose-poke exit).

      We align data to three critical behavioral epochs, stated in the Methods Section 4: Real-time dopamine recordings and analysis - FSCV measurement and analysis: action-cue onset, nose-poke exit, and magazine arrival. Each alignment addresses a specific aspect of our research question:

      Action-cue onset alignment captures VMS dopamine dynamics when animals have explicit knowledge of trial type and the required action. This allows us to characterize how dopamine evolves following correct action selection, which is central to our research question about how dopamine differs during successful Go versus No-go action execution, as well as no overt action (Free) in the long-task variant.

      Nose-poke exit alignment captures dopamine dynamics at the moment animals initiate movement. The reviewer suggests that exiting for Go versus No-go trials is "very different and not comparable" because subsequent actions differ (two lever presses vs. direct reward retrieval). However, this is precisely our experimental manipulation: we compare dopamine signals when animals exit the nose-poke port to perform different actions. This comparison is both valid and necessary to address our research question (does dopamine encode action initiation?).

      Magazine arrival alignment captures dopamine dynamics at reward approach. This allows us to differentiate between spatial proximity to reward from other concepts including temporal proximity (how soon is reward) and action requirements.

      We acknowledge that we cannot identify the precise moment of decision formation. In fact, the precise moment of decision formation is irrelevant for our question. However, our aim is to characterize VMS dopamine dynamics during successful action execution for rewards, following cues associated with specific actions.

      (6) The voltammetry analysis did not appear to test the hypotheses the authors outlined in the intro. All comparisons were done within task variants (DA dynamics in Go vs. No Go trials, aligned to different task events), but there were no comparisons across task variants to determine if the DA signal differed in cued vs self-initiated trials.

      Our aim was to investigate whether VMS dopamine signals consistently differed between action initiation and action suppression during reward pursuit. To test this within-variant contrast (Go > No-go), we manipulated how reward pursuit is initiated (self- vs cue-initiated) and the “effort” requirements (short vs long task variants), and our results show that they did not affect the differential between Go and No-go.

      The consistent Go > No-go dopamine that we observed across all task variants, together with the consistent increase during magazine approach, supports our conclusion that VMS dopamine integrates motivated action and reward.

      The reviewer suggests that we should have compared dopamine signals across self- vs cue-initiated task variants. We acknowledge this is an interesting, but entirely different, question and have addressed it in the Discussion section. We note that differences in controllability altered the time course of increased VMS dopamine, presumably by triggering earlier positive RPEs in Cue-initiated tasks as compared to Self-initiated tasks (illumination of the nose-poke light being the earliest predictor of reward). However, since our primary research question relates to the difference in VMS dopamine between action initiation and suppression, our results show that this difference was unaffected in two variants (short and long), strengthening our conclusions about the relationship between action and reward-related dopamine signaling.

      (7) Classification of No Go behaviors was interesting, but was not well integrated with the rest of the paper and was underdeveloped. It also raised more questions for me than answers. For example:

      (a) Was the behavior classification consistent across rats for all No Go trials? If not, did the DA signal change within subjects between biting vs digging vs calm?

      (b) If "biting rats" were not always biting rats on every No Go trial, then is it fair to collapse animals into a single measure (Figure 3C).

      (c) Some of the classification groups only had 2 or fewer rats in them, making any statistical comparison and inference difficult.

      Behavioral classification for each rat was consistent across trials (i.e., 100%, see Author response image 5). Upon reviewing the consistency of classifications within individual animals, we found that “Biting” animals exhibited biting behavior across the majority of their No-go trials. Only one animal (in the Self-initiated Go/No-go "short" variant) showed mixed classifications across trials, occasionally exhibiting digging or calm behavior. For all other animals, the predominant behavioral classification was highly consistent within subjects across sessions. The occasional trials where “biting” animals did not bite were too infrequent to permit meaningful within-animal comparisons. Therefore, we believe collapsing animals by their predominant behavioral phenotype in Figure 3C is appropriate and accurately represents stable individual differences in No-go response strategies.

      As stated in the Methods Section 6: Statistical Analysis - Clustering No-go behaviors and regrouping animals (last-line), we specifically avoided between-subjects statistical comparisons for groups with n≤2, as this would be inappropriate (see Author response image 5), and reported qualitative observations only. These exploratory findings at the individual-trial level suggest behavioral heterogeneity during No-go trials, that others may use for future investigation, but do not form primary conclusions.

      The behavioral classification is integrated with our central findings on VMS dopamine encoding spatial proximity. Our results demonstrate that individual variation in action suppression strategy, in particular Biting behaviors, consistently manipulates the timing of max dopamine release during subsequent reward approach, but not during the action itself (Figure 3b-c). This links our observations of action-dependent dopamine (Figure 2c) with spatial reward approach (Figure 2e). We believe that our findings shed new light into the understanding of how action modulates reward-related dopamine dynamics at the individual level. This has been discussed in Discussion section: Dopamine dynamics are linked to motivated action.

      Author response image 5.

      Rats predominantly stick to a particular strategy to perform No-go trials. Each bar represents an individual animal, and colours represent the % of each classification type.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures: It would be helpful to have more panel labels on the figures - 1C, for example, labels 4 different dataset panels, similar to a few other cases. This would really help to improve readability, as there are a lot of task types, trial types, behavioral measures, and labels to sift through. The figures overall are very busy, and it is a bit of a challenge to process the tasks and trial comparisons, as well as interpret what the quantification insets mean.

      We thank the reviewer for their suggestion. We have added more panel labels to Figure 1 to improve the ability to follow with the main text. We hope that the reviewer finds it acceptable.

      (2) Figures: A bit more specifically, in Figure 2, the boxplot insets are pretty hard to see, and it's not clear what scale they are on or what data they reflect. Similarly, it's unclear what the horizontal bars reflect in terms of which conditions are being compared. Why are box plots used for some comparisons, and why are some comparisons based on time series bootstrapping, but others are not clear? I would consider broadly reworking this figure and its description for clarity. Figure 3, by comparison, is easier to understand - the quantifications are clearer and labeled.

      We have now labeled the comparisons being depicted by the horizontal bars in Fig 2c. We have also clarified the boxplot analysis in Fig 2d and Fig 2e by adding 'Max dopamine’ labels to the figure, and have made these analysis methods more explicit in the figure legend.

      We used time-series bootstrap analysis to identify when dopamine signals diverged between trial types, which depicts dopamine differences during distinct action requirements. The latency-to-max quantification provides a summary measure to test specific encoding hypotheses (temporal vs. spatial proximity to reward).

      (3) Broadly, more clarity on the FSCV analysis is warranted.

      (a) Targeting: It looks like the dataset contains a mix of medial shell and mostly core accumbens placements. The paper treats VMS as a uniform dopamine region, but it is more standard to separate core and shell (and also other parts of the shell) into subregions. Many of the reported encoding profiles here are known to differ across the accumbens. So, some consideration of this seems appropriate - a minimal signal-behavior comparison for the shell vs core subgroups, for example.

      We thank the reviewer for raising this point. Upon careful re-examination of our histological analysis, we identified an error that occurred when we made the overlay of electrode placements across rostrocaudal planes (to project placements onto a single plane for the sake of simplicity; Fig 2a): We incorrectly assigned some recordings to the nucleus accumbens shell. We corrected this error, which shows that the vast majority of recordings were in the core, with only 2 animals in the shell (black stars). We have adjusted Fig 2a to reflect this, and added an anatomically more complete illustration of electrode placements across rostrocaudal planes as Supplementary Figure S8.

      While we acknowledge reported differences between core and shell dopamine in some contexts, the small number of shell placements precludes meaningful statistical comparison. However, to determine whether average dopamine concentrations differed between core and shell during the action-cue period, we have plotted the values in Author response image 6. The fact that shell data mostly falls centrally into the overall core-data distribution suggests no consistent difference in dopamine release.

      Furthermore, in our experience (and that of colleagues (personal communication)) with appetitive operant tasks, core and shell FSCV dopamine signals do not substantially differ for action-selective encoding and reward approach. Given the sample distribution and our focus on general VMS function in Go/No-go behavior, we believe pooling these regions is appropriate.

      Author response image 6.

      Average dopamine release during the action cue in nucleus accumbens core (circles) and shell (stars) animals showed no distinct separation between regions. Each symbol represents an animal.

      (b) Design: In my understanding of the design, the main distinction between the short and long task variants is a 2-second versus a 3-second required nose poke hold. 3 seconds here is "long" and more "difficult". I'm not sure I agree that a 1-sec distinction really reflects a difference in task difficulty or effort. Can the authors point to a past paper that demonstrates this variation is sufficient to engage a behavioral difference and/or a neural encoding difference? Broadly, some justification of the validity of this manipulation is needed, I think.

      While we lack direct citations for this specific manipulation, we believe that the 2s and 3s hold requirements represent meaningful differences in difficulty based on our extensive rat-behavior experience and behavioral evidence in Author response image 2.

      Importantly, the difficulty of No-go trials does not stem merely from the required time to hold their snouts in the nose-poke port, but from suppressing the motivational/Pavlovian bias to approach reward-associated cues. In No-go trials, subjects must suppress the prepotent tendency to immediately approach reward-related stimuli and instead maintain active suppression of this approach behavior. Even the 2s hold is challenging as animals tend to perform better on Go trials compared to No-go trials (Fig 1b and 1c), demonstrating the inherent difficulty of response suppression even at the shorter duration. The additional 1-second substantially increases this demand, as it represents a 50% increase in hold duration.

      Our training data clearly demonstrate the difficulty in reaching task criterion when increasing the required action (for Go and No-go trials) from 2s to 3s. Across four cohorts of animals trained in Go/No-go task variants, one cohort that was trained first in the 2s task variant, and the remaining three cohorts were trained in the 3s task variant of Go vs No-go. Animals required 41.1 ± 7.9 (n = 27) sessions to learn the 3s hold (approximately 8 weeks), versus

      18.5 ± 7.6 sessions for the 2s hold (approximately 4 weeks; mean ± SEM). This indicates that despite only a 1-second difference in required action performance, animals needed more than double the number of training days to reach criterion, clearly indicating differential effort demands.

      (c) Figure 2 results: the authors state that because there is a greater DA signal to Go vs No Go cues in all the task variants, this means that controllability of reward pursuit and increased task effort do not affect VMS dopamine. But the magnitude of the signals looks different across the task variants - it looks clearly stronger overall in the self-initiated tasks, for example. Given that dopamine signals are not compared across task variants (I think the tasks are all between-subjects?), I don't think the above conclusion is justified.

      We respectfully clarify that our conclusion does not claim controllability and effort have no effect on dopamine magnitude, but rather that these manipulations do not affect the action-selective difference in dopamine (Go > No-go). Our central finding is that the relative difference between Go and No-go remains consistent across all task variants (within-subjects comparison).

      We did not perform across-variant comparisons of absolute dopamine magnitudes because that was not our primary research question. Our focus was to understand whether dopamine differs between action initiation and suppression, and whether this difference can be modulated by controllability or effort.

      We acknowledge the reviewer’s observation that absolute magnitudes appear larger in self-initiated vs cue-initiated task variants. We believe that this likely reflects differences in RPE timing rather than controllability per se: in cue-initiated tasks, the nose-poke light provides an early trial-start signal, distributing RPE temporally across the trial. In self-initiated tasks, trial-initiation and action requirements are temporally integrated. Though understanding how controllability affects absolute dopamine magnitude is an interesting question for future work (e.g., using sophisticated regression-based encoding models), it was beyond the scope of our current investigation, which focuses on action-selective encoding.

      (4) Broadly, I don't think these data, as shown, support the conclusion "dopamine reflects action initiation but not controllability or effort" without more analysis and additional context.

      We have revised the wording of our conclusions throughout the manuscript to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). It is now “Reward-related dopamine depends on action initiation irrespective of controllability and effort”.

      (a) Figure 3 - more description of the classified behaviors would be helpful for interpreting this part of the data. When are the behaviors occurring - during the hold cue? Or is the classification related to what they do immediately after holding? Or something in between> I guess I'm not sure what digging and biting are in the context of a nose poke hold. As described, it's not clear what the signal differences relate to - movement differences? Generally, it's not clear what to make of the behaviors. They seem to emerge spontaneously, but it's not clear whether the specific actions mean anything, so it's a bit difficult to know what to glean from the dopamine is greater during "biting". It's a very different movement pattern, so perhaps this result relates to that, rather than task engagement or motivational drive per se?

      We thank the reviewer for the comment. We have added relevant information in the figure caption and in the Results section to clarify that classified behaviors occurred during the action-cue period (for No-go trials, the hold cue; Figure 3a caption). In addition, we have included Videos to better depict the classified No-go behaviors during the action-cue period.

      We agree with the comment that these classified behaviors, such as biting, seem to emerge spontaneously. Our interpretation of these behaviors is that they may represent the motivational state of each subject. Most importantly, whereas the behavioral differences occurred during the action-cue period (while animals had to suppress actions and stay within the nose-poke port), the difference in VMS dopamine was only observable after this behavior was completed. Thus, the movement pattern per se is likely not relevant to the dopamine release occurring after its completion. This has been discussed in the Discussion section: Dopamine dynamics are linked to motivation action.

      Based on our videos, it appears as though Digging could be perceived as more vigorous (i.e., more general movement in the nose-poke port). However, we did not observe more dopamine during the action-cue period of Digging trials as compared to Biting trials. Furthermore, more vigor during the action-cue period (e.g. Digging trials) did not result in more dopamine during the reward approach period. Together, the data suggest that another process may underlie the large increase in VMS dopamine in Biting trials during reward approach, such as varying attribution of incentive salience.

      (b) In some cases, but not all, dopamine measurement comparisons are done on a total trial basis, and in others, it seems to be subject averages. It's not clear why different approaches are used for different parts of the data. But also, for the trialwise analysis, what statistical steps were taken to incorporate the subject as a random factor in the analysis? If that is not done, then a trial-wise analysis artificially increases the power for the stat (n=trial#).

      We used different analytical approaches depending on sample size and data structure. To compute differences in Go vs No-go dopamine within each animal, as intended by our experimental design, we performed subject-level comparisons (Figure 2).

      For the behavioral classification analysis (No-go, Figure 3), we performed trial-level analyses to increase the statistical power and better characterize this unexpected and interesting phenomenon. We explicitly chose not to perform subject-level group comparisons because

      (1) some groups had only n=2-3 animals, making subject-level statistics underpowered, and (2) behavioral classifications were highly stable within individual animals (see Author response image 5). We acknowledge that formal between-group comparisons (across subjects) are underpowered due to small n, but the stability of within-subject No-go behavioral strategy and qualitatively distinct VMS dopamine profile suggest that these differences may be biologically meaningful and worthy of future investigation in larger samples. We have made these limitations more explicit in the Results.

      This relates to Figure 3, where all trial data are shown next to individual subjects - the subject-wise group comparisons are between 2-5 or so rats, which is quite low. In Figure 2, a subject n of 27 is listed, so it's not clear why this analysis is on such a small set of rats. Generally, it's not clear how many rats/subjects are in each data bit. The methods say only 5 rats are in the long self-initiated task, but 15 in the cue-initiated task. Clarity in all this is needed, including in the figures/captions.

      We have now added detail about the statistical test performed in Figure 3’s caption:

      “After action-cue offset: No-go (trials)’: Individual No-go trials classified by No-go behavior, and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” We have also added in a table in Methods (Table 1), showing the number of trials in each No-go classification, and animals regrouped based on their predominant No-go strategy (see Methods section: Statistical analysis - Clustering No-go behaviors and regrouping animals).

      For the small sample sizes based on the regrouping of animals based on their predominant No-go strategy, we have now added in the caption, “... Rats classified based on their predominant No-go strategy, with no statistical tests performed.”

      (5) I'm also a little confused about the paper's narrative that the dopamine data reflect spatial (but not temporal) proximity to reward - it seems that this conclusion is based on the dopamine signal peaking at magazine entry, but that is different, I think, than a spatial signal per se (space is not manipulated in this study). I think more analysis of the signals during the magazine approach behaviors would be helpful, and possibly comparing rewarded vs unrewarded approaches. The emphasis, including in the title, that a major take-home of the data is that dopamine encodes reward proximity, is not really borne out by the current analyses. Reward expectation is not manipulated independently of the approach action, so it's hard to pin this on "space" vs "reward is soon". This is admittedly a general complexity in characterizing dopamine ramps.

      As the reviewer notes, we acknowledge that 'spatial proximity' and 'reward is soon' are challenging to fully dissociate in appetitive approach paradigms. However, we believe that our data and new additional analyses, which is now included in Results: Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards and Supplementary Figures S10-12, provide compelling evidence that VMS dopamine primarily reflects spatial proximity to the expected reward location, rather than temporal proximity to reward delivery.

      Evidence against full temporal encoding:

      (1) Max VMS dopamine occurred up to 3s before reward delivery in Free trials (Fig 2e, open circles vs triangles). Furthermore, when we calculated the max values of individual trials of realigned traces, maximum dopamine does not consistently coincide with reward delivery across trial types (new Supplementary Figure S11).

      (2) If dopamine encoded temporal proximity from the earliest reward-predictive cue, we would expect consistent accumulation from cue onset (action cue for self-initiated, nose-poke light for cue-initiated). However, realigned trials also did not show a consistent accumulation around these events (new Supplementary Figure S12).

      Evidence for spatial encoding:

      (1) Max dopamine consistently occurred when animals arrived at the magazine across all trial types and task variants, regardless of when reward was actually delivered (Fig. 2e).

      (2) Across individual trials, max dopamine values concentrated around the magazine-panel, with a striking accumulation when animals were in close proximity to it (new Supplementary Figure S10) showing a distance-dependent distribution.

      (3) Assessment of unrewarded magazine approaches during the intertrial interval (ITI) revealed no increase in dopamine release (Fig 2f), indicating that VMS dopamine requires task-relevant reward expectation.

      (6) Examples of the DLC workflow in a supplement would be appropriate. Also, video examples of the 3 kinds of behaviors from the clustering analysis could be useful for understanding what they are/what they mean.

      We have added supplementary figures showing the DeepLabCut workflow (Supplementary Figure S9) and Videos 1-3 demonstrating the three behavioral classifications (biting, digging, calm) during No-go trials.

      (7) Referencing/scholarship: I would suggest broadening the citation pool for the paper to include more older work that has established the notion that dopamine signaling and the accumbens act as a motivation-action interface, as this has been a longstanding notion since at least the 1980s. There is also a sizable literature on dopamine signaling of effort, some of which would be appropriate to cite here.

      We thank the reviewer for the recommendation and have now expanded our citations to include foundational literature to work from the 1980s-90s. These can be found in the Introduction, Results, and Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 353-357- This came as a surprise, as the relevant results are only featured in supplementary information. These should be moved to the main manuscript. As both the biting behavior and faster lever-press completion lead to larger peak dopamine, does this represent response vigor?

      We appreciate the reviewer’s interest in these data. However, we believe that the data that we report in “Supplementary Fig 7: Quartile analysis of Go trials shows coordinated changes in last lever-press timing and VMS dopamine, does not warrant movement to the main manuscript.”

      The purpose of the lever-press timing analysis was to demonstrate that maximum dopamine release coincides with the moment animals arrived at the magazine, rather than with the action period itself. By sorting Go trials based on last lever-press latency, we show temporal coordination between action completion and dopamine timing but critically, dopamine peaks after action completion, not during it.

      This temporal dissociation argues against a 'response vigour' interpretation. If dopamine encoded motor vigour, we would expect the signal to coincide with or precede the vigorous action. Instead, both the lever-press data (Supplementary Figure S7) and the classified No-go trial data (Figure 3, biting behavior) show that changes in max dopamine occur after the actions themselves, during the subsequent approach to reward.

      Together, these findings demonstrate that VMS dopamine reflects spatial approach to the reward location following action completion, not the vigour of the required actions per se. The lever-press analysis serves as supporting evidence for this temporal relationship but does not introduce a novel finding that warrants main figure emphasis. We mentioned this temporal coordination in Results to ensure readers are aware of the converging evidence while maintaining focus on our central findings regarding action initiation versus suppression.

      (2) Line 13- the experimental work cited refers to midbrain dopamine neurons, rather than dopamine release within the VMS. Please correct.

      We thank the reviewer for pointing out this error. We have now corrected the citation to refer to studies measuring striatal dopamine rather than midbrain dopamine neurons.

      (3) Line 166 - Shouldn't this say "consistently delayed for No Go trials"? Looks like peak dopamine occurs later for these trial types.

      This statement refers to traces aligned to nose-poke exit (not action-cue onset). The peak Go dopamine occurred later than other trial types following exit (Green arrows in Figure 2d).

      (4) Lines 391-392 - however however

      Thank you for the comment, we have adjusted the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

      Thank you for the favorable assessment of our manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOGdomain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor remarks:

      (1) A major new experimental finding of this paper is the affinity of TOG domains, which is more than an order of magnitude lower (10 nM) than previous measurements from the same lab (~200 nM). The authors attribute this change to ionic strength differences between buffer conditions, citing the lab's previous work (Ayaz et al., 2014). This argument left me contemplating what the buffer conditions are in both experiments, and I wonder if other readers would feel the same. After going down the rabbit hole, I believe the difference in ionic strength is ~2.3 fold, and at least on the back of my envelope, this works out beautifully with the measured differences in affinities. A short version of this argument may strengthen the manuscript.

      This is a good comment. We should have been clearer about the different buffer conditions. The revised manuscript now explicitly states how the two buffers in question differ in pH and ionic strength. (Page 8, ‘Tubulin binds rapidly …’ section). We tried to perform comparative measurements of TOG:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is stated in the revised manuscript (Page 9, final paragraph before the ‘Unifying measurements …’ section. Along the lines of the reviewer’s ionic strength calculation, and consistent with the increase in affinity we observed with lower ionic strength, we now also state that prior measurements from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with higher ionic strength (100 mM KCl vs 200 mM KCl).

      (2) I am wondering if there may be an alternative explanation to tubulin binding by TOG being the kinetically rate-limiting step for polymerase function:

      TOG + Tubulin ⇌ TOG:Tubulin (fast binding rate, high-affinity binding)

      TOG:Tubulin + MT_end → TOG:MT (tubulin is incorporated into MT, fast transfer rate)

      The binding rate is 3/s, and the transfer rate is 5/s.

      I was wondering if the following step should be considered, which involves a conformational change of tubulin (e.g., straightening) TOG:MT → TOG + MT (ratelimiting straightening and unbinding of TOG from the lattice).

      This is an interesting thought that highlights a gap in the understanding of microtubule dynamics.

      Presumably, the affinity of TOGs for straight tubulin is practically zero for the purpose of this discussion, as there is no lattice binding, which means unbinding is likely very rapid; however, straightening may be the rate-limiting factor here.

      In theory, straightening should also be rapid; however, we lack measurements of how fast or slow this step occurs within the context of a TOG domain, which presumably skews the process towards curved tubulin.

      We agree (based on prior observations) that the affinity of these TOGs for straight tubulin is negligeable in this context. There is much less data about the timescale of tubulin straightening, with or without a TOG domain bound, or even about how tubulin interactions with the microtubule end affect the balance of preferred conformations and/or the rate of conformational change. It’s an extremely interesting topic. Because the straightening process the reviewer envisions is zero-order, the transfer rate in our model could in principle reflect slow straightening (in this view the ‘delivery’ step would need to be very fast, i.e. not rate-limiting). Because there is so little data about this, and because there are not yet methods to study or perturb the timescale of straightening on the microtubule, we prefer not to engage too deeply. We added a sentence to acknowledge this alternative possibility in the revised manuscript (bottom of Page 4 and top of Page 5).

      A hypothetical Stu2, when bound to the microtubule end and with the TOG domain not disengaged from tubulin, would not permit the processivity of that molecule or the binding of a new molecule.

      To emphasize the importance of unbinding, when it is not efficient, as reported for the T238 mutant that results in Stu2 lattice binding (Geyer et al., 2018), the polymerase becomes inefficient.

      The mechanism of polymerase processivity has not been conclusively determined (the Geyer et al. 2018 eLife paper took a step in that direction, though). The model used in this paper is only concerned with how many polymerases are at the microtubule end at steady-state (as opposed to how long a particular polymerase acts before dissociating), so while we appreciate and are interested in these questions, we think it would be better to leave them for future work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al. investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers, and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed-to-open transitional state in which Stu2 unfurls and loads two longitudinally associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Thank you for the nice summary and favorable comments.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model, which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or the use of a third independent assay to measure affinities, would have added rigor.

      This is a good question that was also raised by reviewer #1. We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is now stated in the revised manuscript (page 9, final paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been independently shown to depend on ionic strength in a way that seems consistent with what we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, final paragraph before the ‘Unifying measurements …’ section).

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Thank you for the push to be more explicit about the two contrasting models. We made small changes to the introduction (top paragraph on page 3) and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (page 12, penultimate paragraph of the main text).

      The ‘transfer’ reaction is treated as irreversible (analogous to catalysis by an enzyme), so the biochemical model we use for the polymerase cannot account for polymeraseinduced microtubule depolymerization at low to zero tubulin concentration. We added text to state that the model is limited to the growth reaction (page 4, last paragraph) but otherwise prefer to not engage too deeply in questions about the reverse reaction.

      These polymerases are thought to be plus-end specific because of the domain organization of the protein: TOGs bind tubulin such that the N- to C-terminal polarity of the TOG corresponds to the plus- to minus-end polarity of the tubulin, and the basic region used to make a ‘slippery’ connection to the microtubule is located C-terminal to the TOGs. These two factors mean that it is only at the plus-end that TOGs can engage αβ-tubulins with the basic region contacting surfaces ‘deeper’ in the polymer. We added text about these issues, citing prior work, to the legend of Figure 5 (page 11). We chose to not elaborate much since it is not the primary focus of the paper.

      Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributed to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends, where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a step-wise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase, which bears similarity to Ena/VASP, and suggest a convergent strategy for cytoskeletal polymerases.

      Thank you for the good summary and favorable comments.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across the cytoskeletal fields.

      Thank you for these comments.

      Weaknesses:

      The study is without major weaknesses, but there are several minor weaknesses worth noting. One is that the final model provides new details regarding the Stu2 mechanism, but does not provide a major new advance in our understanding of how the polymerase works. For example, the discussion does not clearly argue for whether the new results and model rule out either of the prior models. This appears consistent with the 'concentrating reactants' model, but does it clearly rule out the 'polarized unfurling' model?

      Thank you for pointing out what in retrospect was an obvious ‘loose end’ in our discussion. The other referees raised the same point. We made small changes to the introduction and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (top paragraph of page 3 and new penultimate paragraph of the manuscript on page 12).

      A second minor weakness is that the comparison to Ena/VASP is not developed at a deep level based on the final model. I found these ideas exciting and want more critical consideration here, but perhaps it is better suited for a commentary piece to follow.

      We appreciate the enthusiasm and understand where this comment is coming from. Because there has been a fair amount of recent movement in the understanding of TOG domains and what they can do, and because some of the mechanistically interesting parallels entail speculation, we agree with the referee’s suggestion that a future commentary will provide a better venue.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Other minor remarks:

      (1) Figure 2 C is missing the label for what should probably be TOG1*-TOG2.

      Fixed

      (2) Figure 5, lower left, is oddly cropped, showing residuals that are slightly distracting from the beauty of the model.

      Apologies that the figure did not look good in the initial submission. We adjusted it and it looks much better in the revised submission.

      (3) The dynamics assay buffer composition stated in the protein purification Method section is not the same as the PEM buffer used for the dynamics assay. And both are different from BRB80, which, with the chambers, are rinsed. This may very well be accurate, but it raises the question of why not stick to one version, as they are virtually the same.

      Thanks for asking these questions, and sorry for the confusion. First, we should have used different names for the (barely) different buffers. This has now been corrected. Second, the PIPES concentration was not 90 mM, it was 100 mM as in our prior work and this discrepancy failed to get caught in proofreading. Why the other small differences? It’s a good question. The differences reflect an arbitrary decision made at the beginning of the work, there is not a deeper rationale.

      (4) Out of curiosity: Why 90 mM PIPES and not 80?

      Why not 80 mM PIPES? This is just a historical difference. The early measurements of yeast microtubule dynamics (e.g. Gupta … Himes MBoC 2002 and Bode … Himes, EMBO Rep 2003) that partly inspired us to use yeast as a model system used 100 mM PIPES as the working concentration, and we never deviated from that.

      (5) Please state the source of PIPES.

      Sorry for the oversight, we have added the source of PIPES (it is Millipore Sigma P6757).

      Reviewer #2 (Recommendations for the authors):

      (1) Page 4, last paragraph, the authors call out "Fig 1C" which I believe should be "Fig 1D".

      We fixed this, thank for catching this error

      (2) Figure 2C: The authors subtract basal tubulin polymerization (growth rate) from the rates measured in the presence of Stu2 constructs. One assumption in doing this is that Stu2 polymerization activity does not compete for the ability of tubulin (not bound to Stu2) to polymerize on the plus end. I think this is a logical assumption, but it would be beneficial for the authors to state this assumption.

      We said this explicitly in the revised submission (first full paragraph on page 7), borrowing from the reviewer’s phrasing.

      (3) Figure 2C: the label for the last bar is missing - presumably: " Stu2 (TOG1*-TOG2)".

      Fixed.

      (4) Figure 2D: Many of the KM values determined are at the border for points measured, or in one case, beyond the concentration of tubulin sampled. In this regard, the authors should discuss how well the fitted curves correlated with their data. Also, as Vmax and KM are calculated, it would be beneficial if another panel were produced (e.g., Figure 2E) in which the data were presented as a Lineweaver-Burk plot. Doing so, the authors would be able to test their enzymatic model by doping the system with their nonpolymerizable tubulin mutants, which should yield competitive inhibitor behavior but not change Vmax.

      Thanks for pointing this out, we should have been more explicit about this point. We incorporated into the results section an explicit statement about this limitation (first full paragraph on page 7). The suggestion to use blocked mutants and Lineweaver-Burke plots is an interesting one that we hope to pursue in future work using blocked or other mechanism-specific mutants. But we think to do so is complicated enough to be beyond the scope of the present study.

      (5) Figure 3: The authors quantitate the amount of Stu2-GFP fluorescence at microtubule plus ends using line scan analysis of the kymograph. Since the kymographs are processed images, it is more appropriate to integrate intensity from the original frames collected using a circular area. i.e., ID points on the kymograph, and return to the respective position in the corresponding frame to calculate background-subtracted GFP intensity at the plus end.

      This is a fair point. We chose to stick with the kymographbased analysis because the symmetry of the point-spread function and lack of rapid variation in GFP and/or background intensity means that the kymograph analysis is adequate for the intensity-based comparison we were doing.

      (6) Page 7, last line, the authors call out "Fig 2D" which I believe should be "Fig 1D".

      Sorry for the error, we have corrected it.

      (7) Figure 4A: The authors discuss the "sortase epitope," but technically, an epitope is the binding site specifically for an antibody, not to be used in general terms for proteinprotein interaction sites. As such, the authors should describe this as the "sortase recognition sequence" or something similar.

      Thank you for noticing this; we had indeed used ‘sortase recognition sequence’ elsewhere in the paper but we did not catch this instance during proofreading. We have now used that same language in the legend for Fig. 4.

      (8) Page 8, the authors state "KM is approximately equal to Kt/Kon and KM negligeable," but I think they mean "...and KD negligeable".

      Thank you for noticing this typo. We corrected it.

      (9) Page 9, first paragraph last sentence: the readership would be aided by modifying the sentence as follows (adding "Kon" and "Kt"): " ... must be kinetically limited by either the rate of TOG:tubulin binding (Kon) or by the rate of TOG-mediated transfer to tubulin to the microtubule end (kt)."

      Very good suggestion, we implemented it.

      (10) Page 9, second paragraph, the authors call out "Fig 2D", but perhaps they intended to call out "Fig 1D"?

      Sorry for the error, the reviewer is correct and we fixed this.

      (11) Page 9, second paragraph: The authors mention that the transfer rate of tubulin to the plus end via a TOG domain is close to the transfer rate of free tubulin to the growing plus end. Can the authors expand on why they are mentioning this comparison?

      Thanks for the push to be clearer about this. The basic idea is that each ‘delivering’ TOG contributes 50% of the background (uncatalyzed) polymerization rate. So the presence of multiple TOGs (in a single polymerase or from multiple end-resident polymerases) can substantially increase the rate of polymerization. We added brief text to try to make this clearer (first paragraph on page 10).

      (12) Page 9, second paragraph: "TOG-TOG2 polymerases" would be better phrased as "TOG2-TOG2 dimeric polymerases". Noting as well that "2" is missing from the first "TOG".

      This is indeed better phrasing and we have adopted it (also corrected the missing ‘2’) (first paragraph on page 10).

      (13) Page 13, BLI methods: The authors should list the final pH for the PIPES buffer (was it pH 6.9 as in the polymerization assay?).

      Sorry for the oversight, we have added the pH and it was indeed 6.9.

      (14) Page 13, BLI methods: What is "LR1-457"?

      LR1-457 is lab-notebook-speak that did not get purged in editing; it refers to the polymerization-blocked tubulin mutant that also carried a sortase recognition sequence. We replaced ‘LR1-457’ with more evocative phrasing and corrected another typo we found there.

      (15) In Ayaz et al. (eLife, 2014) Stu2 TOG1 and TOG2 affinities for tubulin were measured using AUC, for which the fitted curves appeared to correlate with the data quite well. As the authors note, the values were KD = 70 nM and 160 nM, respectively. This contrasts with the BLI measurement for TOG2-tubulin (~10 nM), which suggests that at least one of the experiments was off the mark - or, as the authors do note, that different buffer and salt condition was used could account for the differences, but that the BLI conditions align with the polymerization conditions (though not exactly) and thus are more appropriate to use. In a supplemental discussion, the authors should run the AUC values through the equation for their model and state what types of differences these values could imply for Stu2 mechanism. If the differences are significant for the Stu2 model derived, the authors should give thought as to whether a third assay should be employed to determine TOG-tubulin affinity. Based on the BLI reagents, it appears the authors would be well-positioned to conduct an assay using SPR. As a potential alternative, the authors could repeat the BLI experiment using the buffer conditions from the Ayaz et al., AUC work (25 mM Tris pH 7.5, 1 mM MgCl2, 1 mM EGTA, 100 mM NaCl, 20 μM GTP) - noting that BSA and Triton X-100 may need to be added as well. If the authors are able to replicate the ~160 nM affinity for TOG2:tubulin, this would be a reasonable way to bootstrap to the conclusion that the BLI is measuring affinity correctly and that the current PIPES-based BLI experiments yielded accurate data.

      We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems. Unexpected challenges and personnel turnover made this slower than anticipated. The reviewer’s suggestions are completely reasonable, but ultimately it was not possible for us to get side-by-side results for TOG:tubulin affinity using the same measurement technique, whether it was biolayer interferometry, analytical ultracentrifugation, or isothermal titration calorimetry. For each technique one or the other buffer caused aggregation, non-specific binding, or some other artifact that prevented analysis. The fact that we were unable to compare the buffer conditions in this way is stated in the revised manuscript (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been shown to depend on ionic strength in a way that is consistent with the changes we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added some text to address the comment about affinity and whether/when the shuttle model would hold (first paragraph on page 11).

      (16) A sentence or two in the discussion, relating how their data aligns (or not) with the Nithiantham model would be beneficial, especially as discussing the two models in the introduction was a central point.

      We completely agree and have now added a paragraph to the discussion to explicitly address the two models and how are or are not supported by the new model and observations (page 12, penultimate paragraph of the main text).

      (17) Discussion: The model in Figure 5 depicts Stu2 engaged with the microtubule, perhaps using its basic linker region (?). The authors could note this in the figure caption for 5A. The authors do not discuss the basis for plus-end polymerization activity versus polymerization activity at both the plus and minus ends. Do the authors propose that this is due to differential Kt values for the two ends and/or differential localization via the basic region to the two ends?

      Thanks for bring this up. The plus-end selectivity of these polymerases is thought to result from the polarity of TOG:tubulin engagement and the positioning of TOG domains relative to the basic region that provides ‘slippery’ binding to the microtubule lattice. We have partially addressed these issues in the legend to Figure 5 (page 11).

      (18) Brouhard (Cell, 2008) demonstrated that XMAP215 can catalyze the depolymerization of GMPCPP microtubules when no free tubulin, or very low levels of free tubulin, are present. This is interesting in that it indicates that the enzymatic activity is reversible. Can the authors comment on how their model would behave in the low-tozero free tubulin concentration regime? Would a different model have to be invoked?

      This is an interesting comment. Because the enzyme-like model treats the transfer step as irreversible, the model cannot account for the kind of ‘depolymerase’ activity Brouhard and others have noted. A more general model that could also encompass the depolymerase activity at low-to-no free tubulin would need to explicitly model the step(s) involved in microtubule association and dissociation. These steps remain a major open question in the field and while it would be quite interesting, trying to address this in a model is beyond the scope of what we can confidently do given the data we have. To be more explicit about this assumption/limitation, we now point this out in the results section where the model is introduced (page 4, last paragraph).

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1D, legend. "...and a transfer rate constant kf that describes how fast...". Should kf be replaced with kt?

      We made this correction, thanks for pointing the problem out

      (2) Figure 2C. The x-axis label under the blue bar is missing. Also, I find the arrows to the left of the bars confusing and unnecessary.

      We fixed the legend problem. We sympathize with the dislike of the arrows but respectfully prefer to keep them in the hopes of avoiding confusion about the fact we are fitting ‘growth rate attributable to Stu2’, not simply growth rate. The figure legend has been expanded to hopefully smooth this over.

      (3) Figure 3. The kymographs are convincing, but it may be helpful for future studies to state here what the polymerization rates are for 0.6 µM and 1.4 µM yeast tubulin. These values are probably different than what one might expect for mammalian tubulin at those concentrations, and the authors could simply determine them from the slopes in the kymographs.

      Good suggestion, we added the growth rates to the legend (as the reviewer expected, they differ from expectations based on mammalian tubulin).

      (4) Results, page 9, line 11: "...yielded a value of 9.6 nM...". Should this be 8.9 nM, which is that value stated in Figure 4C?

      Actually, these different values are correct. We just wanted to point out that whether we used response amplitudes or measured on- and off-rates, we get very similar values for K<sub>D</sub>. We changed wording to hopefully make this clearer: “Calculating the dissociation constant K<sub>D</sub> from the measured rate constants (K<sub>D</sub> = k<sub>off</sub>/k<sub>on</sub>) instead of from the amplitudes yielded a value of 9.6 nM, in good agreement with the amplitude-based determination of 8.9 nM.” (page 9).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. Author Response:

      We are grateful for the careful and extensive reviews, and are pleased that the reviewers found the work of broad interest to sensory processing. Please find our proposal for revision based on public reviews:

      Reviewer 1

      1) Request for dose-response curve for DL-TBOA and leak current or RMP. We can provide this, at least for the initial phase of the curve relevant to the concentrations used for synaptic experiments. Prolonged exposure to higher concentrations leads to very large cationic currents (through AMPAR) which appear to be damaging to membrane integrity.

      2) We will increase the N for glial vs neuronal block with the blockers we already used; this seems more practical than doing new experiments with different concentrations of TFB-TBOA. 

      3) We can include data to test the effect of blockers or small depolarizations on excitability.

      Reviewer 2

      1) We differentiated experiments with “25-50 uM” from 200 uM DL-TBOA because the higher concentration clearly led to massive AMPAR activation and depolarization block, as shown in Fig 1. We then chose lower concentrations to minimize background current while allowing glutamate build-up during exocytosis.  We felt we were clear on this point. 

      As to reversibility and “off target effects” like synaptic changes, we will provide this information. See also response to Reviewer 1, comment 1.

      2) See response to Reviewer 1, comment 3.

      3) We are certain that increasing stimulus strength increases the number of stimulated fibers, and this is well accepted. The stimulus electrode is placed in the auditory nerve root, well away from recorded cell and synapses, minimizing current spread to synapses. We can compare PPR for weak and strong stimuli in our current dataset to confirm no effects on release probability. As to variations in the intrinsic properties of myelinated auditory nerve fibers and their sensitivity to stimulation, there is no information about this, and do not understand how it would be relevant, particularly in as much as we report a negative result: no difference in blocker effect with small or large numbers of fibers active. The 3 main types of myelinated auditory nerve fiber, Type 1a,b,c, are known to respond to different sound thresholds, but that is a synaptic issue in the inner ear, and apparently not related to the myelinated axon bundle.

      4) We appreciate the reviewer's caution about a role for neuronal transporters and will revise accordingly.  We cited molecular evidence for expression of subtypes in the pre and postsynaptic neurons and in glial cells. Of course, given how ubiquitous such expression is across the brain, we suspect the kinds of experiments we provided offer more direct evidence for function.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      We thank the reviewers for their positive comments on our manuscript.

      Weaknesses:

      (1) A highly dense presentation of the simulated models and undefined symbols makes it hard for readers outside the modelling community to follow the biological message. An illustration of the models, accompanied by some explanations and references to the main equations and parameters discussed in this paper, would make the first section much more straightforward.

      Thank you for this feedback. To clarify our methods, we have added Figure 7, which illustrates the dynamics of the point neurons and their synapses. We have also added explanations and definitions of variables right where they appear. These variables were previously defined only in a table on a different page.

      We also moved the methods section to the back of the paper, as is common in many modern manuscripts. We hope that relegating method details to the end makes the manuscript more accessible.

      (2) This methodology appears to be a brute-force approach, requiring millions of simulations to tune 32 parameters in a network of 500-700 cells. It isn't scalable. Moreover, the authors did not use cross-validation, which, with a relatively low increase in computational cost, would provide a quantitative measure as to how well it generalizes; this combination raises doubts about both scalability and reliability.

      Scalability is indeed a key challenge of SBI methods. Amortized neural posterior estimation (NPE) is a brute-force approach in that it samples solely from the prior distribution, which is extremely wide. Many of these samples are therefore not very informative for the biologically plausible dynamics we are interested in, which is a downside of amortized NPE. However, amortized NPE is extremely scalable because once the estimator is trained, it can estimate the parameter distribution of any given output dynamic. We tried to build an amortized NPE for our simulator, but simulation-based calibration (a method to validate posterior estimates using additional simulations) showed that the estimators were unreliable.

      Sequential NPE is not a brute-force approach because it samples from posterior estimates, which are narrower than the prior. Because the amortized NPE failed, we use sequential NPE to create the two estimators for the baseline and the hyperexcitable condition described in the paper. While millions of prior samples are used to generate the initial posterior estimate, which is then sequentially refined, the sequential refinement requires only 80,000 additional simulations. This requires a significant amount of computational resources, which is why we consider the results worth reporting, but we make the simulator, the simulation results, and the trained estimators available, so other researchers can use or train their own estimators without running millions of simulations. We hope our rewrites make the advantages and disadvantages of the approach clearer.

      Regarding reliability and cross-validation, we agree that our initial submission has fallen short. We presented results from only one density estimator per condition, which we considered sufficient given the large number of samples. In the revised version, we present the key results from two additional density estimators trained on partially new training data (Figure 4).

      (3) Several parameters remain so broadly distributed after fitting that the model cannot say with confidence which specific changes matter. Therefore, presenting them as "compensatory levers" is somewhat questionable.

      It is indeed difficult to determine which changes matter because of the simulator’s complexity. Especially the marginal correlation coefficient (Figure 2 C) are small and the pairwise histograms are broad (Figure 2A), because all other parameters are unconstrained. But the conditional correlation coefficients are larger (Figure 2D) and narrower (Figure 2B). We have added the histograms in Figure 2B in the revised version to highlight the difference. We cannot provide a definitive threshold for correlation coefficients to discriminate between important and unimportant mechanisms. Therefore, compensatory mechanisms discovered with SBI should be validated mechanistically, as we do in Figure 5.

      (4) Every conclusion is drawn from simulated data; without testing the predictions on recordings, we have no evidence that the proposed interventions would work in real neural tissue. Because today we cannot diagnose which of the three modelled pathological regimes is actually present in vivo, the paper's recommendations cannot yet be used to guide therapy.

      This is indeed an unfortunate drawback of our current work. We are working to apply this approach to constrain microcircuit simulators with data from epilepsy patients. But that work is currently ongoing and will not fit into the present manuscript.

      Recommendations for the authors:

      Beyond the issues I wrote above, which are methodological, I would like to raise my concern about the way this manuscript is written:

      We highly appreciate this editorial feedback on clarity and style. Such feedback is rare and we have worked to address each point to improve the manuscript.

      (1) Paragraphs - several paragraphs start with: "To quantify/identify/find specific compensatory mechanisms of hyperexcitability with simulation-based inference". It is a good idea to orient the reader with the specific goal of each section, but it is not helpful to repeat the overall message of the paper in every paragraph. Several paragraphs open with "However," or "Additionally,". Please restructure sentences so that connectors appear after a clear topic sentence.

      We have done major rewrites to improve the readability of our manuscript. We start paragraphs with more specific context sentences, rather than the broad research goal, and also made paragraphs much shorter with clearer main messages.

      (2) Section 1: The way you present NPE, it would seem like it's specific to neuroscience (and it's not). The paragraph starting at line 26 is not clear. Please revise it. Line 29 - missing a "." before the next sentence begins. Avoid phrases like "for the longest time".

      We now stress that NPE, like SBI, is used across scientific domains.

      (3) Section 2 was tough to read. Please present each equation in its own numbered display, followed immediately by a plain-language explanation of every symbol and parameter. Provide an illustrative diagram: a small schematic of the AdEx neuron, synaptic connections, and the three perturbations. Even a simple block figure will orient nonexperts. Keep critical methodological decisions (priors, summary statistics, simulation length) in the main text, but move voluminous tables of parameter bounds, learning rates, and hardware specs to the supplement. Remove mentions of which Python functions you used. Readers care about algorithmic choices, not function names. Please reserve specific code references for the GitHub README.

      We have added schematic panels at the beginning of Figures 3 & 4 and added Figure 7, which illustrates the neuron and synapse models of the simulator. We also made major rewrites to the methods section to remove programmatic implementation details and define variables where they appear.

      In general, I think it would be a good idea to have an editor to polish syntax, verb tense consistency, and punctuation. A thorough language edit will improve the readability and impact of this manuscript.

      We have attempted to improve the points raised by the reviewer. In particular, we have carefully rewritten verb tense and punctuation throughout the revised manuscript.

    1. Author response:

      We are very happy that our work was positively received by the reviewers and editors and we are looking forward sharing our results via eLife. 

      We have added more details on the generation of the Tribolium brainbow-lines and we have submitted the respective plasmids to Addgene and give the respective IDs. Some additional minor changes were done to make the text more clear. 

      Public Reviews: 

      Reviewer #1 (Public review): 

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both. 

      Strengths: 

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain. 

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells. 

      Weaknesses: 

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects. 

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single cell labelling”, which admittedly limits both precision and use.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function. 

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an e ect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately. 

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question and a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative. 

      Indeed, we do not reach single cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content. 

      We also think that combining our transgenic line with dopamine-expression was su icient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review): 

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically. 

      Strengths: 

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism. 

      Thank you for this encouraging comment. 

      Weaknesses: 

      No weaknesses were identified by this reviewer. 

      Comments: 

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect: 

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." 

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ 

      such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons  https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/ 

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest. 

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal directed navigation, respectively."

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of Toothy, and for recognizing it as a potentially valuable resource for standardizing dentate spike (DS) analysis across labs. We are especially glad that the reviewers found the manuscript clear and easy to follow, judged the detection and classification algorithms to be appropriate and well-validated, and appreciated the tool's graphical user interface (GUI) based, pip-installable design for lowering the barrier to entry for DS analysis.

      We also understand the concerns raised. Most importantly, we will resolve the data-ingestion and classification errors that reviewers encountered and release an updated version of Toothy that we have verified end-to-end across input formats and datasets. Alongside this, we will provide downloadable demo dataset(s) spanning multiple file formats, probe types, and recording conditions, so that users can confirm a correct installation and see how these cases differ.

      To make the pipeline more transparent, we will add a section describing what Toothy does between user steps, including the rationale for decisions users cannot change, such as detection from a single representative channel. We will also expand the documentation of parameter choices with supporting citations and alternatives, and surface this guidance within Toothy where feasible, consistent with our aim that the tool not function as a black box.

      We will clarify Toothy's scope and current limitations. Recordings with irregular spatial sampling (e.g., tetrodes) are supported for detection but not for CSD-based DS-type classification, which requires a laminar probe spanning approximately the hippocampal fissure to the hilus; we will state this explicitly and evaluate adding an optional waveform-based classification mode (Santiago et al., 2024) to extend type classification to such recordings. We will also add data-quality checks (including sampling rate and inter-electrode spacing) that warn users when a recording may not support reliable results.

      Finally, we will situate Toothy among existing open-source toolboxes, describing how it differs, extends beyond, and interoperates with them, and we will add a comparison of Toothy's outputs to previously published analyses while being explicit about the limits of such comparisons. We will of course also address the remaining technical clarifications and figure edits raised by the reviewers.

      We are confident that addressing these points will make Toothy clearer and more useful to the hippocampal community.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the efforts of the reviewers, which have provided insightful and helpful comments to improve the manuscript. The feedback touches upon a number of topics, focusing on clarification or justification of experimental techniques and on understanding the mechanism by which P. aeruginosa detects HOCl. All reviewers raised the issue of how HOCl activates fro expression, including whether free or protein-bound methionine, cysteine, or other HOCl byproducts induce this expression. For the upcoming revision, we plan to perform experiments that address this issue and will discuss potential mechanistic models in light of the new data. In addition, we plan to perform additional experiments to address a reviewer’s concerns regarding the dependence of the fro response on HOCl production by neutrophils. The revision will correct imprecise statements pointed out by reviewers, and address all remaining issues requiring clarification or further discussion, including the range of HOCl sensitivity, relationship between HOCl and flow sensitivity, and justification for testing the fro response to nitric acid.

      We have completed a number of experiments and responded thoroughly to reviewer comments below. We thank the reviewers again for their details comments, which suggested additional experiments and interpretations that have resulted in significant additional insight into the potential mechanism of HOCl sensing and its relevance with neutrophils.

      Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during the flow of fluids and appears to be regulated by the sigma factor FroR and its antisigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cell invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow, and the data were presented clearly.

      We greatly appreciate the efforts of the reviewer and thank them for their succinct summary of the manuscript.

      Weaknesses:

      (1) Lines 60-62: Some of the authors' conclusions are not supported by the data and thus appear unfounded. One example: "we determine that fro upregulation.....These data suggest a novel mechanism..." Their data do not show that MSR upregulation is a direct effect of FroABCD. Instead, it could be possible that the FroR sigma factor also controls the expression of msr genes, which would be independent of froABCD.

      We thank the reviewer for pointing out this important distinction. We have clarified in lines 63-65 in the clean version of the revision that MSR upregulation depends on FroR rather than FroABCD.

      (2) The authors show increased fro transcription both in neutrophils and macrophages; however, the two types of immune cells differ quite dramatically with respect to myeloperoxidase activation and HOCl production.

      Neither has this been discussed nor considered here.

      We agree that the distinction between the cell types is important and have added a brief description of the differences in respiratory bursts and ROS production between the two cell types and our justification for focusing on neutrophils in lines 102-103. We think it’s very interesting that Fro appears to be activated by macrophages, which are not associated with HOCl production on their own. We think that it would be interesting to identify what is inducing Fro in macrophages in future work.

      (3) With respect to the activation of fro expression upon challenge with conditioned media from stimulated neutrophils, does the conditioned media contain detectable amounts of HOCl? Do chloramines, which are byproducts of HOCl oxidation with amines, also stimulate expression?

      This is an excellent question that addresses which molecules Fro is responding to from neutrophils. We have performed additional experiments that confirm that PMA-stimulated neutrophils produce HOCl (Fig. S2) through the use of a commercial hypochlorite sensor assay, which claims high specificity for detecting HOCl. We further confirmed that this production is inhibited by pretreatment with the MPO-specific inhibitor 4ABAH. These data support the interpretation that the Fro response to stimulated neutrophils requires MPO activity, of which the major product is HOCl. We have described this in lines 117-127.

      Our data does not exclude the possibility that other MPO products could activate Fro expression. We were unable to obtain a reliable source of the major secondary MPO product, taurine chloramine, for our experiments, unfortunately. We believe that understanding the potential for secondary products to activate the response is an important and interesting question that can be explored in a future study. We have added a discussion of this in lines 367-376.

      (4) A better control to prove that this fro expression is indeed induced by HOCl in activated neutrophils would be to conduct the experiments in the presence of a myeloperoxidase inhibitor.

      We thank the reviewer for raising this point. We have performed the suggested set of experiments and found that indeed, the pre-treatment of neutrophils with MPO inhibitor 4-ABAH prior to PMA stimulation suppresses the activation of fro (Fig. 2D and Figure 2-figure supplement 1-2). The results are discussed in lines 117-127.

      (5) The work was conducted with two different P. aeruginosa strains (i.e. AL143 and PAO1F). None of the figure legends provides details on which strain was used. For instance, in line 111, the authors refer to Figure S1B for data that I thought were done with PAO1F, while in 154, data were presented in the context of the infection model, which was conducted with the other strain.

      We thank the reviewer for pointing this issue out. We have ensured that strain names appear in all the revised figure legends. To clarify, only mouse experiments and a related RT-qPCR assay used strain PAO1F due to prior IACUC approval of this strain and its use in previous publications.

      (6) It would be good if immune cell recruitment at 2hrs and 20hrs PI could be quantified.

      We previously quantified neutrophil recruitment at the site of corneal abrasion at 24 hours using the same conditions and strains (Ratitong, B. et al., J. Immun, 2022). While we do not have immune cell recruitment data for the 2 hr and 20 hr time points, the previous data show significant neutrophil recruitment near the latter time point, which is consistent with the interpretation that fro expression is activated by stimulated neutrophils. We have discussed this in lines 212-216.

      (7) The conclusions of Figure 4 are, in my opinion, weak (line 187-188; "It is possible that ....."). These antioxidants likely quench the low amounts of NaOCl directly. This would significantly reduce the NaOCl concentrations to a level that no longer activates expression of fro. There is no direct evidence provided that oxidized methionine induces fro expression. Do the authors postulate that this is free methionine, or could methionine and/or cysteine oxidation in FroR increase the binding affinity of the sigma factor to the promoter? Another possibility is that NaOCl deactivates the anti-sigma factor. None of these scenarios has been considered here.

      We acknowledge that our model of HOCl sensing was unclear and thank the reviewer for their insight. This critique is echoed by reviewer #2 in comment 3 as well. We recently found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all known P. aeruginosa anti-sigma factors (Appendix 2—Table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we have proposed an alternative model in which HOCl or secondary RCS molecules are detected through their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-279 and lines 363-367.

      (8) Line 184: The reaction constants of HOCl with Cys and Met are similar.

      We thank the reviewer for pointing out this important clarification. We have revised the sentence to accurately reflect this in lines 243-244.

      (9) Treatment with 16 uM NaOCl caused a growth arrest of ~15 hrs in the WT (Figure 5A), whereas no growth at all was recorded with 7.5 uM in Figure 3A.

      We thank the reviewer for catching this. We have determined that the concentration of NaOCl in the reagent used for this particular experiment was lower than expected, thus requiring a higher concentration to achieve growth inhibition. We have repeated the experiment with new reagent and find that the results (now Figure 6A) are similar to the previous experiment but at a lower concentration of 4 micromolar, consistent with the concentration found to be sub-inhibitory in Figure 3A.

      (10) The concentration range of NaOCl causing fro expression is extremely narrow, while oxidative burst rapidly generates HOCl at much higher concentrations. This should be discussed in more detail.

      We appreciate the reviewer’s comment, which is related to reviewer #2’s comment #9. We have clarified the reported production rates of HOCl, which far surpass the bacterial MIC. After greater consideration, we believe secondary HOCl products including taurine chloramine could have a more significant role in vivo. While this molecule is less potent than HOCl, it is longer-lived, retains bactericidal activity, and retains the ability to oxidize methionine. We have discussed this in lines 377-405.

      Reviewer #1 (Recommendations for the authors):

      (1) Some statements in the text don't match the data shown in the Figures. For instance:

      (a) Figure 2B shows ~65-fold fro expression, but the text states: "...increased expression of fro expression by 30fold..."(line 88).

      The YFP/mCherry value of PMA-stimulated is 67.4 in the Figure (now Figure 2C). The fold-change is computed relative to unstimulated conditioned medium (third column, which has a value of 2.2), which is a 30-fold change. We have added a citation in the main text to the Source Data, which provides these raw values, to help clarify the computation for readers, and added in the legend that the value for unstimulated is greater than 1.

      (b) Figure 3B shows ~35-fold fro expression at 1 uM NaOCl, but the text states: "...NaOCl increased fro expression by up to 74-fold..."(line 110).

      We have clarified that the increase is relative to untreated, added that untreated value is below 1 in the caption, and provided a citation in the main text to the Source Data, which contains the raw values. The change is measured relative to untreated, for which the YFP/mCherry value is 0.48. The value at 1 uM is 35.7, giving a 74-fold change.

      (2) Line 229: While the ∆froR strain was sensitive to HOCl, the strain was tolerant". Please revise.

      We have corrected this typo (now lines 328- 329). We meant to convey that growth was not entirely inhibited in the froR strain.

      Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promoter in an mCherry-expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post-infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. The addition of HOClquenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveal genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article, and some of the models proposed by the authors seem far-fetched, as outlined below:

      We greatly appreciate the reviewer’s efforts and thank them for highlighting strengths and areas for improvement.

      (1) In line 76 the authors claim "Relative to P. aeruginosa that were incubated in host cell-free media, P. aeruginosa in close proximity to human neutrophils or that were engulfed in mouse macrophages appeared to increase fro expression (Fig. 1C)". Counting bacterial cells in Figure 1C shows that 1 in 17 bacteria (5.8%) induce the froA-promotor in media in the absence of immune cells, while 4 in 72 bacteria (only 5.5%) do the same in the presence of neutrophils. Contrary to the authors' claims, it appears that P. aeruginosa actually decreases fro-expression in close proximity to neutrophils. There is a slight increase in fro-expression in bacteria co-incubated with macrophages (3 in 21, or 14.3%). A more rigorous statistical analysis might substantiate the authors' claim, but, as is, the claim "neutrophils increase fro expression" is untenable.

      We believe the images alone do not give an adequate representation of the data and have quantified a larger portion of the data, which has been added as Figure 1D. The quantification supports the original claim that fro expression is increased during co-incubation with macrophages and neutrophils. Since there was not sufficient statistical sampling to distinguish engulfed P. aeruginosa from free ones, this part of the claim has been removed from the text (updated in lines 80-84).

      (2) The authors should explain the rationale behind some of the chemicals used. Why did they use nitric acid? Especially at these high concentrations, a strong acid such as nitric acid might have a significant influence on the medium pH. I understand that the medium is phosphate-buffered, but 25 mM nitric acid in an unbuffered medium would shift the pH well below 2. Similar considerations apply to hydrochloric acid and sodium hydroxide.

      We thank the reviewers for pointing out the need for this clarification. We have updated Figure 3D with a lower concentration of NaOH at 1 uM, which is the same concentration as NaOCl that activates fro expression. Due to the high buffering capacity of our medium, a high concentration of 6 mM NaOH was needed to induce a discernible change in pH and this high concentration of NaOH had no obvious effect on growth. Neither 1 uM nor 6 mM NaOH produced a change in fro expression, consistent with our previous findings that the effect is not due to sodium ions or higher pH. These updated findings are described in lines 181-190.

      Since the effect of chloride is already controlled for using NaCl and the concentration of HCl used was not sufficient to cause a significant change in pH in the buffered medium, we have removed the HCl group from the data.

      We have clarified that nitric acid was used because it is a strong oxidizer that is not found in neutrophils and that concentrations used were near the minimal inhibitory concentrations (lines 145-147 and lines 191-198). We acknowledge that the growth inhibition from HNO<sub>3</sub> could be due to pH or oxidation. However, since no change in fro expression was observed at concentrations approaching the inhibitory concentration, we did not address the potential effects of low pH from nitric acid on fro expression.

      (3) In line 187, the authors state that "It is possible that oxidized methionine increases fro expression" and they suggest a model to that effect in Figure 5D. It is unclear why the authors singled out methionine sulfoxide, since a number of other things get oxidized by HOCl. In line 184, the authors state, in the same vein, that "HOCl oxidizes methionine residues 100-fold more rapidly than other cellular components". The authors should state which other cellular compounds they are referring to. Certainly not cysteine and other thiols, which react equally fast and are highly abundant in the cell: P. aeruginosa contains 340 µM GSH, 140 µM CoA-SH (https://doi.org/10.1074/jbc.RA119.009934) plus free cysteine and cysteines in proteins (based on codon usage, 1.34% of amino acids in proteins are cysteine, while methionine is only slightly more present at 2.10%, although a number of starting methionines are removed from mature proteins).

      We acknowledge that our HOCl sensing model had been vague and unclear and thank the reviewer for their insight. This critique is echoed by reviewer #1 in comment 7 as well.

      Our initial suggestion that methionine sulfoxide was sensed was motivated by the observation that methionine sulfoxide reductases are upregulated by HOCl. However, we have revised this based on feedback from reviewers and further consideration of chlorine redox chemistry. Interestingly, we found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all the P. aeruginosa anti-sigma factors (Appendix 2—table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we propose a model in which HOCl or secondary RCS molecules are detected by their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-278 and lines 362-366.

      (4) Overall (and this is probably not addressable with the authors' data), some very interesting questions remain unanswered: what is the molecular mechanism of fro-induction? How is the FroR/FroI system modulated by HOCl? Does the system sense free or protein-bound methionine-sulfoxide? Are certain methionine residues in these proteins directly oxidized by HOCl? Many "HOCl-sensing" proteins are also modified at cysteine residues or amino groups; could those play a role? And lastly: what is the connection between shear/fluid flow and HOCl, or are these totally separate mechanisms of fro-induction?

      We thank the reviewer for raising these excellent mechanistic questions. Issues relating to HOCl sensing are addressed in the preceding comment.

      Regarding the connection to shear sensing, Padron et al., 2023 found that the detection of flow in P. aeruginosa can be attributed to chemical transport, in particular to H<sub>2</sub>O<sub>2</sub> that was present in growth media. Based on the same principle, we expect Fro to be upregulated in flow at much lower concentrations than those observed in stationary fluids. The activation of Fro and the effects of HOCl would thus be expected to be flow-sensitive. We have commented on this important factor in the discussion in lines 406-417.

      Reviewer #2 (Recommendations for the authors):

      (1) To address 1, the authors could evaluate the microscopic images in the same manner in which they evaluated the other microscopic images, as, for example, presented in Figure 2B or Figure S1B.

      See response to Weakness point (1).

      (2) To address 2, please explain the choice of the chemicals (why nitric acid?), but also provide the pH of the media with those high concentrations of strong acids and bases, and interpret them in light of the permissible pH range for P. aeruginosa growth. More sensible controls might be a lower NaOH concentration in the range that would be reached through the amount of NaOH in the NaOCl stock at the highest NaOCl concentrations used. As for the acids, I don't see a reason to use these acids at these high concentrations. Please explain.

      See response to Weakness point (2).

      (3) To address 3: The authors could specify their methionine-sulfoxide model a bit more, so that testable hypotheses can be developed. If the authors think the FroR/FroI system senses free methionine sulfoxide, they or others could add methionine sulfoxide to the medium and check induction. If they think specific methionine residues in these proteins are oxidized, they could provide evolutionary evidence of conserved methionine residues. Or, based on a structure or structural prediction, they (or others) could mutate methionine residues, e.g., at the protein's surface or potential protein/protein-interaction sites and assess the effect on HOCl-based activation. Or they could consider other amino acids known to be highly reactive towards HOCl and mutate those in a future study.

      See response to Weakness point (3).

      Further comments:

      (4) The headline of the figure legend of Figure S1 seems incomplete. Please mention the flow experiments shown in Figure S1A.

      We have updated the title of this figure, which now appears in the eLife format as Figure 1 – figure supplement 1.

      (5) What is the difference between the data presented in Figure 4C (bars "UTR" and "NaOCl") and the same bars in Figure S2A? Is this redundant or a re-plot?

      In this revision, Figure 4C has become Figure 5B and Figure S2A is now Figure 5—figure supplement 1. Only the 1 uM NaOCl condition is replotted. We have described this in the legend for Figure 5—figure supplement 1.

      (6) Line 228: hpd is more likely a gene of the aromatic amino acid catabolism.

      We thank the reviewer for pointing this out. We have removed the ‘branched’ descriptor in this sentence, now in lines 325-326.

      (7) Line 229: "While the ΔfroR strain was sensitive to HOCl, the strain was tolerant (Figure 5A)". Please clarify. Which strain was tolerant? WT?

      We have corrected this typo (now line 328-329). We meant to convey that growth was not entirely inhibited in the froR strain.

      (8) Line 247: "Activated neutrophils produce HOCl concentrations as high as 50 µM [24]." This "50 µM" number is often quoted; the citation trail typically leads to Weiss et al. 1982 (https://doi.org/10.1172/JCI110652). However, a more factually correct statement based on that paper would be "2 x 10^6 neutrophils, activated with 30 ng/mL PMA at 37C in 1 mL of Dulbecco's buffer can produce around 50 nmol HOCl per hour". In the particular reference 24, Dybpukt et al used a methodology similar to Weiss et al., and here around 50nmol were produced by the same number of cells in 30 min in response to 100 ng/mL PMA. Please clarify accordingly.

      We thank the reviewer for bringing this to our attention. We have altered the language to indicate that the production rate is for this specific set of parameters. Related to this, Reviewer #1 Comment #10 requested a more detailed discussion of the significance of the Fro response, since it is at much lower HOCl concentration than produced by neutrophils. We have discussed this in lines 377-405.

      (9) Line 275: "which is strain PA14 strain". Please clarify.

      We have fixed this error and entered it into the Key Resource Table.

      (10) Line 311: "Cultures containing densities below 10 P. aeruginosa per frame were concentrated using a syringe filter with 0.2 or 0.8 µm pore sizes (Millipore, Burlington, MA)." Isn't that a bit concerning when testing the induction of an operon that is supposedly activated by shear through fluid flow? Did the authors convince themselves that this procedure does not induce fro?

      We do not expect the filtering procedure to cause changes in gene expression because cells are imaged immediately after filtering. Nonetheless, we performed additional experiments (Fig 3F and 2D) entirely without concentrating cells and found that YFP/mCherry levels were consistent with previous data in Figure 3. This rationale has been added to the Methods section under the “Fluorescence and Phase Contrast Microscopy” section.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

      We appreciate the reviewer’s careful evaluation of our manuscript, and we agree that (as with any experimental study) there is still a lot to do to fully understand the mechanisms underlying the linkage between sleep and memory consolidation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependent manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

      We thank the reviewer for their point of view. We have faithfully reported all experimental observations acquired from our enhanced TRIC-LUC dataset as objectively as possible. We note that the reviewer postulates a particular expected activity shift (reduced PAM-α1 activity and elevated DPM activity after memory training), yet it is not clear why. As our data show, this is a highly connected microcircuit and it is not easy to predict how it might change. Our data show that there is change and dismissing our empirically measured results purely based on a theoretical expectation is not justified. The brain operates as an intricately interconnected network; neural activity dynamics cannot always be simply inferred from static circuit polarity. Progress in deciphering neural circuit function relies on iterative rounds of experimental testing. Like most neuroscience investigations, the present study cannot resolve every open question, and we explicitly acknowledge several unresolved directions worthy of future exploration in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      We appreciate the reviewer’s suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      Regarding the question of whether sleep fragmentation of this magnitude per se is critical for long-term memory consolidation, we would like to highlight relevant published evidence. Prior work in rodents and human (Bonnet and Arand, 2003; Van Someren et al., 2015; Ramesh et al., 2012; Baud et al., 2014) and our earlier study (Liu et al., 2019) have demonstrated that alterations in sleep architecture, independent of changes in total sleep amount, can affect multiple physiological processes, including memory. Given this existing supporting evidence, we respectfully argue that this does not constitute a weakness of the present manuscript.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      The concern raised by the reviewer that certain interpretations remain speculative represents a common situation in most published research. This interesting direction warrants further investigation in future work, but falls outside the scope of the present study. In the revised manuscript, we have elaborated on the differences observed between these two GAL4 drivers. We have also conducted additional experiments to investigate a well-characterized memory-related recurrent loop of PAM-α1 neurons in sleep regulation. We respectfully note that no single study can comprehensively address all outstanding questions.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      We appreciate this careful comment on our schematic model in Figure 12. This diagram aims to summarize the key findings obtained in the present study while also explicitly laying out unresolved questions that await future investigation. In our view, including open, outstanding questions in the working model does not undermine the interpretation of our existing experimental results. In the revised manuscript, we have modified the corresponding text to clarify this point and distinguish firmly between conclusions supported by our data and tentative components requiring follow-up validation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca<sup>2+</sup> signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      We fully acknowledge the inherent limitations of the TRIC-LUC reporter system, as pointed out by the reviewer. Every experimental tool comes with characteristic strengths and drawbacks. Although TRIC-LUC suffers from temporal delays, our experiment does not aim to capture acute, immediate effects; instead, it examines long-term dynamics of neuronal activity. To date, within Drosophila neurobiology, no superior technique is available for non-invasive long-term monitoring of neuronal activity in freely behaving flies. We share the hope that new tools capable of reporting neuronal activity in real time will be developed and applied, which will facilitate deeper mechanistic understanding of neuronal dynamics.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

      We appreciate this suggestion. First, our memory paradigm relies on reward-based associative learning, and starvation is required for flies to express robust memory, so to do this would require a completely new experimental set up. Second, our core findings demonstrate that transient perturbations of this neuronal circuit trigger sleep disturbances and memory deficits specifically under starvation conditions. Therefore, measurements of neural activity under fed conditions are not directly relevant to the central conclusions of the present study. We agree that related experiments on other behaviors under fed states constitute an interesting direction and could be pursued in future investigations.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca<sup>2+</sup> signals. I am confident that these questions can be addressed in future studies.

      We greatly appreciate the reviewer’s positive evaluation of our work and the thoughtful suggestions regarding future directions. As acknowledged in the manuscript, our study centers on identifying a shared circuit that coregulates sleep and memory processes. We have performed preliminary investigations of downstream circuitry, and our results indeed support the idea that sleep and memory can be modulated independently. Importantly, we have avoided drawing definitive conclusions that memory deficits arise purely from impaired sleep-dependent consolidation. Instead, we emphasize the existence of a common circuit mechanism governing both processes.

      We also thank the reviewer for drawing attention to dopamine receptor-mediated cAMP signaling and Ca<sup>2+</sup>dynamics. This observation constitutes an additional finding that requires more extensive mechanistic follow-up. Given the scope of the current work, we have not dedicated a separate discussion section to dissecting this pathway. This promising line of inquiry will be pursued in our future research.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further issue, apart from the reverse figure 13 are not found in the main text as indicated.

      We thank the reviewer for this careful check. We have performed a full-text search for Figure 13 throughout the revised manuscript and found no relevant citation. We have also carefully cross-checked all figure numbering and confirm that all figure labels are accurate in the current version.

      Reviewer #2 (Recommendations for the authors):

      In Fig. 8B, some individual GCaMP measurements show values below −100% ΔF/F₀. Under the stated definition, ΔF/F = (Fn - F0) / F0, values below −100% would require Fn to be negative. Since raw fluorescence intensity cannot be negative, values below −100% require further explanation. The authors should clarify whether Fn represents raw fluorescence or processed fluorescence, and whether the plotted traces underwent any normalization, subtraction, detrending, or transformation beyond the stated formula.

      We greatly appreciate the reviewer’s rigorous scrutiny of our data and the valuable question raised. Our fluorescence signals were calculated using the standard formula ΔF/F = (Fn − F0)/F0, where Fn = F_ROI − F_background, and all calculations were implemented accordingly. We have carefully revisited all raw imaging datasets and identified the source of the issue. During initial data processing, we retained all acquired recordings without excluding samples exhibiting focal plane drift. This drift occasionally yielded negative values for Fn (F_ROI − F_background). Beyond the five DPM cell bodies from four brains in Figure 8B highlighted by the reviewer, we further detected six additional DPM cell bodies from four brains in Figure 2A affected by the same artifact. We have now excluded these drifting preparations, regenerated all corresponding plots, and updated the statistical analyses in the revised manuscript. For transparency, we upload both the raw and processed datasets as supplementary materials to clarify this point.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The goal of the study was to address the question of the degree to which social position in a group is a stable trait that persists across conditions. Reinwald et al. use a custom-built cage system with automated tracking and continuous testing for social dominance that does not require intervention by the experimenter. Remixing of individuals from different groups revealed that social position was rather stable and not really predictable from other measures that were taken. The authors conclude that social position is multifaceted but dependent on characteristics like personality traits.

      Strengths:

      (1) Reductionistic, highly controlled setting that allows for the control of many confounding variables.

      (2) Very interesting and important question.

      (3) Confirms the emergence of inter-individual behavior-driven differences in inbred mice in a shared environment.

      (4) Innovative paradigm and experimental setup.

      (5) Fresh perspective on an old question that makes the best use of modern technology.

      (6) Intelligent use of behavioral and cognitive covariables to generate a non-social context.

      (7) Bold and almost provocative conclusion, inviting discussion and further elaboration.

      We thank Reviewer #1 for this constructive and balanced evaluation of our work, and for highlighting both the conceptual importance of the question and the strengths of our automated, highly controlled approach.

      Weaknesses:

      (1) Reductionistic, highly controlled setting that blends out much of the complexity of social behavior in a community.

      Our goal in developing the NoSeMaze was to provide an enriched yet standardized environment that allows animals to interact freely in a complex setting where arena geometry and access contingencies are held constant across groups. As the Reviewer also notes as a strength, this highly controlled setting with minimal experimenter interference enables us to minimize confounds and provides reproducibility across rounds and groups. We now clarify this explicitly and frame the design as a trade-off between ecological complexity and experimental controllability.

      The NoSeMaze is designed as an open system that can incorporate additional configurations. In this first study with the system, we intentionally used a design that focused on single-sex groups without mating, no intruders or external threats, stable environmental conditions, and adult mice. This allowed us to establish social behavior under one defined condition. From here, future studies can add certain levels of ecological and social complexity to progressively understand their impact on specific behaviors and group dynamics.

      Accordingly, we have revised the manuscript to clarify the scope and the role of social and environmental conditions to the here observed phenomena. We also describe how future studies with systematically modified conditions can be used to understand how they change behaviors. We further replaced potentially over-broad terms (e.g., “naturalistic,” “real-world”) with more precise wording throughout.

      We modified the following sections:

      Abstract

      We removed “… within naturalistic mouse groups.” (ll. 52-54) and changed the sentence to “The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      We changed “… in naturalistic groups.” (l. 38) to “… in larger male mouse groups.” and “… from naturalistic tube competitions …” (ll. 41-42) to “… from incidental competitions in the integrated tube tests …”. We also deleted “naturalistic” in l. 50.

      Discussion (ll. 691-701)

      “…The NoSeMaze aims to increase environmental complexity and group dynamics while retaining experimental control, enabling longitudinal high-dimensional phenotyping of individuals. At the same time, it is a controlled laboratory group-housing habitat optimized to capture a subset of the determinants of social complexity present in natural communities. The strength of this design lies in the continuous, observer-independent observation of complex behaviors in defined environmental and social contexts. Accordingly, we interpret our findings as applying to the social contexts and environmental conditions tested here. Building on this, its modular design allows for introducing additional environmental and social factors like stressors, mating behavior, or resource competition in the future to understand their respective impact in modifying social behaviors.”

      Conclusion

      We changed “… complexity of real-world behavior …” (ll. 729-730) to “… complexity of group behavior in semi-naturalistic conditions ...”

      (2) The motivation to enter the test tube is not "trait" (or at least not solely a trait) but the basic need to reach food and water; chasing behavior would be less dependent on this stimulus.

      Tube traversals may reflect different motivations, including routine movement between compartments to access food, the water lickport, the open arena, or the housing area. Nevertheless, we do not interpret tube-entry motivation itself as a ‘trait’. Importantly, our hierarchy readout does not quantify which animals enter the tubes, nor is it confounded by tube-entry frequency itself (cf. ll. 527-532). Rather, it captures the consistent outcomes of incidental dyadic competitions once two animals meet in the tube (push vs. retreat), aggregated across many interactions. Thus, the stable signal we report lies in repeatable competition outcomes, not in traversal propensity. Consistent with this interpretation, the overall number of tube competition events was not associated with social rank, arguing against systematic competition avoidance by low-ranking animals. In contrast, chasing is a voluntarily initiated, asymmetric interaction between an initiator and a recipient and therefore adds a social-action component beyond incidental access-linked encounters, providing a complementary readout of social behavior.

      We also noted that the term ‘trait’ may not be the optimal description in this context and replaced it throughout the ms. with more precise wording, including “internalized social rank” and “propensity to chase”.

      Results

      We added the following paragraph (ll. 524-532):

      “Tube crossings in the NoSeMaze are motivated by the intent to eat, drink, sleep, or socialize. Accordingly, competitions within the tube arise by chance when two animals enter from opposite sides at the same time. Social rank therefore captures the consistent outcomes of repeated incidental competitions (push versus retreat). This measure is not confounded by differences in tube engagement, as social rank was neither associated with participation in tube competitions (ρ = 0.11, p = 0.139; Fig. 6C, Supplementary Fig. S12C) nor with the overall number of tube detections when controlling for chasing (Spearman’s partial correlation between detection count and z-scored David’s score, corrected for the fraction of active chases: ρ = 0.022, p = 0.758).”

      Discussion

      We changed the following paragraphs:

      Line 608-611

      “Importantly, participation frequency and differences in tube entry time did not confound the resulting social ranks, underscoring the robustness of the automated incidental rank assessment.”

      Lines 653-662

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Dominance is only one aspect of sociality, social structure is reduced to rank. The information that might lie in the chasing behavior is not optimally used to explain social behavior beyond the rank measure.

      In this manuscript, we focused on the relationship between social rank derived from incidental tube competitions and chasing. We agree that social rank derived from tube competitions captures only one aspect of social structure. We have now clarified this point more explicitly in the revised manuscript. Here, our specific aim was to determine how chasing relates to competition-derived social rank and whether it provides information beyond rank in these mouse groups.

      We also refer to another manuscript dedicated to additional aspects of social structure captured in the video data, including approach and interaction behavior as well as social clique formation. There, these measures are again examined in relation to chasing and social rank. We also plan future studies leveraging chasing in the NoSeMaze as a readout for neurophysiological investigations.

      In the present study, social rank is based on dominance and subordination in incidental competitions in the integrated tube test. We treat chasing as a distinct, volitional social dimension rather than redundant “rank information”. Importantly, our chasing analyses add beyond social rank information in three ways: (1) structural asymmetry (initiator vs recipient roles) not captured by symmetric social rank measures; (2) elite-centric reciprocal dynamics rather than broad top-down enforcement; and (3) context dependence, with social rank–chasing coupling strengthening when group transitivity is lower. We revised the Discussion/Conclusion accordingly and additionally note that other aspects of social organization (e.g., affiliative bonding and higher-order network structure such as clique/rich-club organization; Nelias et al., 2025) require complementary measures (e.g., video-derived interaction networks) and are not the scope of the present study.

      To avoid ambiguity, we also added the following operational definitions in the Introduction and Results:

      Introduction

      Lines 90-94

      “In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors.”

      Results

      Line 306-308

      “Here, social hierarchy refers to the group-level structure reconstructed from incidental competitions in the integrated tube tests, social rank to an individual’s position within that structure, and social position to social rank together with chasing.”

      We additionally clarified the distinction between chasing and social rank in the following sections.

      Discussion (ll. 632-664)

      “In more constrained or despotic conditions, chasing can serve as a unidirectional, dominance-related behavior directed at subordinates [9,43,44]. The larger groups observed in the NoSeMaze reveal a more nuanced role for chasing behavior. Chasing levels are individually stable, but the expression and meaning of chasing are context-sensitive. Chasing was neither broadly distributed nor consistently directed down the social hierarchy. Instead, it was initiated by a small subset of individuals – primarily those occupying high competition-based social ranks – and frequently occurred reciprocally within this group, suggesting intra-elite social dynamics rather than broad dominance enforcement. Rather than solely serving to impose social hierarchy, chasing appeared to function as a means through which individuals with high social rank monitor, negotiate, and maintain their relative standing within the top tier. Notably, the identity of frequent chasers remained stable over time and persisted across changing group compositions, indicating that the propensity to initiate chases reflects a consistent individual-level tendency rather than a purely situational response.

      However, its coupling to social rank depended on the group’s hierarchy structure. In groups with less clearly defined social hierarchies (i.e., lower transitivity), active chases aligned more strongly with social rank, suggesting that mice in less structured groups rely more on proactive signaling to clarify social rank. Indeed, the top-ranked mice in these groups exhibited relatively high levels of active chases. This context-sensitivity highlights chasing’s dual role: it serves as a tool for negotiating social rank among mice at the upper end of the hierarchy, and additionally functions to establish or reinforce hierarchical clarity when social structures are ambiguous. These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure. These findings reveal a novel aspect of social dynamics: chasing is not merely a dominance display but a flexible context-dependent mechanism shaped by both individual disposition and group-level social structure.”

      Conclusion (ll. 718-723)

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization. Chasing behavior played a dual role, reflecting a stable individual propensity while also adapting to group-level structure.”

      (4) Focus on rank bears the risk of overgeneralization for readers not familiar with the context.

      As already discussed above (see also point 2 and 3), we have sharpened the framing to reduce the risk of overgeneralization. Throughout the manuscript, we now refer to tube-derived social rank explicitly as a dominance-subordination-related axis of social organization rather than a comprehensive measure of “sociality”, and we also avoid the term “traits”, but prefer the use of “internalized social rank” or “propensity to chase” when discussing stability across rounds and contexts. Also, we removed “personality-like” as the scope of the study is to set specific behaviors in relation and not to enter the field of mouse personality classification.

      We clarified the definition of social rank in this manuscript in the Introduction (ll. 90-94) and the beginning of the Results (ll. 306-308, see also point 3).

      Abstract

      Lines 41-46

      “… Across more than 4,000 mouse-days, hierarchies derived from incidental competitions in the integrated tube tests were non-despotic, transitive, and stable even when group compositions changed. This stability supports an internalized component of competition-based social rank. Chasing was also stable across contexts. Notably, chasing was concentrated among high-ranking individuals, consistent with ongoing negotiation of social rank among individuals at the upper end of the hierarchy.”

      Lines 50-54

      “In summary, high-dimensional tracking with the NoSeMaze reveals that social position in mice is multifaceted and shaped by stable dimensions of individual behavior that persist across changing social contexts. The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      Discussion (ll. 622-626)

      “This temporal and contextual stability supports interpreting social rank as a stable, internalized characteristic of individuals. In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes.”

      Changes in the Conclusion starting l. 709 as highlighted in answer to the above point 3.

      (5) Conclusion only valid for the reductionistic setting, in which environment, social and non-social changes only within narrow limits, and in which the mouse population does not face challenges

      Our conclusions are bounded by the conditions tested, namely an enriched but stable environment, controlled group composition and remixing, and the absence of explicit ecological challenges such as resource scarcity or predators. As described above in the answer to point 1, we now state explicitly that our conclusions apply to the social-context variation tested here (controlled group reshuffling) under otherwise stable conditions, and we frame resource-competition and environmental-stressor manipulations as future variation to test of how this manipulation affects the here described behaviors.

      We accordingly changed the Discussion as already highlighted in point 1 above and added ll. 693-704.

      (6) Animals are not naive at the beginning of the experiment, but are already several weeks old.

      Animals entered the study as adults and were continuously group-housed (3-5 mice/cage) before entering the NoSeMaze. To reduce effects of initial apparatus novelty, mice underwent two NoSeMaze habituation sessions prior to data collection (each several hours). We now clarify these points explicitly in the Methods and scope the inference accordingly. One of our future steps is to extend this framework to earlier developmental stages (e.g., adolescence/weaning) to capture full lifespan trajectories. This however first required establishing the approach in adult mice under controlled conditions as done in the present study.

      Methods (ll. 739-747)

      “A total of 79 adult male homozygous OXTRfl/fl mice (B6.129(SJL)-Oxtrtm1.1Wsy/J, RRID: IMSR_JAX:008471, Jackson Laboratory) backcrossed > F10 to C57BL/6J background (Charles River, Sulzfeld) were used for the experiments. Animals entered the NoSeMaze as adults (see Supplementary Table S2 for ages at NoSeMaze entry across rounds) and had been continuously group-housed after weaning (3–5 mice/cage), i.e., they were not developmentally or socially naïve at study onset. Of the 79 mice, 26 mice were injected six weeks before the start of the experiment with an AAV expressing Cre recombinase (rAAV1/2-CBA-Cre) into the AON pars centralis to induce bilateral OXTR deletion (OXTRΔAON). The remaining 53 animals received an AAV expressing only dTomato (rAAV1/2-CBA-dTomato). …”

      Discussion (ll. 629-631)

      “… A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      In summary, this is a wonderful study, but not one that is easy to interpret. The bold conclusion is valid only within the constraints of the study, but nevertheless points in an important direction. The paradigm is clever and could be used for many interesting follow-ups.

      To define social position as a personality trait will elicit strong opposition and much debate; the nuances of the paper might be lost on many readers and call for the (re)-consideration of many concepts that are touched. I find this attitude a strength of the paper, but the approach bears the risk of misunderstanding.

      We thank the Reviewer for the helpful comments. As detailed above, we tightened terminology and framing throughout the revised manuscript by defining stability explicitly in terms of repeatability across rounds and reshuffled groups within this paradigm, and described tube-derived social rank as one dimension of social behavior. We hope these edits preserve the conceptual message while reducing the risk of misunderstanding.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:

      The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:

      The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:

      The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:

      The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      We thank the Reviewer for their evaluation of our manuscript. We appreciate their recognition of the NoSeMaze as an innovative and conceptually integrated platform, the scale and longitudinal rigor of the dataset, and the clarity and appropriateness of the analytical framework.

      Weaknesses:

      (1) Conceptual Novelty and Prior Work:

      While the study is carefully executed and methodologically innovative, several of its core findings reaffirm concepts already established in the literature. The emergence of stable, transitive social hierarchies, the persistence of individual differences in social behavior, and the presence of non-despotic social structures have all been previously reported in mice, including under semi-naturalistic conditions (e.g., Fan et al., 2019; Forkosh et al., 2019). Although this work extends those findings with greater behavioral resolution and scale, the manuscript would benefit from a clearer articulation of what is genuinely novel at the conceptual level, beyond the technological advance.

      We agree with the Reviewer that transitive dominance hierarchies and stable inter-individual differences have been demonstrated previously in mice, including in semi-naturalistic settings. Our intent was therefore to address two more specific conceptual questions that prior work typically has not tested in an integrated way:

      (1) Whether an individual’s social position generalizes across distinct social contexts created by systematic changes in group composition rather than reflecting stability that is only observable within a fixed group;

      (2) How the two frequently employed social dominance metrics tube competition-based social rank and chasing relate to each other in unperturbed larger groups.

      Specifically, the conceptual advance is enabled by (A) continuous, handling-free estimation of competition-based social rank from incidental tube contests over weeks, (B) a repeated group-reshuffling (“accelerated longitudinal”) design that explicitly tests whether an individual’s social rank generalizes across distinct social contexts, and (C) parallel, continuous measurement of chasing and non-social reinforcement-learning behaviors in the same individuals and environment.

      This integration also lets us dissociate dominance-related social dimensions (competition-based social rank vs chasing) and test their relation to individual styles in non-social reinforcement-learning. Importantly, these questions are addressed in larger intact societies (9–10 mice), where hierarchy structure is shaped by more complex network-level dynamics than in smaller groups.

      We revised the Introduction and Discussion to foreground these conceptual points and to position them more explicitly in the context of prior works.

      Introduction

      We adapted the following section (ll. 76-98) and integrated the suggested references:

      “… In mice, small groups tend to form highly despotic hierarchies [24-26], whereas larger groups exhibit more complex structures [27]. These patterns suggest a strong influence of emergent group-level dynamics on social structure [12,27]. Yet, animals do not enter social groups as blank slates [28-30]; stable latent factors in the individual may also contribute to hierarchy formation [13]. Thus, it remains unclear to what extent an individual’s social position is internalized and persists across different social contexts [31], or instead is primarily an emergent property of group-level dynamics. Disentangling these possibilities requires experimental conditions that allow unperturbed, continuous tracking of all individuals in sufficiently large groups, together with systematic changes of group composition to modulate social context.

      While stable hierarchies and behavioral identity domains have been described previously in semi-naturalistic settings [9-13], many studies quantify these features either within fixed group compositions or in separate assays. Here, we therefore use systematic group reshuffling in 10-member societies to directly test whether individual differences in social rank and chasing persist across distinct social groups. In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors. By continuously measuring social rank, chasing, and reinforcement-learning behavior in the same individuals, we further test how chasing and social rank are related to each other and how they relate to individual styles in non-social reinforcement learning.

      To enable these tests, we developed the Non-invasive Sensor-rich Maze (NoSeMaze).”

      Discussion

      We added a brief framing statement at the beginning of the Discussion (ll. 577-583) to clarify the study’s conceptual novelty:

      “Ecologically enriched, yet experimentally controlled assessments allow us to study behavioral individuality and social structure in group-living animals over extended timescales. Here, we show that individual mice carry stable, individual-specific, and multi-faceted profiles of social position and cognitive styles across changing group contexts. Our approach goes beyond prior work by testing cross-context stability under repeated, systematic group reshuffling in larger mouse societies, while measuring competition outcomes, chasing, and reinforcement-learning behavior in parallel. …”

      (2) Role of OXTR Deletion:

      The inclusion of the OXTR manipulation feels somewhat disconnected from the manuscript's central aims. The effects were minimal and transient, and the authors defer full interpretation to a separate study.

      We appreciate the Reviewer’s point and agree that the OXTR<sup>ΔAON</sup> manipulation can appear secondary to the manuscript’s central aims. We included OXTR<sup>ΔAON</sup> because oxytocin-dependent social recognition memory was hypothesized originally to impact potentially also learning an individual’s position in social hierarchy networks. Even though the effects were small and transient, we nevertheless believe it is relevant to report them, also in relation to the more profound effects of the genetic manipulation reported in a related manuscript (Nelias et al., bioRxiv 2025, 10.1101/2025.08.26.672298). Therefore, we explicitly account for genotype as a covariate in the analyses such that the main conclusions do not depend on this manipulation.

      To improve coherence, we have added one sentence in the Introduction motivating why OXTR<sup>ΔAON</sup> was included, and finally more clearly signposted that deeper mechanistic interpretation is beyond the scope of the present manuscript.

      Introduction (ll. 115-120)

      We added one short sentence introducing the rationale behind the perturbation:

      “… and (3) determine whether social rank, chasing, and non-social reward-seeking behaviors represent stable individual characteristics or dynamic features across time and changing group composition. As a secondary analysis, motivated by oxytocin’s established role in social recognition memory [32,33], we also tested whether OXTR deletion in the anterior olfactory nucleus produces detectable shifts in rank dynamics.”

      Results

      Lines 161-164

      “… This manipulation was included as a secondary biological perturbation. The primary analyses and conclusions focus on the platform and cross-context stability, and genotype is treated as a covariate unless stated otherwise. …”

      Lines 451-454

      “… As a secondary analysis, we tested whether OXTR<sup>ΔAON</sup>, which impairs de novo social recognition memory required for social clique formation in this cohort [36], also affects the measures reported here. Consistent with largely internalized features, OXTR<sup>ΔAON</sup> produced only transient effects. …”

      Discussion (ll. 594-605)

      “This study focused primarily on the relation of social rank and chasing. We however also considered their relation to additional variables including the loss of oxytocin receptors in the olfactory cortex in the adult (OXTR<sup>ΔAON</sup>), involved in de novo social recognition learning. The propensity to chase was largely unaffected by OXTR<sup>ΔAON</sup>. Mice carrying OXTR<sup>ΔAON</sup> displayed a transient reduction in social rank during the first week that normalized thereafter. This transient effect contrasts to the persistent impairment by OXTR<sup>ΔAON</sup> in forming higher-order social bonds that enable membership in stable cliques, as identified by video tracking of self-paced interactions in the same cohort [36]. Together, these findings suggest that OXT-dependent olfactory learning is critical for the formation of social context-dependent higher-order bonds, but plays a limited role in shaping hierarchy-related behaviors.”

      (3) Scope Limitations (Sex and Age):

      The study is limited to male mice, and although this is acknowledged, the title and overall framing imply broader generalizability. This sex-specific focus represents a common but problematic bias. Additionally, results from the older mouse cohort are under-discussed; if age had no effect, this should be explicitly stated.

      We thank the Reviewer for this relevant point. The study is limited to male mice. We therefore revised the title and the abstract and strengthened the limitations to make the sex-specific scope explicit.

      We additionally note ongoing work extending the same framework to female groups, where we find similar hierarchy structure and cross-context stability in the NoSeMaze in preliminary unpublished data.

      Regarding age, our design included two separate cohorts of different adult age ranges (young and older adult animals), but age was not the primary experimental factor. As detailed in our response to Reviewer #3 (point 1), we now quantify age structure explicitly and test age effects using a decomposition that separates between-group age differences from within-group age variation (mean age per group and each animal’s deviation from that mean). We also recomputed stability estimates with age-adjusted ICC models, and the resulting ICCs are highly similar to the original estimates (cf. new Supplementary Table S4), indicating that the reported metrics’ stability is not driven by age differences across groups.

      Title

      We change the title from “Individual differences drive social hierarchies in mouse societies” to “Individual differences drive social hierarchies in male mouse societies”

      Abstract

      We also added male in the abstract (ll. 37-38):

      “The interaction of these behaviors in the shaping of social position in larger male mouse groups remains largely unknown.”

      Discussion (ll. 613-617)

      “… While this study focused on male mice, in which social hierarchies are best established [9], future work is needed to explore sex-specific expressions of social structure and their neurobiological underpinnings in female groups. We therefore restrict our interpretation to male mice. Critically, male social ranks were robustly maintained within the same group over time. …”

      For a more detailed integration of age in the Methods, Results, and Discussion, we kindly refer to the reply to Reviewer #3, point 1.

      (4) Ambiguity of Dominance as a Construct:

      While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear. As in prior work (e.g., Varholick et al., 2019), dominance rank here shows only weak associations with physical attributes (e.g., body weight), cognitive strategy, or neuromodulatory manipulation (OXTR deletion). This recurring pattern, where rank metrics are reliably established yet poorly predictive of other behavioral or biological traits, raises important questions about what such measures actually capture. In particular, it challenges the assumption that outcomes in paradigms like the tube test or chase frequency necessarily reflect dominance per se, rather than other constructs.

      We thank the Reviewer for this clarifying point. We agree that stable social rank metrics do not necessarily imply a complete or unitary measure of “dominance.” In the revised manuscript, we therefore clarified that, in this study, social rank is operationalized as consistent competitive outcomes in incidental tube-test encounters in the NoSeMaze (see also Reviewer #1, points 2-4).

      Our data indicate that this competition-based social rank is related to, but not identical with, other social behaviors such as chasing. Likewise, body weight significantly contributes to social rank, but explains only part of the variance. We therefore do not interpret weak or partial associations with other variables as invalidating the social rank measure. Rather, we interpret them as indicating that social position is multidimensional and only partially captured by any single assay.

      In independent subsequent studies that are currently in preparation or revision, we observed that heterogeneity in the neurobiology and response to challenges was best predicted by the competition-based social rank, also compared to the other behaviors assessed here. While these observations are beyond the scope of the present manuscript, they support our view that incidental tube competition and the resulting dominance-subordination structure may provide a biologically informative measure of one important dimension of social position. At the same time, the observation here and in many previous studies that factors such as body weight explain only a limited portion of the variance remains important for our understanding of these constructs.

      We have revised the manuscript accordingly to make this distinction more explicit. In particular, we now define more precisely the terms social hierarchy, social rank, and social position in the context of this manuscript (see also reply to Reviewer #1, point 3; ll. 90-94 in the Introduction and ll. 306-308 in the Results), and we use these terms more consistently throughout. We also revised the text to avoid overstating the meaning of “dominance” where the data support a more specific interpretation.

      Finally, we now emphasize more clearly both the strength and the limitation of the present approach. The NoSeMaze allows these relationships to be assessed continuously in a minimally perturbed group-housing ecology, laying ground for future incorporation of further variables and dimensions to capture how they shape social organization.

      Specifically, we added text in the Discussion and Conclusion to clarify that competition-based social rank and proactive chasing represent separable dimensions related to dominance and subordination, but do not exhaust sociality or individuality, and that future work should integrate additional measures such as affiliative behavior and higher-order network structure.

      Discussion

      Lines 624-6631:

      “… In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes. The tube-derived social rank describes here a dimension of individual social behavior. The capacity to quantify stable individual differences across changing social contexts highlights the value of the NoSeMaze for lifespan-oriented studies of behavioral individuality. A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      Conclusion

      Lines 718-722:

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization.”

      Reviewer #3 (Public review):

      Reinwald et al. present the NoSeMaze, a semi-natural behavioral system designed to track social behaviors alongside reinforcement-learning in large groups of mice. Accumulating more than 4,000 days of behavioral monitoring, the authors demonstrate that social rank (determined by tube competitions) is a stable trait across shuffled cohorts and correlated with active chasing behaviors. The system also provides a solid platform for long-term measurements of reinforcement learning, including flexibility, response adaptation, and impulsiveness. Yet, the authors show that social ranking and chasing are mostly independent of these cognitive traits, and both seem mostly independent of oxytocin signaling in the AON.

      Strengths:

      (1) The neuroethological approach for automated tracking of several mice under semi-natural conditions is still rare in social behavioral research and should be encouraged.

      (2) The assessment of dominance by two independent measures, i.e., spontaneous tube competitions and proactive chasing, is innovative and valuable.

      (3) The integration of a long-term reinforcement-learning module into the semi-natural system provides novel opportunities to combine cognitive traits into social personality assessments.

      (4) The open-source system provides a valuable resource for the scientific community.

      Limitations:

      (1) Apparent ambiguity and inconsistency in age structure and cohort participation across rounds, raising concerns about uncontrolled confounds.

      (2) Chasing behavior appears more stable than tube-test competitions (Figure 4D vs. Figure 3D), which challenges the authors' decision to treat tube competitions as the primary basis for hierarchy determination.

      We thank the Reviewer for the evaluation of our work. We have addressed the limitations raised below with additional analyses, clarifications, and corresponding manuscript revisions.

      Major concerns:

      (1) Unclear and inconsistent handling of age groups and repeated sampling. The manuscript repeatedly refers to "younger" and "older" adults, but it is unclear whether age was ever controlled for or included in models. Some mice completed only one round, others 2-5 rounds, without explanation of the criteria or balancing.

      We thank the Reviewer for this clarifying comment. We now explicitly quantify participation in the different rounds in new Supplementary Table S3. Importantly, our primary stability analyses are implemented using variance-component mixed models (REML) that naturally handle unbalanced repeated-measures data. This is explained in more detail in the Methods section (ll. 1003-1111), where we added a statement that the LMEs are well suited for unbalanced repetitions.

      To further address unbalanced participation, we additionally performed a conservative sensitivity analysis restricted to a balanced subset, including only sessions 1 and 2 and only mice with observations in both sessions. For the key social measures, ICC estimates were highly similar in the full dataset and in the balanced subset (e.g., z-scored competition David’s score, ICC across cohorts: 0.55 without age adjustment vs. 0.56 in the balanced first-two-session subset; active chasing: 0.74 vs. 0.72; being chased: 0.61 vs. 0.60; Supplementary Table S4), indicating that unbalanced participation did not inflate the stability estimates.

      We also addressed age structure explicitly. Although age was not the primary experimental factor, the inclusion of two age cohorts allowed us to assess whether the observed behaviors and their interrelations were robust across most of the adult lifespan (cf. new Figure 1). We therefore decomposed age into a between-group component (mean age per group) and a within-group component (each animal’s deviation from its group mean) and included these terms in the relevant LME models. Tube-based dominance rank (David’s score, z-scored) showed no age effect (p<sub>age, cond.</sub> = 0.63), whereas chasing metrics showed modest age associations (cf. new Supplementary Table S5). This however only indicates that chasing was associated with age to some degree. More importantly, recomputing all stability estimates using age-adjusted ICC models yielded nearly identical ICCs for the core social measures (new Supplementary Table S4), indicating that the reported stability was not affected by age differences across groups.

      In addition, we revised the study design schematic (new Fig. 1; cf. Reviewer #1, Recommendations for the authors) to depict the separate age cohorts and the reshuffling procedure more clearly.

      We adapted the following sections accordingly.

      Methods

      Lines 980-985

      “… Most mice (n = 68) participated in at least two NoSeMaze rounds with reshuffled group members. Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems. We examined the stability of social and reward-seeking metrics by correlating values from the first and second round (Spearman’s correlation, cf. Fig. 5). These round-1-to-round-2 correlations use one paired observation per mouse and are therefore not inflated by mice contributing >2 rounds.”

      Lines 991-993

      “The ICC treated mouse identity as a random intercept and NoSeMaze group (i.e., round-specific social group) as a random effect, with repetition included as fixed effect (i.e., stability across groups while holding repetition means constant).”

      Lines 1003-1011

      “Variance components for ICC estimation were obtained from LMEs fit by restricted maximum likelihood (REML), which yields less biased variance-component estimates and is well suited for unbalanced repeated-measures designs (i.e., different numbers of rounds per mouse). To assess potential confounding by age structure, age was decomposed into a between-group component (group-mean age) and a within-group component (each animal’s deviation from its group mean) and included as covariates. ICCs were recomputed in age-adjusted models (see Supplementary Table S4). Finally, we performed a conservative sensitivity analysis restricted to the first two sessions per mouse and to mice with observations in both sessions (“balanced first2”), to additionally account for unbalanced participation structure.”

      Results

      Lines 167-173

      “… The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (Fig. 1C, see Supplementary Table S2). Mice lived in groups of 9-10 for multiple rounds in the NoSeMaze, with different group members in each round (Fig. 1D, see Supplementary Table S1). This allowed us to test which individual behaviors were stable across different group compositions. …”

      Lines 441-450

      “Because the number of rounds in the NoSeMaze was unbalanced between animals (Supplementary Table S3) and age varied across groups (Supplementary Table S1-2), we also performed robustness checks. Stability estimates changed only minimally when recomputed in age-adjusted ICC models (between-group mean age and within-group age deviation; see Methods) and when restricting analyses to a balanced first-two-session subset (sessions 1–2 only; mice with both sessions) (Supplementary Table S4). Mixed models indicated that some chasing and reinforcement-learning measures showed modest age- and/or session-related shifts in absolute levels (Supplementary Table S5), but importantly, these did not affect the observed stability patterns.”

      (2) Stability of chasing appears stronger than the stability of tube competitions. Figure 4D shows highly consistent chasing behavior across weeks, while Figure 3D shows weaker and more variable correlations for tube-based David scores. This is also evident from Figure 5A-B,D. Thus, it appears that chasing, which serves to quantify dominance in similar semi-natural setups, may be a more reliable and behaviorally meaningful measure of dominance than the incidental tube competitions.

      Indeed, active chasing showed higher cross-round correlations than tube-derived social rank (e.g., R1–R2 Spearman ρ = 0.75 vs 0.57; ICC<sub>across cohort</sub> 0.74 vs 0.55, see Fig. 5 and new Supplementary Table S4). Importantly, this does not contradict the central finding that tube-derived social rank is stable across time and across remixed groups. Rather, it may highlight that chasing and social rank capture different aspects of dominance-subordination-related behavior with different statistical properties. Chasing reflects an individual’s propensity to actively initiate interactions (a strongly expressed, asymmetric behavior), which can be highly consistent across contexts. By contrast, social rank is a relational measure inferred from symmetric dyadic win–loss outcomes based on incidental competitions in the integrated tube tests and can vary with the specific set of competitors and interaction opportunities in each reshuffled NoSeMaze group, while still remaining substantially stable overall.

      Accordingly, we do not interpret the higher repeatability of chasing as evidence that it is the “better” hierarchy measure. Tube competitions yield symmetric dyadic outcomes that directly support formal hierarchy reconstruction (David’s score/Elo, transitivity, steepness) and show convergent validity with traditional tube testing. Chasing, in contrast, is asymmetric and volitional and, in our data, is concentrated in the upper social ranks and modulated by group-level social hierarchy structure, consistent with rank negotiation/signaling rather than a mechanism that assigns a full ordering to all individuals. These different properties lead to very different “win-lose” relationships when comparing tube competition events to chasing events that we specifically illustrated in Supplementary Fig. S13. We therefore revised the Results and Discussion to frame tube-derived social rank and chasing as complementary social dimensions.

      Results (ll. 402-413)

      “… Specifically, stability was high for the tube competition-based David’s score (Fig. 5A, ρ = 0.57, p < 0.001), as well as for the fraction of active chases (Fig. 5B, ρ = 0.75, p < 0.001) and of times being chased (Fig. 5C, ρ = 0.54, p < 0.001). Notably, the fraction of active chases showed slightly higher across-round stability than tube-derived David’s score and the fraction of being chased. This is in line with active chases capturing an individual propensity to initiate this behavior, whereas competition-based social rank is a relational measure that is also influenced by the set of competitors and interaction opportunities in each reshuffled group. Nonetheless also social rank and being chased were overall stable. Across the full series, ICCs for these social measures were also in the good–excellent range, indicating high across-round stability (Fig. 5D, ICC = 0.548 to 0.890).”

      Discussion

      Line 653-664 (see also Reviewer #1, point 2)

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Unbalanced participation across rounds compromises stability analyses. Stability analyses (e.g., ICCs, round-to-round correlations) assume comparable sampling across individuals. However, some mice contribute 1 round, others 2, 3, 4, and even 5 rounds. This imbalance may inflate stability estimates or confound group reshuffling effects, and the rationale for variable participation is not explained.

      We thank the Reviewer for raising this point. Indeed, participation was unbalanced across rounds (see the same Reviewer #3, Major concerns 1 and new Supplementary Table S3), mainly due to missing data from occasional technical failures of the RFID detectors or the reinforcement learning water port during acquisition in some groups (for details, see Supplementary Table S2). We added this rationale behind variable participation to our Methods section (for details, see Major concern 1, ll. 981-982, “Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems.”)

      We now quantify round participation (new Supplementary Table S3) and directly address potential bias from unequal sampling in two ways. First, the round-1-to-round-2 (R1–R2) stability correlations use one paired observation per mouse and therefore are not inflated by mice contributing more than two rounds. Second, ICCs were estimated from REML variance-component mixed models that account for unbalanced repeated-measures by design and use all available observations. To further rule out inflation from unequal sampling or non-random missingness, we additionally report a conservative sensitivity ICC restricted to a balanced subset including only each mouse’s first two observed sessions and only mice with both sessions (“first2-balanced”). Full-sample ICCs and first2-balanced ICCs were highly similar (new Supplementary Table S4), indicating that participation imbalance did not affect the stability estimates. For details on the changes made in the manuscript, see Reviewer #3, Major concern 1.

      Recommendations for the authors: 

      Editor's notes:

      Should you choose to revise your manuscript, if you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      Readers would also benefit from noting that the mice were male in the abstract.

      We have revised the manuscript accordingly and now report exact p-values values wherever possible for the key results in the main text. Because full statistical reporting for every analysis in the main text would substantially interrupt readability, we provide the complete statistical details in a Supplementary Excel File (Supplementary Material – Systematic Statistical Reporting), including sample sizes, degrees of freedom, the number and type of permutation tests, exact p-values, and 95% confidence intervals. We also provide Extended Data Sheets for all linear-mixed effects models that account for covariates such as age and round of participation in the NoSeMaze (cf. Reviewer #3, Major Concern (1) for the additional analyses). To guide readers to these resources, we now explicitly refer to the Supplementary Material – Systematic Statistical Reporting at several points in the manuscript:

      Results

      Lines 271-274

      “Detailed statistical reporting for all analyses, including n, degrees of freedom, exact p-values, and 95% confidence intervals, is provided in the Supplementary Material - Systematic Statistical Reporting. Extended Data includes additional linear mixed-effects models controlling for potential confounding variables.”

      Lines 430-431

      “All statistical details are provided in the Supplementary Material – Systematic Statistical Reporting.”

      Additional references to the supplementary statistical reporting were inserted at lines 252-254, 562-563, and 567-568.

      Methods

      Lines 962-965

      “Full details on the statistical tests, including n, degrees of freedom, exact p-values, and 95% confidence intervals, are provided in the Supplementary Material – Systematic Statistical Reporting, as well as in the Extended Data for the LMEs accounting for different covariates.”

      We also revised the abstract and the title to explicitly state that the mice were male.

      Manuscript changes:

      Results, figure legends, and supplementary tables expanded to include full statistical reporting; abstract revised to specify male mice.

      Reviewer #1 (Recommendations for the authors):

      I would recommend being much more explicit about the reductionistic nature of the study and how the limitations are turned here to an advantage, while at the same time acknowledging the challenges of extrapolating beyond these boundaries. Most importantly, dominance should be positioned more clearly and cautiously within a framework of social behavior (and social structure) in general. The authors include cognitive tests, etc., to generate context, but this context is dependent on the same circumstances that possibly contribute to the social structure. The information that lies in the chasing behavior as an additional measured variable might be used better to provide more context.

      We thank the Reviewer for these recommendations. We have better clarified the framing to make explicit that the NoSeMaze is a controlled laboratory group-housing system. The specific conditions are now discussed in more detail in relation the observed social behaviors. We describe the tube-derived social rank as one dimension of social organization, and proactive chasing as another one. This study presents a first necessary step to understand their shared and distinguishing features. We also elaborated the Discussion to better clarify the trade-off between ecological complexity and experimental control, and to highlight that the value of the NoSeMaze for continuous, observer-independent phenotyping under standardized conditions. We now clearly state the importance to vary conditions in the system to see how the social behaviors and also their relation to non-social features changes depending on context conditions. For detailed changes, see our responses to Reviewer #1, points (1), (3), (4), and (5), and Reviewer #2, Weakness (4).

      Manuscript changes:

      Abstract, Introduction, Discussion, and Conclusion revised to clarify scope, construct interpretation, and the complementary roles of competition-based social rank and chasing.

      The precision of the description of the experimental design should be improved. When were the cohorts mixed, or did they stay separate? Which animals were old, which were young? This remained a bit confusing.

      We revised the presentation of the study design accordingly. Specifically, we clarified that the younger and older adult cohorts were run as separate experimental series and that group reshuffling occurred within, but not across, these cohorts. We also revised the study schematic in Fig. 1 and the corresponding description in the manuscript to more explicitly depict cohort structure, timing of NoSeMaze rounds, and between-round reshuffling (details are provided in the Supplementary Tables 1-3). For new age-related analyses and additional robustness checks, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Figure 1 and its legend were revised to depict cohort structure, NoSeMaze rounds, and between-round reshuffling more explicitly. In addition, the Results subsection “Ecological longitudinal assessment in the NoSeMaze” were updated to clarify that the younger and older adult cohorts were run separately, that reshuffling occurred within but not across cohorts, and which animals belonged to each age-defined cohort. We also now cross-reference Supplementary Table S2 for round-specific age information and cohort composition.

      Results

      Lines 167-170

      “The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (see Supplementary Table S2).”

      It is also not fully clear across which groups stability measures were obtained: are these across all groups or within the subgroups with a given characteristic?

      We thank the Reviewer for highlighting this point. We now state explicitly that stability measures were computed across all eligible rounds, using mixed-effects models that account for repeated observations of the same mouse and for round-specific group membership. We also added a conservative sensitivity analysis restricted to a balanced first-two-session subset. Full and balanced-subsample stability estimates were highly similar, indicating that the main conclusions are robust to the participation structure. For details, see our response to Reviewer #3, Major Concerns (1) and (3).

      Manuscript changes:

      Methods and Results revised to clarify the level of analysis and model structure, as well as new supplementary robustness tables (Supplementary Tables 3-5) added.

      Minor point: In Figure 3, 18 groups are mentioned, but 19 are shown.

      We thank the Reviewer for this clarification. The apparent discrepancy arose because the group labels in Figure 3 follow the numbering of all experimental groups, whereas only groups with available tube-competition data are shown in this panel. Accordingly, group 16 is absent because no tube data were available for that group (see Supplementary Table S2), and groups 20 and 21 are likewise not included for the same reason. We have revised the figure legend and corresponding text to make this explicit and to avoid the impression of a numbering inconsistency.

      “Fig. 3: Social rank derived from incidental competitions in the integrated tube tests of the NoSeMaze.”

      “C, Box plots of metrics characterizing social hierarchy for 18 groups, including transitivity, steepness, stability, and uncertainty-by-repeatability (for details, see ‘Source Data’). Group labels correspond to original experimental group IDs. Only groups with available tube-competition data are shown. Therefore, numbering is non-consecutive (e.g., groups 16, 20, and 21 are absent; see Supplementary Table S2).”

      Reviewer #2 (Recommendations for the authors):

      (1) To better distinguish this study from previous literature, we recommend incorporating a more focused discussion (or adding to the intro) of how the findings advance our understanding of social hierarchy beyond prior works.

      We have sharpened the conceptual framing in both the Introduction and Discussion. In particular, we now distinguish more explicitly between the well-established observation that hierarchies form and remain stable within fixed semi-naturalistic groups, and the more specific question addressed here: whether an individual’s social position generalizes across changing social contexts created by repeated group reshuffling. We added additional references and also emphasize that the present study integrates continuous measurements of competition-based social rank, chasing, and reinforcement-learning features in the same individuals and environment. For details, see our response to Reviewer #2, Weakness (1).

      Manuscript changes:

      Introduction and Discussion revised to foreground conceptual novelty relative to prior semi-naturalistic work.

      (2) Given the weak associations between dominance rank and other traits such as body weight, cognitive performance, and oxytocin receptor manipulation, we suggest further clarifying what is being captured by these measures.

      We agree and have clarified this point throughout the manuscript. We now define tube-derived social rank explicitly as an operational measure based on repeated competitive outcomes from incidental dyadic tube tests, highlighting it as one dimension of social behavior.

      We also make clearer that proactive chasing is not redundant with rank, but instead captures a distinct, partly dissociable dominance-related interaction mode. For details, see our response to Reviewer #2, Weakness (4).

      We understand the question on the meaning of social rank as the correlations to body weight and cognitive performance are only punctual. We would like to mention here already that the competition-based social rank turns out to be a strong predictor of individual reactivity in a series of challenges. These works are in currently in preparation for publication and will make the relevance of these measures more clear. Related to this, we find in these studies that chasing and competition-based social rank predict different behavioral and neuronal aspects of individual reactivity.

      Manuscript changes:

      Abstract, Results, Discussion, and Conclusion revised to clarify construct interpretation and multidimensionality.

      (3) We recommend modifying the title and abstract to more clearly reflect the male-only design of the study. In addition, please indicate whether any age-related differences were observed. If age had no measurable effect, this should be stated explicitly to justify the combination of age groups.

      We revised the title and abstract to make the male-only design explicit. We also now report age-related analyses directly in the manuscript. Briefly, age was modeled by separating between-group age structure from within-group age variation. Some measures showed modest age associations in mean level, but most importantly, age-adjusted stability estimates for the core social metrics were highly similar to the original estimates, indicating that the reported stability is not driven by age differences across groups. For details, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Title and abstract revised. Methods, Results, and supplementary robustness analyses expanded to report age effects explicitly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examined whether infraslow fluctuations in noradrenaline and in heart rate are coupled and how they are affected by sleep transitions. The authors used the fluorescent NA biosensor GRAB-NE2m in the medial prefrontal cortex of mice to record extracellular NA while also recording EEG and EMG during sleep-wake episodes. They also analyzed previously published human data to reproduce relationships they found between sigma power and RR intervals in mice.

      Strengths:

      This is an impressive study with significant strengths, as it involves a rich set of data that includes not only observations of associations between heart rate and noradrenergic dynamics but also optogenetic manipulation of the locus coeruleus. Human data is presented to show parallels in the association between sigma power during sleep and phasic heart-rate bursts.

      We thank Reviewer #1 for their thoughtful, detailed, and constructive evaluation of our manuscript. We appreciate their recognition of the strengths of the study, particularly the integration of noradrenergic recordings, optogenetic manipulation, and cross-species analyses. We are especially grateful for the reviewer’s careful attention to clarity, experimental interpretation, and control comparisons. The comments have helped us sharpen the framing of our hypotheses, clarify causal claims, improve statistical reporting, and better explain our closed-loop approach and heart rate analyses. We have addressed each point in detail below and believe that the revisions substantially strengthen the manuscript.

      Weaknesses:

      (1) Language could be clearer and more precise. As detailed below, in both the introduction and the discussion, the way the hypotheses and study objectives are described could use some revision to be more precise and accurate.

      Thank you for this helpful comment. We have sharpened the description of the study objectives, hypotheses, and interpretation of the findings to better distinguish between what was directly tested, what was inferred, and what remains speculative. We revised the language throughout these sections to improve clarity, accuracy, and overall readability

      (1A) In the introduction on p. 4: The overarching question is framed as "could the peripheral autonomous systems be a read-out of the central LC-NE system and thus be a biomarker of memory consolidation and LC dysfunction?" This gives the impression that the LC function would be the main influence on peripheral autonomous systems. There are, of course, many influences on peripheral autonomous systems, so it would be advisable for the authors to be more specific here about what signal(s) in particular would be predicted to be sensitive markers of LC function.

      Thank you for this important point. We agree that heart rate reflects the integrated output of multiple autonomic mechanisms and should not be interpreted as being exclusively driven by LC activity. Cardiac dynamics arise from the balance between sympathetic and parasympathetic influences, which themselves are regulated by several central and peripheral systems. In addition, recent work shows that the infraslow oscillations observed during NREM sleep are not restricted to norepinephrine alone but also involve other neuromodulatory systems, including acetylcholine and serotonin (e.g., Teng et al., PNAS 2025, Kjaerby et al., iScience, 2026). Our intention was therefore not to imply that the LC is the sole driver of peripheral autonomic dynamics. We have revised our overarching questions to make them more specific: Please see new text below:

      (Introduction, page 4/5). “The sympathetic and parasympathetic autonomic nervous system are involved in HRV, which is conventionally analyzed across three primary frequency bands: high frequency (HF) HRV, low frequency (LF) HRV, and very low frequency (VLF) HRV (Berntson et al., 1997). Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by HRV under different physiological states or does LC–HR coupling scale differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1B) In the discussion on p. 12: "In this study, we leveraged real-time measurements of mPFC NE levels and HR measurements from EMG recordings in mice to investigate the causal link between the two variables with high temporal resolution in freely moving sleeping mice, with similar inspection in humans." To test the causal link between mPFC NA levels and HR measures, the study would manipulate NA levels just in the mPFC and not elsewhere in the brain. However, in this study, the manipulation occurred in the LC, and so there would be broad cortical changes in NA levels. Thus, it could be that LC activity causes HR changes via a non-PFC pathway.

      We thank the reviewer for this important comment. Indeed, mPFC NE is merely a readout of LC activations and we expect that NE in other brain regions would show the same patterns. Indeed, mPFC NE is not expected to provide any causal link to heart rate. We have revised added a sentence to the results section and also changed the initial summary part of the discussion to reflect this better.

      (Results, page 6). “mPFC was selected as a representative cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (Discussion, page 14/15). “Variability in HR is a non-invasive biomarker of autonomic nervous system function and is frequently disrupted in ageing and Alzheimer’s disease. Here, by combining real-time measurements of mPFC NE dynamics with simultaneous HR recordings in freely sleeping mice, we demonstrate that HR closely tracks the infraslow phasic activity of the LC–NE system.”

      (2) Comparisons with the control condition need further development.

      (2A) While the authors did include a key YFP control condition, in the main text no direct statistical comparison between the closed-loop optogenetic stimulation (ChR2) condition and the YFP control condition was reported. (It was reported in Supplementary Figure 2c-d.) Instead, in the main text, the authors only reported that the effects of stimulation were significant in the closed-loop condition and not in the control. However, that is not the same as demonstrating that the two conditions significantly differed from each other, and it is the direct test that is important for the conclusions, so it seems important to include this result in the main presentation.

      We thank the reviewer for this important point and agree that direct statistical comparisons between ChR2 and YFP conditions are important for interpretation. These comparisons were performed and are shown in Supplementary Figure 2c–d, but we acknowledge that this was not sufficiently emphasized in the main text. We are now more clearly referring to this comparison in Result section:

      (Results, page 9). “The magnitude of pre-stimulation NE descent and post-stimulation NE ascent was reduced as the thresholds increased, indicating less pronounced NE dynamics as LC stimulation became more frequent (Fig. 2e, for direct comparison with YFP control, see Suppl. Fig. 2c-d).”

      Our rationale for prioritizing the within-animal threshold comparisons in the main figure was that the central experimental question concerned how progressive shifts in infraslow NE oscillatory frequency influence the NE–HR relationship. Because variability in viral expression levels (both NE sensor expression and LC opsin expression) introduces substantial between-animal variability, we considered within-animal comparisons across threshold conditions to provide the most informative representation of how changes in LC-driven NE dynamics alter cardiac responses. That said, we agree that highlighting the direct ChR2 versus YFP comparison is important for the overall interpretation. As mentioned, the reference to Suppl. Fig. 2c– d, where the between-group analyses are visualized are now clearly referred to.

      (2B) In addition, the authors should address the issue that the pre-stimulation NE was consistently significantly lower in the YFP condition than in the ChR2 condition (see Supplementary Figure 2c), which is a potential confound.

      We thank the reviewer for bringing up this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design and appreciate the opportunity to clarify this aspect of the experiment.

      In the ChR2 condition, animals were exposed to repeated optogenetic LC activation designed to mimic progressively faster infraslow NE dynamics. Such repeated stimulation is expected to produce a gradual elevation in tonic NE levels across the recording session, which explains the higher pre-stimulation baseline relative to YFP controls. We acknowledge that elevated tonic NE levels could introduce additional physiological effects. For example, higher NE tone would be expected to increase α2mediated autoinhibitory feedback on LC neurons and presynaptic NE release. Within the LC itself, we expect that optogenetic stimulation would largely override such effects due to the strong Na+-mediated depolarization induced by ChR2 activation. However, NE release in downstream regions such as the mPFC may be influenced to some extent by elevated tonic noradrenergic tone. Importantly, such feedback mechanisms would likely also occur under physiological conditions characterized by elevated LC activity, such as stress or sleep fragmentation. Because the goal of our stimulation paradigm was to model progressively faster infraslow NE dynamics under physiologically relevant conditions, we believe this feature of the manipulation may in fact increase the translational relevance of the model. We specifically address this elevation in Figure 3, where we show that very rapid stimulation regimes are accompanied by signs of compensatory cardiovascular regulation, likely reflecting baroreceptor-mediated responses to sustained increases in heart rate.

      We agree that the precise contribution of elevated tonic NE to the overall manipulation cannot be fully disentangled in the present study. We therefore avoid overinterpreting these effects and have instead added text to the manuscript acknowledging this consideration without extensive speculation.

      (Results, page 9) “Since pre-stimulation NE baseline levels progressively became higher in the ChR2 condition compared with YFP controls (Fig. 2c+e, Supplementary Fig. 3a-f), it demonstrates that they arise from stimulation-dependent modulation of noradrenergic tone rather than nonspecific signal drift. As a result, the pre-stimulation state at higher thresholds differed between ChR2 and control conditions, which should be considered when interpreting the immediate effects of laser stimulation across thresholds.”

      (2C) Direct comparison of the strengths of correlations shown in Figure 2h vs. Supplementary Figure 2f should be included. Currently, we see relatively weak correlations in both ChR2 and YFP conditions, and it is not clear if the relationships differ in the control. It seems they are still present in the control condition but weaker which would contradict the apparently broad claim on p. 7 that "No such effects were present in the control condition" (it is not entirely clear whether this claim refers to all effects discussed in the figure or just a subset - this language should be clarified).

      We thank the reviewer for this important comment and agree that the original wording could be interpreted as implying a complete absence of an NE–RR relationship in the YFP condition. To address this concern, we directly compared the strength of the NE– RR relationship between ChR2 and YFP animals. Please see Supplementary Figure 3h.

      Using a linear mixed-effects model that accounted for repeated measurements within animals, we found that the slope of the NE–RR relationship was significantly steeper in ChR2 animals than in YFP controls (all thresholds: slope difference = 1.915, p < 0.0001; thresholds −15, −10, and −5 only: slope difference = 2.367, p < 0.0001). Consistent with this result, comparison of Pearson correlations using Fisher's r-to-z transformation also indicated significantly stronger coupling in ChR2 animals than in YFP controls (see Statistics in the Supplementary File).

      These analyses demonstrate that an inverse relationship between NE and RR is present under physiological conditions in YFP animals, but that optogenetic LC activation substantially strengthens this coupling. We have revised the corresponding text to clarify this:

      (Results, Page 9): “An inverse NE–RR relationship was present in both ChR2 (Fig. 2h) and YFP (Suppl. Fig. 3f) groups but was significantly stronger in ChR2 animals than in YFP controls (Suppl. Fig. 3h).”

      (2D) Did the YFP controls vs. ChR2 animals show any differences in the number of NA states that triggered stimulation in the closed-loop system? With ChR2 animals, stimulation changes NA, which could change future triggering. In YFP animals, nothing changes NA (other than natural fluctuations), so the dynamics of stimulation timing could diverge between groups in a way that complicates interpretation. Specifically, if ChR2 stimulation raises NA and prevents future threshold crossings, ChR2 animals may end up receiving fewer subsequent stimulations than YFP animals (or a different temporal clustering). If the number or pattern of stimulation differed in two groups, it would be important to have a yoked control where matched animals get the same stimulation pattern but not triggered by their own NA.

      We thank the reviewer for this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design.

      Importantly, stimulation triggering was based on relative declines in NE fluorescence calculated against a rolling 2-minute baseline, rather than absolute NE levels. Thus, as tonic NE levels gradually increased in ChR2 animals, the threshold adapted accordingly, reducing the likelihood that elevated baseline NE alone would prevent future triggering. Instead, stimulation continued to occur when NE declined relative to the recent baseline, thereby preserving the infraslow closed-loop structure.

      The number of stimulation events across thresholds is already reported in the manuscript (Methods, p. 30 and corresponding figure legends), but we have now indicated more clearly in the result section where to find the information:

      (Results, Page 9). “Mean traces of NE and RR were aligned to LC stimulation onset (Fig. 2c-d, for number of laser stimulations see Fig. 2 legend or Methods).”

      (Methods, Page 31). “For the LC activation-related analysis, 108 events were found for Threshold -15 (11 of these being YFP), 260 events for Threshold -10 (55 of these YFP), 777 events for Threshold -5 (296 of these YFP), 1,444 events for Threshold 0 (510 of these YFP), and 1,148 events for Threshold 5 (377 of these YFP) across ten animals (four being YFP).”

      While the total number of events was lower in YFP animals, this is expected in part due to the smaller group size (4 YFP vs. 6 ChR2 animals included in this analysis). Furthermore, because ChR2 stimulation increased NE levels by design, more pronounced subsequent declines in NE may have modestly facilitated additional threshold crossings.

      Nevertheless, the overall temporal structure of the stimulation paradigm remained comparable across groups, and YFP animals underwent the same closed-loop stimulation protocol. Our primary comparison was mainly based on shifts in within-animal NE oscillatory frequency and how this change would impact the connection to HR. Thus, we believe the present control condition appropriately addresses the central question of whether optogenetic LC activation are able to conduct a continuum of NE oscillations.

      (3) Some more discussion/explanation of the rationale for the closed-loop approach and how it influences how we should interpret the results could be useful. For instance, currently, it is not clear whether LC stimulation needs to be timed after an NA dip to yield the effects seen.

      We thank the reviewer for pointing this out. The rationale for the closed-loop LC stimulation approach was to test whether the relationship between infraslow LC–NE dynamics and heart rate is maintained only under physiological infraslow conditions or whether it breaks down when the rhythm becomes progressively faster, as occurs during sleep fragmentation and other high-arousal states. Specifically, we asked whether heart rate continues to track LC–NE fluctuations as the infraslow rhythm shifts toward higher frequencies, thereby assessing its utility as a potential biomarker of disrupted restorative sleep.

      To address this while preserving the intrinsic temporal structure of infraslow LC activity, we implemented a closed-loop strategy in which stimulations were triggered following defined declines in the NE signal, using a rolling preceding 2-minute window as baseline. This allowed LC activation to occur during the descending phase of the endogenous infraslow cycle, maintaining its physiological phase structure while systematically increasing its effective frequency. By progressively relaxing the decline threshold, stimulations were triggered earlier in the cycle, thereby compressing the infraslow period in a controlled manner.

      Importantly, the intention was not to test whether LC stimulation specifically needs to occur after an NE dip to elicit the observed effects. Rather, triggering stimulation during the decay phase provided a way to accelerate the infraslow rhythm without disrupting sleep through indiscriminate stimulation. This enabled us to examine whether the coupling between LC–NE dynamics and heart rate remains stable under increasingly rapid infraslow regimes. Our results indicate that this relationship weakens at higher infraslow stimulation frequencies, suggesting that heart rate reliably reflects physiological LC–NE oscillations but becomes less tightly coupled when the rhythm is compressed beyond its normal range.

      We have clarified this rationale in the revised manuscript by adding the below section in the result section.

      (Result, page 8/9). “This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (4) The section on heart rate decelerations is hard to follow. In particular, I was not sure how to interpret Figure 3f-j. For Figure 3f, what does the middle line represent? The laser onset or the max RR value after laser onset? What is the baseline that is used to correct the values to obtain amplitudes? If it is the whole period before the maximal RR value or the laser onset, wouldn't baseline values differ significantly across conditions and so potentially account for differences seen between conditions in the reported HR decelerations? Larger HR decelerations may be seen in conditions with higher HR simply as a regression to the mean phenomenon.

      We thank the reviewer for this feedback and agree that additional clarification of Figure 3f–j is warranted.

      For Figure 3f, the central line represents the peak RR value (maximal heart-rate deceleration) identified within the 2–7 s window following laser onset, rather than the laser onset itself. We realize this was not sufficiently clear and have revised the figure and corresponding Results text to clarify this point.

      Regarding baseline correction, RR amplitudes were calculated as the difference between the RR at peak deceleration and the mean RR during the 8–10 s period preceding the RR peak, as described in the manuscript. Thus, the baseline was defined locally for each event and was not based on the entire pre-laser period or stimulation onset. We chose this approach to account for shifts in baseline heart rate across conditions and to capture the relative magnitude of the deceleration response rather than absolute RR values.

      We appreciate the reviewer’s point regarding potential regression-to-the-mean effects, particularly in conditions with higher baseline heart rates. This is an important consideration. However, because the amplitude measure was baseline-corrected on an event-by-event basis, we believe the reported differences are unlikely to be explained solely by higher pre-stimulation heart rate. At the same time, our findings clearly show that elevated baseline heart rate influence the dynamic range of deceleration responses. Our findings show that under physiological conditions with elevated HR, larger heart-rate fluctuations would also contribute to increased HRV, which is often interpreted positively, despite potentially reflecting fragmented or dysregulated sleep states in this context.

      To improve readability, we have revised the figure and associated text.

      (Results, page 10). “To quantify the HR decelerations that happened after LC activation, we took the maximal RR value (so slowest HR) 2-7 s after LC stimulation and baseline corrected the value to the mean RR during the 8–10 s period preceding the RR peak to obtain their amplitude.”

      (5) The findings regarding LC suppression could be further clarified.

      (5A) Page 8: "observed a response in NE decline" - please be more precise. Did NE decline more or less?

      We thank the reviewer for this suggestion and agree that the original wording was imprecise. To clarify the direction and nature of the response, we have revised the text to state:

      (Results, page 11). “...we observed a gradual NE decline sustained throughout the laser period that was not observed in the YFP condition…”

      (5B) It would be helpful to also show the correlation between NE and RR in the control (YFP) condition and whether there were any differences between YFP and Arch conditions (Figure 4e).

      We thank the reviewer for this suggestion. We have now added the corresponding YFP correlation to Figure 4e. In the YFP group, the relationship between NE and RR showed a similar negative trend but did not reach statistical significance (p = 0.053). To directly assess whether the NE–RR relationship differed between Arch and YFP animals, we performed both a linear mixed-effects analysis and a Fisher r-to-z comparison.

      Neither analysis revealed a significant difference between groups. The linear mixed-effects model showed no significant Group × NE interaction (slope difference = 0.994, p = 0.51), indicating that the NE–RR coupling was not altered by LC suppression. Similarly, Fisher's r-to-z comparison found no significant difference between the correlations (p = 0.42).

      We believe this result is consistent with the relatively modest nature of the LC suppression paradigm. While Arch stimulation produced a clear reduction in NE levels, it did not induce a large shift in the overall NE–RR relationship. Instead, the data suggest that heart-rate responses remain coupled to noradrenergic fluctuations under both physiological conditions and during mild LC suppression. We have added the YFP data and clarified this interpretation in the revised manuscript.

      (Results, page 12): “A similar negative relationship was observed in YFP controls (Fig. 4e), and the strength of the NE–RR association did not differ significantly between Arch and YFP animals (Suppl. File, Statistics), suggesting that LC suppression did not substantially alter the underlying coupling between these measures.”

      (5C) This sentence took me multiple readings to understand - it would be helpful to rewrite to make it clearer: "indicating that, while HR generally did not respond strongly to LC suppression, the variability in RR responses was dependent on NE changes to the suppression (Figure 4e)."

      We agree with the reviewer that this phrasing is hard to understand and we have optimized for better clarity. Please see new version below:

      (Results, page 11/12). “Notably, despite the absence of a robust group-level HR effect, NE and RR responses remained negatively correlated across trials (Fig. 4e), indicating that HR dynamics continued to track the magnitude of noradrenergic suppression at the individual-response level”.

      (5D) The two colors in Figure 4 are similar and hard to distinguish.

      We agree that the colors are hard to separate and have altered them to make them easier to separate.

      (5E) The correlations shown in Figure 4j seem to be driven by just two of the cases. Are the effects significant when outliers are removed?

      We thank the reviewer for raising this point. To assess whether the observed correlation in Figure 4j was disproportionately driven by a small number of data points, we performed a formal outlier analysis using the ROUT method (Q = 1%). This analysis did not identify any statistical outliers in the dataset. Therefore, we did not have an objective basis for excluding any observations from the analysis.

      (5F) Page 10: Were there any differences in memory performance between the Arch and YFP conditions?

      We thank the reviewer for this question. The memory experiments were based on a previously published dataset (Kjaerby, Andersen et al., 2022), in which the primary objective was to assess the effect of LC suppression on sleep spindle dynamics and memory consolidation. In the present study, we performed an additional analysis by extracting heart-rate (RR) measures from these recordings to evaluate whether cardiac responses could serve as a biomarker of LC-mediated noradrenergic regulation.

      However, due to technical limitations (EMG recording often suffers from noise) in extracting reliable RR signals from all animals in this dataset, the number of subjects available for this secondary analysis was reduced. All these considerations are described in Methods/Mice. As a result, we were not sufficiently powered to perform a direct statistical comparison of memory performance between Arch and YFP groups based on RR measures alone. Instead, we examined whether RR responses to LC suppression predicted behavioral performance across animals. When pooling Arch and YFP conditions, we observed a correlation between the magnitude of the RR response and subsequent memory performance, suggesting that heart-rate dynamics reflect noradrenergic modulation relevant for memory consolidation.

      To avoid overinterpretation, we have therefore limited our conclusions to reporting this association rather than making direct group-level comparisons between Arch and YFP animals, and we have clarified this point in the revised manuscript.

      (Results, page 13): “Interestingly, across pooled Arch and YFP animals, larger RR increases following LC suppression were associated with better subsequent memory performance. Additionally, RR and NE responses to LC suppression were negatively correlated indicating that animals showing stronger NE reductions also exhibited larger RR changes. Together, these findings suggest that heart-rate dynamics covary with noradrenergic responses during sleep and may reflect physiological processes relevant for sleep-dependent memory consolidation.”

      (5G) Page 10: "We found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power." It should be noted in the text that the correlation was not specific to sigma (it was also seen for theta and beta, Figure 4i).

      We agree with reviewer that this should be highlighted. We have changed the sentence:

      (Results, page 12). “Furthermore, we found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power; similar correlations were also observed in the theta and beta frequency ranges (Fig. 4h-i).”

      (6) It is not clear which of the sigma power and RR interval findings do/do not exactly line up between the mice and humans. It could be helpful to have a table comparing them. For instance, was the finding in humans that pre-HRB sigma power was positively associated with slowing in heart rate after the HRB also seen in mice? Was there evidence in mice (as seen in the human sample) that sleep-dependent memory improvement was associated with pre-HRB sigma power?

      We thank the reviewer for this thoughtful comment and agree that the cross-species comparisons could be communicated more clearly. Our intention was not to imply exact one-to-one correspondence between all mouse and human findings, but rather to examine whether central–autonomic coupling surrounding phasic heart-rate events shows conserved features across species while acknowledging species-specific physiological differences.

      Importantly, the mouse and human analyses were designed to address related but not identical questions. In mice, we leveraged optogenetic LC suppression to probe a more causal relationship between noradrenergic activity, heart-rate slowing, spindle-related dynamics, and memory consolidation. Previous work using this dataset demonstrated that 2-minute LC suppression robustly enhances spindle density and that spindle enhancement correlates with improved memory performance. In the present study, we therefore asked whether heart-rate slowing covaries with this LC-mediated spindle/memory relationship, supporting HR as a potential biomarker of these restorative processes.

      By contrast, causal manipulation of LC activity is not feasible in humans. Instead, we focused on naturally occurring HRBs as putative downstream signatures of phasic LC– NE activity, motivated by our mouse findings that NE increases precede HR accelerations. We observed conserved coupling between sigma activity and HR dynamics across species, although the temporal profile differed, with sigma activity occurring closer to the HRB in humans than in mice. These temporal differences may reflect species-specific differences in cardiac and sleep physiology.

      Regarding the reviewer’s specific questions, the positive relationship between pre-HRB sigma power and post-HRB heart-rate slowing was tested in humans, where this metric showed the strongest relationship to behavioral outcome. We did not directly test the same measure in mice because, unlike humans, HR recovery following HRBs did not show a pronounced baseline shift (Fig. 5b), limiting the interpretability of this comparison. Similarly, we did not directly correlate pre-HRB sigma power with memory performance in mice, as the more causal LC suppression paradigm already demonstrated a spindle–memory relationship in this species and was the focus of our mechanistic analysis.

      To reduce confusion, we have revised the Results conclusion to make a clearer overview:

      (Results, page 14): “In conclusion, mice and humans displayed evidence of conserved autonomic-central coupling, reflected in coordinated HR and sigma power dynamics surrounding phasic cardiac events.

      However, the temporal relationship between these events differed across species, likely reflecting differences in sleep and cardiovascular physiology. In mice, causal manipulation of LC activity demonstrated that heart-rate dynamics covary with LC-mediated noradrenergic and spindle-related processes linked to memory consolidation. In humans, sigma power preceding HR bursts was associated with both post-HRB heart-rate slowing and sleep-dependent memory improvement, suggesting that autonomic–central coupling surrounding HR events may provide a translational marker of restorative sleep processes (Fig. 5n).”

      (7) Page 18: It is not clear if the sex of mice was balanced across controls and optogenetics groups.

      We thank the reviewer for this important comment and agree that the sex distribution should be reported more clearly. These experiments relied on the availability of animals from the heterozygous TH-Cre transgenic line, and given the relatively small cohort sizes, perfect balancing across sex and experimental groups was not always feasible.

      For the LC activation experiments, the sex distribution was: YFP: 2 male / 2 female; ChR2: 4 male / 2 female. For the LC suppression experiments, the distribution was: YFP: 4 female; Arch: 3 female / 1 male.

      Although the groups were not perfectly sex balanced, we had no strong reason to expect robust sex-dependent differences in the physiological effects of these optogenetic manipulations, particularly given the relatively strong and acute nature of the intervention. At the same time, we acknowledge that the present study was not powered to assess sex as a biological variable, and subtle sex-dependent effects therefore cannot be excluded. To improve transparency, we have now clarified the sex distribution in the Methods/Mice section.

      Reviewer #2 (Public review):

      Summary:

      The major part of this study reproduces previously published findings in both mice and humans and provides incremental analyses on these findings. In essence, the work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep. It further confirms previous findings that coordination depends on noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      The authors successfully replicate key previously reported phenomena across both mice and humans. Confirmatory studies and demonstrations of reproducibility are essential for progress in neuroscience. To maximize their value, such studies should clearly acknowledge their confirmatory nature and carefully situate what, in their view, are novel results, going beyond existing literature.

      Weaknesses:

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      In the present manuscript, several citations of literature on the work they reproduce lack precision or completeness, which reduces transparency and obscures how the reported findings relate to previously established results.

      We thank Reviewer 2 for the thoughtful comment regarding positioning of our findings relative to the literature, and caution in mechanistic interpretation. In response, we have revised the Introduction, Results, and Discussion to more clearly acknowledge foundational studies in this area and to better clarify how the present work extends beyond them.

      We agree that prior work has demonstrated infraslow coupling between sigma activity, norepinephrine (NE) dynamics, and heart rate (HR), and has established a role for the locus coeruleus (LC) in coordinating these oscillations. However, cardiac measures in these studies were typically treated as secondary observations rather than as primary experimental targets. A central goal of the present study was therefore to provide a systematic and mechanistically grounded characterization of NE-mediated HR dynamics during sleep across multiple timescales, including infraslow oscillations, sleep–wake transitions, and causal manipulations of LC activity.

      Importantly, we also aimed to relate infraslow HR fluctuations to the very-low-frequency (VLF) component of heart rate variability (HRV), which remains comparatively under-characterized and mechanistically unresolved in the clinical HRV literature. By linking LC activity, NE dynamics, and HR fluctuations across behavioral states, our findings provide a biologically grounded framework that may help explain this component of HRV.

      A second major objective of the study was translational. Because direct LC recordings are not feasible in humans, we asked whether cardiac dynamics alone could reflect the infraslow, memory-consolidating potential of sleep and thus serve as a noninvasive biomarker. By directly manipulating LC activity and demonstrating corresponding changes in HR dynamics, our results strengthen the mechanistic rationale for using HRV—particularly its VLF component—as an accessible proxy of LC-dependent sleep physiology.

      We therefore respectfully disagree with the suggestion that the present study does not provide novel insight. Rather, the revised manuscript now more clearly emphasizes that our contribution lies in (i) systematically characterizing NE-dependent HR dynamics across sleep states, (ii) linking these dynamics to the poorly understood VLF component of HRV, and (iii) establishing a causal and translational framework for using cardiac measures as markers of LC-mediated sleep processes.

      We hope the reviewer finds that the revised Introduction and Discussion better highlight both the existing literature and the specific advances provided by the present work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I have been convinced by discussions about the replicability crisis that it should be a standard practice to share data upon publication in a publicly accessible online repository such as Open Science Framework or OpenNEURO. Making data available can increase the impact of the research study. Simply stating "data available upon request" as done in the current draft is not sufficient as, unfortunately, when data are not shared upon publication in a public repository, it can be impossible to gain access by request from the researchers - the majority of requests from other researchers to obtain data are not complied with (e.g., Vanpaemel, Vermorgen, Deriemaecker, & Storms, 2015; Wicherts, Bakker, & Molenaar, 2011).

      We fully agree with the reviewer about the importance of data sharing for transparency, reproducibility, and maximizing the impact of research. Consistent with these principles, we will make the human dataset publicly available on the Open Science Framework upon publication and have added the link to this repository in the manuscript under Data Availability (https://osf.io/g6emj/). With respect to the mouse dataset, we respectfully note that this dataset is currently the subject of multiple planned and ongoing analyses that extend beyond the scope of the present manuscript. Releasing these data publicly at this stage could compromise these efforts and lead to potential misinterpretation prior to completion of the full analytic pipeline. For this reason, we believe that it would not be appropriate to share the mouse data in a public repository at this time. However, we remain committed to transparency and will make the mouse data available upon reasonable request during this period, with the intention of publicly releasing the dataset once the planned analyses are complete.

      (1) p. 4: There is some orphan text at the top of the page ("marker of Alzheimer's disease. Furthermore, maintaining LC neural density prevents neurodegeneration (13). Given the reported reduction in HRV in aging and Alzheimer's disease, suppressed").

      Thank you. We have removed the orphan text.

      (2) p. 22: What does NFR refer to?

      Novel-to-familiar ratio. We have removed the abbreviation from the main text. It is now only used in the figure.

      (3) I might have missed this information, but it was not clear to me how the epochs to be analyzed were selected, and for the mice, how much wake vs. sleep they included.

      We thank the reviewer for this comment and apologize that the epoch selection criteria were not sufficiently clear. They were mentioned under Methods. For the optogenetic analyses in mice, epochs were selected based on NREM sleep including microarousals to ensure that physiological responses were evaluated within stable sleep conditions while allowing for natural brief interruptions.

      For the LC activation (ChR2) experiments, stimulation epochs were included only if NREMinclMA began at least 30 s before laser onset and continued for at least 30 s after laser onset. For the LC suppression (Arch) experiments, the criterion was similarly ≥30 s of NREMinclMA prior to laser onset, but extending ≥60 s following laser onset to accommodate the longer suppression response profile.

      Thus, analyses were intentionally restricted to sleep periods, and wakefulness was not included except where it emerged naturally as an outcome of the manipulation or transition under investigation. We have now clarified these inclusion criteria in the Methods/ Event marker selection section to improve transparency.

      (4) In the Supplement, Figure 3 appears before Figure 2.

      Thanks. We have corrected it.

      Signed, Mara Mather

      Reviewer #2 (Recommendations for the authors):

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      There are three major directions in which this study would need to be revised.

      Part 1 - Literature citations of both mouse and human literature need to be revised, and citations placed in a manner that accurately reflects what has been previously done. In detail:

      (1.1) The statement on p. 4 regarding "the extent to which phasic infraslow NE fluctuations ...is not well understood" disregards a previous publication in which closed-loop optogenetic stimulation of LC was already shown to regulate HR variations on the infraslow time scale (10.1016/j.cub.2021.09.041).

      We thank the reviewer for this important point. The study by Osorio-Forero et al. (2021) was already cited as ref. 4 in our original manuscript and in the discussion, we specifically highlighted this study: ‘These findings build on Osorio-Forero et al. (35) as well as other studies’; however, we agree that our wording did not sufficiently emphasize its key finding that closed-loop optogenetic manipulation of LC activity can coordinate infraslow heart-rate fluctuations and spindle clustering during NREM sleep.

      We have revised the relevant paragraph in the introduction to explicitly acknowledge that ref. 4 demonstrated causal coordination between LC activity, sleep spindle dynamics, and heart-rate fluctuations. We have also refined our statement of the knowledge gap to clarify which novel questions our study addresses. We believe these revisions more accurately position our work within the existing literature and clearly distinguish our contributions from prior studies.

      We have updated the introduction in several places to address these comments:

      (Introduction, page 4). “It has previously been reported that HR fluctuates at similar infraslow frequencies as NE fluctuations and sigma power in mice (Lecci et al., 2017; Osorio-Forero et al., 2021) and that phasic HR fluctuations correlate with infraslow changes in pupil diameter (Carro-Domínguez et al., 2025), a proxy for changes in NE levels (Murphy et al., 2014; Reimer et al., 2016). Importantly, optogenetic manipulation of the LC has demonstrated that infraslow LC activity coordinates sleep spindle clustering and heart rate fluctuations during NREM sleep (Osorio-Forero et al., 2021). Together, these findings support a functional coupling between the central LC–NE system and peripheral cardiac dynamics, further supported by findings that HR increases accompany MAs during NREM sleep (Carro-Domínguez et al., 2025;

      Osorio-Forero et al., 2025).”

      (Introduction, page 4/5). “Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by heart-rate variability under different physiological states or does LC–HR coupling scales differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1.2) In this same published paper, optogenetic stimulation of LC was already used to "determine the causal relationship..." (p.6). However, the authors do not cite these data.

      We thank the reviewer for this comment. As noted in our response to Comment 1.1, ref. 4 was already cited in the Introduction, and we have now revised that section to more explicitly emphasize this paper. In addition, we have modified the wording in the Results section.

      (Results, page 8). “After finding the inverse correlation between NE and RR in natural sleep transitions, we next sought to further characterize the causal influence of LC activity on HR dynamics, using a closed-loop optogenetic approach to modulate NE oscillatory frequency during NREM sleep.”

      (1.3) The authors' speculation about LC-induced sympathetic and parasympathetic actions is premature: this study does not provide pharmacological experiments in this direction. However, two published studies implied a parasympathetic mechanism linked to infraslow fluctuations of LC activity (10.1016/j.cub.2021.09.041, 10.1016/j.cub.2017.12.049). These findings should be appropriately cited. While HR decelerations may reflect compensatory autonomic responses, there is no direct evidence in the present study that these effects are sympathetically mediated. An alternative, and equally plausible, interpretation is enhanced parasympathetic activity. This distinction is particularly important given that the observed increases in mean HR during LC stimulation cannot distinguish between reduced parasympathetic activity and increased sympathetic drive.

      We thank the reviewer for raising this important point and agree that the current study does not provide direct mechanistic evidence to disentangle sympathetic versus parasympathetic contributions to LC-mediated heart rate regulation. This was not the intention of our study; rather, the relevant discussion section was meant to provide mechanistic interpretations and hypotheses based on the observed physiology. To avoid overstating our conclusions, we have revised the wording to more clearly emphasize the speculative nature of these interpretations. Furthermore, we have added the suggested references demonstrating that muscarinic blockade reduces heart rate and pupil fluctuations during sleep, which support the possibility of a parasympathetic contribution to infraslow LC-related dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) and also influences sympathetic output through direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014), there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018), led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition.”

      (1.4) It is not clear why prefrontal NE signals were associated with HR fluctuations. Literature evidence indicates other brain areas that are functionally more directly linked to autonomous fluctuations.

      We thank the reviewer for raising this point. We do not speculate that mPFC NE activity is causally linked to heart rate fluctuations or that the mPFC directly mediates the observed autonomic dynamics. Instead, mPFC NE signaling was used as an experimentally accessible readout of LC activity. Importantly, accumulating evidence suggests that infraslow NE oscillations are globally coordinated phenomena that are expressed across multiple brain regions during sleep, making mPFC NE a valid proxy for LC-driven neuromodulatory state dynamics. We have clarified this in the result section:

      (Results, page 6). “mPFC was selected as a cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (1.5.a) The lack of effect of Arch-inhibition of LC on infraslow NE signals is concerning. Prior work showed that bilateral LC inhibition does affect NE signals and also infraslow sigma power fluctuations (10.1016/j.cub.2021.09.041, 10.1038/s41593-024-01822-0). This discrepancy should be explicitly acknowledged and discussed on p.9.

      We thank the reviewer for highlighting this important point. To clarify, we did observe a robust effect of LC inhibition on noradrenergic signaling, with clear suppression of the NE signal following Arch-mediated LC inhibition (Fig. 4c), aligned to laser onset. In addition, LC suppression increased neuronal synchronization, including enhanced sigma power relative to YFP controls (Fig. 4h), consistent with previous reports showing that reduced LC activity promotes synchronized sleep-related oscillations. Thus, we do not interpret our findings as indicating an absence of LC suppression effects.

      The apparent discrepancy relates specifically to the absence of a statistically significant group-level shift in infraslow NE or HRV power during NREM sleep, rather than the efficacy of the manipulation itself. We note that our inhibition paradigm was intentionally mild, consisting of repeated 2-minute suppression periods separated by 4-minute intervals, and was designed to introduce subtle shifts within physiological ranges rather than globally reorganize infraslow sleep structure. Accordingly, we consider the immediate NE and heart-rate responses to LC inhibition to be the most sensitive physiological readouts of LC-mediated regulation in this context.

      We also note that the studies cited by the reviewer used different suppression paradigms, including more frequent manipulations. Thus, while our suppression scheme did not result in detectable infraslow reorganization, we do not believe they reflect an ineffective LC suppression as we demonstrated a clear NE reduction in response to time-locked LC suppression.

      We have added a sentence to the result section to explain the lack of effect infraslow power:

      (Results, page 12). “This likely reflects the relatively mild and intermittent LC suppression paradigm, which was designed to remain within physiological ranges and therefore did not globally reorganize infraslow sleep dynamics.”

      (1.5.b) Additionally, the authors state in the Discussion that they "find no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection is more strongly associated with sympathetic activity rather than parasympathetic inhibition". However, the results primarily demonstrate an absence of modulation in mean HR, while preserving a significant relationship between RR intervals and stimulation. This indicates a modulation of heart rate variability, even in the absence of changes in average HR. Notably, such variability-related effects may fall outside the VLF range and could instead involve higher-frequency components.

      We thank the reviewer for this important clarification. We agree that our original wording may have conflated the absence of modulation in mean HR with the absence of autonomic modulation more generally. To address this point, we revised the Discussion to emphasize that LC activity may influence VLF HRV through sympathetic activation and/or indirect modulation of cardiac vagal activity, even in the absence of robust changes in average HR. We have softened our previous interpretation that the LC-HR relationship is primarily sympathetic in nature and instead discuss a more nuanced interaction between sympathetic and parasympathetic influences on HRV dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018) led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition. Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of non-linear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      (1.6) The relationship between sigma power fluctuations and HR is different in humans than in mice. This has been shown before (10.1126/sciadv.1602026, 10.1038/s41593025-02159-y). This work should be mentioned on p. 11.

      We already acknowledge prior studies demonstrating species differences in the relationship between sigma power fluctuations and heart rate in the discussion section. However, to accommodate the reviewer’s comment and improve clarity for the reader, we have now also added the suggested references to the Results section, where the relationship between sigma power fluctuations and HR is first discussed.

      (Results, page 14). “These temporal differences may reflect species-specific physiology differences in cardiac timescales, which has also been previously reported (Bergel et al., 2025; Carro-Domínguez et al., 2025; Lecci et al., 2017).”

      (1.7) Correlations between the strength of infraslow sigma power fluctuations and memory consolidation have been published and should be discussed (10.1126/sciadv.1602026). It is surprising to see that correlations with learning in humans are done using pre-HRB sigma peaks rather than heart rate. This is a measure that is very close to the one used by Lecci et al.; this similarity should be clearly acknowledged. The way the data are currently presented limits this manuscript's novelty, also in its translational aspect.

      We agree that the work by Lecci et al. (2017) established an important relationship between the association of infraslow sigma power fluctuations and memory consolidation, which is highly relevant to our findings.

      Importantly, the underlying infraslow fluctuations in neuromodulatory tone are increasingly recognized as key regulators of sleep spindle dynamics (sigma power), including from our own previous work demonstrating that direct manipulation of locus coeruleus–norepinephrine infraslow rhythms alters spindle organization and sleep continuity. Thus, our findings are conceptually aligned with prior studies linking sigma fluctuations to memory consolidation.

      However, we would like to clarify an important distinction in our translational approach. While Figure 4 demonstrates that heart rate dynamics during sleep can predict memory performance in mice, Figure 5 was designed to address the translational potential of these findings in humans, where direct neuromodulatory readouts are not readily accessible. Here, we deliberately focused on heart rate bursts and their associated sleep dynamics as a clinically tractable physiological measure.

      We acknowledge that the pre-HRB sigma increase may appear conceptually similar to the measure used by Lecci et al.; however, our approach is not equivalent. Rather than selecting spindle or sigma peaks themselves, we aligned analyses to heart rate accelerations and examined the robust upregulation of sigma activity preceding these events. In this framework, sigma activity serves as a physiological readout linked to autonomic dynamics, rather than being the primary anchor of analysis. We chose this measure because heart rate bursts are influenced by multiple physiological factors, and the associated sigma dynamics provided the clearest and most robust relationship with memory outcomes in the human dataset.

      We have revised the Discussion to more clearly acknowledge the similarity to prior work. We believe our findings extend prior observations by providing evidence that sleep-related heart rate fluctuations may serve as a non-invasive readout of the memory-preserving function of sleep.

      (Discussion, page 18). “Previous work in humans demonstrated that the strength of infraslow sigma oscillations correlates with sleep-dependent memory consolidation in humans (Lecci et al., 2017).”

      (1.8) Conclusions as to whether sigma fluctuations might be slightly slower and less powerful in mice are not justified. More work is required to determine which infraslow manifestations are most useful for cross-species comparisons. Moreover, little is currently known about LC activity in human sleep. A careful look into how pupil diameter correlates with sigma power should provide clues for further discussion (see Carro-Dominguez et al). A detailed study of infraslow fluctuations in human sleep should also be discussed https://doi.org/10.1101/2024.11.06.620875.

      We thank the reviewer for this thoughtful comment. We agree that our original phrasing suggesting that sigma dynamics in mice may be “slower and less powerful” than in humans was overly interpretive. We have removed this sentence from the Results section. In addition, we have expanded the Discussion to more thoroughly integrate recent human literature on infraslow sleep dynamics.

      (Discussion, page 19). “Importantly, infraslow fluctuations of sigma power in human sleep have received growing attention. Recent work demonstrates that the infraslow fluctuation of sigma power segments N2 sleep into functional phases associated with arousal and memory-related sleep markers (Dimitriades et al., 2024). Complementary findings using pupillometry show that pupil diameter fluctuates on similar infraslow timescales during NREM sleep and is inversely related to spindle clustering, providing indirect evidence that arousal-related noradrenergic dynamics shape human sleep microstructure (Carro-Domínguez et al., 2025). However, LC activity during human sleep remains inferred rather than directly measured, and systematic perturbation studies linking LC output to spindle–autonomic coupling in humans are currently lacking. Together, these observations underscore both the promise and the current limitations of cross-species comparisons of infraslow sleep dynamics.”

      (1.9) Citation of literature should be as explicit as possible. Referring to "many studies rely on plasma levels..." while including some that actually did real-time fiber photometric measures is misleading.

      We thank the reviewer for this suggestion and agree with the reviewer about the importance of accurately representing prior studies. We have now updated the discussion to clarify this.

      (Discussion, page 15). “Many studies have linked HR to NE (Fawaz and Simaan, 1963; Sundaram et al., 1991; Tanoue et al., 2022; Watson et al., 1979). Many rely on plasma levels of NE, which, while linked to central NE (Gurguis and Uhde, 1998), has low temporal resolution, making causal interpretations harder. In recent years, the use of biosensors and fibre photometry allows for very reliable estimate of the temporal dynamics of NE changes making association to HR more precise (Osorio-Forero et al., 2021).”

      (1.10) Regarding the discussion on the baroreflex: The emphasis is placed predominantly on sympathetically mediated effects. However, the description of the baroreflex loop is incomplete, as it overlooks the substantial contribution of parasympathetic modulation. In particular, heart rate adjustments within the baroreflex are primarily mediated by parasympathetic mechanisms.

      We agree with reviewer that this important notion should be added. We have rephrased the discussion as below:

      (Discussion, page 20). “These neurons suppress the activity of the rostral ventrolateral medulla, ultimately resulting in reflex parasympathetic activation with sympathetic inhibition lowering the HR (Aicher et al., 2000; Lanfranchi and Somers, 2002).”

      (1.11) Regarding the interpretation of HRV analysis, the discussion places disproportionate weight on sympathetic modulation in the interpretation of HRV metrics. This framing is inconsistent with recent conceptual clarifications, including a recent Nature Reviews Cardiology article by Menuet et al. (10.1038/s41569-02501160-z), which cautions against simplistic low-frequency/high-frequency (LF/HF) interpretations of autonomic balance.

      We thank the reviewer for this thoughtful comment. We agree that HRV frequency bands should not be interpreted as exclusive markers of specific autonomic branches. In the original manuscript, we cited Billman (2013), which challenges the validity of LF/HF as a measure of sympatho-vagal balance, to acknowledge these conceptual limitations. However, we recognize that some of our phrasing, particularly in the section discussing compensatory HR decelerations, may have implied branch-specific dominance.

      We have updated the Discussion section as follows:

      (Discussion, page 17). “Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of nonlinear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      We have also replaced the title of the Discussion section title “Locus-coeruleus-mediated sympathetic control drives compensatory heart rate decelerations” with “Locus-coeruleus activation drives compensatory heart rate decelerations via autonomic feedback mechanisms”

      (1.12) In particular, the manuscript attributes VLF power primarily to sympathetic activity, despite evidence that very-low-frequency RR-interval oscillations are strongly dependent on parasympathetic integrity. Notably, parasympathetic blockade has been shown to nearly abolish VLF oscillations in humans (Taylor et al., Circulation, 1998; doi:10.1161/01.CIR.98.6.547).

      We thank the reviewer for highlighting this important point. It was not our intention to imply that VLF HRV is driven primarily by sympathetic activity or to disregard the important role of parasympathetic integrity in shaping VLF oscillations. Our intention was to present the multifactorial and incompletely resolved nature of the VLF component; however, if our wording can be interpreted otherwise, we agree that clarification is warranted.

      In response, we have revised the manuscript to better reflect the current understanding of VLF physiology. Specifically, we now make clearer that, while VLF HRV remains less mechanistically defined than HF and LF HRV, substantial evidence supports a strong parasympathetic contribution to the expression of VLF oscillations. We have incorporated the reviewer-suggested references within this comment and others and clarified this point throughout the revised manuscript.

      (Discussion, page 16). “While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998).”

      (Discussion, page 17). “Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014).”

      (1.13) Moreover, based on work by Armour (2003) and Kember et al. (2000, 2001), the VLF rhythm is thought to emerge from stimulation of afferent sensory neurons within the heart, further arguing against a purely sympathetic interpretation. Together, these findings indicate that the discussion overemphasizes sympathetic mechanisms and underrepresents the contribution of parasympathetic and afferent cardiac pathways to HRV, particularly in the VLF range.

      We thank the reviewer for this important point. Our response to this comment is largely aligned with our response to comment (1.12). In the revised Discussion, we now more clearly acknowledge that VLF oscillations likely arise from more complex cardioautonomic interactions than simple sympathetic and parasympathetic innervation alone. At the same time, we have intentionally avoided an extensive discussion of these mechanisms, as our experimental design does not directly address these pathways.

      Part 2 - There are a number of conceptual issues that need more careful elaboration:

      (2.1) A highly problematic point throughout this study is the choice of AUCs, notably of NE signals and RR intervals, rather than the signal amplitudes. It confuses the correlations to events of different durations, such as MAs and wakefulness. It is not possible to draw conclusions of the kind "cardiac rhythm are tightly coupled with the infraslow phasic NE ..." because such statements ignore that AUCs conflate amplitude and duration of a signal.

      We thank the reviewer for this comment and agree that the distinction between amplitude- and duration-related signal features is important. We deliberately chose AUC measures because our intention was to capture the overall physiological response over time, including both the magnitude and temporal evolution of the signal, rather than relying solely on a single peak value. In this context, AUC provides information about the shape and sustained nature of NE and RR changes within a defined time window, which we considered particularly relevant for temporally dynamic responses.

      At the same time, we appreciate the reviewer’s concern that AUC may conflate response amplitude and duration, particularly when comparing events of different lengths such as microarousals and wakefulness. To minimize this issue, the AUC windows were intentionally kept relatively short and fixed, thereby limiting the influence of prolonged wake episodes or differences in transition duration on the measure. Thus, our intention was not to quantify the total duration of awakenings, but rather the immediate physiological response profile surrounding the event.

      To address this concern more directly, we compared amplitude- and AUC-based measures across all mice. The results are now shown in Suppl. Figure 1g. We found a strong correspondence between amplitude and AUC measurements for NE signals, indicating that the observed relationships are not dependent on the choice of metric. A similar, albeit weaker, relationship was observed for RR responses. This likely reflects physiological constraints on heart-rate dynamics, where the initial heart-rate acceleration is relatively similar across vigilance-state transitions (Figure 1e), while the duration of the response differs substantially. As a result, amplitude measures may underestimate differences between transitions, whereas AUC better captures the extent of the cardiac response. Importantly, direct comparison of NE and RR amplitudes still revealed a significant relationship, supporting the overall conclusion that cardiac and noradrenergic responses are coupled. However, this relationship was weaker than that observed using AUC measures, suggesting that incorporating temporal aspects of the response captures additional biologically relevant information.

      In addition to the figure, we have revised the Results section to include these considerations:

      (Results, page 7): “These comparisons were performed using area under the curve (AUC) estimates of NE and R-R responses. Importantly, comparing AUC and peak amplitude measures showed a similar overall relationship, and the coupling between NE and RR remained significant when only peak amplitudes were considered (Suppl. Fig. 1g), indicating that the findings are not solely driven by the response duration aspect of AUC. The somewhat stronger relationship observed with AUC-based measures may reflect rapid saturation of heart-rate responses across vigilance-state transitions, making response persistence an informative component of the physiological signal.”

      (2.2) A next problematic aspect of the study is the analysis of noradrenergic signals and HR at transitions (e.g., from NREM sleep to sleep wakefulness or to microarousals). The abstract does not mention these data, leaving open how they fit into the paper's message. There are also two problems with it: a) Noradrenaline levels increase with wakefulness, as do many other neuromodulators. This is not novel. Furthermore, why use an AUC measure for 0-25 s when MAs last only 5 or 15 s? b) biosensor signal comparisons are difficult to make for state transitions, because blood flow changes and modifies the fluorescent signal.

      We thank the reviewer for these comments and appreciate the opportunity to clarify the rationale and interpretation of these analyses.

      Regarding the inclusion of vigilance-state transitions, our intention was not to claim novelty in the observation that NE levels increase during wakefulness. Rather, these analyses were included to provide an additional physiological context in which to examine the coupling between NE and heart-rate dynamics. Specifically, the transition analyses allowed us to determine whether graded changes in NE across sleep-to-wake transitions were mirrored by corresponding RR changes (Fig. 1d–e) and whether these responses covaried (Fig. 1f), thereby strengthening the overall conclusion that cardiac dynamics track noradrenergic signaling across naturally occurring sleep-state fluctuations.

      Regarding the use of AUC measures, we deliberately chose this metric because it captures the overall physiological response over a defined time window, including both magnitude and temporal evolution, rather than relying solely on a peak value. In this context, AUC was intended to reflect differences in the overall response profile, for example that NE and RR responses during microarousals may return more rapidly toward baseline than during sustained wakefulness. To minimize the influence of differing event durations, the analysis window was kept fixed (0–25 s) across all transition types. Importantly, we directly compared AUC- and amplitude-based measures and found that they produced largely similar relationships (Suppl. Fig. 1g). The coupling between NE and RR responses remained significant when only peak amplitudes were considered, indicating that the observed relationship is not solely driven by response duration. However, the relationship was somewhat stronger with AUC-based measures, likely because heart-rate responses rapidly saturate across vigilance-state transitions, making response persistence an informative component of the physiological signal. As mentioned in the previous comment, we have added a Figure and new result text to highlight this.

      Regarding the concern about blood-flow related artifacts in fluorescent biosensor signals, we agree that hemodynamic contamination is an important consideration for neuromodulator recordings and applies broadly to biosensor-based measurements, including analyses of infraslow fluctuations. To minimize this issue, ΔF/F calculations were performed using the isosbestic control channel, which serves to correct for movement- and hemodynamic-related signal fluctuations. Furthermore, hemodynamic artifacts typically occur rapidly at state transitions, whereas GRAB-NE signals display slower dynamics. If uncorrected hemodynamic contamination strongly influenced the signal, we would expect abrupt signal distortions tightly aligned to arousal onset, which was not evident in our recordings. While we cannot completely exclude residual hemodynamic influences, we do not believe they account for the graded NE responses observed across vigilance-state transitions or the corresponding relationship with RR dynamics.

      (2.3) Wordings such as 'extent of arousal' to compare MAs and wakefulness are problematic. Microarousals and wakefulness are qualitatively different behaviorally, physiologically, and in terms of neuromodulatory conditions

      We thank the reviewer for the helpful comment. Our intention was not to imply that microarousals and wakefulness are qualitatively identical states differing only in magnitude. Rather, we used arousal in the broader neurophysiological sense, referring to the degree of activation of central arousal systems, with wakefulness representing part of this continuum. However, we recognize that in the sleep field, arousal is often used more specifically to describe brief EEG desynchronization events and that we furthermore use micro-arousals as a broad term for these sleep arousals, which may make our wording confusing. To avoid ambiguity, we have revised the manuscript to replace this terminology with vigilance state, sleep–wake state, or state transitions, depending on the context.

      (2.4) To interpret linear correlations between datasets, even for the ones with Rsquare values < 0.5, as 'predictive' represents an overextended interpretation that is not supported by available evidence. This concern is aggravated due to the use of AUCs that confound amplitudes and time courses. This is in particular the case for Figure 5d, 5k, or 5m…

      We thank the reviewer for this important comment. We agree that the term predictive may overstate the interpretation of these correlations, particularly given the modest R² values in some analyses. Our intention was to highlight an association between heartrate and NE-related measures rather than imply strong predictive performance or causality. We have therefore revised the wording throughout the manuscript to avoid predictive language and instead refer to these relationships as associations or correlations.

      Regarding the use of AUC, we selected summary metrics based on the physiological characteristics of the signal of interest. In cases where responses were characterized primarily by rapid shifts to a new level, amplitude measures were used. In contrast, AUC was chosen when both the magnitude and duration of the response were considered physiologically relevant.

      Part 3. A substantial number of experimental and analytical points require clarification. Here is a list of a few examples; many observations noted here apply equally to other figure panels.

      (3.1) Many figure panels leave it open about whether averages or representative data are shown, how many animals are included, and what kind of measures are plotted.

      We thank the reviewer for this comment and agree that clarity in figure presentation is important. In response, we have carefully revised the figure legends throughout the manuscript to more explicitly state whether data shown are representative examples or group averages, clarify the number of animals included in each analysis, and specify the measures being plotted. We have also added “data are shown as mean ± SEM” where this information was previously not explicitly stated and clarified when n refers to the number of animals. We believe the revised figure legends now provide clearer guidance for interpretation, and further statistical details are available in the accompanying statistics table. It should also be noted that number of animals used for all experiments can be found in the Method section ‘Mice’.

      (3.2) In case experiments were done in a paired manner (e.g., the LC stimulations), individual data points should be shown connected for the different conditions.

      We thank the reviewer for this suggestion. While we agree that connecting individual data points is valuable for paired experimental designs, this visualization is not appropriate for the LC stimulation analyses presented here. Specifically, each condition reflects pooled stimulation events selected across animals rather than a single summary value per animal that can be directly matched across columns. As such, individual points in one condition do not map one-to-one onto points in the next condition, making connected visualizations potentially misleading.

      Importantly, although the data are presented as pooled event-level measures, the paired structure of the experiment was accounted for in the statistical analyses, such that repeated measurements within animals and the paired nature of the design were included in the relevant comparisons.

      (3.3) Numerous analyses involve heart rate measures from the neck EMG during wakefulness. However, it is not specified how, in this case, RR peaks could be detected within the high-activity EMG.

      We thank the reviewer for this question. The procedure for heart-rate detection during wakefulness is described in detail in the Heart rate detection Methods section. Briefly, RR intervals were extracted from preprocessed neck EMG recordings using a previously validated approach for mouse sleep studies. To minimize contamination from movement-related EMG activity during wakefulness, R-peaks were not detected within periods extending 100 ms before and 250 ms after detected movement, as these segments were considered too noisy for reliable peak detection. Heart-rate estimates during these excluded periods were subsequently interpolated using surrounding valid RR intervals to preserve temporal continuity and enable analysis of HR dynamics before and after movement episodes.

      (3.4) Figure 1b: PSD for RR intervals. The supplementary figure says that 11-minutelong NREMS or 5-minute-long NREMS periods were used. These are very rare events in mice, for which the average bout duration is around 2 min and the cycle length is 10 minutes. The methods do not explain how these bouts were chosen, how many of them were included, and why shorter bouts were not analyzed. Single cases or means?

      We thank the reviewer for this important point and agree that additional clarification was warranted. The selection of long NREM (including microarousals) periods was motivated by methodological considerations related to spectral analysis of very slow oscillations rather than by an assumption that these bout lengths are representative of average NREM duration in mice. Because our analysis focused on very-low-frequency dynamics (~1 cycle every 50 s), sufficiently long continuous recordings are required to reliably estimate power at these frequencies and avoid fragmentation-related edge effects that disproportionately affect shorter bouts.

      For this reason, shorter NREM episodes were excluded from the PSD analysis, as they do not provide sufficient duration to robustly capture slow-frequency components. We initially compared PSD estimates using NREM periods of at least 11 min (allowing ~10 VLF cycles) and 5 min duration to assess whether the shorter 5 min recordings introduced bias (Supplementary Fig. 1f). While longer periods resulted in higher overall power estimates, the frequency distribution remained highly similar between conditions. We therefore selected 300 s (5 min) as the inclusion criterion for the remainder of the study, as periods exceeding 10 min are uncommon in mice and would substantially limit analyses across experimental paradigms.

      To further reduce bias related to bout duration, PSD estimates were weighted by NREM episode length, as longer bouts showed systematic effects on power estimates. For Figure 1, the analysis included 144 NREM-with-MA bouts across 7 animals, and data shown represent group means rather than single examples.

      This description is also included in the Methods section/Data Analysis.

      (3.5) Figure panel 1c,d: In Panel c, what is plotted?

      As stated in the figure legend, it is the cross-correlation between NE and R-R. We have added more description in the figure legend.

      A cross-correlation between two signals? If yes, how were these signals chosen per transition? Looks rather like they plot some time course across a transition. What is time point 0?

      We thank the reviewer for pointing out that this analysis was insufficiently explained. The analysis shown represents a cross-correlation between the NE and RR signals, performed to assess their temporal relationship across different vigilance-state transitions. Specifically, for each transition type, NE and RR signals were extracted within the corresponding time windows and cross-correlated to determine the strength and timing of their interaction.

      In this context, time point 0 (lag = 0) represents perfect temporal alignment between the two signals. Positive or negative lags indicate whether changes in one signal systematically precede or follow changes in the other. Due to methodological differences between the rapid electrical heart signal and the slower fluorescent NE signal, we intentionally avoided overinterpreting fine temporal lead–lag relationships. Rather, the aim of this analysis was to characterize the overall interaction between the two signals across transitions.

      Because the relationship between NE and RR was predominantly inverse, negative cross-correlation values indicate that increases in NE are associated with decreases in RR (i.e., faster heart rate), and vice versa. Thus, this analysis served primarily to confirm and extend our other findings by quantifying the interaction between NE and heart-rate dynamics across vigilance-state transitions.

      We have revised the descriptive sentence in the Results section to improve clarity and have added a more detailed description of the signals included directly in the figure panel.

      (Results, page 7). “Here, cross-correlation analysis revealed a predominantly negative relationship between NE and RR across vigilance-state transitions, indicating that increases in NE were associated with reductions in RR (i.e., faster heart rate; Fig. 1c).”

      (Figure 1 legend). “Cross correlation (how strongly and at what temporal offset the two signals covary) between NE and RR during transitions”.

      In Panel d, what is measured here? Is the time point of the dotted line a NA trough or the moment of a transition? Show the data with connected lines. Heart rate calculation during wakefulness?

      We thank the reviewer for these questions and apologize that this was not sufficiently clear. In panel d, the dotted line indicates the NE trough, not the moment of a vigilance state transition, as specified in both the figure and figure legend. The analysis is aligned to detected NE troughs during sleep and examines the subsequent physiological dynamics, including transitions into wakefulness.

      We have updated the sentence in the result section to make it a bit more clear:

      (Results, page 7): “To explore the NE-RR relationship across sleep-wake transitions, we examined four progressive sleep-to-wake transitions using the preceding NE trough as the time stamp (time 0):…”

      Regarding the suggestion to connect data points, as noted in an earlier response, these analyses are based on event detection, where multiple events contribute from each animal. Thus, each condition represents pooled events across animals rather than a single matched value per animal, making connected-line visualizations inappropriate and potentially misleading. Importantly, the paired structure of the experimental design was accounted for in the statistical analyses.

      Heart-rate detection during wakefulness was usually not possible due to movement artefacts in the EMG. In the detection, we excluded periods with movement. Our EMG-based detection approach is described in the Heart rate detection Methods section. Briefly, movement-contaminated periods were excluded from R-peak detection (100 ms before and 250 ms after detected movement), and RR intervals were subsequently interpolated using surrounding valid values to preserve temporal continuity. Because analyses were aligned to NE troughs occurring during sleep, we limited the temporal window to avoid excessive contamination from movement-related noise associated with subsequent wakefulness. We have slightly revised the text to make these points clearer.

      Panel C is a cross correlation between NE and RR from the same traces included in the mean traces. x=0 is the NE through like the other figures.

      Panel D is the mean traces of NE during transitions with the dotted line at x=0 being the NE though

      Yes, but x = 0 for the cross-correlation is not the NE trough. It says something about how aligned NE and R-R are in time (see response further up in (3.5)).

      (3.6) The sigma band should ideally be chosen between 10-15 Hz for better consistency with the literature.

      We thank the reviewer for this suggestion and agree that consistency in the definition of frequency bands is important for comparison across studies. The sigma range used in the present manuscript was selected to match our previous publications and analyses, thereby allowing direct comparison with our earlier findings on LC-mediated regulation of sleep spindles and infraslow sleep dynamics. Maintaining the same band definition also ensured consistency across the datasets analyzed in this study.

      We acknowledge that having a shared definition of sigma power would improve alignment and facilitate comparisons across laboratories. At present, there remains some variability in the exact frequency boundaries used for spindle and sigma analyses across studies and species, although there are ongoing efforts within the sleep field to improve standardization. Importantly, we do not expect that modest adjustments of the sigma-band boundaries would materially affect the conclusions of the present study, as the spindle-related activity of interest lies well within the selected frequency range and the observed effects are broad rather than restricted to a narrow frequency bin.

      Moving forward, we aim to follow emerging consensus recommendations where appropriate. To clarify this point for readers, we have added a statement in the Methods section explaining that the sigma band was chosen to maintain consistency with our previous publications.

      (Methods: EEG power, page 29). “The sigma band was defined as 8 - 15 Hz to maintain consistency with our previous publications. Modest differences in sigma-band boundaries are not expected to affect the main conclusions.”

      (3.7) What are 'extreme LC stimulation frequencies'. The only information available is that stimulations were done for 2s at 20 Hz.

      We thank the reviewer for pointing out that this wording was unclear. By “extreme LC stimulation frequencies”, we did not refer to the within-stimulation pulse frequency (which remained constant at 20 Hz for 2 s across all conditions). Rather, we referred to the effective frequency of LC activation at the infraslow timescale, which was progressively increased through the closed-loop stimulation paradigm.

      Specifically, stimulations were triggered when NE levels crossed increasingly permissive thresholds during the descending phase of the endogenous NE signal. As thresholds increased over time (from −15 ΔF/F (%) to +5 ΔF/F (%)), stimulations occurred progressively earlier in the infraslow cycle, thereby compressing the oscillatory period and increasing the effective frequency of LC recruitment while preserving the endogenous temporal structure of NE dynamics.

      Thus, “extreme stimulation frequencies” refers to the highest rate of repeated LC activations achieved through the closed-loop paradigm, where stimulations became increasingly frequent at the infraslow level rather than changes in the 20 Hz pulse train itself. To avoid confusion, we have revised the wording throughout the manuscript to refer more explicitly to faster infraslow LC activation frequencies or increased infraslow stimulation frequency.

      We have added more information in the result section to highlight this better.

      (Results, page 8). “We employed a closed-loop paradigm, where LC stimulations (2 s 20 Hz (10 ms) blue laser pulses with a light intensity of 5 mW) were triggered when NE levels fell below increasing thresholds (-15, -10, -5, 0 and 5 ΔF/F (%), Fig. 2a-b, Methods). This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (3.8) Could the discrepancy between panels 4c, left and right, be due to limited sample size? The whole figure lacks indications of sample numbers, making interpretation difficult.

      We thank the reviewer for this comment. We assume the reviewer is referring to the apparent discrepancy between the NE response and heart-rate response following LC suppression (Fig. 4c–d), where LC inhibition induced a clear reduction in NE levels, whereas mean heart-rate responses were less pronounced.

      The figure legends state sample sizes and event numbers. Specifically, for these analyses, n = 8 animals (4 Arch, 4 YFP) were included, comprising 48 Arch events and 38 YFP events, our interpretation is that this discrepancy reflects a biological observation rather than a failed manipulation. Specifically, LC suppression robustly reduced NE levels, confirming the effectiveness of the optogenetic intervention, whereas HR did not exhibit a similarly consistent group-level response. However, as highlighted by the correlation analyses, variability in RR responses remained associated with the magnitude of NE suppression, suggesting that heart-rate dynamics still reflected noradrenergic modulation at the individual-response level despite the absence of a strong mean effect.

      (3.9) It would be great if Figure 4 j could be more explicitly illustrated. For example, behavioral traces that lead to higher NFR and corresponding changes in RR AUC should be shown for animals with large and small effect sizes. The sample size seems excessively low. Can this explain the difference in slopes compared to Figure 4e, right panel?

      We thank the reviewer for this suggestion. To improve the interpretation of Figure 4j, we have now added representative examples illustrating animals with high and low memory performance and their respective RR responses following LC suppression.

      We agree that the sample size for this analysis is limited. As noted in the Methods, Figure 4j is based on a secondary analysis of a previously published dataset (Kjaerby, Andersen et al., 2022), where heart-rate measures were retrospectively extracted from EMG recordings. Due to noise-related limitations in RR detection, reliable cardiac measures could not be obtained from all animals, reducing the number of subjects available for this analysis. For this reason, we have deliberately avoided direct statistical comparisons between Arch and YFP animals and instead limited our conclusions to the observed association between RR responses and memory performance across animals.

      Regarding the difference in slope compared with Figure 4e, the two analyses are based on different levels of aggregation and address different questions. Figure 4e examines the relationship between NE and RR responses across individual LC suppression events, resulting in multiple observations per animal. In contrast, Figure 4j uses a single mean RR response and a single behavioral outcome per animal. Furthermore, Arch and YFP animals were pooled in Figure 4j to maximize statistical power and because the dataset was not sufficiently powered for direct group comparisons. Consequently, the slopes are not expected to be directly comparable between the two figures.

      (3.10) It would be important to show anatomical validation of viral expression in THcre animals and optic fiber positioning.

      We thank the reviewer for this comment. Anatomical validation of viral expression in TH-Cre animals and optic fibre placement was performed for these experiments and has been reported previously in Kjaerby et al. Nature Neuroscience paper, from which this dataset was derived. Specifically, viral targeting and fibre positioning were histologically verified as part of the original experimental validation. This is mentioned in the Method section ‘Surgery’: Viral expression and injection sites were validated through immunostaining of perfused brain slices from the experimental animals (see Kjaerby et al. (3) for more information).

    1. Author Response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

      Thank you for this excellent summary. We recommend one change: removing the word “simultaneous”. It could perhaps be replaced with “concurrent” or simply omitted. Tuft dendrites and somata were imaged on alternating days, and most readers will probably interpret “simultaneous” as implying a faster, interleaved sampling rate.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      (1.1) In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      We thank the reviewer for their insightful summary of the significance of the differences that we identified in the task-related activity of tuft dendrites and somata.

      Weaknesses:

      (1.2) At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      We thank the reviewer for this feedback. We absolutely agree that detailed computational models of each proposed computation will be very valuable and constitute an important follow-up to this work. We hope to collaborate with theorists to take that next step. Each possible computation noted by the reviewer reflects distinct differences that we observed in the task-related activity of tuft dendrites and somata. They are not mutually exclusive hypotheses to explain the same phenomenon. As such, we think they are best addressed independently in future modeling work. By making the data and a concise description of the main findings available immediately, we hope to allow computational experts in each of these areas to take advantage of the results of this study without delay.

      (1.3) My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      We thank the reviewer for raising this point. We agree that our use of projections along coding directions (CDs) defined at each time point is a less conventional use of coding directions, although nearly identical calculations have been previously used to assess population-level selectivity and code stability in this task (Chen et al., 2017; Yang et al., 2022). As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task-dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task-variable. Finally, because the primary goal of the study was to identify any differences between tuft dendrite and somatic encoding, we think that calculating the population selectivity at each timepoint gives readers a less biased view of the selectivity of the two compartments, whereas calculating a CD over a single arbitrary time window could conflate differences in dynamics with differences in selectivity.

      We agree that calculating the CD at each timepoint makes it hard to see where the code is stable and where it is rotating, and thus how a downstream neuron might “read out” the signal. To provide this information, we have added new panels to the supplement showing the correlation of CDs across time (Figure 4 - figure supplement 1B,D). We also now provide this information for CR-CA in Figure 5—figure supplement 2A (the plots previously presented in 2A were the correlations of CR with CA, rather than CR-CA; an error that has been fixed). The following changes were also made to the Results section to clarify this issue:

      “From the linear model, we calculated coding directions (CDs) at each timepoint that maximally separated Stimulus, Choice, and Outcome activity (Figure 4F; Figure 4—figure supplement 1A) and estimated the direction and selectivity along each dimension across time (Figure 4—figure supplement 1B-E; see Methods). Allowing CDs to rotate in time (see Figure 4—figure supplement 1B,D), although unconventional, ensured that comparisons of population selectivity across the two compartments were not biased by the selection of an arbitrary CD time window.”

      We also identified a mistake in the description of CD orthogonalization in the Methods, which has been corrected as follows:

      “For each timepoint, each selectivity CD was then orthogonalized with respect to the other two selectivity CDs by a QR decomposition in which that selectivity CD was last in the order.”

      (1.4) A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      We agree with the reviewer that by the end of the late period, performance on the right side (Figure 1I) has still not returned to pre-shift levels. This may reflect mice not fully mastering the task, as the reviewer suggests, or it may reflect that after the shift, the right port is substantially more difficult to reach than the left port. Unfortunately, because of high animal-to-animal and lick-to-lick variability, plotting the post-shift lick angle at a finer level of granularity is not statistically informative.

      With regard to the nature of the learning in our task, we selected “skill learning” as the best description of the motor learning task based on distinctions between adaptation and the learning of motor skills by Krakauer et al., 2019 and Heald et al., 2021. Conceptually, the difference is whether an existing motor controller memory is simply updated with new parameters, or whether the motor context has changed sufficiently that a distinct motor controller memory (which can still use parts of previous memories) is formed. In our task, after the port shift the left port forms an obstacle to reaching the right port. This obstacle was simply not present before the shift. Before the shift, ports were approximately equidistant from the mouth and easily avoided given the port separation and tongue width. Thus, avoiding an obstacle would presumably not be part of the initial motor controller memory and a distinct memory would need to be constructed.

      We agree with the reviewer, however, that given that we do not have fine-timescale dynamics of behavioral changes in response to the shift, and did not conduct other experiments (such as returning the ports to their original location) that would typically be conducted to identify “adaptation-like” or “skill learning-like” dynamics, we cannot empirically distinguish between the two. We now clarify in “Study limitations” that we call the studied behavior “skill learning” based on the nature of the task, but that our behavioral analysis cannot distinguish between adaptation and skill learning:

      “We refer to the behavioral paradigm as motor “skill learning” strictly based on the nature of the task. After the shift, mice must avoid a new obstacle close to the mouth (i.e., the left port), which we assume requires the formation of a distinct motor controller memory and therefore would be considered skill learning (Krakauer et al., 2019). However, we did not confirm that the mice exhibited specific behavioral characteristics of skill learning and it is possible that other kinds of motor learning (e.g., motor adaptation) were dominant.”

      (1.5) This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

      We understand how this could cause confusion. “Sensory” was meant to denote the nature of the differences in external events between the trial types used to calculate selectivity, not to imply that the activity was necessarily selective for detailed features of the cues outside the context of the task. Previous work in ALM cortex has labeled this selectivity direction as “stimulus” (Yang et al., 2022; Chen et al., 2024) to better emphasize that it is simply defined by the external cue. Where appropriate, we have revised the article to use “stimulus” or “instructional cues” in place of “sensory” for clarity and to better conform with convention.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      (2.1) The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      We thank the reviewer for their appreciation of the study design, imaging methods, and scholarship of the article.

      Weaknesses:

      (2.2) It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      We thank the reviewer for emphasizing the need to discuss this potential issue. For the following reasons, it is likely that the vast majority of dendrites we imaged in layer 1 originated from layer 5 ET neurons. First, as the reviewer notes, the provided images suggest enrichment in layer 5. This enrichment likely reflects the fact that most L6 CT neurons in motor and premotor cortex send denser projections to other thalamic nuclei than to VM thalamus (Winnebust et al., 2019, Cell), where we targeted our retrograde-Cre injections. Second, L6 CT neurons are predominantly untufted (Ledergerber and Larkum, 2010, J. Neurosci.), including in motor and premotor cortex (Peng et al., 2021, Nature; Ichikawa, 2025, Front. Neuroanat.). A recently identified subclass of L6 CT neurons in secondary motor cortex has dense projections to VM thalamus, but this class also appears to extend minimal dendrites into L1 (Li et al., 2024, bioRxiv). Nonetheless, we did not label post-hoc tissue collected from imaged mice with markers of precise laminar boundaries, and thus cannot definitively rule out the possibility that dendrites from a subclass of L6 CT neurons with tuft dendrites were also imaged. We have added the following paragraph to the “Study limitations” section to make readers aware of these issues:

      “L5 ET neurons in premotor cortex elaborate extensive tuft dendrites in L1, whereas Layer 6 (L6) corticothalamic (CT) neurons are predominantly untufted (Jiang et al., 2020; Peng et al., 2021). Thus, although we cannot rule out the possibility that dendrites from a subclass of L6 CT neurons were also sampled, it is likely that the vast majority of dendrites we recorded in L1 originated from L5 ET neurons.”

      (2.3) The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      We thank the reviewer for requesting this useful additional information.

      In all cases, the model was retrained for each dendritic or somatic imaging session. Denoising improved segmentation consistency, as measured by comparing segmentations of individual sessions from the same animal. This is now specified in the Methods as follows:

      “The DeepInterpolation model was trained on each imaging session prior to denoising of that session. Denoising prior to NMF-based segmentation resulted in more robust and consistent dendrite segmentation than NMF-based segmentation without prior denoising (0.79 +/- 0.01 ⍴ vs. 0.46 +/- 0.01 ⍴; mean of the max Spearman correlation of components across sessions; random subsample of N = 3 mice, 15 sessions, 400 components).”

      With regard to how denoising impacts activity traces, examples were shown in Figure 2I, K. To provide more quantitative information to the reader, we calculated estimates of the power and reliability of the spectral content of dendrite activity traces extracted with or without denoising. These data are now shown in the new panel, Figure 2 - figure supplement 2G. The power spectral density of the denoised activity and the estimated reliable power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C). Some frequencies beyond this point have been suppressed beyond what would be expected due to photon shot noise (as estimated by the replicate coherence-weighted PSD, or “recoverable” PSD). Further characterization of the precise nature of the suppressed high-frequency information – which could be suppressed artifacts (e.g., fast brain motion) or lost signal detail (i.e., GCaMP8m rise kinetics) – is beyond the scope of this paper.

      Details of the PSD calculations have been added to the Methods, and the following statement has been added to the Results: “Power spectral density of the denoised traces and the coherence-weighted power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G; Methods), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C).”

      (2.4) The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      Preparatory selectivity and ramping activity can be seen in Figure 3H, Figure 6I, and Figure 5 – figure supplement 1B. We note that in ALM cortex, the ramping mode explains a minority of the total variance (~17%, Yang et al., 2022), but it can appear particularly prominent in projections along certain fixed CDs.

      (2.5) It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      We agree with the reviewer that being able to compare differences in signals between the dendrites and somata of the same cells would be very valuable. However, reliable tracing of tuft dendrites to somata from in vivo 2P anatomical imaging requires extremely sparse labeling, such that very few neurons are recorded per animal (Kerlin et al., 2019, eLife; Otor et al., 2022, Science). As stated in the “Study limitations” section of the Discussion, we suspected (correctly) that some task-related selectivity (i.e., selectivity for corrective action) would be sparsely represented in the dendrites, and thus adopted a labeling and image processing strategy that allowed us to record from many dendrites per animal. This strategy necessarily comes at the expense of generating a labeling density that precludes reliable tracing of tuft dendrites to their respective somata based on 2P morphology alone. As discussed in our response to reviewer comment 2.2 and a new paragraph of “Study limitations,” substantial contamination of the dendrite recordings by dendrites of L6 CT neurons is highly unlikely. Future studies could use simultaneous functional imaging across large volumes combined with activity-based segmentation or post-hoc high-resolution imaging of tissue sections registered to in vivo 2P imaging to accomplish both high-throughput dendritic imaging and reliable tracing.

      We thank the reviewer for pointing out that we could discuss selective backpropagation as a potential mechanism more explicitly. Our results are consistent with previous studies of L5 tufts in vivo (Francioni et al., 2019, eLife), including in ALM cortex (Maristany de las Casas et al., 2026, Science), that reported that rates of multi-branch calcium transients in the tuft dendrites of L5 neurons are lower than somatic spike rates. As discussed in “Study limitations,” there is not a clear approach in our data to determine the precise nature of the events underlying the calcium transients we measured in the tuft dendrites. Selective backpropagation of action potentials is certainly one possibility and we agree that recent research from Dr. Adam Cohen’s group should be discussed. We have added the following to the Discussion:

      “Based on previous calcium imaging of L5 tufts in ALM cortex of mice engaged in similar tasks (Kerlin et al., 2019; Maristany De Las Casas et al., 2026), we suspect that most of the activity we measured was coincident with global tuft or hemi-tree events, as well as somatic spiking. Recent in vivo voltage imaging in the hippocampus has also indicated that most spikes in distal dendrites start as bAPs that have been selectively amplified (Wu et al., 2026; Lee et al., 2026).”

      (2.6) The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      We thank the reviewer for giving us the opportunity to clarify this issue. As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task variable in the population activity. Thus, the analyses are not at odds with the nature of the measurements in the study.

      Nevertheless, it is true that by recalculating the CD at each timepoint, our selectivity projections do not provide the same information as conventional projections along a fixed CD, which can indicate where the selectivity code is stable and where it is changing. To provide this information we have added new panels to the supplement showing the correlation of selectivity CDs across time (Figure 4 - figure supplement 1B, D).

      (2.7) This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

      We agree with the reviewer. As noted by the reviewer in comment (2.1), we combined a number of approaches in an innovative manner to explore how tuft dendrite activity differs from somatic activity at the population level during motor learning. These measurements provide the necessary foundation for future mechanistic studies and we think it is appropriate to share them at this stage of investigation and in the format of this article.

      Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      We thank the reviewer for highlighting interesting findings in the paper and their assessment that it provides a “useful foundation for future mechanistic studies”.

      Weaknesses:

      No major weaknesses were identified.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A very minor suggestion: it would be useful to mention the model organism in the abstract or title.

      (1.6) Thank you for catching this. We have added the model organism to the abstract as follows:

      “Using longitudinal two-photon calcium imaging, we investigated sensorimotor encoding in the apical tuft dendrites and somata of L5 extratelencephalic (ET) neurons in the frontal cortex of mice during learning of a discrete change to a cued dexterous action.”

      Reviewer #3 (Recommendations for the authors):

      Major:

      (3.1) Lines 197-199: It is unclear why the authors conclude that somas have stronger representations of choice and task outcome. In Figure 4G, there is no significant difference between dendrites and somas for Choice or Outcome coding selectivity. The differences in the Sensory/Choice and Sensory/Outcome ratios shown in Figure 4H,I are likely explained by stronger Sensory selectivity in dendrites (Fig 4g), rather than by stronger Choice or Outcome encoding in somas.

      We agree with the reviewer’s interpretation of the data. The statement at 197 - 199 was meant to reflect relative selectivity, but it was imprecise. We have replaced that sentence with the following, more precise sentence:

      “Somatic activity also encoded these features, but the representation of the stimulus was weaker – and the representation of action timing was stronger – than in the tuft dendrites.”

      (3.2) Figure 4B: The authors realign FL-associated IRFs to GO-cue timing using the mean FL latency for each trial type and animal. Because FL timing is jittered across trials and may differ between CL and CR trials, this could smear the realigned traces and complicate the interpretation of contact-associated activity. The authors should consider using trial-by-trial FL timing for realignment or quantify the impact of FL-timing variability on the resulting traces.

      We aligned average GO- and contact-IRFs in Figure 4B so that comparisons of their magnitudes could be drawn from the same time window.

      With regard to jitter across trials, we think the reviewer may have misinterpreted how the mean IRFs in Figure 4B are calculated. The contact-IRF, by its nature, is calculated once per animal and trial type with respect to FL timing and shifted once based on mean FL latency. There is no smearing due to trial-to-trial FL timing.

      With regard to systematic differences in FL timing across animals and CR vs. CL, the reviewer is correct that this could – in theory – smear the realigned mean contact-IRF shown in Figure 4B. However, differences in mean FL latency across animals and trial-types are small compared with the long-timescale contact-IRFs. Thus, the non-realigned (i.e., always FL-aligned) mean contact-IRF looks nearly identical to Figure 4B just globally offset in time, as shown in Author response image 1:

      Author response image 1.

      Since this is nearly identical to data already presented in Figure 4B, we do not think it is necessary to include it in the revised article. However, we have added the following to the Methods:

      “Population averages of contact-IRFs that were not shifted prior to averaging were nearly identical (excluding the overall temporal shift; data not shown), indicating that pooling of mean IRFs across animals and trial types produces minimal smearing of the final population IRF.”

      (3.3) Figure 5A: Are CA trials specific to motor learning, or do they reflect a corrective lick toward the alternative port after an unrewarded lick? An analysis of the second lick on left-error trials or pre-shift right-error trials could help distinguish whether correction licking reflects a general decision change after failed reward, or a motor-command correction specific to post-shift motor learning. The authors should also report the prevalence of CA versus AP trials and clarify whether these trial types are behaviorally distinct.

      We thank the reviewer for highlighting the need to emphasize that CA trials reflect a distinct behavior related to reaching the displaced port.

      By definition, CA trials started as Motor Error trials and thus reflected a corrective lick toward the same port after an unrewarded lick. Almost all first contact licks on Motor Error trials were well outside the distribution of correct left licks both pre- and post-shift (Figure 1 - figure supplement 1B,D), consistent with the interpretation of this first lick as directed toward the right port. Thus, we see no evidence suggesting that CA trials involve a decision change. CA trials are exceedingly rare pre-shift, because Motor Error trials are rare pre-shift (Figure 1I, only ~5% of all right trials).

      With regard to other error types before the shift, most expert-trained mice did not immediately sample the other port with a second lick after an unrewarded lick. They usually either stopped licking immediately or licked the unrewarded port multiple times before switching ports. When port switches occurred pre-shift, timing was highly variable across mice and trials. Even on rewarded trials, some mice would “check” the unrewarded port after consuming the reward, as can be seen in Figure 3I, J. All of these behaviors are clearly distinct from the stereotyped second lick that occurred on CA trials after the shift. We agree that the prevalence of CA and AP trials, as well as the prevalence of immediate port alternation, should be reported, and we have added that information to the article as follows:

      “On Correction Attempted (CA) trials, the first lick made contact with the incorrect port, and the mouse chose to direct a second lick toward the correct port (Figure 5A; prevalence: 54% of motor error trials). We interpreted these licks as a corrective action, because the tongue exit angle shifted further toward the correct target (Figure 5B). Abandoned Port (AP) trials were the same as CA trials, except the mouse either did not make a second attempt or the second lick was directed toward the incorrect port (Figure 5A; prevalence: 46% of motor error trials).”

      (3.4) The classification of pre-shift errors into motor and decision errors is not clear. If error-trial exit angles follow a unimodal distribution (Figure 1- Figure Supplement 1C), then the distinction between motor and decision errors may not be behaviorally well separated. The authors should explain how these categories are validated and whether conclusions depending on this classification are robust to alternative definitions.

      We do not conclude that motor errors and decision errors are distinguishable pre-shift. Pre-shift licks were classified into motor error and decision error categories only to demonstrate that the boundary we established for classifying post-shift licks classifies extremely few (~5%, Figure 1I) pre-shift licks as motor errors. No conclusions were drawn from comparisons between pre-shift licks classified as decision errors and those classified as motor errors. The categorization is defined by the distribution of exit angles pre-shift and validated by the bimodal distribution of exit angles on error trials post-shift. To improve clarity regarding our classification of pre-shift errors, we have added the following to the Results:

      “Exit angles after the shift exhibited a bimodal distribution across error trials (Figure 1G,H; Figure 1—figure supplement 1C,D), supporting this distinction in error type. The frequency of licks classified as motor errors on right-cued trials increased significantly after the shift (median pre-shift 0.06, median post-shift 0.42, p < 0.001; Figure 1I; Figure 1—figure supplement 1C,D), reflecting the new challenge of avoiding the left lickport. In contrast to after the shift, exit angles on error trials before the shift were unimodal (Figure 1—figure supplement 1C). These errors were classified based on the fixed CB in order to demonstrate that very few pre-shift licks qualify as motor errors (Figure 1I), and not to suggest that tongue trajectories before the shift are behaviorally well-separated.”

      (3.5) Figure 1- Figure Supplement 1D, post-shift decision errors: Are these truly decision errors? The lick angles appear similar to those observed before the shift, suggesting that these trials may reflect execution of a "default" or "uncertain" lick trajectory rather than an incorrect choice under the new contingency.

      The post-shift exit angles on right-cued decision error trials (Figure1 - figure supplement 1D, grey) are similar to the lick angles on correct left-cued trials pre-shift (Figure 1 - figure supplement 1A, red) and clearly different from the correct right-cued trials pre-shift (Figure 1 figure supplement 1A, blue). Thus, to the extent that the animal’s intention can be measured from lick trajectory, it was targeting the incorrect (left) port. It is also true that it may still target the previous location of the left port (a “default” left trajectory), but because the decision error makes precise targeting irrelevant to the task outcome (it is easy to reach the left port after the shift), we do not designate it as a joint decision error and motor error. As to whether the deliberative process leading to this action is somehow cognitively distinct from other behaviors typically labeled as decision errors or incorrect choices, we cannot say.

      Minor:

      (3.6) Vocabulary consistency: soma vs somata.

      When data are shown for, or derived from, multiple somata, we use “somata”. When data are shown for an individual soma (such as in a panel with data from a single example soma), we use “soma.” We could not find any use of “somas,” which would indeed be inconsistent.

      (3.7) Figure 1C: I am not sure why the lick trajectories do not depict the tongue exiting the mouse. What time window is shown? Why does it look like the trajectories are shifted to the left?

      We thank the reviewer for identifying this issue. The definition of the location labeled “mouth” was accidentally omitted. The lick trajectories in Figure 1C do depict the tongue tip once it became visible to the cameras. Jaw opening and shifting partly determined the location where the tongue became visible in the videography. These movements varied from mouse to mouse and trial to trial, so exit angle was measured from the approximate midpoint between the temporomandibular joints, which is the grey point in 1C. We have fixed the captions and Methods to precisely define this location. With regard to the appearance of a slight leftward shift in the trajectories, this reflects how the tongue exits the mouth and how the tongue tip curves downward as the tongue approaches the port.

      (3.8) Figure 1- Figure supplement 1: it could ease the comparisons to report population statistics, such as median, from panel A to panel B and D, population statistics from B to D.

      Thank you. We have added these statistics to the Figure 1 - figure supplement 1 caption.

      (3.9) Choice boundary (CB) should be defined in line 100, not 110.

      Thank you. We have fixed this.

      (3.10) Line 109: claim not supported by referenced figure (Figure 1 - Figure Supplement 1). Lick angle histogram to the right port, pre-shift does not overlap substantially with lick angle to the left port, post-shift.

      We thank the reviewer for the opportunity to clarify this. We agree that Figure 1 - figure supplement 1 is not sufficient to support the claim. First, we want to make clear that Figure 1 - figure supplement 1 does not contradict the claim. The new location of the left port can obstruct the tongue during right-cued licks, regardless of the distributions of left licks pre- or post-shift. Second, to confirm that the new location of the left port would obstruct a substantial fraction of pre-shift right-cued lick trajectories, we measured the minimum distance between tongue trajectories and the post-shift location of the left port. Of pre-shift right-cued exit trajectories, 30 +/- 5% came within 1.25 mm – half of the combined tongue width (1.5 mm) and port width (1 mm) – of the port center.

      To make this claim more precise, we have changed the statement as follows:

      “Thus, on right-cued trials, mice continuing to follow the pre-shift motor plan would be biased to more frequently contact the new left port location (Figure 1E,F; 30 +/- 5% of pre-shift trajectories came within a tongue-width of the new location) and receive punishment (i.e., timeout).

      (3.11) Line 113: claim not supported by referenced figure. Figure 1G does not display error trials.

      We have changed the line to refer to “both correct and error trials”, such that reference to Figure 1G is also appropriate.

      (3.12) Figure 2 - Figure Supplementary 3 & method: how is noise estimated?

      Thank you. The following has been added to the Methods:

      “For Figure 2 - figure supplement 3, noise was estimated as the square-root of the geometric mean of the Welch power spectrum in a high-frequency band (0.25–0.5 times the frame rate; Giovannucci et al., 2019).”

      (3.13) Figure 2D: Was imaging during the shift epoch always performed in dendrites? If so, could the imaging schedule bias comparisons between dendritic and somatic activity during learning, especially given that mice show behavioral learning between early and late post-shift sessions (Figure 1J)?

      No, imaging during the shift was not always performed in the dendrites. The following has been added to the Methods to make clear that the post-shift data reflect dendritic and somatic imaging conducted on the day of the shift with roughly similar frequency:

      “For Figure 5 and Figure 6, which make comparisons between dendritic and somatic activity during the post-shift period, 67% of animals providing somatic data (4 of 6 mice) underwent somatic imaging on the day of the shift and 80% of mice providing dendritic data (8 of 10 mice) underwent dendritic imaging on the day of the shift.”

      (3.14) Lines 163-164, "we observed that the onset of tuft activity was consistently time-locked to the GO cue (vertical green line; Figure 3B). This was in contrast to somatic activity, which had more variable timing (Figure 3E)." The authors cite panels B and E in support of this point, but these appear to be example ROIs. It would be helpful to clarify how representative these examples are, since the corresponding population summaries in panels G and H do not make the effect immediately apparent.

      These examples are representative, as supported by the population summary of activity time-locked to the GO-cue versus port contact in Figure 4B.

      (3.15) Figure 4B: It could be useful to add the lick traces here as well. To allow the reader to have an idea of contact timing with respect to the Go cue and compare the sustain response with the licking pattern.

      We understand how this could be helpful. However, since these exact traces are already present in Figure 3I,J, we think that adding them to Figure 4 is unnecessary and would add complexity to an already very busy figure.

      (3.16) Figure 5E: Why are the imaging sessions labeled 0 and +1 rather than 0 and +2? Are the dendritic and somatic imaging not alternated?

      Yes, imaging was not alternated for 3 of the 22 mice. We have clarified this in the Methods, as follows:

      “Somatic and dendritic imaging sessions alternated every other day (19 of 22 mice), except for 3 mice in which only one compartment was imaged daily (dendrite-only: 2 mice, soma-only: 1 mouse). The exceptions were due to brain curvature or the angle of the coverslip with respect to the brain, such that only one compartment could be imaged and the other compartment was underneath skull regrowth or dural thickening that made high-quality imaging impossible.”

      (3.17) Figure 6C, legend: I suppose the authors meant "remapping", not "Post-shit SI distribution" for the description of the right column.

      Thank you. We have fixed this label.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a useful finding on the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using technically sophisticated spatio-temporal calcium recordings to confirm that AVP influences α and β cells differently depending on glucose concentrations. While the study's methods - particularly the calcium imaging techniques and peptide ligand design targeting V1b receptors - are strong, the reviewers were concerned about several aspects of the experimental design. However, the results on βcell responses are incomplete and insufficient to support the manuscript's claims, especially due to the high variability of islet responses and lack of mechanistic and functional (hormone release) data. There are also concerns about the possibility of off-target effects and incomplete receptor specificity, noting that the study would have been significantly strengthened by inclusion of signaling pathway interrogation, hormone output assays, genetic validation (e.g., β cell-specific deletion of V1br), and receptor localization, although the work will still be of interest to researchers studying islet physiology in the context of health and diabetes.

      We sincerely thank the reviewers and editors for their thorough evaluation of our manuscript and their recognition of its technical strengths, including the advanced spatio-temporal calcium imaging and the rational design of selective V1b receptor ligands. We appreciate their acknowledgement of the study’s relevance for understanding AVP effects in a physiologically intact islet context and their positive assessment of our methodological rigour and innovation. The reviewers’ constructive feedback has helped us clarify the boundaries and intent of our study, which focuses on the glucose- and context-dependent modulation of α- and β-cell activity, rather than exhaustive molecular dissection.

      While the reviewers rightly emphasize the importance of receptor specificity and downstream signaling validation, we respectfully suggest that some of their concerns may reflect a lingering bias toward reductionist frameworks. Our interpretation is rooted in the emerging understanding that β-cell behaviour is largely defined by dynamic intercellular interactions within the islet collective, rather than by static gene expression or receptor localization alone (Jin et al., 2025; Korošak et al., 2021; Rutter et al., 2024). Recent studies have demonstrated that roles such as “leader” or “hub” β cells are transient and emergent, governed more by timing, environment, and local network structure than by fixed molecular identity (Postic et al., 2023; Gosak et al., 2018).

      This has profound implications for how we interpret cell responsiveness to agents like AVP: what appears as biological variability may in fact reflect context-sensitive transitions within a non-linear, self-organizing system (Stožer et al., 2021). Hence, we chose to focus on functional collective dynamics using intact pancreatic slices, rather than isolated cell models which fail to preserve the essential network architecture of islets. Although the addition of genetic models or isolated receptor measurements would strengthen receptor-specific conclusions, we argue that such approaches alone cannot resolve the physiological complexity of a system where function arises from cell–cell communication and spatiotemporal context.

      Indeed, the lack of direct correlation between receptor transcript abundance and functional outcomes has been noted in prior studies, reinforcing the view that function cannot be strictly predicted by molecular presence (Rutter et al., 2024). As articulated in our manuscript, the islet behaves as a sensory collective (Fancher & Mugler, 2017), where emergent patterns— not static cell identity—determine behaviour. This perspective aligns with broader shifts in biology away from strict genetic determinism toward causal emergence and collective agency (Ball, 2023; Levin, 2021).

      We therefore believe our study contributes not only new pharmacological insights but also a conceptual reframing of how AVP responses should be interpreted in a complex organ like the pancreas. We have added new data addressing reviewer suggestions—such as glucagon secretion assays, clarifications on the role of forskolin, and an analysis of event timing—that further support our conclusions. We also expanded the discussion on how islet variability is functionally meaningful, not just noise, and explained why β-cell responses to AVP must be interpreted within this probabilistic framework.

      We agree that future work should include receptor-specific knockouts and more direct signaling pathway assays, but these would need to be designed with careful consideration of the islet’s dynamic topology and the emergent nature of β-cell roles. In this light, we see our study not as the final word, but as a necessary systems-level foundation for more targeted interventions. We thank the reviewers again for their careful critiques and hope that our response clarifies both the rationale and scope of our work. Our revisions aim to enhance the paper’s clarity while maintaining its commitment to an integrative, physiology-rooted approach.

      We thank the reviewers and editors for their thoughtful and constructive assessment of our work. We are especially grateful for their recognition of the study’s technical strengths, including the use of spatio-temporal calcium imaging in intact pancreatic tissue and the strategic development of receptor-selective peptide ligands. We also appreciate their acknowledgement that our study contributes to the understanding of glucose-dependent AVP effects in islet physiology. The reviewers’ concerns regarding variability, receptor specificity, and functional validation helped us further clarify the scope and context of our study.

      We respectfully submit that some reservations stem from a reductionist framing that may not fully account for the collective behaviour of islets. As we and others have shown, β-cell function arises from emergent, self-organizing network dynamics, not just from static gene expression or receptor abundance (Jin et al., 2025; Korošak et al., 2021; Postic et al., 2023). In this view, pharmacological heterogeneity across islets is not simply noise or experimental inconsistency, but a signature of dynamic attractor states within the islet network (Stožer et al., 2021). Because an islet functions as a coupled system, most response variability originates from its emergent collective behavior, which eclipses variability in receptor expression or metabolic state.

      For this reason, even single-islet receptor quantification or ATP measurements would provide limited explanatory power: it is the state of the network—not absolute receptor levels—that determines whether a perturbation elicits activation or inhibition. As we illustrate in our graphical abstract, a single islet tested repeatedly under identical glucose conditions can yield divergent responses, simply because it occupies different dynamic states. These findings are in line with systems biology and network science approaches, which have revealed that cell function, especially in the β-cell collective, cannot be fully understood through reductionist parameters alone (Gosak et al., 2018; Ball, 2023).

      We have included glucagon secretion assays and new analyses to address key reviewer suggestions. Still, we chose not to pursue extensive knockouts or cAMP imaging, as these would require a different experimental scope and could risk disrupting the very dynamics we aim to understand. Likewise, while direct measurements of V1bR or IP3R expression would add molecular detail, they are not definitive without network context. The bell-shaped AVP dose-response curve and its explanation through IP3R inactivation are supported by prior studies; we invoke this mechanism not speculatively, but because it provides the most parsimonious explanation for the glucose-dependent shift in β-cell responsiveness.

      We also clarify that our study does not aim to resolve every mechanistic detail, but rather to offer a systems-level insight into how AVP modulates islet dynamics across varying glucose and cAMP contexts. The implications extend beyond AVP pharmacology, suggesting that perturbations to β-cell function must be understood within a probabilistic, state-dependent framework (Fancher & Mugler, 2017). This resonates with emerging concepts in cell physiology that emphasize causal emergence and local agency over static molecular determinism (Levin, 2021; Rutter et al., 2024).

      In summary, we see our work as part of a necessary shift in perspective—from linear receptor-function models to context-sensitive dynamic systems. We are grateful for the opportunity to revise our manuscript in response to insightful feedback and hope our clarifications and new data will strengthen its impact for the islet research community.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors confirmed earlier findings that AVP influences α and β cells differently, depending on glucose concentrations. At substimulatory glucose levels, AVP combined with forskolin - an activator of cAMP -did not significantly stimulate β cells, although it did activate α cells. Once glucose was raised to stimulatory levels, β cells became active, and α cell activity declined, indicating glucose's suppressive effect on α cells and permissive effect on β cells. Under physiological glucose levels (8-9 mM), forskolin enhanced β-cell calcium oscillations, and AVP further modulated this activity. However, AVP's effect on β cells was variable across islets and did not significantly alter AUC measurements (a combined indicator of oscillation frequency and duration). In α cells, forskolin and AVP led to increased activity even at high glucose levels, suggesting that α cells remain responsive despite expected suppression by insulin and glucose.

      Experiments with physiological concentrations of epinephrine suggest that AVP does not operate via Gs-coupled V2 receptors in β cells, as AVP could not counteract epinephrine's inhibitory effects. Instead, epinephrine reduced β cell activity while increasing α cell activity through different G-proteincoupled mechanisms. These results emphasize that AVP can potentiate αcell activation and has a nuanced, context-dependent effect on β cells.

      The most robust activation of both α and β cells by AVP occurred within its physiological osmo-regulatory range (~10-100 pM), confirming that AVP exerts bell-shaped concentration-dependent effects on β cells. At low concentrations, AVP increased β cell calcium oscillation frequency and reduced "halfwidths"; high concentrations eventually suppressed β cell activity, mimicking the muscarinic signaling. In α cells, higher AVP concentrations were required for peak activation, which was not blunted by receptor inactivation within physiological ranges.

      Attempting to further dissect the role of specific AVP receptors, the authors designed and tested peptide ligands selective for V1b receptors. These included a selective V1b agonist; a V1b agonist with antagonist properties at V1a and oxytocin receptors; and a selective V1a antagonist. In pancreatic slices, these peptides seem to replicate AVP's effects on Ca<sup>2+</sup> signaling, although responses were highly variable, with some islets showing increased activity and others no change or suppression. The variability was partly attributed to islet-specific baseline activity, and the authors conclude that AVP and V1b receptor agonists can modulate β cell activity in a statedependent manner, stimulating insulin secretion in quiescent cells and inhibiting it in already active cells.

      We applaud the reviewer to capture the essence of work in their introduction.

      Strengths:

      Overall, the study is technically advanced and provides useful pharmacological tools. However, the conclusions are limited by a lack of direct mechanistic and functional data. Addressing these gaps through a combination of signaling pathway interrogation, functional hormone output, genetic validation, and receptor localization would strengthen the conclusions and reduce the current (interpretive) ambiguity.

      Thank you!

      Weaknesses:

      (1) The study is entirely based on pharmacological tools. Without genetic models, off-target effects or incomplete specificity of the peptides cannot be fully ruled out.

      We partially agree with this comment and acknowledge that genetic models would provide a valuable complementary approach to address possible off-target effects or incomplete peptide specificity. However, genetic models also have important limitations, particularly when the aim is to resolve subtle, population-level physiological differences in beta cell activity. We therefore used pharmacological tools at different concentrations to test whether the observed effects were concentration-dependent and consistent with the expected receptor-mediated actions. An advantage of the pancreatic slice preparation is that it preserves much of the native tissue environment and allows pharmacological manipulation within concentration ranges closer to in vivo efficacy, thereby reducing the likelihood of nonspecific effects. To compensate for the lack of genetic models, we now emphasize the collective activity analysis as an additional strength of the study and have clarified this limitation in the revised manuscript.

      (2) Despite multiple claims about β cell activation or inhibition, the functional output - insulin secretion - is weakly assessed, and only in limited conditions. This aspect makes it very hard to correlate calcium dynamics with physiological outcomes.

      We agree that the functional output needed stronger support and have therefore expanded the hormone secretion experiments. While the effects of AVP and its analogues were tested during a stable plateau phase in the Ca<sup>2+</sup> imaging experiments, this phase provides only a narrow dynamic range for insulin release measurements in mouse slices. We therefore added a sequence of stimulations on the same slices, using 8 mM glucose and 500 nM forskolin, with glucose lowered to a non-stimulatory range between different AVP concentrations. These new experiments better define how AVP-dependent changes in Ca<sup>2+</sup> dynamics translate into insulin secretion under conditions with a broader secretory dynamic range. The new insulin and glucagon secretion data have now been added to the manuscript as Figure 5, and the text has been revised accordingly.

      (3) Insulin and glucagon secretion assays should be provided; the authors should measure hormone release in parallel with Ca2+ imaging, using perifusion assays, especially during AVP ramp and peptide ligand applications.

      We added insulin and glucagon secretion assays for AVP ramp to Figure 5.

      Additionally, there is no standardization of the metabolic state of islets. The authors should consider measuring islet NAD(P)H autofluorescence or mitochondrial potential (e.g., using TMRE) to control for metabolic variability that may affect responsiveness.

      We agree that standardization of the metabolic state of the islets would further strengthen the interpretation of the responsiveness data. We attempted to address this experimentally, but the results were inconclusive and therefore not included in the manuscript. Based on our previous unpublished observations, NAD(P)H levels appear to be significantly higher and less variable in islets within tissue slices than in isolated islets, suggesting that the slice preparation may better preserve the native metabolic state. However, we acknowledge that this remains an important limitation and we now indicate that in the manuscript. Additional experiments will be required to establish a robust and standardized approach, for example by combining NAD(P)H autofluorescence and/or mitochondrial potential measurements with Ca<sup>2+</sup> imaging.

      (4) There is a high degree of variability in response to AVP and V1b agonists across islets (activation, no effect, inhibition). Surprisingly, the authors do not fully explore the cause of this heterogeneity (whether it is due to receptor expression differences, metabolic state, experimental variability, or other conditions).

      This is a well-taken point and has indeed been one of the major bottlenecks in interpreting the results of this study. We agree that the variability in responses to AVP and V1b agonists may reflect several factors, including receptor expression, metabolic state, experimental conditions, and differences in the functional state of individual islets. However, our data also suggest that the beta cell population within an islet should be considered as a dynamic, non-linear system, in which even small differences in initial conditions or collective state can result in qualitatively different outcomes, including activation, no apparent effect, or inhibition. In this framework, the response to AVP is not determined by receptor expression alone, but by the current physiological context of the islet network. This is also why we believe that pharmacological tests are most informative when interpreted within a defined functional state rather than as isolated receptor-specific readouts. As indicated in the graphical abstract, apparently similar islets may occupy different dynamic states and therefore respond differently to the same Gq/PLC/IP3R stimulus. We have now expanded the discussion to make this interpretation more explicit and to acknowledge that receptor expression, metabolic variability, and experimental factors remain possible contributors that will require further targeted studies.

      The following text has been added to expand the discussion:

      “The heterogeneous responses observed across different islets, where some showed increased activity while others showed no detectable change or inhibition, could intuitively be attributed to variability in V1b receptor expression or signaling capacity among β cells. Such an explanation would be consistent with differences in receptor density, coupling efficiency to Gq proteins, or downstream signaling components such as PLC or IP<sub>3</sub> receptors. However, our data suggest that receptor-level variability alone is unlikely to fully explain the observed response spectrum, and that the current functional state of the islet collective must also be considered. The islet behaves as a non-linear dynamic system in which the same molecular perturbation can produce different functional outcomes depending on the current state of the β-cell collective. In such systems, cells or cell populations do not occupy a single deterministic activity state, but rather move within a landscape of possible states, with perturbations shifting the probability distribution of transitions between them. This concept is well established in dynamical systems approaches to biological cell-state transitions, where attractor landscapes, noise, and signaling inputs determine the probability of moving between alternative functional states rather than enforcing a single fixed output.

      In this framework, AVP and V1b receptor-selective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that βcell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone. Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (5) There is no validation of V1b receptor expression at the protein or mRNA level in α or β cells using in situ hybridization, immunohistochemistry, or spatial transcriptomics.

      We agree with the reviewer that spatial validation of V1b receptor expression is important for interpreting the cellular targets of AVP signaling in the islet. We have therefore added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      (6) AVP effects are described in terms of permissive or antagonistic effects on cAMP (especially in relation to epinephrine), but direct measurements of cAMP in α and β cells are not shown, weakening these conclusions. The authors should use Epac-based cAMP FRET sensors in α and β cells to monitor the interaction between AVP, forskolin, and epinephrine more conclusively.

      We agree that direct measurements of cAMP dynamics in alpha and beta cells would provide a more conclusive assessment of the interaction between AVP, forskolin, and epinephrine signaling. We attempted to address this experimentally; however, within the time domain of the Ca<sup>2+</sup> oscillations analyzed here, the temporal resolution and robustness of currently available cAMP readouts were not sufficient to resolve these interactions reliably. Even at slower time scales, cAMP sensor signals can be difficult to interpret quantitatively and may be overinterpreted if not tightly linked to the functional readout. We have therefore moderated the wording of the manuscript and now describe the proposed permissive or antagonistic interaction between AVP/V1b and cAMP-dependent signaling as an interpretation supported by the pharmacological Ca<sup>2+</sup> response patterns, rather than as a directly demonstrated cAMP mechanism. We now explicitly acknowledge in the limitations that most experiments were performed under cAMP-permissive conditions, which increases sensitivity for detecting AVP-dependent modulation but complicates the separation of direct beta cell effects from intra-islet interactions. Future studies using optimized cell-type-specific Epac-based sensors will be required to resolve this interaction.

      (7) Single-islet transcriptomics or proteomics (also to clarify variability) should be provided to analyze receptor expression variability across islets to correlate with response phenotypes (activation vs inhibition). Alternatively, the authors could perform calcium imaging with simultaneous insulin granule tracking or ATP levels to assess islet functional states.

      We agree that single-islet transcriptomics, proteomics, or simultaneous metabolic readouts could provide useful complementary information, particularly for describing molecular variability across islets. However, we do not think that differences in receptor expression or ATP levels alone are sufficient to explain the diversity of response phenotypes observed here. Our interpretation is that the beta cell population behaves as a collective dynamic system, in which the same input can lead to different outcomes depending on the current state of the network and its local physiological context. In such a system, AVP/V1b signaling does not necessarily impose a single deterministic response, but changes the probability distribution of accessible states, including activation, inhibition, or no detectable response. Theoretically and partially confirmed by the preliminary data, even the same islet exposed repeatedly under apparently identical conditions could be expected to display different responses if it occupies a different position within this dynamic state space at the time of stimulation. This concept is summarized in the graphical abstract and is central to our interpretation of the pharmacological data. We have therefore clarified in the Discussion that receptor expression, ATP levels, and other molecular parameters may modulate the response landscape, but are unlikely to fully define the observed functional phenotype without considering the collective dynamics of the islet.

      Added to Discussion section: “In this framework, AVP and V1b receptorselective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that β cell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone (63). Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (8) While the study implies AVP acts through V1b receptors on β cells, the signaling downstream (e.g., PLC activation, IP3R isoforms involved) is simply inferred but not directly shown.

      We agree that downstream signaling was not directly resolved at the level of PLC activation or specific IP3R isoforms. However, we did not infer Gq/PLC/IP3R involvement solely from AVP pharmacology, but used ACh as an independent Gq-coupled receptor reference stimulus in the same pancreatic slice preparation. With ACh concentration ramps, we could reproduce both activation and inactivation patterns observed with AVP/V1b stimulation, supporting the interpretation that these responses arise from modulation of the Gq-dependent Ca<sup>2+</sup> signaling axis.

      In addition, in prelilminary expriments we could observe that inhibition of Gq activity with YM254890, as well as interference with IP3R-dependent signaling using Xestospongin C, diminished the response, although not completely. This incomplete suppression is important, because it suggests that beta cell Ca<sup>2+</sup> homeostasis and collective islet activity are not controlled by a single linear pathway, but by partially redundant and context-dependent mechanisms. We have therefore revised the manuscript to state more cautiously that our data support the involvement of Gq/PLC/IP3R-dependent signaling, while acknowledging that direct measurements of PLC activity and IP3R isoform-specific contributions remain outside the scope of the present study.

      (9) The interpretation that IP3R inactivation (mentioned in the title!) underlies the bell-shaped AVP effect is just hypothetical, without direct measurements. Assays in β (and/or α)-cell-specific V1b KO mice and IP3R KO mice must be provided to support these speculations.

      We agree that the involvement of IP3R-dependent signaling should be stated with appropriate caution. However, the concept of IP3R inactivation as a mechanism contributing to bell-shaped Gq-dependent Ca<sup>2+</sup> responses is not purely hypothetical, since IP3R inactivation has been directly demonstrated in previous studies and provides a parsimonious explanation for the shift from activation to suppression at higher AVP concentrations. In the present study, this interpretation is further supported by the glucose dependence of the AVP concentration-response relationship, where different stimulatory glucose conditions shift the apparent efficacy peak.

      We also agree that cell-specific V1b receptor and IP3R knockout experiments would be valuable future approaches. In this respect, we have obtained preliminary results from a small sample of IP3R triple-knockout mice, which cannot yet be fully included because they are part of an ongoing collaboration. In these experiments, supraphysiological AVP concentrations did not produce the IP3R-like beta-cell response pattern observed in controls, namely reduced halfwidth and increased frequency, whereas alpha cell stimulation was preserved similarly to WT slices.

      At the same time, we believe that definitive knockout experiments must be carefully designed, because the beta cell population behaves as a dynamic collective system in which the response to AVP depends on the current functional state of the islet, glucose context, and intercellular coupling. We therefore now present IP3R inactivation as a strongly supported mechanistic interpretation rather than as a directly proven mechanism in this study, and we explicitly acknowledge that cell-specific V1b and IP3R genetic models will be required to fully resolve this pathway.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Drs. Kercmar, Murko, and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper falls short of supporting this claim.

      Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations of this effect.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      We thank the reviewer for their detailed input and support to increase the impact of our work and our understanding of important cellular processes overall. We have considered their suggestions carefully to further expand the strenghts of our approach and analysis.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Thank you!

      Weaknesses:

      (1) The introduction is long and summarizes a substantive body of literature on AVP actions on insulin secretion in vivo. There are a number of possible explanations for these observations that do not directly target islet cells. If the goal is to resolve the mechanistic basis of AVP action on alpha and beta cells, the more limited number of papers that describe direct islet effects is more helpful. There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is a) not expressed by beta cells and b) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript).

      We thank the reviewer for this important comment and agree that the literature on AVP actions in vivo is complex, with several possible sites of action outside the islet. We have therefore revised the Introduction to make the rationale more focused and to better separate systemic effects of AVP from studies addressing direct actions on pancreatic islet cells. At the same time, we chose not to restrict the Introduction only to the alpha cell V1bR literature, because one of the aims of the manuscript is precisely to address why AVP effects on insulin secretion have remained difficult to interpret across experimental contexts.

      Our results fully confirm a central aspect of the study cited by the reviewer, namely that V1bR activation robustly stimulates alpha cell Ca<sup>2+</sup> activity under non-stimulatory glucose conditions, and that 10 nM AVP does not produce a uniform activation of beta cell Ca<sup>2+</sup> activity. In fact, in a substantial fraction of beta cell populations, 10 nM AVP failed to activate oscillations, consistent with the view that alpha cells are the more sensitive and more direct cellular target of AVP/V1bR signaling. However, we do not think that the available transcriptomic evidence is sufficient to categorically exclude V1bR expression or functional relevance in beta cells. Re-analysis of the published dataset, together with more recent datasets and our newly added RNAscope data, supports a higher relative expression of V1bR transcripts in alpha than in beta cells, but does not justify treating beta (or non-alpha) cell expression as absent.

      We have therefore revised the manuscript to avoid overstating beta cell V1bR expression as a major isolated claim. Instead, we now present the data as evidence that AVP/V1bR signaling acts most prominently through alpha cells, while beta cell responses emerge in a concentration-, glucose-, and statedependent manner within the intact islet. This interpretation is consistent with the reviewer’s concern that 10 nM AVP preferentially activates alpha cells, but it also accommodates our observation that beta cell collective activity can be modulated under defined pharmacological and metabolic conditions. We believe that this is an important distinction, because the absence of a uniform beta cell Ca<sup>2+</sup> activation at one AVP concentration does not exclude beta cell modulation by AVP/V1bR signaling within the intact islet network. The Introduction and Discussion have been revised accordingly to clarify that our study does not simply challenge the alpha cell V1bR model, but expands it by examining how AVP-dependent alpha cell activation, possibly lower beta-cell receptor expression, and collective beta cell dynamics interact in the native pancreatic slice preparation.

      (2) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/dataghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet, but also the lack of expression of V1bra, V1brb, and Oxtr in beta cells. Instead of the detailed list of expression of these 4 receptors elsewhere in the body, it would be more directly relevant to set up their pancreatic slice experiments to summarize the known expression in pancreatic islets that is publicly available. It would also have helped ground the efforts that involved the generation of the V1aR agonist and V2R antagonist, which confirm these known AVP/OXT receptor expression patterns.

      We thank the reviewer for pointing us more directly to the publicly available islet expression datasets. We agree that the expression of AVP/OXT receptors in purified alpha, beta, and delta cells provides an important reference frame for interpreting our pharmacological data, and we have revised the manuscript to summarize these islet-specific datasets more directly rather than emphasizing receptor expression in other organs. These data support the absence or very low expression of V2 receptors in islet endocrine cells and confirm that V1b receptor expression is substantially enriched in alpha cells compared with beta cells.

      At the same time, as outlined in our response above, we do not think that the currently available transcriptomic datasets are sufficient to categorically exclude low-level V1bR transcript expression or functional relevance in beta cells within the intact islet. For this reason, we added independent RNAscope validation to assess V1bR transcripts in the pancreatic slice preparation. These data confirm stronger V1bR expression in glucagon-positive alpha cells, while also showing a broader expression pattern within the islet and pancreas.

      We have also revised the rationale for the pharmacological experiments using V1aR- and V2R-directed tools. We now present these experiments not as evidence for unexpected receptor expression, but as functional controls that are consistent with the known AVP/OXT receptor expression patterns in pancreatic islets. This better aligns the manuscript with the existing transcriptomic literature while preserving the main physiological question of the study: how AVP/V1bR-dependent signaling reshapes alpha-cell activity and beta-cell collective dynamics in intact pancreatic tissue.

      (3) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated a) indirectly, downstream of alpha cell V1br or b) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest that they are not mediated by the same receptor.

      We agree with the reviewer that the absence or very low abundance of V1bR transcripts in beta cells in published transcriptomic datasets would not invalidate the observation that AVP modulates beta-cell Ca<sup>2+</sup> activity. It does, however, raise the important question of whether this modulation is mediated indirectly through alpha-cell V1bR activation, through V1bR expression in beta cells that is difficult to resolve transcriptomically, or through another mechanism. To address this more directly, we have now added RNAscope data, which confirm relatively stronger V1bR transcript enrichment in glucagon-positive alpha cells, but also show a broader V1bR transcript signal within the islet and pancreas. Thus, while our data support alpha cells as the dominant V1bR-positive endocrine population, they do not support a strict absence of V1bR-associated signaling capacity in the beta cell compartment.

      We also agree that different peak efficacies in alpha and beta cells could be interpreted as evidence for distinct receptors or indirect mechanisms. However, we favor a different interpretation: the apparent efficacy of AVP depends strongly on the physiological state in which the cells are tested. This is particularly evident in beta cells, where the AVP efficacy peak shifts with glucose concentration, suggesting that the beta-cell response is shaped by the metabolic and Ca<sup>2+</sup>-handling context rather than by receptor occupancy alone. In this framework, the same V1bR/Gq-dependent input can generate different downstream Ca<sup>2+</sup> outcomes in alpha and beta cells because the two cell types operate in different dynamic regimes.

      We have therefore revised the manuscript to acknowledge this dilemma more explicitly. We now state that beta cell effects of AVP could include indirect alpha cell-dependent components, but given the magnitude and statedependence of the beta cell Ca<sup>2+</sup> response it is unlikely to be driven by alpha cell activation. Instead, our preferred interpretation is that AVP/V1bR signaling acts within the intact islet as a context-dependent perturbation of the collective beta cell Ca<sup>2+</sup> system, with IP3R-dependent mechanisms being modulated by glucose-dependent changes in beta cell excitability and intracellular Ca<sup>2+</sup> handling.

      (4) The rationale for the use of forskolin across almost all traces is unclear. It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct all studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMPgenerating stimuli. The permissive actions by AVP that are cited are on hormone secretion, which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium response is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e., the cAMP stimulated by it in beta cells. However, the design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon, particularly upon co-stimulation with AVP, which permits glucagon release by activating a calcium response in alpha cells. This glucagon could then activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments in Figures 1, 2, and 4 should have been completed in the absence of cAMP first.

      We agree with the reviewer that the use of forskolin needs to be explained more clearly, and we have revised the manuscript accordingly. Our rationale was based on the established permissive role of cAMP in AVP-dependent endocrine responses, but we acknowledge that this does not necessarily imply that the upstream V1bR/Gq-mediated Ca<sup>2+</sup> response itself is cAMP-dependent. The reviewer is also correct that forskolin elevates cAMP broadly and therefore may affect both alpha and beta cells, including the possibility that AVP-enhanced alpha cell activation and glucagon release secondarily influence beta cell activity.

      In fact, our initial experiments were performed without forskolin and revealed an important difficulty: stimulatory glucose alone can increase cAMP levels to a variable extent, as also supported by our previous work on epinephrine signaling, thereby shifting the apparent peak efficacy of AVP stimulation. Thus, forskolin was originally used to reduce this variability and create a more defined cAMP-permissive background in which alpha and beta cell responses could be compared in the same slice. However, we agree that this design works against isolation of beta cell-autonomous AVP effects.

      Within a scope of another study we have done an independent series of more focused experiments using GLP-1 receptor stimulation, which preferentially increases cAMP signaling in beta cells compared with the broad cAMP elevation produced by forskolin. We have clarified that the modulation of the AVP-dependent pathway by GLP-1 and related ligands at largely supports beta cell-autonomous AVP effects. It is part of ongoing work and will be reported independently, because a full mechanistic dissection of cAMP–AVP interactions goes far beyond the scope of the present study.

      (5) It is unexpected that epinephrine in Figure 2 does not activate the alpha cell calcium? A recent paper from the same group (Sluga et al) shows robust calcium activation in alpha cells in a similar prep by 1 nM epinephrine, which is similar to the dose used here.

      We thank the reviewer for pointing this out, but we would like to clarify that epinephrine did significantly activate alpha-cell Ca<sup>2+</sup> activity in our experiments, as shown in Fig. 3F. This result is consistent with our previous study by Sluga et al., where low nanomolar epinephrine robustly activated alpha cell Ca<sup>2+</sup> signals in the pancreatic slice preparation. The main point of the present comparison was therefore not that epinephrine is inactive in alpha cells, but that AVP produces a substantially stronger and reproducible alpha cell Ca<sup>2+</sup> response under comparable experimental conditions. This is also consistent with the data of van der Meulen et al., supporting the view that AVP/V1bR signaling is a particularly potent activator of alpha cell activity. We have revised the text to make this comparison clearer and to avoid the impression that epinephrine failed to activate alpha cells in our preparation.

      (6) Figure 8 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this comparison with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 4A)? I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual, and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose, irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 5B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 8).

      We agree that the interpretation of low-pM AVP effects requires caution, particularly when compared with reported V1bR binding affinities. The apparent discrepancy between Fig. 4A and Fig. 8 most likely reflects differences in experimental design, stimulation context, and readout sensitivity. In Fig. 4A, we assessed acute AVP effects under conditions in which beta cell Ca<sup>2+</sup> responses are relatively threshold-dependent and where low AVP concentrations produced little or no activation. In contrast, Fig. 8 analyzes prolonged beta cell population dynamics during sustained stimulation with 8 mM glucose, a physiological stimulatory context in which even weak modulatory inputs may become detectable at the level of collective Ca<sup>2+</sup> activity.

      Importantly, the strongest and statistically significant effect was observed at 100 pM AVP, while lower pM concentrations showed only a trend. We therefore do not interpret the low-pM range as evidence for robust direct pharmacological activation of beta cell V1bR. Rather, these data suggest that AVP may exert permissive or modulatory effects within an already active beta cell network, where glucose-dependent excitability, receptor-effector coupling, and Ca<sup>2+</sup> amplification mechanisms can enhance the apparent efficacy of weak inputs. This interpretation is consistent with the known permissive role of AVP in endocrine responses, where AVP may not act as a primary activator alone but can increase the efficacy of other physiological stimuli.

      We have also clarified that sustained 8 mM glucose alone does not account for these effects, since Suppl. Fig. 1 shows no comparable time-dependent progression of Ca<sup>2+</sup> behavior under sustained glucose stimulation alone. Thus, we now present the low-concentration AVP effects as subtle, contextdependent modulation within the physiological stimulatory range, rather than as evidence for direct beta cell activation at concentrations below the expected receptor affinity range.

      Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      We would like to thank the reviewer for recognizing the potential of our emerging area of research.

      Weaknesses:

      The conclusions are only modestly supported by the data and lack experimental depth and rigor. The rationale for only conducting studies at high cAMP conditions is not entirely clear and limits the conclusions that can be made. The use of Ca2+ is helpful, but it is a surrogate for hormone secretion. Additional measurements of hormone secretion are needed to enhance the robustness of these conclusions. Consideration of paracrine effects between alpha- and beta-cells is only superficially made and is likely essential in the context of the experimental design. For instance, there is clear literature that alpha-cells secrete several factors that work in paracrine interactions on beta-cells and autocrine actions back on alpha-cells. Conducting these studies in a high cAMP context only completely overlooks these interactions, skewing the interpretations made by the investigators. Finally, the clarity of the experiments and results could be significantly enhanced.

      We thank the reviewer for this balanced assessment and for emphasizing several issues that are central to the interpretation of our study. We agree and now explicitly state in the Limitations, that Ca<sup>2+</sup> oscillations are a surrogate readout for hormone secretion and that currently used stimulation protocols are not optimized to directly quantify the relationship between Ca<sup>2+</sup> dynamics and secretory output. To address this limitation, we have now expanded the functional part of the study by adding complete insulin and glucagon secretion measurements during AVP concentration ramps. These new data provide a stronger functional framework for interpreting the Ca<sup>2+</sup> imaging results, while also clarifying that Ca<sup>2+</sup> activity and secretion cannot be assumed to correlate linearly under all stimulation protocols.

      We have also revised the rationale for the high-cAMP experimental condition. The original aim was to reduce variability arising from glucose-dependent endogenous cAMP signaling and to study alpha and beta cell responses in a common permissive background. However, we agree that broad forskolin stimulation complicates the interpretation of cell-autonomous versus paracrine mechanisms. Independent experiments within a scope of another study demonstrate that using GLP-1 co-stimulation, which provides a more beta-cell-oriented cAMP-permissive condition and supports the interpretation that AVP can modulate beta-cell collective activity in a manner that is not solely secondary to alpha-cell activation.

      We fully agree that paracrine interactions within the islet are physiologically important and must be considered, particularly in intact pancreatic slices. Nevertheless, the rapid onset of the AVP effects observed in beta cell Ca<sup>2+</sup> activity argues against a mechanism mediated predominantly by slower indirect paracrine loops. High AVP concentrations, as shown in Fig. 5, significantly shorten the intervals between Ca<sup>2+</sup> events in both alpha and beta cells, but that the activity of the two cell populations remains largely noncoordinated. This temporal dissociation does not exclude paracrine modulation altogether, but it argues against a simple alpha-cell-driven explanation for the beta-cell response.

      We have revised the manuscript to state these points more clearly and to moderate conclusions where the data support modulation rather than definitive cell-autonomous receptor action. We believe that the added secretion experiments and alpha/beta event-timing analysis substantially strengthen the physiological interpretation of the study, and we thank the reviewer for raising these issues.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The paragraph discussing the benefits of slice physiology over islets is not reflective of how most - if not all of your colleagues who do islet experiments conduct these. Many labs have reported for years high-quality GSIS experiments, synchronous calcium responses, and a plethora of studies detailing the mechanism of hormone and neurotransmitter actions using islet models, and have done so well. Slice physiology is a unique and helpful model that can have advantages over other models. This particular reviewer uses both models in their lab and each has benefits and - inevitably - drawbacks. Many of the possible drawbacks cited for islet studies apply equally to slices, including the possibility of altered gene expression, lack of innervation, and circulation. Added drawbacks are the exposure to higher levels of pancreatic enzymes from the slice, which require co-culture with enzyme inhibitors.

      We agree with the reviewer and have revised these limitations accordingly. Our intention was not to imply that isolated islet preparations are generally inferior, since they have provided a highly productive and rigorous experimental platform for GSIS, synchronized Ca<sup>2+</sup> dynamics, and mechanistic studies of hormonal and neurotransmitter regulation with standardized protocols with all their positive and negative sides. We modified the presentation of pancreatic slices as a complementary model with specific advantages, particularly preservation of local tissue architecture, while also acknowledging their limitations. The revised text therefore avoids a comparative hierarchy between slices and isolated islets and instead emphasizes that both models have distinct strengths and drawbacks depending on the experimental question.

      (2) If you want to demonstrate direct actions on beta cells, deconstructing the islet would be a better way to go. Less complicated, not more. Dissociated beta cells, instead of slices, were used just to prove or disprove the hypothesis of direct beta cell effects of AVP.

      We agree that dissociated beta cells can be a useful reductionist model to test whether AVP is capable of acting directly on individual beta cells. However, this approach would also remove the collective beta-cell activity that is central to the physiological question addressed in the present study. Since our data indicate that AVP effects emerge within the intact islet as rapid and extensive changes in coordinated Ca<sup>2+</sup> dynamics, dissociation would not necessarily provide a more informative model for understanding these responses. The fast onset and magnitude of the beta cell response argue against a predominantly indirect non-autonomous mechanism, and this interpretation is further supported by the GLP-1 co-stimulation experiments, which are more consistent with beta cell-autonomous modulation. We have therefore clarified in the revised manuscript that dissociated-cell experiments would be valuable for a narrowly defined receptor-cell autonomy question, but would not resolve the collective islet dynamics that are the focus of this work.

      (3) If you want to sustain the claim of beta cell expression of V1br, you would have to demonstrate this far more directly by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by highly selective, well-validated pharmacology. This should include a demonstration of Gaqdependence in isolated beta cells.

      We have added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      Minor

      (1) O'Carroll et al. should be cited in the context of islet permissive actions of AVP/cAMP. PMID: 18434353, although that paper offers no evidence that the AVP-dependent potentiation of insulin release is mediated directly by beta cells. It does confirm dependence on PKC.

      We agree and have added O’Carroll et al. in the revised manuscript in the context of AVP/cAMP-dependent permissive actions on islet hormone secretion. We therefore use it as support for the broader concept of AVP-dependent amplification of secretion in a permissive signaling context, rather than as direct evidence for beta-cell-autonomous V1bR signaling.

      (2) Figure 2 E-H, glucose concentration mislabeled.

      We thank the reviewer for pointing this out. The glucose concentration label has been clarified: panels E–H show pooled data from separate experiments performed at 8 mM glucose, whereas panel D shows a representative experiment performed at 9 mM glucose.

      (3) The insulin secretion in 4E is difficult to interpret without a low-glucose control. If this is hard to do in a slice preparation, a separate static islet secretion experiment would help here. The possibility that the inhibition of insulin secretion traces back to the activation of delta cells by AVP could be considered - I struggle to come up with a plausible mechanistic explanation why AVP (which activates calcium in alpha and in beta cells in the presence of 8 mM G plus forskolin according to your data) would inhibit insulin secretion.

      We agree that the original insulin secretion experiment was difficult to interpret without a clearer low-glucose reference condition. To address this, we have added new insulin release experiments in which glucose was lowered to a non-stimulatory range between AVP concentrations, followed by sequential stimulation with 8 mM glucose and 500 nM forskolin in the same slices. These new data provide a broader dynamic range for assessing insulin secretion and allow a more direct comparison between AVP-dependent Ca<sup>2+</sup> modulation and secretory output.

      We also agree that AVP-dependent inhibition of insulin secretion requires careful interpretation. One possible explanation is not simply activation of delta cells, but a failure of the beta cell collective to maintain coordinated activity at very high AVP concentrations. In the Ca<sup>2+</sup> imaging data, high AVP concentrations increase activity in many beta cells, but numerous cells within the islet fail to keep pace with the collective oscillatory rhythm, leading to fragmented and less synchronized population activity. Thus, despite increased frequency of Ca<sup>2+</sup> oscillations in the islet, the integrated beta cell output may become less efficient for insulin secretion. We have added this interpretation to the revised manuscript and now discuss delta cell activation as a possible contributing mechanism, but not as the primary explanation supported by our current data.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function while devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. Some experimental presentation and design concerns should be addressed to support the stated conclusions.

      We thank the reviewer for their supportive comments. We agree that portions of the manuscript were not fully clear and that some aspects of the experimental presentation and design required clarification.

      To address this, we have revised the manuscript to make the experimental logic more transparent. In particular, we now explain more clearly the rationale for the screening strategy, including the use of histidinol as a canonical GCN2 activator, latrunculin A as a modulator of PPP1R15A-mediated eIF2α dephosphorylation, and tunicamycin as a PERK-dependent ER-stress control. We also clarify why a submaximal concentration of histidinol was used: this was intended to reveal compounds that enhance ISR signalling when GCN2 is partially activated.

      We have clarified the use of the two ATF4 reporter systems. The ATF4–NanoLuc reporter was used for sensitive primary screening, whereas the ATF4–luc2 reporter was used as a more stringent orthogonal assay to prioritise robust ISR activators. We now state explicitly why some initial hits were not retained after testing in the second reporter line, and why the NanoLuc system was subsequently used again for mechanistic experiments.

      Finally, we have revised the presentation of the orthogonal validation steps to make clearer how they support the stated conclusions. These include assays designed to distinguish GCN2-dependent ISR activation from indirect activation through ER stress, additional analysis of GCN2 dependence, and clearer interpretation of biochemical and docking data.

      We hope that these revisions address the reviewer’s concern that the experimental design and data presentation needed to be made clearer in order to support the conclusions.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Zhu, Emanuelli, and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and, in particular, GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      The rationale for the screen is not exploited in the results (e.g., pathogenic GCN2 mutants), and lots of cell-based read-outs are not endogenous.

      We thank this reviewer for their positive assessment of the work. We address the three major concerns in turn below.

      Major points

      (1) Regarding the justification of the work. Since the authors justify the screen for GCN2 activators with loss-of-function mutants associated with diseases, it would be of interest to evaluate whether the best compounds identified in the study are indeed able to prompt activation of those mutants (or at least of the most prevalent). This approach could actually go in parallel with the docking experiments carried out in the last figure of the manuscript, where mutants could be modelized as well.

      To address this point, we tested whether the lead compounds could activate disease-associated GCN2 variants linked to pulmonary veno-occlusive disease. In contrast to GCN2iB, the new compounds did not activate these variants. We now state this explicitly in the manuscript, thereby clarifying that although the compounds identify a new mode of GCN2 activation, they do not rescue the pathogenic GCN2 variants tested here.

      Results

      “Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (2) The compounds are only tested using « artificial » proximal signaling outputs. It would be interesting to evaluate whether the best identified compounds are capable of prompting endogenous eIF2alpha phosphorylation in cellular models.

      We thank the reviewer for this suggestion. Detecting eIF2α phosphorylation following activation of GCN2 is technically challenging and typically produces weaker signals compared to activation of other ISR kinases, such as PERK (e.g. by thapsigargin). For this reason, many studies rely on downstream reporter assays to monitor GCN2 activity. To address the reviewer’s concern, we have now included an orthogonal readout of ISR activation by assessing global translation using a puromycin incorporation assay. Using this approach, we show that compound 20 significantly reduces translation, and importantly, this effect is attenuated in GCN2-deficient cells, supporting a GCN2-dependent mechanism.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      (3) Other GCN2 activators (other than GCN2iB, e.g., HC-7366) were recently identified. In this context, it would be of interest to carry out a small benchmarking study to evaluate how the compounds identified in the current study perform against the previously identified molecules.

      We thank the reviewer for this suggestion. In response, we obtained HC-7366 and assessed its activity alongside our compounds in the CHO ATF4::NanoLuc reporter assay. In this system, compound 20 demonstrated greater potency than HC-7366 (see reviewer figure below). However, we note that HC-7366 showed relatively limited activity in CHO cells in our hands, despite previously reported strong effects in other cellular systems and in vivo models. This context-dependent activity makes direct benchmarking difficult. Accordingly, we have included this comparison in the revised manuscript and discuss this limitation in the Discussion.

      Discussion

      “We also evaluated the reported GCN2 activator HC-7366 in our CHO ATF4::NanoLuc reporter system. In this context, HC-7366 showed limited activity relative to compound 20, despite its reported efficacy in other cellular systems and in vivo (data not shown). This highlights potential context dependence in small‑molecule activation of GCN2 and limits direct cross-study comparison.”

      Author response image 1.

      ISR activation by compound 20 and GC-7366 in CHO cells

      Normalised fold-change in ATF4 signal in CHO ATF4::NanoLuc reporter cells treated for 19 hours with Compound 20 or HC-7366. DMSO was used as vehicle control. (representative experiment, mean ± SEM, n=3 technical replicates).

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high-throughput screen for small-molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20), is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each binds in the ATP-binding pocket of GCN2, and that at least in vitro, C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      We agree that GCN1-independent activation suggests a potentially distinct mechanism of action. While we are currently unable to define the mechanistic basis underlying the GCN1-independence of compound 20, prior work provides some relevant context. Recent studies have shown that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions, notably in the context of the GCN2 E26A mutant [Carlson, 2023]. This observation raises the possibility that, under certain conditions, engagement of the kinase domain, potentially via the ATP-binding pocket, may bypass the requirement for GCN1. However, in our system, we did not observe GCN1-independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow or context-dependent window for such activity, or differences between wild-type and mutant GCN2. These findings suggest that GCN1-independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds, although further work will be required to define the underlying mechanism for compound 20. We have added the following to the main text:

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      Strengths:

      Of the 3 compounds identified by the authors, C20 is the most interesting, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T in Figure 4, and in contrast to other GCN2 activators) but also for its potency. In in-cellulo assays, compound 21 appears as more of an ISR enhancer than an activator per se, and although compound 18 and compound 21 lead to upregulation of the ISR targets (Figure 2), that degree of upregulation is probably not significantly different from that induced by those compounds in Gcn2-/- cells. For C20, the effect appears stronger (although it is unclear whether the authors performed statistical analysis comparing the two genotypes in Figure 2D). In Figure 3, only C20 activates the ISR robustly in both CHO and 293T. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation, and could be exploited therapeutically.

      Prompted by this suggestion, we assessed C20 in two additional commonly used cell lines: human colon carcinoma HCT116 cells and African green monkey COS7 cells. C20 showed no activity in these models. In contrast, primary mesothelioma cells (Mesobank T12) exhibited robust PPP1R15A induction in response to the compound. The following text has been added to the manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses… [data not shown].”

      Weaknesses:

      There are some limitations to the existing work. As the authors acknowledge, they do not use any of the compounds in animals; their in vivo efficacy, toxicity, and pharmacokinetics are unknown. But even in the context of the in cellulo experiments, it is puzzling that none of the three compounds, including C20, has any effects in HeLa cells when Neratinib does. It's beyond the scope of this paper to address definitively why that is, but it would at least be reassuring to know that C20 activates the ISR in a wider range of cells, including ideally some primary, non-immortalized cells. In addition, the ISR is a complex, feedback-regulated response whose output varies depending on the time point examined. The in cellulo analysis in this paper is limited to reporter assays at 18 hours and qRT-PCR assays at 4 and 8 hours. A more extensive examination of the behaviour of the relevant ISR mRNAs and proteins (eIF2, ATF4, CHOP, cell viability, etc.) for C20 across a more extensive time course would give the reader a clearer sense of how this molecule affects ISR output.

      We thank the reviewer for this insightful suggestion. To address the need for a more comprehensive assessment of ISR signalling, we have extended our analysis across a broader time course and incorporated additional functional readouts. Using the ATF4-NanoLuc reporter, compounds 20 and 21 exhibit peak ISR activation at approximately 6-8 h in wild-type cells. In parallel, we assessed global mRNA translation using puromycin incorporation and found that compound 20, but not compound 21, induces a progressive reduction in translation over this period, which is dependent on GCN2. While we agree that direct measurement of upstream ISR markers such as eIF2α phosphorylation can be informative, detection of GCN2-mediated eIF2α phosphorylation is technically challenging and often less robust than activation of other ISR kinases (e.g. PERK). For this reason, we have prioritised orthogonal downstream functional readouts, including reporter activity and translational output, to capture ISR pathway engagement. These additional data provide a clearer picture of the kinetics and functional consequences of compound-induced ISR activation and have been incorporated into the revised manuscript.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      I also find it a bit strange that the authors describe C20 as "demonstrat(ing) weak inhibition of ... PKR" - the measured IC50 is ~4 μM, which is right around its EC50 for GCN2 activation. This raises the confounding possibility that C20 would simultaneously activate GCN2 while inhibiting PKR. While perhaps inhibition of PKR is not relevant under the conditions when GCN2 would be activated either experimentally or therapeutically, examining in cells the effects of C20 on GCN2 and PKR across a dose range would shed light on whether this cross-reactivity is likely to be of concern.

      We thank the reviewer for highlighting compound 20 as the most interesting lead compound and for recognising its apparent ability to activate GCN2 independently of GCN1. The reviewer identified several limitations relating to cell-type specificity, the temporal behaviour of ISR activation, and possible PKR cross-reactivity.

      In response to the concern about cell-type specificity, we tested compound 20 in additional cellular models. Compound 20 did not induce detectable ISR signalling in HCT116 or COS-7 cells under the conditions tested, consistent with the reviewer’s observation that activity is not universal across cell types. However, we observed induction of PPP1R15A in primary Mesobank T12 mesothelioma cells. We have therefore revised the manuscript to present compound 20 activity as cell-context dependent rather than broadly generalisable across all cell types.

      To address the reviewer’s concern that the ISR output was examined only at limited time points, we extended the time-course analysis for compounds 20 and 21. Using live-cell ATF4::NanoLuc reporter measurements, both compounds showed maximal reporter activation at approximately 6-8 hours. We also measured translational output by puromycin incorporation and found that compound 20, but not compound 21, caused a progressive GCN2-dependent reduction in translation. These data provide a clearer view of the kinetics and functional consequences of compound 20-mediated ISR activation.

      The reviewer also noted that compound 20 inhibits PKR in vitro at concentrations close to those required for GCN2 activation in cells. We agree that this is an important potential liability. Because PKR signalling was not robustly or reproducibly inducible in our CHO-based reporter system, we were unable to perform a reliable cellular dose-response analysis of PKR engagement in the present study. We have therefore revised the Discussion to acknowledge kinase cross-reactivity, including possible PKR inhibition, as an important limitation and an issue for future development of this chemical series.

      Finally, we have moderated our mechanistic interpretation of compound 20. Although the data support direct engagement of GCN2 and suggest a mechanism distinct from canonical GCN1-dependent activation, we now discuss GCN1-independent activation more cautiously and in the context of prior reports that GCN2iB can display GCN1-independent activity under specific experimental conditions.

      Discussion

      “While the functional relevance of PKR inhibition in our cellular systems is uncertain, these observations highlight the potential for kinase cross-reactivity, which will be important to address in future studies.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):<br /> (1) The description of the chemical screen for Gcn2 activators is not sufficiently clear and detailed. a) Briefly and early on provide the rationales (modes of action) for using histindinol and latruculin A. Explain further the rationale in Figure 2, outlining the purpose for the combined compound + submaximal dose of histindinol.

      The text has been amended.

      Results

      “Histidinol activates GCN2 by inhibiting histidyl‑tRNA synthetase, leading to the accumulation of uncharged tRNAHis. This uncharged tRNA binds to GCN2, relieving its autoinhibition and activating the kinase (35). Latrunculin A sequesters G‑actin, thereby inhibiting PPP1R15A activity (36, 37). Tunicamycin inhibits protein glycosylation in the endoplasmic reticulum (ER), resulting in activation of PERK (38).”

      “This submaximal concentration was used to allow detection of compounds that enhance ISR signalling when GCN2 is partially activated.”

      b) What is unique about the second CHO: ATF4-luc2 reporter line? Why do only 89 out of the original 130 compounds induce the ISR in this line versus the original CHO: ATF4-Nanoluc cell line? This is confusing for the reader about how compounds were triaged for characterization.

      The ATF4‑Nanoluc and ATF4‑luc2 reporter lines differ only in the luciferase used, but this has important practical consequences. The Nanoluc reporter is substantially more sensitive, so it was used for the primary screen to detect even weak ISR activation. The luc2 reporter has lower sensitivity and a narrower dynamic range, making it a more stringent orthogonal assay. As a result, not all hits from the Nanoluc screen (130 compounds) reproduced in the luc2 line; the 89 compounds retained are those that robustly activate the ISR under these more stringent conditions. This step was therefore used to prioritise stronger, more reproducible activators for downstream characterisation.

      Results

      “While primary screening was performed in ATF4‑Nanoluc lines for maximal sensitivity, hits were subsequently re-tested in a second CHO ATF4::luc2 reporter line as a more stringent orthogonal assay to prioritise robust ISR activators. Of the 130 hits identified in the sensitive Nanoluc screen and passing early toxicity assessment, 89 were confirmed in the luc2 assay, consistent with enrichment for higher-amplitude ISR activators under more stringent detection conditions.”

      c) The study uses a second CHO reporter line in the flow scheme (CHO:ATF4-luc2) and then switches back to an ATF4-Nanoluc line to establish GCN2 dependence. What is the rationale for switching back to the original reporter line?

      The luc2 reporter line was used as a more stringent, orthogonal validation step to prioritise robust ISR activators. For subsequent mechanistic studies, including assessment of GCN2 dependence, we returned to the ATF4‑Nanoluc line because its higher sensitivity and simpler single‑reagent assay format are better suited to multi‑point measurements and comparative analyses. In effect, the luc2 reporter was used for triage, whereas the Nanoluc system was retained for mechanistic characterisation and downstream screening.

      Results

      “In subsequent mechanistic studies, the ATF4‑Nanoluc reporter was again used to take advantage of its higher sensitivity and simpler assay format for multi‑condition comparisons.”

      d) The rationale for the first orthogonal screen described in the results section to identify inducers of ER stress is not clearly explained. The compounds were already determined to be dependent on GCN2 prior to this test, and one would have thought that this criterion would have covered ER stress and alternative eIF2 kinase activators.

      We agree with the reviewer that, in principle, establishing GCN2 dependence should reduce the likelihood of capturing compounds acting through alternative eIF2α kinases. However, we performed this orthogonal ER stress screen to address two practical considerations. First, high‑throughput screening is inherently prone to false positives, as it is typically conducted at a single concentration and time point, and compound libraries may contain degraded or chemically inconsistent material. We therefore used a lower‑throughput, more controlled ER stress assay with freshly sourced compounds and additional readouts (e.g. CHOP and XBP1) to improve confidence in the hits. Second, despite prior evidence of GCN2 dependence, ER stress signalling via PERK converges on the same downstream endpoints: eIF2α phosphorylation and ATF4 induction. We therefore wished to explicitly exclude compounds that activate the ISR indirectly via ER stress. In practice, this proved important, as the orthogonal assay did identify compounds that induced ER stress, which we subsequently excluded from the lead set.

      Results

      “Although hits were prioritised for GCN2 dependence, we performed an additional orthogonal screen to exclude compounds that activate the ISR indirectly via ER stress, which converges on the same downstream outputs.”

      (2) A major point of the manuscript is that there is GCN1 independence for the small molecule activation of GCN2, and this has not yet been reported. One report for this GCN1 independence is reference 9 [Carlson … Wek 2023] (Figure 5). In this report, low doses of GCN2iB that can activate GCN2 (although by the present manuscript at much lower levels than the identified new compounds) induce ATF4 expression in cells expressing an E26A mutant of GCN2 that is suggested to negate GCN1 binding and enhancement of GCN2 activity. Halofuginone induction of ATF4 expression was thwarted by the GCN2 E26A mutant.

      We thank the reviewer for highlighting this important point. We agree that Carlson et al. (2023) suggest that, under certain conditions, GCN2iB can activate GCN2 independently of GCN1 using the E26A mutant. In our experiments, however, we did not observe GCN1‑independent activation with GCN2iB under the conditions tested, i.e. similar low doses. One possible explanation is that the GCN1‑independent activity reported by Carlson et al. occurs only within a narrow concentration range; their observations were made at very low compound concentrations, whereas higher concentrations may engage additional regulatory mechanisms. In our study, we used concentrations optimised for robust ISR activation, which may mask such effects. We also note that the E26A mutation (E18A in yeast), originally identified by two‑hybrid analysis, disrupts the GCN2-GCN1 interaction but may not completely eliminate all modes of functional coupling under all conditions. Taken together, these observations raise the possibility that GCN1‑independent activation represents a context‑dependent mechanism that may be unmasked only under specific experimental conditions or by particular classes of compounds.

      We have revised the Discussion to acknowledge this prior report explicitly and to clarify how our findings relate to it.

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      (3) The authors state that the ISR was exaggerated in Ppp1r15a KO cells. It would be helpful to include statistical analyses to support this statement.

      Thank you for enabling us to be more precise. New text added:

      Results

      “Activation of the ISR by tunicamycin was exaggerated in the Ppp1r15a<sup>-/-</sup> cells owing to their defective dephosphorylation of eIF2a (wild type vs Ppp1r15a<sup>-/-</sup>, p<0.05).”

      (4) The results state that Chop and Ppp1r15a mRNAs were measured following 4 hours of treatment with compound 18, 20, or 21, but Figure 2D shows treatment from 0 to 8 hours? It appears that the compound still induces these mRNAs in GCN2 KO cells, possibly with delayed kinetics. A lengthened time course study would help determine if this is indeed the case.

      We thank the reviewer for this careful observation. To address the reviewer’s point regarding delayed or GCN2‑independent signalling, we have extended our analysis using compounds 20 and 21, which are more tractable experimentally (solubility). Using the ATF4‑Nanoluc reporter, both compounds show peak ISR activation at ~6–8 h in wild-type cells over an extended time course. In parallel, functional readouts of mRNA translation (puromycin incorporation) demonstrate that compound 20, but not 21, induces a progressive, GCN2‑dependent reduction in translation over this period. These clarify the temporal aspects of signalling by these two compounds.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      Legend

      “Supplementary Figure S1. Kinetics of responses to compounds 20 and 21

      (A-B) Wild-type CHO cells stably expressing the ATF4::nanoLuc-PEST reporter were treated with Nano-Glo and either (A) 13mM compound 20 or (B) 13mM compound 21. Median bioluminescence (fold change normalised to DMSO control) ± 95% confidence. Representative experiment (n=4 technical repeats). (C-F) Representative immunoblot of lysates from wild-type or Eif2ak4<sup>-/-</sup> CHO cells treated with 10μM compound 20 or 7.5μM 21 for the indicated times. Immediately before harvesting, cells were treated with 10μg/mL puromycin to label newly synthesised polypeptides. “-“ indicates cells not incubated with puromycin. “U” cells were treated with puromycin but without test compound. “CHX” represents the cycloheximide control (100μg/mL). Molecular size in kDa. (E-F) Quantification of puromycinylated proteins normalised to GAPDH. Mean ± SEM. CHO WT (black) and Eif2ak4<sup>-/-</sup> cells (turquoise. N = 4 independent experiments. Two-way ANOVA with Šídák's multiple comparisons test; ***: p ≤ 0.001.”

      (5) In the section describing the differences between cell lines in the ability of compounds to induce the ISR, this is difficult for the reader to interpret, as no controls are included. How does histidinol (or other canonical inducers of the ISR) behave in the three reporter assays (CHO, 293T, and HeLa)?

      As requested, we now provide ATF4::Nanoluc reporter activation (3mM, 20 hours because of this drug’s slow kinetics)

      Results

      “To benchmark ISR activation in these models, each cell type was treated with 3mM histidinol (Figure S2). Reporter activation was most robust in CHO cells, followed by 293T cells, then HeLa cells.”

      Discussion

      “Moreover, histidinol-induced ISR activation showed a clear hierarchy across cell lines, with CHO cells being the most responsive and HeLa cells the least.”

      Legend

      “Supplementary Figure S2. Cell-type differences in response to histidinol

      Fold-change of ATF4::NanoLuc reporter signal in HEK293T, HeLa and CHO cells transiently transfected with reporter and treated for 20 hours with 3mM histidinol. Fold-change calculated relative to vehicle control. Mean ± SEM).”

      We thank the reviewer for this important question. However, we respectfully disagree that a direct correspondence between the concentrations required for target engagement in the BRET assay and for ISR activation in functional assays should necessarily be expected. BRET (including NanoBRET) is a target engagement assay that measures compound binding to the protein in intact cells, typically by competition with a labelled tracer, and thus reports on apparent intracellular affinity and occupancy rather than downstream biological effect (Robers 2019, PMID 30519940). By contrast, ISR activation is a functional readout that reflects amplification through signalling networks, and can be influenced by multiple additional variables including pathway non-linearity, feedback, and kinase regulation. Consequently, it is well established that potencies derived from target engagement assays do not always align with those measured in functional assays. For example, intracellular kinase profiling studies using NanoBRET have demonstrated systematic potency offsets between binding/engagement measurements and downstream cellular activity, arising from factors such as intracellular ATP competition and pathway context (PMID Capener 2026, PMID 41495225). More generally, target engagement assays provide a quantitative measure of binding, whereas functional assays measure biological outcome, and these readouts need not coincide because they capture distinct aspects of a compound’s mechanism of action. Accordingly, we interpret our BRET data as evidence of direct interaction with GCN2 in cells, rather than as a predictor of the concentration required to activate the ISR. The observation that higher concentrations are required in the BRET assay is therefore not unexpected and does not argue against a requirement for kinase-domain engagement in ISR activation. Instead, it reflects the different mechanistic endpoints captured by the two assay formats.

      We will clarify this point explicitly in the revised manuscript.

      Results

      “The concentrations required to detect target engagement in NanoBRET assays did not directly mirror those required for ISR activation, reflecting the distinction between ligand binding and downstream pathway output.”

      (7) In Figure 5E, the authors suggest that compounds 18 and 20 are non-competitive inhibitors of GCN2 since the Vmax increases with increasing ATP concentration. What is the Km for ATP in the absence or presence of compound 18 or 20? It would be helpful to include progress curves as supplementary data to support the Vmax plots in Fig. 5E. Consider providing more specific units (currently arbitrary units) for the y-axis.

      We thank the reviewer for this insightful comment and agree that our original wording overstated the mechanistic interpretation of these data. In particular, the use of the term “non‑competitive” is not well supported by the current analysis and may be misleading, especially given that our data are consistent with binding within or proximal to the ATP-binding pocket. We have therefore revised the text to remove this designation and instead describe the data more conservatively in terms of changes in apparent Vmax, without assigning a specific inhibition mechanism. With respect to kinetic analysis, we agree that full determination of K<sup>m</sub> values and inclusion of progress curves would provide a more rigorous mechanistic interpretation. However, given the primary focus of this manuscript on identifying and characterising small‑molecule activators of GCN2 in cells, we believe that a detailed steady‑state kinetic analysis would be beyond the scope of the current study. We have therefore moderated our conclusions accordingly and now present these data as preliminary kinetic observations rather than definitive evidence of inhibition modality.

      Results

      “Compounds 18 and 20 altered the apparent kinetic parameters of GCN2, including an increase in the observed V<sub>max</sub>; however, these data do not allow assignment of a specific inhibition modality.”

      (8) The full-length GCN2 assay presented in Figure 5G appears to be unresponsive to uncharged tRNA, a known regulator of GCN2. The statement that compound 20 induces eIF2 phosphorylation to a greater extent than tRNAs is true for the in vitro assay, but arguably is because the in vitro assays do not recapitulate the in vivo arrangement.

      We thank the reviewer for this comment. We respectfully disagree that the assay is unresponsive to uncharged tRNA. In our hands, GCN2 does exhibit activation in response to tRNA; however, the magnitude of this effect is modest (~4‑fold) compared to the substantially stronger activation observed with compound 20 (~40‑fold). As a result, the tRNA response can appear compressed when both are plotted on the same scale. We agree with the reviewer that the in vitro assay does not fully recapitulate the in vivo regulatory environment, where factors such as GCN1 and ribosome association are known to potentiate GCN2 activation. This limitation likely explains the relatively weaker response to tRNA under our assay conditions and was a key motivation for incorporating cellular assays in our study. Interestingly, the marked difference in activation magnitude between tRNA and compound 20 in vitro raises the possibility that compound-mediated activation may, at least in part, bypass regulatory features that normally constrain GCN2 activity in a GCN1‑dependent manner. While we have not directly tested this hypothesis, we will temper the wording and include this as a speculative point in the Discussion.

      Results

      “Of note, uncharged tRNA produced a modest (~4‑fold) activation of GCN2 under these conditions, whereas compound 20 induced substantially greater (~40‑fold) activation.”

      Discussion

      “The markedly greater activation observed with compound 20 compared with uncharged tRNA in vitro raises the possibility that such compounds may partially bypass regulatory constraints on GCN2 activation, including those normally mediated by GCN1.”

      (9) Using purified GCN2 kinase domain at low ATP concentrations (10 μM), compounds 18 and 20 were shown not to inhibit GCN2 up to concentrations of 3 μM. In previous assays, much higher concentrations of compounds 18 and 20 were used to inhibit GCN2. Why were different concentrations used? This makes this interpretation of this data difficult for the reader to draw conclusions.

      We thank the reviewer for this comment and agree that the use of different concentration ranges across assays may not have been sufficiently clear. The kinase‑domain assay performed at low ATP (10 μM) was specifically designed to assess whether compounds 18 and 20 have a propensity to inhibit GCN2 under conditions that sensitise detection of ATP‑competitive effects and facilitate comparison with related eIF2α kinases. This assay was therefore optimised for detecting inhibition, rather than activation. In contrast, the higher concentrations used in other experiments were selected to robustly measure ISR activation in cellular or full‑length protein contexts, where higher compound exposure is required to observe downstream signalling outputs. These two assay systems therefore address distinct mechanistic questions—targeting inhibition under controlled biochemical conditions versus activation in more complex functional settings—and are not directly comparable in terms of concentration–response relationships. We will revise the manuscript to clarify this distinction and to emphasise that the kinase‑domain assay was not intended to define the activation potency of the compounds.

      Results

      “This assay was performed at low ATP concentrations to sensitise detection of ATP-competitive inhibition and was not optimised to detect compound-mediated activation of GCN2.”

      (10) In silico docking studies support the binding of compound 20 in the ATP-binding pocket of the GCN2 kinase domain. How is this compatible with the stated ATP non-competitive mechanism?

      We agree with the reviewer that our previous description was misleading. The designation of compounds 18 and 20 as “ATP non‑competitive” is not supported by the available data and is inconsistent with the docking results suggesting binding within the ATP‑binding pocket. We have therefore revised the manuscript to remove this terminology and to describe the kinetic behaviour more cautiously, without assigning a specific mode of inhibition.

      (11) The study does not appear to feature biological assays demonstrating the effects of GCN2 activation. For example, does compound 20 reduce translation or growth of cells in a GCN1/GCN2-dependent manner, or do the compounds overcome PVOD mutations akin to the authors' Hum Mol Genet 2024 Aug 18;33(17):1495-1505 article?

      We thank the reviewer for this suggestion. We agree that defining downstream biological consequences of GCN2 activation is an important goal. We did explore this using several cellular systems; however, these effects were context-dependent and not consistently observed across models. Specifically, compounds did not induce a detectable ISR in HCT116 or COS‑7 cells under the conditions tested, despite responsiveness of these systems to canonical activators such as histidinol. In contrast, we did observe induction of PPP1R15A in primary mesothelioma cells, indicating that biological responses can be elicited in certain cellular contexts. We also tested whether these compounds could rescue disease-associated GCN2 variants linked to PVOD, as previously reported for GCN2iB, but did not observe activation of these mutants. These findings suggest that while the compounds robustly activate GCN2 signalling in reporter assays, downstream biological outputs are context-dependent and may require specific cellular conditions or co-factors. Given the variability across systems, we have limited our conclusions to ISR activation and have not generalised broader biological effects. We will clarify this point in the revised manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses. Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (12) For the control of neratinib and other activators linked with ATP binding that are suggested to be dependent on GCN1 in Fig. 4, include reporter induction by drug treatment in Gcn2-/- cells. It would be helpful to be clear in the list of compounds between those suggested to be direct activators versus those that may create stress that leads to GCN2 activation.

      We thank the reviewer for this suggestion. In response, we have performed additional reporter assays in both WT and GCN2<sup>-/-</sup> cells to assess the dependence of drug-induced ISR activation on GCN2. These experiments reveal clear GCN2 dependence for reporter induction upon treatment with sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, as signal is markedly reduced in GCN2⁻/⁻ cells compared to WT.

      In contrast, dovitinib and AZD1775 show less clear dependence, with relatively low reporter signal even in WT cells (notably lower than observed in Fig. 4C), limiting interpretation. Interestingly, dabrafenib induces stronger reporter activity in GCN2<sup>-/-</sup> cells than in WT, indicating that its effects are independent of GCN2 and may reflect activation of alternative stress or signalling pathways.

      Results

      “To further assess the mechanism of compound-induced ISR activation, we evaluated reporter responses in GCN2-deleted cells (Supplementary Figure S3). Several compounds, including sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, showed reduced reporter activity in GCN2-deficient cells, consistent with GCN2-dependent activation. In contrast, dovitinib, AZD1775 and dabrafenib produced weaker or inconclusive responses even in the paired wild-type lines, limiting analysis. These data allow us to distinguish compounds consistent with direct or GCN2-dependent activation from those more likely to induce ISR indirectly through cellular stress upstream of GCN2. These findings support a distinction between compounds that activate the ISR through GCN2-dependent mechanisms and those that likely act indirectly via alternative stress pathways.”

      Legend

      “Supplementary Figure S3. GCN2-dependence of ISR activation by putative GCN2 agonists

      Normalised fold-change in ATF4 signal in CHO WT (purple) and Gcn2-/- (blue) ATF4::NanoLuc reporter cells treated for 19 hours with a panel of ATP-competitive kinase inhibitors reported to activate GCN2 (at 1 and 3µM, sunitinib used at 3 and 10µM, AZD1175 used at 0.3 and 1µM, gefitinib and erlotinib used at 3 and 10µM). DMSO was used as vehicle control. (n=3; mean ± SEM).”

      Minor Comments:

      (1) In the introduction, the authors state that "Several type 1 and 1.5 kinase inhibitors can activate GCN2 at low concentrations while inhibiting at higher concentrations". What is the evidence that type I inhibitors can activate GCN2?

      Thank you for this query. We believe our initial phrasing was open to misinterpretation. The new wording is

      Introduction

      “Several kinase inhibitors classified as type 1 or type 1.5 with respect to their canonical targets can activate GCN2 at low concentrations while inhibiting it at higher concentrations; however, their binding mode to GCN2 remains undefined.”

      (2) A CHO:ATF4-Nanoluc translation reporter screen was used to screen 123K compounds and divide them into a pool of inhibitors and a pool of activators. The criteria used to make this distinction are not sufficiently described in the manuscript, and screen results are not supplied as supplementary data.

      We thank the reviewer for highlighting the need for greater clarity regarding the screening criteria and reproducibility. Compounds from the primary screen (BioAscent library of 123,222 drug-like compounds performed by the ALBORADA Drug Discovery Institute) were classified based on Z-score thresholds, with activators defined as those with Z-score > 3. To assess robustness, the primary screen data were re-analysed independently. In an initial analysis of 121,000 compounds (excluding plates failing quality control), 6,521 compounds met the activator threshold. Of these, 6,461 overlapped with the original hit list. The small number of discrepancies included: (i) compounds absent from the analysed dataset (e.g. originating from failed plates), and (ii) compounds with Z-scores close to the threshold (typically between −3 and −3.02), consistent with minor analytical variation (e.g. rounding). Overall, these analyses show a high degree of concordance in hit identification, with differences restricted to borderline cases near the selection threshold. We have clarified these criteria below.

      Results

      “Compounds were classified based on Z-score thresholds derived from the primary screen, with activators defined as those with Z-score > 3. An initial analysis of 121,000 compounds, excluding plates failing quality control, identified 6,521 activators, of which 6,461 overlapped with the subset selected for follow-up screening. Minor discrepancies were restricted to compounds absent from the analysed dataset (e.g. originating from failed plates) or those with Z-scores close to the threshold (3 to 3.02), consistent with limited analytical variation. Using this approach, we assembled a subset of 6,461 compounds enriched for potential ISR activators and screened these in 384-well format at 10 µM for 16 hours.”

      Reviewer #3 (Recommendations for the authors):

      (1) The legend to Figure 4 should read "Compound 20 displays GCN1 independence", not "dependence".

      Thank you for spotting this error. We have made the correction.

      (2) The text describes Figure 2D as examining 4h, but it also examines 8h.

      We have amended the text to:

      “After validating their effects using the assays described above, we next confirmed activation of the ISR by these compounds at the transcriptional level at 4 and 8 hours.”

      (3) Unless I'm misreading Table S1, compound Z134826202 is listed as activating the ISR, but the authors describe it in the text as inactive.

      Thank you. We have corrected this error.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study would be greatly strengthened by testing key predictions made using perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

      We thank eLife for this nice assessment. We indeed believe that the presented machine learning approach will be useful for many types of research questions, and this was the main aim of this manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study only provides limited new insight into chromatin dynamics.

      We respectfully disagree with this conclusion. To the best of our knowledge, our study is the first to provide any insights into chromatin dynamics upon serum starvation. Previous studies are solely based on studies in fixed cells, and the dynamics have remained unexplored. Moreover, for example the notion that homologous chromosomes show differential dynamic behavior is novel and will likely have implications and relevance to many chromatin-based processes beyond the example studied here.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      First, we would like to point out that the other reviewer found our machine learning analysis pipeline a major strength of our manuscript. Indeed, analyzing single features and assessing their impact on the studied phenomenon could have been achieved relatively easily by conventional analysis. However, this analysis would have ignored the interactions (some of which were not intuitively obvious) between different features and thereby limited the knowledge gain from the experiment.

      Unbiased analysis of the interactions between the different features would have been already very difficult and time-consuming with conventional approaches. We believe that our analysis pipeline, especially with the Shapley values, addresses the key issue of combinatorial explosion prominent to multiparametric data, such as imaging data, and helps the researcher to navigate complex datasets.

      (3) There are several specific technical points:

      (a) It was not clear what the CRISRP-Sirius probes actually labelled. The chromosome 1 sgRNA sequence is provided, but I could not find information as to which region(s) of the chromosome are actually labelled (size, location, etc.).

      We have added a schematic as Supplementary Figure 1A to show the region of the chromosome that is labelled. In addition, the target sequence, together with the relevant references can be found in the Materials and methods (page 16). Please see also below Reviewer #1 (Recommendations for the authors) point 4a.

      (b) The authors visualize a relatively small region of chromosome 1 but make conclusions regarding the entire chromosome. Additional probes on the same chromosome should be used.

      Related to this point, the discussion of why the authors are unable to reproduce the prior findings of relocation of chromosomes 10, 13, and X is not satisfying. It would be worth comparing the FISH-based painting of entire chromosomes, which generated the results suggesting relocation of these chromosomes, with the point-labelling method used here.

      We agree that our approach to labeling chromosome 1 is very different than the FISH-based probes utilized before. However, we also feel that we discuss this aspect, and the difference between our and previous results, which may also stem from the used cell model, in quite a detail in the first paragraph of the results (page 4). Also, we are very careful throughout the manuscript to indicate that here we study the dynamics of a specific chromosome loci, not the entire chromosome, and have further amended the text to emphasize this. In the future, it would be very interesting to study the dynamics of also other loci of chromosome 1. As indicated also below in response to reviewer 2, we have failed to identify further gRNAs that would reliably and reproducibly label further chromosome 1 loci, suggesting that we would need to change the labeling system entirely. Unfortunately, this is not in the scope of this manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 1.

      (c) The study lacks controls. Since in their hands chromosomes 10, 13, and X do not change position, they should be used as a negative control in all experiments demonstrating a shift in the location of chromosome 1.

      We disagree that our study lacks controls, since we use telomeres as controls throughout the manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 2,3.

      (d) I did not find information about the spatial or temporal resolution of the imaging modality. This is important to assess whether the observed changes in position, relative to time, are meaningful.

      To estimate the spatial resolution, we have added new data using fixed cells (Supplementary figure 1E; corresponding text in results on page 5); temporal resolution is indicated in Materials and methods (page 17). Please see also below Reviewer #1 (Recommendations for the authors) point 4d.

      (e) The authors analyze surprisingly early timepoints (up to 40 minutes) of serum starvation. Would these results look different if longer serum starvation timepoints of several hours were analyzed?

      We chose to analyze early time points of serum starvation based on the previous literature reporting the chromosome relocation within the first 15 minutes of starvation. Indeed, the results might look very different later during serum starvation, since we already observe differences between 0-20 min vs 20-40 min into starvation (see for example Figure 5A-D). Analyzing further time points is not in the scope of this manuscript.

      (f) The authors can do a better job of explaining what the biological meaning of the various parameters (DistR, TDist, etc.) they measure is.

      We have amended Table 1 to describe the measured features more clearly. Please see also below Reviewer #1 (Recommendations for the authors) point 4e.

      (g) I did not understand the reasoning for the authors' conclusion of differential behavior of homologues. Please explain this better, or idealy use more direct labeling methods that identify the individual homologues.

      The differential behavior of homologues is best demonstrated in Figure 6H, which shows that in serum-containing media, the peripheral homolog has equal probability of being faster or slower compared to its homolog. However, the distribution changes upon starvation, with the peripheral loci being more frequently the faster homolog. We completely agree that further studies are needed to understand this phenomenon better, but changing the labeling method is not in the scope of this manuscript.

      (h) In many figures, statistical analysis of the data is missing, including, but not limited to, Figures 1B, C, G, Figures 4, 5, 6.

      We have added a Supplementary table to include inferential statistics. See also below Reviewer #1 (Recommendations for the authors) point 4b.

      (i) No information is provided throughout the manuscript as to how many cells were analyzed in each experiment. This should be indicated in every figure legend.

      The number of analyzed loci or nucleus is indicated in every figure. See also below Reviewer #1 (Recommendations for the authors) point 4c.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Livecell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      We appreciate the comment about the readability of our manuscript, and have seriously evaluated this point. We also completely agree that it would be interesting and important to understand the mechanistic and functional basis of the observed changes in chromatin dynamics take place upon serum starvation. However, we feel that it is not in the scope of the present manuscript. See also below Reviewer #2 (Recommendations for the authors) points 3,7-11.

      Recommendations for the authors:

      Reviewing Editor Comments:

      I would like to first offer my congratulations on a very interesting study; second, I would like to encourage you to test a few key predictions using a perturbation experiment. Two reviewers with deep expertise in this area were supportive of the work, and both noted that such an addition would greatly increase the impact and visibility of this work in the field. I welcome a revision that addresses this seminal point. Thank you for sending your work to eLife!

      We thank eLife for the positive assessment. We have aimed to address all of the reviewers comments and suggestions. However, we feel that some of the suggestions are not in the scope of this particular manuscript, since they would require setting up a different chromatin labeling system.

      Reviewer #1 (Recommendations for the authors):

      The following experiments would strengthen the study:

      (1) Please label additional regions on chromosome 1 so as not to rely on a single point to represent the behavior of the entire chromosome.

      This is an excellent suggestion, but unfortunately, despite our extensive efforts, we have failed to identify further gRNAs that would reliably label chromosome loci with the CRISPR-Sirius system. Changing the labeling system is not in the scope of the presented manuscript.

      (2) Please use chromosomes 10, 13, or X as a negative control since these chromosomes do not change position in the authors' hands.

      (3) Please compare the behavior of the homologues to that of either random loci or control loci on 10, 13, or X to assess whether the differential behavior observed for chromosome 10 is a specific effect.

      Related to points 2 and 3, we opted to use telomeres as controls in this study. Throughout the manuscript, the behavior of chromosome 1 loci is compared to telomeres, demonstrating the specific effect of serum starvation on chr 1. For example, Figure 5A and 5B show that when analyzing mean locus displacement, chr1 and telomeres show the opposite behavior.

      (4) In addition:

      (a) Please provide detailed information on the sequence and location of the probes used.

      We have added a schematic showing the location of the probes as Supplementary Figure S1A. In addition, the sequences are indicated in Materials and methods (page 16).

      (b) Please provide a statistical analysis in all graphs.

      To make statistical analysis more comprehensive, we have added supplementary table 1, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. See also below Reviewer #2 (Recommendations for the authors) point 5.

      (c) Please provide throughout the manuscript in each figure legend information as to how many cells were analyzed in each experiment.

      The number of analyzed loci (or nucleus) is indicated in each graph.

      (d) Please provide information on the spatial and temporal resolution of the imaging modality.

      The imaging settings are indicated in Materials and methods, including the temporal resolution of 0.25 frames per second (page 17). To estimate spatial resolution, and especially its relationship with the observed repositioning of the chromosome loci, we performed experiments in fixed cells, using an optically identical set-up as utilized for live imaging. Unfortunately, the microscope utilized for live imaging was taken out of use by the core facility after submission of the original draft of this manuscript, but we used a microscope with essentially a similar set-up. The data from fixed cells is now presented as Supplementary figure 1E and discussed in results on page 5. This analysis indicates that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error.

      (e) Please better explain what the various measured parameters mean in biological terms.

      We have amended Table 1 to provide better explanation of the measured parameters.

      (f) Please add a scale bar to Figure 6I.’

      Scale bar has been added to figure 6I.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) SHAP analysis identified nuclear area (MA) and its change (sA) as the most predictive features of starvation state, while motility features (MD, MaxD, TD) showed strong interactions with nuclear morphology. Discrete features, such as displacement outliers and homolog subclassification by speed/proximity, influenced classification, particularly in MLP models. Could the authors clarify why morphological and motility features act in combinatorial and context-dependent ways? A biological interpretation of this interdependence would strengthen the study.

      Unfortunately, we do not have a good biological interpretation for this. The fact that some interactions are context-dependent indicates that there could be subpopulations of cells/analyzed loci. For example, we found that the predictive value of nuclear area was high in a subgroup of low motility loci (Figure 4F and Supplementary figure 4D-F). We do not believe that adding more speculation would strengthen the study.

      (2) The manuscript shows that chromosome 1 moves toward the periphery within the first 20-40 minutes of serum withdrawal. However, it remains unclear whether the locus eventually "touches" the periphery and whether it subsequently stabilizes or retracts. It would be valuable to compute the time point of minimal nuclear distance and examine whether this is transient or sustained.

      With the experimental set-up utilized here, we imaged the loci for only two minutes at random time point within the first 40 minutes of the starvation. Hence extracting the time point of minimal nuclear distance is not meaningful from this dataset. As we discuss in the manuscript, following the dynamics of the same locus for longer periods of this would be very interesting in the future. However, this is not in the scope of the present manuscript.

      (3) The manuscript is at times difficult to follow, partly because methodological descriptions are highly detailed in the main text. Consider moving more of the methodological content into Supplementary Methods and emphasizing the main results and interpretations in the main text for clarity.

      We have carefully evaluated this point. Most methodological descriptions in the manuscript relate to the machine learning models and their explanation with SHAP. As we feel that this combination is an essential part of the manuscript, and likely the aspect that can have widest impact beyond chromatin dynamics studies, we feel that the background and our reasoning related to the chosen methods are important.

      (4) The distinction between the first 20 minutes and the latter 40-minute window is intriguing. Could these different time scales be paralleled with early versus delayed gene expression responses to serum starvation? A discussion of this temporal connection would add biological depth.

      This is an intriguing idea, and we have added a short note on this in the discussion (page 13). However, as we do not know how the U2OS cells utilized here respond transcriptionally to serum starvation, we are hesitant to speculate too much.

      (5) If the observed interpretations are robust, could this be demonstrated more explicitly through statistical principles or reproducibility tests across independent datasets?

      To provide further evidence of the robustness of our findings, we have 1) added new data to estimate the spatial resolution (Supplementary figure 1E) and 2) expand the statistical analysis as supplementary table 1. Regarding the spatial resolution (see also the response to reviewer 1), our experiments on fixed cells demonstrate that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. To make statistical analysis more comprehensive, we added a supplementary table, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. MW test was used as a non-parametric statistic, with null hypothesis assuming the samples are coming from the same distribution. Null hypothesis was rejected at the p-value below 0.05. In most cases (bold font in table) MW test results were in accordance with the descriptive statistics confirming the conclusions in the study. In case of exceptions (MD, TD for telomeres and TDist), the results were reported, for example, as an “appearing trend” to reflect descriptive statistics while highlighting certainty levels.

      (6) Figure labeling is difficult to follow. Please include abbreviation explanations directly in the figure panels or legends for clarity.

      Abbreviations have been added to figure legends. Adding them to figures themselves would have made the figures too busy.

      (7) The manuscript reports higher displacement at the nuclear periphery. Can the authors explain why displacement amplitudes increase near the periphery and how this relates to nuclear architecture?

      We speculate in the manuscript (results, page 11; discussion, page 14) that actually the lower displacement observed with the central locus may, at least partially, result from anchoring this locus to the nucleolus (Fig 6I). Nevertheless, alternative explanations, such as differences in transcriptional and/or chromatin states may exist (see also the response to point 11), and this is now mentioned in the discussion (page 14).

      (8) How is the movement of chromosome 1 directed specifically toward the periphery, rather than being random fluctuations? This point requires clarification.

      This is an important question, but unfortunately our data does not provide an answer to this, and suggesting any mechanism would be pure speculation. Nevertheless, our results agree with previous studies utilizing fixed cells that also demonstrated movement of chromosome 1 towards nuclear periphery (Mehta et al., 2010), arguing against random fluctuation.

      (9) Only chromosome 1, and not the other tested chromosomes, undergoes this relocalization. Could the authors elaborate on why some chromosomes but not others display this behavior?

      Previous studies (Mehta et al., 2010) utilizing chromosome paints in fixed cells actually show the relocalization of several chromosomes upon serum starvation. The fact that we observed the relocalization of only chr1 loci is likely due to the labeling method and/or the cell model utilized in this study. This is quite explicitly discussed in the first paragraph of results (page 4).

      (10) The magnitude of these movements appears relatively small (0.5 micron). Can the authors discuss whether such small but reproducible displacements are likely to be biologically meaningful in terms of nuclear function or gene regulation?

      At the moment, our experimental set up allows us to analyze the dynamics of only a small portion of chr1, which indeed shows an average 0.5 micron displacement towards the nuclear periphery. Based on the chromosome painting data from fixed cells, the displacement at the level of whole chromosome is significantly larger. As mentioned in the discussion (page 13), the functional implications of radial repositioning of chromosomes upon serum starvation is not known. Therefore further discussion on the relevance of the magnitude reported here would be pure speculation.

      (11) Peripheral homologs of chromosome 1 became faster and more dynamic under starvation. Why might these loci be more prone to movement? Could this be linked to differences in transcriptional activity or chromatin state between central and peripheral homologs?

      At the moment we favour the idea that the central homolog is constrained by its anchorage to the nucleolus (Figure 6I). However, transcriptional activity and/or chromatin state may also play a role, and this possibility is now mentioned in the discussion on page 14.

      Minor points:

      (1) Figure legends use inconsistent capitalization and panel labels. These should be standardized across all figures for better readability.

      We apologize for these inconsistencies, and have aimed to standardize all labeling.

      References

      Mehta, I.S., Amira, M., Harvey, A.J., and Bridger, J.M. (2010). Rapid chromosome territory relocation by nuclear motor activity in response to serum removal in primary human fibroblasts. Genome Biol 11, R5.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We are grateful to all the reviewers for dedicating time to review our manuscript and for providing insightful comments and suggestions. We have revised our manuscript in line with the reviewers' feedback. The major revisions include characterization of Dcp-1 overexpression-induced cell death, demonstration of the involvement of autophagy in Dcp-1 activation, characterization of the interaction between full-length Bruce and cleaved Dcp-1. We have introduced new figures (Figure 1 – figure supplement 1, Figure 2 – figure supplement 2, Figure 3 – figure supplement 1, Figure 5 – figure supplement 1), new panels (Figures 1D, Figure 3C, Figure 4I, J) and a new table (Table S2). The previous Figure 5 – figure supplement 1 has been relocated to Figure 4 – figure supplement 2.

      With all concerns and suggestions from the reviewers addressed, our conclusion—that Bruce suppresses autophagy-regulated caspase activity and wing tissue growth in Drosophila— is now more robustly supported. We are confident that our revised manuscript makes a significant contribution to the fields of cell death, autophagy, and developmental biology, as it provides a new conceptual framework for understanding non-lethal caspase regulation. We remain hopeful that the reviewers will find it suitable for publication in eLife.

      Reviewer #1 (Public review):

      Summary:

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components. Using TurboID-based proximity labeling, they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp1 activation. Using structure-based prediction using AlphaFold3, they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, acts as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy-Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Strengths:

      This is an excellent paper with very good structure, excellent quality data and analysis.

      Weaknesses:

      This reviewer did not identify any weaknesses or recommendations for revision.

      We sincerely thank the reviewer for their highly positive evaluation of our work. We are pleased that the reviewer found the overall structure, data quality, and analyses to be strong, and that they clearly recognized the key findings of our study. No changes to the manuscript were required in response to this review.

      Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce-expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      On the positive side, the study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a.

      Weaknesses:

      The data supporting the Dcp-1/Bruce interaction are not strong, even though the title of this manuscript highlights Bruce. For example, the authors' turboID data does not support Dcp1/Bruce interaction. The case for the interaction is based on a single experiment that overexpresses a truncated Bruce transgene in S2 cells.

      We sincerely thank the reviewer for their constructive and detailed evaluation of our manuscript. We appreciate the positive assessment that our study identifies new Dcp-1-associated factors and provides functional links between Dcp-1 and Sirt1, Fkbp59, and multiple autophagy-related genes, including Debcl, Buffy, Atg2, and Atg8a. We also thank the reviewer for clearly pointing out concerns regarding the strength and interpretation of the evidence connecting Bruce and Dcp-1. In the revised manuscript, we have addressed these concerns in two major ways. First, we provided additional experimental evidence explaining why TurboID-mediated labeling did not identify Bruce. Specifically, we showed that the majority of TurboID-tagged Dcp-1 expressed in wing imaginal discs remains in its full-length form, which is unlikely to engage Bruce. Second, and more importantly, we now demonstrated that endogenously expressed full-length Bruce interacts with cleaved Dcp-1 in wing imaginal discs. These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Detailed descriptions of these experiments and results are provided in the point-by-point responses in the “recommendations for the authors” section. Together, these newly added data substantially strengthen the evidence for the Dcp-1/Bruce interaction and support the focus of the original manuscript title.

      Reviewer #2 (Recommendations for the authors):

      (1) The title of the manuscript highlights Dcp-1/Bruce interaction, even though the evidence there is not strong. The evidence for Dcp-1/Sirt1 and Dcp-1/Fkbp59 is stronger. How about changing the title to highlight these other Dcp-1 interactions?

      We thank the reviewer for the thoughtful suggestion. We agree that several Dcp-1-associated factors identified in our study, particularly Sirt1 and Fkbp59, are supported by functional evidence. Specifically, our data show that Sirt1 and Fkbp59 are required for Dcp-1 overexpression-mediated activation. However, Bruce differs from these factors in both the scope and the nature of its effects on Dcp-1. Bruce is not only shown to specifically suppress Dcp-1 activity, but also to suppress wing tissue growth, indicating a broader physiological role in modulating non-lethal Dcp-1 function. Importantly, we further demonstrate that Bruce can specifically physically interact with cleaved Dcp-1. In addition, in this revised manuscript, we show that using the endogenously mStayGold::V5-tag knock-in-tagged Bruce allele, cleaved Dcp-1, induced by overexpression of Dcp-1::VENUS in wing imaginal discs, can be co-immunoprecipitated with full-length Bruce (new Figure 4I, J). These results support a physical interaction between full-length Bruce and activated Dcp-1 in vivo, consistent with a direct inhibitory role. Based on these findings, we decided to retain Bruce in the manuscript title, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      (2) The case for Dcp-1/Bruce interaction is not strong because the Dcp-1 turboID fails to identify Bruce. In fact, the Dcp-1 turboID approach may not have been effective, as it failed to detect many established interactions, including Diap1 (Wang et al. 1999 PMID 10481910; Tenev et al., 2006 PMID 15580265). The authors may want to comment on this.

      We thank the reviewer for raising this important point. We agree that Bruce, as well as DIAP-1, was not identified in our TurboID-MS labeling dataset (Figure 2C, Table S1). Previous studies have shown that DIAP1 interacts with Dcp-1 and Drice only after exposure of the IAP-binding motif (IBM) at the neo-N-terminus of the large executioner caspase subunit following cleavage (Tenev et al., 2005). Similarly, our co-immunoprecipitation analyses show that Bruce interacts specifically with cleaved Dcp-1, but not with full-length Dcp-1. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs is present in the full-length pro-form (new Figure 2 – figure supplement 2A). Thus, the failure to identify Bruce and DIAP1 by TurboID-MS using full-length Dcp-1 as bait is expected, as this approach primarily labels interactors of the inactive, full-length form of Dcp-1. To evaluate whether our proximity labeling approach was nevertheless effective, we compared our TurboIDMS dataset with a previously published immune-affinity purification (IAP)-MS dataset generated using catalytically inactive, C-terminally V5-tagged Dcp-1 overexpressed in Drosophila 1(2)mbn cells (Choutka et al., 2017). Although the experimental conditions differ in several respects, we observed a substantial overlap between the TurboID-MS-mediated and IAP-MS-mediated interaction lists (new Figure 2 – figure supplement 2B, new Table S2). Importantly, SesB, one of the best-characterized Dcp-1 interactors located in mitochondria (DeVorkin et al., 2014), was also identified in our mass spectrometry dataset (new Figure 2 – figure supplement 2B, new Table S2). Based on these analyses, we now more explicitly describe the experimental context and limitations of the TurboID approach, clarifying that it preferentially labels interactors of full-length Dcp-1 in the revised manuscript. We also incorporate comparisons with prior studies to further support the validity of our mass spectrometry experiments in the revised manuscript

      (3) The best experimental evidence for Bruce/Dcp-1 interaction can be found in Figure 4H. But here, they see a weak interaction only when a truncated Bruce construct is overexpressed in S2 cells. Whether Dcp-1 interacts with Bruce in a physiological setting remains unsupported.

      We thank the reviewer for the important comment. We agree that, in the original manuscript, the biochemical evidence for the Bruce/Dcp-1 interaction relied primarily on experiments using an overexpressed truncated Bruce construct in S2 cells and therefore did not sufficiently establish whether this interaction occurs in vivo, especially in wing imaginal discs. To address this concern, we performed additional experiments to examine the Bruce/Dcp-1 interaction. In the background of the mStayGold::V5-tag knocked-in Bruce allele, we overexpressed Dcp-1::VENUS using WPGal4 driver to induce Dcp-1 activation and tested whether full-length Bruce under endogenous expression interacts with cleaved Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates with cleaved Dcp-1 in vivo and thus support the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      (4) The genetic interaction between Bruce and Dcp-1 is interesting, but the interpretation becomes complicated because Bruce inhibits Reaper, and at the same time, Dcp-1 genetically interacts with Reaper, Hid, and Grim (Figures 1E, F, G). Thus, it remains unclear if the genetic interaction between Bruce/Dcp-1 is due to a direct interaction between Bruce/Dcp-1 or alternatively, because Bruce inhibits Reaper and Grim.

      We thank the reviewer for the comment. The primary function of Reaper, Hid, and Grim (RHG proteins), collectively referred to as IAP antagonists, is to directly interact with inhibitor of apoptosis proteins (IAPs), most notably DIAP-1 (Kornbluth and White, 2005; Ryoo and Baehrecke, 2010), leading to the inhibition of DIAP-1 function. RHG proteins have not been shown to directly inhibit caspases. Because inhibition of RHG proteins results in the stabilization of DIAP-1, it is likely that the effects observed upon RHG gene knockdown are mediated through DIAP-1. Consistent with this idea, overexpression of DIAP-1, while less potent than Bruce, can also suppress Dcp-1 activation (Figure 5B, C). However, we also acknowledge that Bruce suppresses Reaper- and Grim-dependent, but not Hid-dependent, cell death (Vernooy et al., 2002). In addition, Bruce directly targets Reaper through non-lysine ubiquitination, promoting its degradation (Domingues and Ryoo, 2012). Thus, it is possible that Bruce overexpression suppresses Reaper and thereby strengthens DIAP-1 function, which could indirectly contribute to the inhibition of Dcp-1 activation. Nevertheless, because the effect of Bruce overexpression is stronger than that of DIAP-1 overexpression (Figure 5B, C), and together with our physical interaction data of Bruce with cleaved Dcp-1, we propose that Bruce most likely inhibits Dcp-1 directly to attenuate its activation.

      (5) In general, the manuscript could benefit from highlighting the strong data on Sirt1 and Fkbp59, while clearly acknowledging the limitations of the Bruce/Dcp-1 interaction.

      We thank the reviewer for the comment. As described above, in the revised manuscript we now demonstrate that endogenously expressed full-length Bruce physically interacts with cleaved Dcp-1 in wing imaginal discs (Figure 4I, J). These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Based on this evidence, we decided to highlight Bruce in the manuscript, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a nonapoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating nonapoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      We sincerely thank the reviewer for their thoughtful and constructive evaluation of our study. In response to these concerns, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. Detailed explanations and experimental results are provided in the point-by-point responses below. Overall, we believe that these additions strengthen the manuscript by clarifying the dual roles of Dcp-1 in promoting tissue growth at low activity levels while triggering apoptosis when excessively activated.

      Specific points:

      (1) Nature of the wing ablation phenotype

      A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Figure 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      We thank the reviewer for this important point regarding the nature of cell death. We agree that nuclear condensation alone is not sufficient to conclude apoptotic cell death, and we therefore performed additional experiments. First, we performed TUNEL staining and detected robust TUNEL-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1D), supporting apoptotic DNA fragmentation. Second, to directly test whether Dcp-1 overexpression leads to executioner caspase activity in vivo, we used two independent executioner caspase activity probes, GC3Ai (Schott et al., 2017; Zhang et al., 2013) and CD8::PARP::VENUS (Williams et al., 2006). Both probes showed clear executioner caspase activity-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1 – figure supplement 1C–F), demonstrating that Dcp-1 overexpression leads to executioner caspase activity capable of cleaving standard substrates in vivo. In addition, staining with an anti-cleaved Dcp-1 antibody was positive upon Dcp-1 zymogen overexpression (new Figure 1 – figure supplement 1B), indicating that the overexpressed Dcp-1 zymogen is converted into its active form. Consistent with this result, western blot analysis revealed that Dcp-1 zymogen overexpression results in the appearance of a cleaved Dcp-1 (new Figure 1 – figure supplement 1A). Importantly, consistent with our original observation that the wing ablation phenotype is suppressed by expression of the caspase inhibitor p35, we further showed that p35 overexpression completely abolished the appearance of cleaved Dcp-1 in western blot (new Figure 1 – figure supplement 1A), suggesting that Dcp-1 activation is mediated by self-cleavage. Taken together, these new results demonstrate that Dcp-1 zymogen overexpression induces typical executioner caspase activity-dependent apoptotic cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (2) Role of Decay

      In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At a minimum, this point should be discussed explicitly.

      We thank the reviewer for the comment regarding the role of Decay. In our previous study (Shinoda et al., 2019), we demonstrated that both Dcp-1 and Decay promote wing tissue growth in a non-lethal manner. In the present study, however, we focused our analysis on Dcp-1. This decision was based on both technical and biological considerations. From a technical perspective, we had established TurboID knock-in lines and UAS overexpression lines for Dcp-1, Drice, and Dronc, whereas corresponding genetic tools are not available for Decay. From a biological standpoint, Dcp-1 exerts a stronger effect on wing growth than Decay, as shown in our previous work, and exhibits a dual functional spectrum: Dcp-1 promotes tissue growth at low activity levels, whereas excessive activation induces overt cell death. By contrast, Decay has not been implicated in cell death in wing imaginal discs (Kondo et al., 2006). Given that a central aim of the present study was to dissect how executioner caspase activity is differentially regulated to support nonlethal functions versus apoptotic cell death, we therefore focused on the two executioner caspases that are known to participate in apoptosis, Dcp-1 and Drice. We agree with the reviewer that Decay remains an intriguing caspase with largely unexplored physiological roles, and further investigation into its regulation and function will be an important direction for future studies. Importantly, Decay has been shown to mediate Hid-induced cell death in the DIAP1- and apoptosome-independent manner in differentiating photoreceptors and accessory cells of the eye (Leulier et al., 2006). In addition, although not required for cell death, Decay accounts for most of the caspase activity during metamorphic midgut programmed cell death, which is executed by autophagy (Denton et al., 2009). Thus, similar to Dcp-1, Decay might be an executioner caspase that can be regulated independently of the canonical apoptosome-mediated pathway, potentially involving autophagy-Bruce axis, and thereby contributing to the regulation of tissue growth. We have now discussed this point in the revised manuscript.

      (3) Figure 2: Proximity labelling analysis

      The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICEassociated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity would be valuable.

      We thank the reviewer for the comment regarding the TurboID-based proximity labeling analysis and the interpretation of the identified Dcp-1-associated factors. With respect to Bruce, we clarify that Bruce was not identified as a Dcp-1 interactor in the TurboID proximity labeling dataset. Our co-immunoprecipitation analyses in S2 cells indicate that Bruce interacts specifically with cleaved Dcp-1, but not with the full-length, inactive form. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs exists in the full-length pro-form (new Figure 2 – figure supplement 2A). Therefore, the failure to detect Bruce in the TurboID experiment using full-length Dcp-1 as bait is expected, as this approach primarily labels proteins proximal to the inactive form of Dcp-1. To examine the Bruce/Dcp-1 interaction under more physiological conditions, we performed additional in vivo experiments. Using the mStayGold::V5-tag knock-in allele of Bruce, we overexpressed Dcp1::VENUS using WP-Gal4 driver to induce Dcp-1 activation and assessed whether endogenously expressed full-length Bruce associates with Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates selectively with the cleaved, active form of Dcp-1 in vivo, thereby supporting the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      FK506-binding proteins (FKBPs) are a conserved group of proteins known to bind FK506, an immunosuppressive drug. FKBPs contain FK domains, which correspond to peptidyl cis-trans isomerase (PPIase) domains. Drosophila Fkbp59 is an orthologue of the mammalian FKBP4 and FKBP5, both of which possess a C-terminal tetratricopeptide repeat (TPR) domain that functions independently of the PPIase domain by mediating protein-protein interactions. The mammalian orthologues of Drosophila Fkbp59 function as Hsp90 co-chaperones (GharteyKwansah et al., 2018). Importantly, loss of Fkbp59 results in pupal lethality (Iki et al., 2020), which precludes further mechanistic analysis on Dcp-1 activation using adult wing phenotypes. To date, the involvement of Fkbp59 in caspase regulation has not been reported. Given that Fkbp59 functions as a co-chaperone, it may facilitate Dcp-1 activation by promoting proper folding, stability, or subcellular positioning of Dcp-1 or its regulatory factors. Importantly, Dcp1 proximal proteins are enriched in chaperone-related factors, including CCT2, CCT8, Droj2, CG16817, Fkbp59, Sgt1, and nudC; seven out of sixteen identified proximal proteins are chaperone-related. These observations suggest that Dcp-1 activity may be regulated by chaperone proteins or that Dcp-1 activity may be spatially restricted to regions enriched in chaperone machinery. Further analysis of the relationship between Dcp-1 activity and chaperone-related proteins will be important to elucidate the mechanisms and functions underlying non-lethal Dcp1 activation.

      (4) Figure 3: Autophagy-related factors

      Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      We thank the reviewer for the comment. To address whether knockdown of Debcl, Buffy, Atg2, or Atg8a independently affects wing development, we performed RNAi-mediated knockdown of each gene using the WP-Gal4 driver in the absence of Dcp-1 overexpression. Under these conditions, knockdown of Debcl, Buffy, Atg2, or Atg8a did not cause any detectable defects in wing morphology (new Figure 3 – figure supplement 1A), indicating that these autophagy-related factors specifically function to suppress Dcp-1-mediated cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (5) Evidence for canonical autophagy

      The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      We thank the reviewer for the constructive comment regarding the involvement of canonical autophagy. To further strengthen the evidence that autophagy is required for Dcp-1 activation, we examined additional core autophagy-related genes that function at distinct steps of the autophagy process, in addition to the previously tested Atg2, which mediates autophagosomal membrane expansion, and Atg8a, a core component directly associated with autophagosomal membranes. Specifically, we performed knockdown of genes including FIP200/Atg17, which is required for the initiation of autophagosome formation; Atg9, which is required for autophagosomal membrane nucleation; Atg5, which is required for autophagosomal membrane expansion through Atg12-Atg5-Atg16 ubiquitin-like conjugation system; and Stx17, which is required for autophagosome-lysosome fusion (Umargamwala et al., 2024). Because Atg8b is known to be specifically expressed in the male germline and is dispensable for autophagy, at least in fat body cells (Jipa et al., 2021), we did not further examine Atg8b in wing imaginal discs. Using WPGal4 driver, knockdown of each of these genes significantly suppressed Dcp-1-induced wing ablation phenotype (new Figure 3 – figure supplement 1C), supporting a requirement for canonical autophagy components across multiple stages of autophagosome biogenesis in Dcp-1 activation. Importantly, knockdown of these autophagy-related genes alone did not affect wing morphology in the absence of Dcp-1 overexpression (new Figure 3 – figure supplement 1B), as observed previously for Atg2 and Atg8a, suggesting the suppressive effects are specific to Dcp-1 overexpression-dependent cell death. Together, these results indicate that inhibition of autophagy at any of several key steps can suppress Dcp-1-dependent cell death, demonstrating that intact canonical autophagy is required for Dcp-1 activation. In addition, to directly assess autophagy at the cellular level, we monitored autophagosome formation using mCherry::Atg8a reporter. Upon overexpression of Dcp-1::VENUS in the wing pouch region, we observed a clear accumulation of Atg8a-positive puncta in wing imaginal discs (new Figure 3C), demonstrating that Dcp-1 overexpression induces autophagy in vivo. Together, these results provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and is required for Dcp-1-dependent cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (6) Figures 4-5: Functional consequences

      It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      We thank the reviewer for the suggestion regarding the functional consequences of Synr, Debcl, and Buffy on wing size. As requested, we knocked down Debcl or Buffy using WP-Gal4 driver and found that this led to reduced wing size (new Figure 5 – figure supplement 1A), indicating that endogenous Debcl and Buffy promote wing growth potentially through regulating endogenous Dcp-1 activity. We have included the corresponding data in the revised manuscript. Because Synr RNAi did not show any detectable effect on the Dcp-1 overexpression-induced phenotype (Figure 3A, B), we did not further examine the effect of Synr knockdown on wing development alone. Overexpression of Synr was not examined in this study. However, Synr overexpression has previously been reported to induce cell death in wing imaginal discs, resulting in malformed adult wings (Ikegawa et al., 2023), suggesting that increased Synr expression is likely to have deleterious rather than growth-promoting effects. Because Debcl and Buffy are both required for Synr-induced cell death, overexpression of Debcl or Buffy may lead to similar phenotypes. Therefore, we did not test Debcl or Buffy overexpression in the wing imaginal discs.

      (7) Terminology and interpretation of cell death

      Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of nonapoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis; there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      We thank the reviewer for the important comment on terminology and interpretation of the cell death phenotype. As explained in our response to comment #1, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. At the same time, as explained in our response to comment #5, we provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and promotes Dcp-1 activation. We recognized that the phrase “Dcp1 activity-regulating alternative apoptosis signaling pathway” used in Figure 5L could be misleading, as it may imply the existence of an “alternative apoptosis”. To avoid this confusion, we have revised the figure legend to read “autophagy-facilitated alternative caspase activation pathway.”

      Reviewer #3 (Recommendations for the authors):

      Figure 1c should be annotated more clearly so that it is evident that the images shown are grouped by genotype.

      We thank the reviewer for the helpful suggestion. We have added lines to Figure 1C to improve clarity by indicating that the images are grouped by genotype.

      References

      Choutka C, DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2017. Hsp83 loss suppresses proteasomal activity resulting in an upregulation of caspase-dependent compensatory autophagy. Autophagy 13:1573–1589.

      Denton D, Shravage B, Simin R, Mills K, Berry DL, Baehrecke EH, Kumar S. 2009. Autophagy, not apoptosis, is essential for midgut cell death in Drosophila. Curr Biol 19:1741–1746.

      DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2014. The Drosophila effector caspase Dcp-1 regulates mitochondrial dynamics and autophagic flux via SesB. J Cell Biol 205:477–492.

      Domingues C, Ryoo HD. 2012. Drosophila BRUCE inhibits apoptosis through non-lysine ubiquitination of the IAP-antagonist REAPER. Cell Death Differ 19:470–477.

      Ghartey-Kwansah G, Li Z, Feng R, Wang L, Zhou X, Chen FZ, Xu MM, Jones O, Mu Y, Chen S, Bryant J, Isaacs WB, Ma J, Xu X. 2018. Comparative analysis of FKBP family protein: evaluation, structure, and function in mammals and Drosophila melanogaster. BMC Dev Biol 18:7.

      Ikegawa Y, Combet C, Groussin M, Navratil V, Safar-Remali S, Shiota T, Aouacheria A, Yoo SK. 2023. Evidence for existence of an apoptosis-inducing BH3-only protein, sayonara, in Drosophila. EMBO J 42:e110454.

      Iki T, Takami M, Kai T. 2020. Modulation of Ago2 loading by Cyclophilin 40 endows a unique repertoire of functional miRNAs during sperm maturation in Drosophila. Cell Rep 33:108380.

      Jipa A, Vedelek V, Merényi Z, Ürmösi A, Takáts S, Kovács AL, Horváth GV, Sinka R, Juhász G. 2021. Analysis of Drosophila Atg8 proteins reveals multiple lipidation-independent roles. Autophagy 17:2565–2575.

      Kondo S, Senoo-Matsuda N, Hiromi Y, Miura M. 2006. DRONC coordinates cell death and compensatory proliferation. Mol Cell Biol 26:7258–7268.

      Kornbluth S, White K. 2005. Apoptosis in Drosophila: neither fish nor fowl (nor man, nor worm). J Cell Sci 118:1779–1787.

      Leulier F, Ribeiro PS, Palmer E, Tenev T, Takahashi K, Robertson D, Zachariou A, Pichaud F, Ueda R, Meier P. 2006. Systematic in vivo RNAi analysis of putative components of the Drosophila cell death machinery. Cell Death Differ 13:1663–1674.

      Ryoo HD, Baehrecke EH. 2010. Distinct death mechanisms in Drosophila development. Curr Opin Cell Biol 22:889–895.

      Schott S, Ambrosini A, Barbaste A, Benassayag C, Gracia M, Proag A, Rayer M, Monier B, Suzanne M. 2017. A fluorescent toolkit for spatiotemporal tracking of apoptotic cells in living Drosophila tissues. Development 144:3840–3846.

      Shinoda N, Hanawa N, Chihara T, Koto A, Miura M. 2019. Dronc-independent basal executioner caspase activity sustains Drosophila imaginal tissue growth. Proc Natl Acad Sci U S A 116:20539–20544.

      Tenev T, Zachariou A, Wilson R, Ditzel M, Meier P. 2005. IAPs are functionally non-equivalent and regulate effector caspases through distinct mechanisms. Nat Cell Biol 7:70–77.

      Umargamwala R, Manning J, Dorstyn L, Denton D, Kumar S. 2024. Understanding developmental cell death using Drosophila as a model system. Cells 13:347.

      Vernooy SY, Chow V, Su J, Verbrugghe K, Yang J, Cole S, Olson MR, Hay BA. 2002. Drosophila Bruce can potently suppress Rpr- and Grim-dependent but not Hid-dependent cell death. Curr Biol 12:1164–1168.

      Williams DW, Kondo S, Krzyzanowska A, Hiromi Y, Truman JW. 2006. Local caspase activity directs engulfment of dendrites during pruning. Nat Neurosci 9:1234–1236.

      Zhang J, Wang X, Cui W, Wang W, Zhang H, Liu L, Zhang Z, Li Z, Ying G, Zhang N, Li B. 2013. Visualization of caspase-3-like activity in cells using a genetically encoded fluorescent biosensor activated by protein cleavage. Nat Commun 4:2157.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Thank you for reviewing our manuscript and for your constructive comments. Our point-to-point responses are shown below.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

      We thank you for these helpful comments. Based on your suggestions, we revised the presentation of error bars in Fig. 2G and in other panels showing G–V or F–V relationships (Figs. 2J, 3F, 3N, 4C, 4F, 4I, 4L, and 5E; Supplementary Fig. 5D) to make the SEM bars clearer. We also added statistical comparisons of V<sub>1/2</sub> values among the three lamprey KCNQ1 orthologs in Fig. 2G and among truncation-series constructs in Figs. 3F and 3N using one-way ANOVA followed by Tukey–Kramer multiple-comparison tests. For Fig. 4, we added statistical comparisons between KCNQ1 species isoforms expressed with or without PmKCNE0 (Figs. 4C, 4F, and 4I), and between PmKCNQ1 expressed alone or with human KCNE1 or KCNE3 (Fig. 4L), using unpaired two-tailed Welch’s t-tests. These statistical comparisons are included in Supplementary Table 1.

      Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

      We thank you for the positive assessment of our work and for the constructive suggestions. Our point-by-point responses are provided below.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What is the physiological role of this KCNE0 and lamprey KCNQ1 in the lamprey species? While the authors mention that the physiological roles of KCNE0 are the next focus, it is preferable to discuss some of the potential functional significance of this newly characterised KCNE.

      We thank you for this helpful suggestion. We agree that discussing the potential physiological roles of KCNQ1–KCNE0 complexes in lamprey strengthens the manuscript. We have therefore expanded the Discussion to raise the possibility that, given the broad tissue distribution of kcne0 transcripts and the ability of KCNE0 to render lamprey KCNQ1 constitutively active, KCNQ1–KCNE0 complexes may contribute to general ion homeostasis, potentially analogous to the epithelial K<sup>+</sup> recycling function of mammalian KCNQ1–KCNE3, rather than to the highly specialized KCNQ1–KCNE1 function in the mammalian heart and inner ear (page 14, lines 263–267).

      (2) Human KCNQ1 has 676 amino acids, but LcKCNQ1 contains just 507 amino acids. The species-specific regulatory effects of KCNE0 may not only be attributed to KCNE0 itself but might also be influenced by the species of KCNQ1. Some discussion on this possibility will be helpful.

      We thank you for raising this important point. We agree that the species-specific regulatory effects observed in our cross-species pairing experiments are unlikely to be determined by KCNE0 alone and may also be influenced by species-specific features of the KCNQ1 α-subunit. To address this point, we expanded the Discussion to note that the KCNQ1 proteins used in this study vary in amino-acid length, largely reflecting differences in the cytoplasmic C-terminal region, which may affect KCNQ1–KCNE compatibility and thereby influence channel gating and coupling to KCNE subunits (pages 12–13, lines 227–238). We also updated Supplementary Fig. 4 to include LrKCNQ1 and LcKCNQ1, and clarified that the LcKCNQ1 construct used in this study encodes 644 amino acids.

      (3) In Supplementary Figure 3, bands corresponding to LcKCNQ1 (507 amino acids, Supplementary Figure 1) were not seen.

      We thank you for pointing out this potentially confusing point. Supplementary Fig. 3 shows RT-PCR products amplified from tissue cDNA, not full-length amplification of the LcKCNQ1 ORF. As stated in the Methods section (pages 19–20, lines 374–399), the primers used for RT-PCR in Fig. 1F and Supplementary Fig. 3 were different from those used for cloning the full-length LcKCNQ1 cDNA shown in Supplementary Fig. 1. Therefore, a band corresponding to the full-length LcKCNQ1 coding sequence was not expected in Supplementary Fig. 3.

      To clarify this point, we revised the figure legends and indicated the RT-PCR primer-binding sites with orange arrows in Supplementary Figs. 1 and 2.

    1. Author response:

      We thank the reviewers for their constructive and careful assessment of our manuscript. We are encouraged that both reviewers recognised the value of the empirical contribution: the semantic advantage is robust across experiments, the two new experiments are pre-registered, and the drift-diffusion modelling provides an informative decomposition of behavioural performance. At the same time, both reviewers raise an important and convergent point: the manuscript currently places too much interpretive weight on non-decision time and sometimes moves too quickly from decision-model parameters to claims about the representational format of working memory.

      We agree that this aspect of the manuscript should be revised. In the next version, we will substantially soften claims about adaptive reformatting and long-term-memory-like formats. We will instead frame the central contribution more precisely: semantic-category judgements show a reliable advantage at stages preceding evidence accumulation and this advantage is modulated by attentional prioritisation and interference during the maintenance interval. Our data constrain the dynamics with which different kinds of information become available for WM-guided decisions, but they do not, on their own, provide a direct measure of representational format. This hypothesis should be tested in future experiments.

      At the same time, we think the data provide stronger constraints on alternative explanations than the current manuscript makes clear. The reviewers correctly note that non-decision time is not a pure retrieval parameter, as we also note in the discussion. It can include probe encoding, response preparation, motor execution, and other processes. We will therefore avoid more explicitly equating NDT directly with retrieval latency. However, many of the alternatives raised by the reviewers, such as easier question reading or simpler response mapping for semantic probes, predict a relatively fixed semantic–perceptual offset. In our experiments, the probes and response mappings are held constant across attentional conditions, while the semantic NDT advantage changes as a function of whether the relevant item can be prioritised in advance or must be selected/reactivated at test. We will restructure the Results and Discussion to make these condition × feature interactions central to the argument.

      We will also clarify the logic of Experiment 1. We agree with Reviewer 1 that a valid retro-cue likely triggers retrieval or reactivation of the cued item. Our original phrasing, which described the valid-cue condition as reducing retrieval demands, was imprecise. The critical manipulation is better described as shifting item prioritisation/retrieval earlier in the trial. Under valid cueing, the relevant item can be prioritised before the probe appears, whereas under neutral cueing, item selection and access must occur after probe onset. We will rewrite this section accordingly.

      We will also clarify our operationalisation of semantic and perceptual categories. The present contrast is specifically between semantic category information (animate versus inanimate) and perceptual-format information (photograph versus drawing). We agree that the perceptual judgement is still categorical and does not measure fine-grained perceptual fidelity. We will therefore avoid broad claims about semantic structure or perceptual detail in general. However, as pointed in the manuscript, we believe the contrast remains meaningful: the two dimensions are orthogonal within the same stimuli, and previous work using the same feature space showed the opposite ordering during perception (Linde-Domingo et al., 2019), where perceptual-format information was available before semantic-category information. We will move this argument earlier in the manuscript and present it as converging evidence for dissociable access dynamics, while acknowledging that it does not by itself prove representational format.

      In response to the modelling concerns, we will expand the model-validation section. Specifically, we plan to add posterior predictive checks for the reported models, report model comparisons more transparently, clarify when more complex models do or do not provide practically meaningful improvements, and include sensitivity analyses using alternative parameterisations where identifiable.

      We will also make several methodological clarifications. First, because the reanalysis of Kerrén et al. (2022) forms a substantial part of the manuscript, we will add a fuller description of the original task in the main text, including how the probed item was indicated at test. Second, we will rewrite the unclear sentence describing pseudo-random stimulus selection in Experiment 1 and add a control analysis testing whether performance differs when the probed item belongs to the majority versus minority category within the trial. Third, we will clarify the stimulus repetition scheme and discuss possible long-term-memory contributions. Importantly, because semantic and perceptual probes are applied to the same items from the same trials, any repetition history or proactive-interference contribution is shared across the two probe types, although we agree that this should be discussed explicitly.

      Finally, we will revise the broader theoretical framing. We will remove or substantially qualify claims linking the present data directly to episodic memory and imagery. We will also integrate the recent literature suggested by Reviewer 2 on semantic structure, associative relations, long-term-memory contributions to working memory, and boundary conditions for semantic labelling effects. This will allow us to position the study as one piece of a broader literature on how semantic information influences WM performance, rather than as direct evidence for a general representational reformatting mechanism.

      In summary, the revised manuscript will make a narrower but stronger claim: semantic-category information shows a robust pre-accumulation advantage during WM-guided decisions, and this advantage is shaped by attentional prioritisation and interference during maintenance. We will present this as evidence about WM access dynamics and decision components, not as direct evidence that WM representations are transformed into long-term-memory-like formats.

    1. Author response:

      eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

      We would like to counter the final statement. We were very careful in our description of state 2, while we describe it as ‘active’ we clearly explain that the resolution of the reconstruction is not sufficient to define all the classical indicators of a kinase active state; however, the map is consistent with the active conformation, the complex is active in vitro, and the MEK1 variant used is the well known DD mutant that is constitutively active. While the A-loop is not observed, this is in agreement with many crystal structures of other DD mutants. We therefore decided to define this state as ‘active’ as the A-loop of ERK2 approaches the active site, the alpha-C helix has moved in and the A-loop of MEK1 no longer occludes the active site - to clarify the state we refer to the classically active confirmation as ‘fully active’. Our supporting data also show that the complex is highly dynamic during turnover, meaning we have captured MEK1 in a number of conformations on the landscape of an active state – we feel that rather than a limitation, this is an important observation in MAP2K studies. Finally, the determination of an 80 kDa complex by cryoEM to resolutions well below 4 Å is a huge technical achievement allowing the first snapshots of the MEK1-ERK2 complex to be visualised.

      Regarding the A-helix release – our observation is that the helix becomes less folded on binding of substrate. There are many studies, which we cite, that show that destabilising this helix leads to release of the catalytic machinery, see Mansour et al, 1996, Biochemistry, 35, 15529-15536 and Jindal et al 2017 J. Biol. Chem. 292, 18814-18820 for initial studies. Our observation shows that this is linked to substrate binding – a very relevant new insight that demonstrates the importance of this helix, in addition to many previous studies, but links unfolding to substrate recognition for the first time.

      For the mechanism of processive phosphorylation – it has been well established that both processive and distributive mechanisms exist. While the way that a distributive mechanism could work is obvious (complete dissociation of the two proteins), it has not been clear how a MAP2K can remain bound to its substrate and exchange nucleotides. While caution should be employed in interpreting our state 3 structure, it clearly shows what nucleotide exchange when bound to substrate can look like and that this low nucleotide affinity state is linked to disorder in the P-loop, the A-helix and substrate binding via the KIM. We would love to perform experiments that could demonstrate this but cannot at present think of an appropriate method – the reviewers did not suggest a route either.

      Finally, for the cancer-causing mutations – there are many studies demonstrating that the mutations lead to a destabilisation of the A-helix. Our study links this to substrate recognition. While this is inference, it seems justified to describe a link between substrate recognition, A-helix unfolding and disease mutations given the large body of literature describing these events.

      We are currently performing a series of in-cell activity assays that should strengthen our claims regarding the A-helix and other observations in the structure - the histidine interactions in particular.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

      We thank reviewer #1 for in-depth comments and analysis of our manuscript. However, there is a misunderstanding regarding HDX-MS and SAXS data. First, the HDX-MS data were performed on the ADP.AlF<sub>4</sub><sup>-</sup> inhibited complex - this was not made sufficiently clear in the text, and we will amend this. Secondly, we are not trying to support local flexibility observed in the SAXS data with the HDX data. The HDX data support the interactions observed in the cryoEM structure and demonstrate flexibility in the proline-rich loop, the ERK A-loop and unfolding of the MEK1 A-helix. The SAXS data demonstrate that when the transition state complex is formed, the complex is more compact but flexibility within the complex increases - as observed in the dimensionless Kratky plot, supporting our observations in the cryoEM maps. Therefore, the HDX data and SAXS data are separate observations. We thank reviewer #1 for all the comments and will rewrite the manuscript in order to increase clarity as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.

      - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

      We thank reviewer #2 for comments and thorough analysis of our manuscript. We agree that perhaps the abstract should be toned down in terms of claims of an active conformation even though we feel that the combination of the first structure of the MEK1-ERK2 complex combined with MD simulation studies clearly demonstrate how phosphoryl transfer will occur. We are now also performing in-cell assays to support our theory of A-helix regulation. For point 2 we based the order of events on data from Juyoux et al 2023 Science, 381, 1217-1225, where in long-timescale MD simulations and experimentally validated adaptive Markov state model simulations the KIM interaction was the last to dissociate after the alpha-G helix interaction. In our MD simulations of the MEK1-ERK2 complex, the interactions formed by the N-lobe were weaker than those formed by the alpha-G. Indeed, dissociation of the N-lobe was observed in various independent simulations, whereas dissociation of the alpha-G was observed only once.

      Assuming that, as in the MKK6-p38a complex, the association proceeds along the reverse of the dominant dissociation pathway, the simulations suggest that the KIM interaction forms first, followed by the alpha-G and finally the N-lobes. While alternative association pathways are possible, this interpretation is consistent with the MD and in line with the main association and dissociation pathway observed for the MKK6-p38a complex. This is additionally supported by the observation that there is no catalytic activity if the KIM is removed, demonstrating this as the first essential recruitment event. We will expand this section to include our arguments.

      For the second example, we feel this is rather unfair. The statement that if the His-His interaction is absent, catalysis will be prevented is only obvious if one knows about the His-His interaction - this is the first structure showing pathway-specific interactions in the variable loop regions of a MAPK. If it is obvious, why has no one described these residues as important before?

      We have taken on board the comments on the style of the manuscript and will make significant changes as suggested by both reviewers.

    1. Author response:

      Reviewer #1

      (1) Causality and the role of DMS; a specific, falsifiable prediction. We agree that the paper should not leave the causal question implicit, and we will expand the Discussion to state a concrete prediction rather than a general claim of involvement. Briefly, if the accumulation signal we describe carries the animal's intended departure time, then suppressing DMS during patch occupancy should not simply shift exit times in one direction but should degrade their structure: exit-time variability should increase, and exit timing should lose its systematic dependence on reward-rate context and on the time of the most recent reward. A plausible alternative outcome is disengagement from the task altogether, which would be uninformative and would need to be controlled for. The result that would falsify our hypothesis is the one we will state explicitly: exit timing that remains as predictable, and as sensitive to context and reward history, under DMS suppression, as without it. We will also discuss the alternative the reviewer raises, that these signals are inherited from cortical inputs such as OFC or mPFC with DMS acting as a relay, and note that our data cannot presently distinguish this from a locally generated signal.

      (2) Behavior under a normative strategy. This is an interesting question and we will address it in the Discussion. Our expectation, which we will frame as a prediction rather than a result, is that an animal timing from patch entry rather than resetting at each reward would show accumulation that begins at entry and proceeds to a context-dependent threshold at the reward-rate-optimal time, rather than the reward-triggered resets we observe. In this view, the reset structure of the neural signal is a reflection of the behavioral policy rather than a property of the region. We will make clear that this is a testable prediction that our current dataset does not address.

      (3) Do these neurons also track time at the context port? We intend to answer with new analysis, and it converges with Reviewer #2's suggested analysis (below), so we treat the two together there.

      Reviewer #2

      (a) Timing versus movement initiation: activity at the context port. We take this to be a central interpretational concern. We will examine whether the step-like DMS activity extends to the context port, testing the interval between the final context-port reward and departure for the same step-like transitions and accumulation we observe in the time-investment port. We will apply the same comparison to the context-port inter-reward intervals, which addresses Reviewer #1's third point about whether these neurons also track time at the context port.

      We want to flag one feature of the task that bears on how the outcome should be read. The context port is not a timing-free epoch. Its four rewards are delivered at predictable, regularly spaced intervals, and the interval between the final reward and the animal's departure is self-timed. Departure from the context port is therefore also a self-timed action, and observing accumulation there would not by itself indicate that the signal reflects movement initiation rather than timing. What the comparison can inform is whether the accumulation is specific to a decision about when to disengage from a depleting resource, or is a more general feature of self-timed departures. This is a meaningful distinction either way, and one we will report and interpret whichever direction the result falls.

      (b) Operational definition of leaving; the exit-to-entry latency. We agree these details are necessary for interpretation and their absence is our omission. Exit is the final withdrawal from the time-investment port preceding the next context-port visit, and we will make that clear in the revised methods. Mice can and occasionally do re-enter the time-investment port without an intervening context-port visit (especially early in training). Such re-entries are not counted as exits. We will also examine whether the latency between time-investment-port exit and context-port entry relates to the accumulation slope on the corresponding interval, to test whether the neural signal relates to the decision or to the execution of the transition.

      (c) Independence of interval-level fits from the session-level model. This is a fair characterization of the procedure, and we accept the distinction the reviewer draws between a detector of consistency with a session-defined step model and an unbiased test of discrete versus continuous dynamics. We will address it with a held-out validation: estimating each unit's state parameters and transition time from one half of its intervals and testing whether the transition times recovered from the withheld half agree. The discrete-versus-continuous comparison is addressed directly by the ramp simulations under Reviewer #3's point (5) below.

      (d) The inclusion threshold for the accumulation analysis. We will add a supplementary figure reporting the total number of recorded units per session and the fraction classified as step-like, so that the seven-unit criterion can be evaluated against the recorded population rather than in the abstract. Yield varied substantially across sessions, from a handful of units to roughly one hundred, and we will show this distribution directly. We will also report the accumulation results across a range of inclusion thresholds spanning approximately five to eight simultaneously recorded step-like units, so that readers can assess sensitivity to the choice.

      Reviewer #3

      (1) Specification of the task. We agree the task description was insufficient, and we will correct this at the front of the Results and in the Methods. The time-investment port operates on a poke-and-hold basis: the mouse maintains its head in the port and rewards are delivered stochastically over time for as long as it remains, with no requirement to withdraw and re-poke. Reward delivery follows an exponentially decaying rate in time since port entry, with a time constant of eight seconds, integrating to an expected eight rewards of one microliter each (8 µL total) for indefinite occupancy; we will give the explicit function. The reward probability function in the time-investment port is identical across blocks. The high- and low-reward-rate contexts are properties of the context port alone (four rewards over five seconds versus four rewards over ten seconds), and we will make this contrast unambiguous, since it is the manipulation on which the design rests.

      (2) Derivation of the optimum, Figure 1h, and a formal test of overstaying. We appreciate this comment. The optimal residence time in our task is not obtained by the classical Charnov tangent construction. It is computed by explicit maximization of the overall reward rate over the full cycle, following the framework in Sutlief et al. (2025) and shown graphically in Figure 1h. We will expand the legend of Figure 1h so its conventions are stated explicitly, give the reward-rate-maximizing derivation as an explicit equation in the Methods, and reframe the surrounding text around reward-rate maximization as the normative principle, with MVT identified as the special case it is. We will also add the formal statistical comparison of observed residence times against the computed optimum, which the reviewer correctly notes was asserted rather than tested.

      (3) Interpretation of SVM coefficients with correlated predictors. We accept this criticism. We will not rest the ordering of predictors on coefficient magnitudes alone. We will add a variable inclusion-and-ablation analysis, reporting cross-validated classification performance as each predictor is added to and removed from the model, so that the contribution of time since last reward is assessed by its effect on performance rather than by its normalized weight. We will additionally add a complementary analysis of the leave hazard that estimates the contribution of each variable without requiring the classification framework.

      (4) The five-second post-exit window. The reviewer is right that we did not explain this choice, and right that it rests on an assumption. Our reasoning was that a single exit moment per trial leaves the decision boundary badly under-constrained, and that treating the moments immediately following an exit as conditions under which the animal would also have left is licensed by the fact that within-patch reward rate declines monotonically with time, so conditions in the counterfactual continued visit would have been strictly less favorable than those already rejected. We accept that this is an assumption rather than an observation and will state it as such. We will also report the analysis across a range of window durations so that the independence of the result on this choice is visible. To the reviewer's specific questions: both time since entry and time since last reward are computed with respect to the time-investment port throughout, and we will state this explicitly.

      (5) False positive rate against ramps and random walks. We agree this is the more informative validation, and that continuous ramping is the most relevant alternative to ours. We will generate simulated units with continuous ramp-like rate changes, matched to the firing rates and interval structure of our recorded units, and pass them through the identical classification pipeline to obtain false positive rates comparable to the Poisson analysis already reported. We will retain the flat-rate Poisson simulation and present the ramp results as additional panels of the same supplement. We will also explore random-walk dynamics. Together with the held-out validation of transition times described under Reviewer #2(c), this converts the step characterization from a single-null validation into a comparison against the relevant continuous alternatives.

      (6) Specificity of the accumulation signal versus a generic increasing function. This is a valuable challenge and we will address it directly. We will implement the trial-mismatch control the reviewer proposes, randomly reassigning accumulation trajectories to reward-to-exit intervals within session and showing the extent to which predictive performance degrades relative to the correctly matched case.

      We would also note two features of the existing results that speak to this concern, and which we will bring forward in the revision because we did not make them salient enough. First, a signal that merely tracked elapsed time would be expected to shift its starting level as well as its rate across trials with different exit times; instead, the accumulation slope is strongly related to exit time (mean r = -0.551) while the intercept is not (mean r = 0.013), indicating a variable rate from a stable origin. Second, and more to the point, the accumulation arrives at a common level at the moment of exit whether the animal leaves early or late. The rate of accumulation shifts with the animal’s policy on that trial such that the threshold is met at the intended time. 

      What makes this predictor non-trivial is its trial-by-trial correspondence to behavior, not simply that it increases over time. We will make this argument explicitly alongside the shuffle control.

      Summary

      To summarize the planned additions: (i) analysis of step-like activity and its accumulation at the context port, with the interpretive caveat noted above; (ii) validation of step detection against ramping alternatives, together with held-out estimation of transition times; (iii) a trial-mismatch control for the specificity of the accumulation predictor; (iv) an inclusion-and-ablation analysis of the behavioral predictors and a complementary hazard model of the leave decision; (v) reporting of unit yield, step-like fraction, and sensitivity of the accumulation results to the inclusion threshold; (vi) analysis of the exit-to-context-entry latency in relation to the neural signal; (vii) a formal statistical test of overstaying relative to the computed optimum; and (viii) substantial clarification of the task specification, the operational definition of exit, and the derivation of the optimal residence time, including an expanded Figure 1h legend.

      Our aim in the revision is to meet the specificity concern raised in the assessment as directly as the existing data allow, and we hope the revised manuscript will warrant reconsideration of the strength-of-evidence characterization.

      We are grateful to the reviewers for the care evident in their reports, and to you both for handling the manuscript.

      References

      Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136.

      Sutlief, E., Walters, C., Marton, T., & Hussain Shuler, M. G. (2025). The value of initiating a pursuit in temporal decision-making. eLife. https://doi.org/10.7554/eLife.99957.2.

    1. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Recommendations for the authors):

      (1) Interpretation of Syllable-Tracking in the RND Condition:

      The finding of greater syllable-tracking in the LL group compared to the HL group in the RND condition warrants cautious interpretation. Currently, there is no direct statistical evidence demonstrating greater PLV at 4 Hz in the Structured versus Random conditions for either group; readers must infer this solely from numeric differences in Figure S5 B and D. Therefore, while the interpretation on Page 14 (Lines 443-446) "successful segmentation may enhance syllable tracking via top-down predictions of the next syllable" is an interesting speculation, it feels somewhat far-reaching. Additionally, the authors should discuss whether this upregulated syllable tracking in the structured condition (which is specific to the HL group) represents an adaptive or maladaptive response.

      The reviewer correctly highlights the lack of direct  comparison between conditions (RND versus STR). We tempered our claims in the cited paragraph and insisted on the speculative nature of this part of the discussion. We also clarified that, to us, it may represent an adaptive compensatory strategy:

      Page 14, line 441: “Interestingly, our supplementary analyses (Supplementary Material Figure S4-5) suggest that syllable entrainment may be differentially affected in HL versus LL infants, depending on the statistical structure of the input stream (RND versus STR). However, as our experiment was not explicitly designed to test stream effects, these results should be interpreted with caution. Future studies could explore how successful segmentation may enhance syllable tracking via top-down predictions of the next syllable in both LL and HL infants. If confirmed, such a mechanism may improve alignment to syllable onsets, potentially constituting a compensatory process allowed by preserved segmentation abilities.”

      (2) Preservation of Statistical Learning in HL Infants:

      The text added on Pages 17-18 (Lines 562-566) regarding a "heightened dependence on bottom-up mechanisms (in autism)" does not appear to be supported by the data or by theories of implicit statistical learning. Because greater syllable-level entrainment was observed in the LL group than the HL group across both the random and structured conditions, the data actually point toward impaired bottom-up processes. Furthermore, implicit statistical learning typically involves an interplay of both bottom-up and top-down mechanisms; the implicit nature of a task does not guarantee a strictly bottom-up process. Consequently, this interpretation is not entirely convincing.

      We agree with the reviewer that the concepts of “top-down” and “bottom-up” were not fully appropriate to support our point in the cited paragraph. We should have used the concepts of implicit versus explicit learning instead, in line with previous literature suggesting increased reliance on preserved implicit learning in autism to compensate for altered explicit processes. The paragraph was slightly modified.

      Page 18, line 564: “According to these studies, autistic impairments in explicit attentional processes, such as social orienting - which are critical for bootstrapping language acquisition (70) - may result in a heightened dependence on implicit mechanisms, including statistical learning. As previously discussed, preserved word segmentation abilities may further compensate for alterations in lower-level implicit processes, such as syllable tracking.”

      Reviewer #2 (Recommendations for the authors):

      Potential typo on line 199 - I think an apostrophe is needed here.<br /> Potential typo on line 255 - do you mean Central electrodes?

      We addressed the typos spotted by reviewer.

      Line 199: variables’

      Line 255: Centro-frontal electrodes


      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability - particularly in the context of neurodevelopmental conditions - remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables, and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Weaknesses:

      (1) Clarifying longitudinal vs. concurrent associations

      Because the current analytical approach incorporates all time points, including the final visit, it is challenging to determine to what extent the brain-language associations are driven by longitudinal relationships vs. concurrent correlations at the last time point. This does not undermine the main findings, but clarifying this issue could significantly enhance the impact of the individual-differences results. If feasible, the authors might consider (a) showing that a model excluding the final visit still predicts verbal outcomes at the last visit in a similar way, or (b) more explicitly acknowledging in the discussion that the observed associations may be partly or largely driven by concurrent correlations. Either approach would help readers interpret the strength and nature of the longitudinal claims.

      We thank the reviewer for this insightful comment. We agree that distinguishing between longitudinal predictive power and concurrent correlations at the final visit is crucial for clarifying the nature of these brain-language associations. Following the reviewer’s suggestion (a), we re-ran the two critical Partial Least Squares Correlation (PLS-c) analyses by excluding all EEG and behavioral data from the final 18–21 month visit (n = 54 recordings kept) to test whether earlier trajectories still predict the final verbal outcome.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2C–D): The PLS-c restricted to the 3- to 15-month visits still identified a single significant component (p=.001, r=.56, 64.0% explained covariance, Figure 2 -figure supplement 3). Bootstrap ratios (BSR) were: contrast (low vs. high autism likelihood) 4.1; mean age −1.3; contrast*mean-age −2.1; delta-age 10.5; contrast*delta-age 2.0; age<sup>2</sup> −6.0; contrast*age<sup>2</sup> −5.8; and notably verbal outcome 7.1; contrast*verbal-outcome −6.3.

      The latent component and its spatial electrode configuration remain highly consistent with the original analysis (Figure 2C–D). This confirms that excluding the final visit preserves the model’s predictive validity: lower syllable entrainment correlates with poorer verbal outcomes at 18–21 months, particularly in the high-likelihood group.

      (2) Late evoked response to novel words (original analysis on Figure 6): The PLS-c analysis on the ERP late time window (1500–3000 ms), excluding the final visit, also revealed one significant component (p=.002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (part-word versus word) 14.5; mean age -5.2; contrast*mean-age 6.2; delta-age -0.8; contrast*delta-age 3.5; age<sup>2</sup> -0.8; contrast*age<sup>2</sup> -12.1; verbal-outcome 20.1; contrast*verbal-outcome -7.2. The latent component closely mirrored the original analysis (Figure 6), with frontal electrodes contributing negatively and posterior electrodes positively. Minor divergences in age-related parameter contributions were observed, likely due to the absence of 18-21 month timepoints, which previously contributed to the convex/concave shapes of the group age trajectories in figure 6B (left panel).

      Crucially, both models (with and without the final visit) positively predicted verbal outcomes (Figure 6B, right panel, and Figure 6 -figure supplement 1B). However, excluding the final visit reversed the direction of the group*verbal-outcome interaction (from 3 to -7.2): This indicates that after ruling out cross-sectional correlations at 18–21 months, the early predictive value of the late ERP to word novelty is more prominently observed in high-likelihood infants, suggesting that the original result was influenced by concurrent cross-sectional correlations at the final visit. This aligns with the syllable entrainment findings (Figure 2 -figure supplement 3), as both 4 Hz neural tracking and late ERP responses to novelty predominantly predict verbal outcomes in infants at high likelihood for autism.

      We reported these supplementary analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 3 and Figure 6 -figure supplement 1. In general, most of figures that were present in Supplementary materials were moved as figure supplements to enhance readability.

      Page 8 lines 234-240 (pages and lines refer to the reviewed uploaded manuscript): “To rule out the possibility that the association between syllable entrainment and verbal outcome was driven by concurrent measures taken at 18–21 months, we re-ran the PLS-c analysis excluding EEG data from the final visit (n = 54 recordings kept). The resulting latent component remain significant (p = .001) and showed contributions from behavioral and EEG variables that were highly similar to those observed in the previous analysis, with a verbal outcome BSR of 7.1 and a group’verbal-outcome interaction BSR of −6.3 (Figure 2 -figure supplement 3).”

      Page 12 lines 387-394: “As we did for neural entrainment to syllables, we conducted a new analysis on late ERP to word novelty, excluding EEG data from the final visit. This PLS-c yielded one significant latent component (p = .002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1) with globally similar EEG parameter contributions and age trajectory modelling. Verbal outcome still significantly contributed to the latent component (BSR=20.1), with a negative verbal outcome*group interaction (BSR=-7.2). These results suggest that, after ruling out cross-sectional correlations at 18–21 months, the late ERP to word novelty predominantly predicts verbal outcomes in high-likelihood infants for autism.”

      Page 17 lines 547-548: “As with syllable entrainment, the late ERP to novel words primarily predicted verbal outcomes in high-likelihood (HL) infants.”

      Page 18 lines 588-590: “Likewise, the absence of a late ERP orientation response in HL participants may represent an early neural signature of altered attention to novelty that can be used both as a non-invasive predictor of language development and as a potential target for early intervention.’

      (2) Incorporating sleep status into longitudinal models

      Sleep status changes systematically across developmental stages in this cohort. Given that some of the papers cited to justify the paradigm also note limitations in speech entrainment and word segmentation during sleep or in patients with impaired consciousness, it would be helpful to account for sleep more directly. Including sleep status as a factor or covariate in the longitudinal models, or at least elaborating more fully on its potential role and limitations, would further strengthen the conclusions and reassure readers that these effects are not primarily driven by differences in sleep-wake state.

      The reviewer is highlighting here a limitation of our study design that comprised sleeping status that varied from one timepoint to another among participants. To rule out any confounding effect of wake status (coded as a binary variable: sleeping or awake during recording) on analyses comparing groups, a linear mixed-effect model with repeated measures was fitted finding no significant difference between high- and low-likelihood participants (p=.769, reported at page 20, lines 646-647). However, as rightly suggested by the reviewer, this doesn’t prevent from a sleep bias on age trajectories, especially given that sleeping status significantly decreases with age in our sample.

      Including sleep status as a covariate in our analyses, as suggested by the reviewer, would be difficult to implement in our PLS-c methods, since a categorical behavioral parameter that varies within participants is not possible in the models provided by myPLS toolbox.

      As an alternative option, we re-ran all analyses that explored the condition effect on the whole sample within the sleeping participants only (n=25 recordings) to confirm that the same age-trajectories of EEG parameters were highlighted. However, negative results should be interpreted with caution since the sample is small for such a multivariate approach, resulting in modest statistical power.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2A–B): The PLS-c identified one significant component (p <.001, r = .78, 85.1% explained covariance, Figure 2 -figure supplement 2 and Figure 3 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (4hz vs. adjacent frequencies) 30.3; mean age -2.9; contrast*mean-age -2.4; delta-age 3.8; contrast*delta-age 3.4; age<sup>2</sup> -1.1; contrast* −2.5. The spatial distribution of contributing electrodes globally matched that shown in Figure 2A. The high contrast BSR (30.3) confirms robust syllable entrainment in sleeping infants. Critically, the contrast*age<sup>2</sup> parameter contributed negatively to the latent component (BSR = −2.5), confirming that the convex age trajectory of syllabic entrainment (Figure 2B) is also present in the sleeping subsample.

      (2) Word entrainment (1.3 Hz) (original analysis on Figure 3A–B): The PLS-c identified one significant component (p <.001, r = .63, 37.3% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (1.3hz vs. adjacent frequencies) 24.9; mean age -7.5; contrast*mean-age -3.2; delta-age 2.9; contrast*delta-age 0.0; age<sup>2</sup> 1.5; contrast* 1.1. The spatial distribution of significant electrodes partially overlaps with the ones in the original analysis, primarily showing fronto-central positive contribution to the latent component. The high contrast BSR confirms a robust word entrainment in sleeping participants, in line with previous studies (e.g., Flò et al, Sci Rep, 2022). However, the lack of a significant contrast* age<sup>2</sup> suggests that the U-shape age trajectory illustrated on Figure 3 might be modulated by wakefulness or due to a lack of power in the present analysis. A non-significant trend towards a U-shape pattern with a 12-month nadir is visible in sleeping participants, but additional data from sleeping 18-21 months sleeping infants would be required to confirm or refute this trend.

      (3) Early evoked response to novel words (original analysis on Figure 4): The PLS-c analysis on the ERP early time window (0–1000 ms) in sleeping participants revealed no significant component. The absence of early response to word novelty in sleeping participant might account for the lack of response observed in the whole sample, illustrated on Figure 4. To test this hypothesis, we conducted the same PLS-c in awake participants (n=58 recordings), which also yielded no significant latent component. This suggests that the lack of a measurable early response to word novelty observed in the whole sample is consistent across both sleeping and awake infants, and not driven by any of the two subsamples.

      (4) Late evoked response to novel words (original analysis on Figure 5): The PLS-c analysis on the ERP late time window in sleeping participants revealed no significant component. This suggests that sleeping participants might present a reduced or even absent late response to novel words. Given this identified effect of sleep on late ERP response, we reran the PLS-c on the late ERP window using group as contrast (original analysis on figure 6), excluding the sleeping participants to avoid any confounds. This PLS-c revealed one significant component (p = .006, r = .68, 30.5% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (low versus high likelihood) 7.6; mean age -5.9; contrast*mean-age -3.8; delta-age 2.5; contrast*delta-age -3.5; age<sup>2</sup> -2.5; contrast*age<sup>2</sup> -4.9; verbal-outcome 8.5; contrast*verbal-outcome -1.0. Behavioral parameters contribute to this latent component with similar magnitude and polarity as in the original analysis. Electrode contributions are also highly consistent, with frontal negative and posterior positive contributions. This confirms that sleeping participants, despite their potentially reduced late response, did not significantly bias the results presented in Figure 6.

      Author response image 1.

      Late evoked response potential (ERP) to word novelty in awake participants. A. Design and brain saliences derived from the significant latent component. Brain topographies of bootstrap ratios (BSR) are displayed at 250ms intervals. Black dots indicate BSR > 2.3. B. Participants’ brain scores for part-word and word conditions, as a function of age (left panel) and verbal DQ (right panel). For details on brain scores, see Figure 6 -figure supplement 1. Linear fitting is used for illustrative purposes only. HL: high likelihood for autism; LL: low likelihood for autism.

      We reported these analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 2A and Figure 3 -figure supplement 1.

      Page 7, lines 218-224: “Because some infants were asleep during the recording session, particularly at younger ages, we performed a supplementary control analysis restricted to this sleeping subsample (n = 25 recordings, Figure 2 -figure supplement 2). This PLS-c also identified a significant latent component (p < .001, r = .78, 85.1% explained covariance), with a significant contrast effect (BSR = 30.3) and a significant negative contrast*age<sup>2</sup> interaction (BSR = −2.5). These findings confirm that the convex age trajectory observed in the main analysis remains present and observable even in sleeping infants.”

      Page 8, lines 252-259: “We further investigated word entrainment in sleeping participants (n=25), which yielded one significant latent component (p<.001, r=.63, 37.3% explained covariance, Figure 3 -figure supplement 1). Centro-frontal electrode contributed to this component, with a high contrast BSR (24.9), confirming a similar word entrainment pattern in the sleeping subsample. The contrast*age<sup>2</sup> was also positive but not significant (1.1), suggesting a trend toward a U-shape age trajectory with a 12-month nadir in sleeping infants. Additional 18-21 month recording would be required to confirm this trend.”

      Page 11, lines 347-349: “The same PLS-c, conducted separately in sleeping (n=25) and awake subsamples (n = 58), yielded no significant latent component, indicating a consistent absence of early response to word novelty in both sleeping and awake infants.”

      Page 11-12 lines 368-370: “The same PLS-c in the sleeping subsample yielded no significant latent component, suggesting that sleep may reduce or even abolish the late response to word novelty.”

      Page 12 lines 382-385: “Given that no late response was detected in sleeping participants, we re-ran the PLS-c analysis using group as a contrast in the awake subsample (n=58). This yielded one significant latent component (p=.006, r=.68, 30.5% explained covariance), with behavioral and electrode contributions highly overlapping with those in Figure 6.”

      Page 18 lines 590-592: “This potential biomarker might nevertheless be modulated by participants’ sleep status, warranting careful consideration of vigilance state in future studies.”

      (3) Use of PLS-c and potential group × condition interactions

      I am relatively new to PLS-c. One question that arose is whether PLS-c could be extended to handle a two-way interaction between group and condition contrasts (STR vs. RND). If so, some of the more complex supplementary models testing developmental trajectories within each group (Page 8, Lines 258-265) might be more directly captured within a single, unified framework. Even a brief comment in the methods or discussion about the feasibility (or limitations) of modeling such interactions within PLS-c would be informative for readers and could streamline the analytic narrative.

      The reviewer raises a valid concern regarding the capacity of PLS-c to accommodate multi-way interactions among categorical and continuous variables. While PLS-c has no inherent theoretical constraints on the number of predictor terms (they can even exceed the sample size in number), practical limitations arise from model stability and interpretability when the ratio of predictors to sample size becomes excessive. As noted by Geladi and Kowalski (1986), exceeding ~10% of the sample size with predictors increases noise sensitivity and overfitting.

      In our study, the PLS-c analyses already reach this ~10% limit, with a maximum of nine predictors for a sample size of n=83. Attempting to integrate both group and condition as contrasts — along with necessary age parameters to account for developmental trajectories — would result in 12 predictors (or 15 if verbal outcome is included). Specifically, the model would require behavioral terms for Group, Condition, Group*Condition, Mean-age, Group*Mean-age, Condition*Mean-age, Delta-age, Group*Delta-age, Condition*Delta-age, Age<sup>2</sup>, Group*Age<sup>2</sup>, Condition*Age<sup>2</sup>, Verbal-outcome, Group*Verbal-outcome, and Condition*Verbal-outcome.

      Although a unified multivariate model capturing the complex dynamics at play in our sample is theoretically appealing, the substantial risk of overfitting precludes its feasibility. Therefore, we opted to use only one categorical predictor per PLS-c analysis to maintain model parsimony and reliability. However, a larger sample could overcome this limitation, allowing a stable and unified model of longitudinal EEG data that simultaneously captures age trajectories, group, clinical outcome, and condition.

      Reference:

      Geladi, P., & Kowalski, B. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/S0003-2670(00)82582-3

      We added the following comment in the method section:

      Page 24, lines 773-777: “We limited the number of behavioral variables to nine to mitigate noise sensitivity and overfitting risks associated with exceeding the 10% sample size threshold (Geladi & Kowalski, 1986). This limitation precluded the implementation of a single PLS-c model incorporating group, condition (STR vs. RND), age, and their interactions.”

      (4) STR-only analyses and the role of RND

      Page 8, Lines 241-245: This analysis is conducted only within the STR condition. The lack of group difference observed here appears consistent with the lack of group difference in word-level entrainment (Page 9, Lines 292-294), suggesting that HL and LL groups may not differ in statistical learning per se, but rather in syllabic-level entrainment. As a useful sanity check and potential extension, it might be informative to explore whether syllable-level entrainment in the RND condition differs between groups to a similar extent as in Figure 2C-D. In other work (e.g., adults vs. children; Moreau et al., 2022), group differences can be more pronounced for syllable-level than for word-level entrainment. Figure S6 seems to hint that a similar pattern may exist here. If feasible, including or briefly reporting such an analysis could help clarify the asymmetry between the two learning measures and further support the interpretation of syllabic-level differences.

      The reviewer points to the interesting pattern highlighted in supplementary figure S6, suggesting that group differences in syllabic entrainment might be modulated by the structure of the stream (STR versus RND). Such modulatory effect of stream structure on entrainment to syllables has been suggested by many studies, like Moreau et al (2022), as pointed by the reviewer, and seems at play in our sample, as illustrated on supplementary figure S5 (decline in the 4hz PLV that exceeds the size of confidence intervals, ~90 s after STR onset).

      Following the reviewer’s suggestion, we ran a PLS-c testing group effect on 4hz PLVs in each stream:

      (1) in the RND stream: the analysis yields one significant component (p<.001, r=.49, 52.7% explained covariance, Author response image 2A-B). Bootstrap ratios (BSR) are: contrast (low versus high likelihood) 6.9; mean age -1.0; contrast*mean-age 0.1; delta-age 11.8; contrast*delta-age 0.1; age<sup>2</sup> -4.9; contrast*age<sup>2</sup> 4.6; verbal-outcome 9.7; contrast*verbal-outcome -0.6. Interestingly, the model still highlights a strong link between syllable tracking and group, suggesting that RND also discriminate between HL and LL. However, RND syllable tracking doesn’t appear to be linked to group x verbal-outcome as we observed in Figure 2C-D.

      (2) In the STR stream, we obtained one significant latent component (p=.002, r=.51, 57.8% explained covariance, Author response image 2C-D). Bootstrap ratios (BSR) are: contrast 2.2; mean age -1.6; contrast*mean-age -1.1; delta-age 6.1; contrast*delta-age -0.7; age<sup>2</sup> -4.4; contrast*age<sup>2</sup> 0.5; verbal-outcome 8.2; contrast*verbal-outcome -7.3. Here, the strong association between syllable tracking and group x verbal-outcome is similar to the model presented in Figure 2C-D.

      Taken together, these results suggest that the apparent STR/RND dissociation illustrated in Figure S6 might primarily reflect a Group*Verbal-outcome divergence, with syllable tracking in the STR stream being related to verbal outcome mainly in high likelihood for autism.

      Author response image 2.

      Syllable entrainment within RND (A-B) and STR (C-D).

      These results were reported in the revised manuscript in the Result section (Time course of the entrainment along experiment subheader), implying a slight reframing of the result presentation of supplementary analysis S6. Author response image 2 was added in supplementary material as Figure S5.

      Page 10, lines 307-319: “The group, age and verbal outcome parameters were mainly correlated (BSR>2.3) with the neural entrainment occurring~90 seconds after the onset of the STR stream, coinciding with the time participants began tracking word boundaries (Supplementary material, S3). This result suggests that the group differences in syllable entrainment, as shown in Figure 2C-D, as their associations with verbal outcome, are modulated by the structure of the stream (STR versus RND). We ran one additional PLS-c for each stream separately, using group as contrast. In both streams, the PLS-c yielded a significant LC (p<.001 for RND and p=.002 for STR), with a positive group effect (BSR>2.3) in both LC (Supplementary material, S5). Most strikingly, the group*verbal outcome parameter reached significance exclusively within the STR latent component (BSR:-7.3). These results suggest that while syllable tracking is generally decreased in HL infants across both streams, its association with verbal outcome is prominently driven by the stream containing words (STR).”

      Page 14, lines 443-446: “This temporal overlap suggests that successful segmentation may enhance syllable tracking via top-down predictions of the next syllable, improving alignment to syllable onsets in LL infants as well as in HL with better verbal outcome.’

      (5) Multi-speaker input and voice perception (Page 15, Lines 475-483)

      The multi-speaker nature of the speech input is an interesting and ecologically relevant feature of the design, but it does add interpretive complexity. The literature on voice perception in autism is still mixed: for example, Boucher et al. (2000) reported no differences in voice recognition and discrimination between children with autism and language-matched non-autistic peers, whereas behavioral work in autistic adults suggests atypical voice perception (e.g., Schelinski et al., 2016; Lin et al., 2015). I found the current interpretation in this paragraph somewhat difficult to follow, partly because the data do not directly test how HL and LL infants integrate or suppress voice information. I think the authors could strengthen this section by slightly softening and clarifying the claims.

      We acknowledge the reviewer’s concern regarding the potential ambiguity in the cited paragraph. To address this, we have revised the text to explicitly clarify the aims of our study and its design. Furthermore, we now emphasize the speculative and post-hoc nature of the hypotheses and interpretations presented, thereby ensuring transparency regarding the limitations of our findings.

      Page 16 lines 520-530), as follows: “HL infants, on the other hand, did not show this transient disruption. In this group, word entrainment remained stable over time. To account for this unexpected finding, we followed up on the post-hoc hypothesis proposed above: a reduced sensitivity to social and vocal cues observed in HL infants may have spared segmentation abilities by limiting the interference introduced by speaker variability. If this post-hoc hypothesis holds true, LL and HL infants would differ not in their intrinsic ability to learn statistical regularities per se, but rather in how they integrate or suppress competing cues (such as speaker changes) during the segmentation process. It is important to note, however, that the present study was not designed to isolate and evaluate the specific impact of speaker changes on word segmentation. Consequently, this interpretation remains speculative, and additional research is required to further address this question.”

      (6) Asymmetry between EEG learning measures

      Page 16, Lines 502-507 touches on the asymmetry between the two EEG learning measures but leaves some questions for the reader. The presence of word recognition ERPs in the LL group suggests that a failure to suppress voice information during learning did not prevent successful word learning. At the same time, there is an interesting complementary pattern in the HL group, who show LL-like word-level entrainment but does not exhibit robust word recognition. Explicitly discussing this asymmetry - why HL infants might show relatively preserved word-level entrainment yet reduced word recognition ERPs, whereas LL infants show both - would enrich the theoretical contribution of the manuscript.

      We concur with the reviewer’s observation that our findings imply a theoretically significant double dissociation between HL and LL groups, specifically concerning the asymmetries between word-level neural entrainment and word recognition mechanisms. We believe this point was partly addressed in the subsequent paragraph, where we stated that “in contrast” to LL, HL infants “showed no clear ERP difference between novel and familiar triplets”, while “both groups showed similar word neural entrainment during learning”. We further explored potential explanations for this apparent dissociation, such as a possible deficit in novelty orientation that may be specific to HL infants and unrelated to statistical learning itself. We cited Liu et al (2023) as a reference showing the dissociation between mechanisms underlying implicit versus explicit traces of statistical learning. We acknowledge that we can discuss more in depth the potential preservation of statistical learning in HL infants. We have incorporated the following discussion in the reviewed manuscript, supported by relevant references:

      Pages 17-18, lines 562-566: “Interestingly, this dissociation between spared implicit versus impaired explicit statistical learning in autism has been previously discussed in the literature (Zwart et al, 2018, Kissine, 2021). According to these studies, autistic impairments in top-down attentional processes, such as social orienting — which are critical for bootstrapping language acquisition (Kuhl, 2007) — may result in a heightened dependence on bottom-up mechanisms, including implicit statistical learning.”

      References:

      Zwart, F.S., Vissers, C.T.W.M., Kessels, R.P.C. and Maes, J.H.R. (2018), Implicit learning seems to come naturally for children with autism, but not for children with specific language impairment: Evidence from behavioral and ERP data. Autism Research, 11: 1050-1061. https://doi.org/10.1002/aur.1954

      Kissine, M. (2021). Autism, constructionism, and nativism. Language 97(3), e139-e160. https://dx.doi.org/10.1353/lan.2021.0055.

      Kuhl, P.K. (2007), Is speech learning ‘gated’ by the social brain?. Developmental Science, 10: 110-120. https://doi.org/10.1111/j.1467-7687.2007.00572.x

      References:

      (1) Moreau, C. N., Joanisse, M. F., Mulgrew, J., & Batterink, L. J. (2022). No statistical learning advantage in children over adults: Evidence from behaviour and neural entrainment. Developmental Cognitive Neuroscience, 57, 101154. https://doi.org/10.1016/j.dcn.2022.101154

      (2) Boucher, J., Lewis, V., & Collis, G. M. (2000). Voice processing abilities in children with autism, children with specific language impairments, and young typically developing children. Journal of Child Psychology and Psychiatry, 41(7), 847-857. https://doi.org/10.1111/1469-7610.00672

      (3) Schelinski, S., Borowiak, K., & von Kriegstein, K. (2016). Temporal voice areas exist in autism spectrum disorder but are dysfunctional for voice identity recognition. Social Cognitive and Affective Neuroscience, 11(11), 1812-1822. https://doi.org/10.1093/scan/nsw089

      (4) Lin, I.-F., Yamada, T., Komine, Y., Kato, N., Kato, M., & Kashino, M. (2015). Vocal identity recognition in autism spectrum disorder. PLOS ONE, 10(6), e0129451.https://doi.org/10.1371/journal.pone.0129451

      Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants with increased likelihood of autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards in the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Weaknesses:

      While the statistical analyses are rigorous, a few of the components of the models are not clearly defined, and some corrections and thresholds for significance warrant further justification. Further, a few stimuli and participant details that could influence results are not specified. It is not clear whether all participants came from majority French-speaking families; differences in the amount of French language exposure (compared to other languages that may be spoken by a participant's family) could influence results. The standardized volume of the stimuli is also not included. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

      We thank the reviewer for these remarks.

      Regarding the amount of French exposure: while all participants were raised in primarily French-speaking environments (i.e., French as the dominant language at home and daycare), the parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. We did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism. The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      Regarding the volume of stimuli, they were played at 50cm distance with an intensity of 75dB. Both considerations have been included in the new version of the manuscript. In general, we moved most of the figures present in Supplementary material to figure supplements to improve readability.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Comments:

      Figure 6: The figure caption is not complete (there is no description for the right half of panel B).

      We thank the reviewer for this observation, figure 6 caption has been completed.

      Reviewer #2 (Recommendations for the authors):

      Broadly speaking, I would recommend reducing the number of abbreviations in this article, and I would recommend that the authors take care as to where these abbreviations are being introduced. Many of the abbreviated terms are defined in the Materials and Methods section, which is presented after the abbreviations are used in the main results.

      We acknowledge that our extensive use of abbreviations compromises the readability of the manuscript. Consequently, we have removed the following abbreviations:

      - SL (replaced by statistical learning)

      - LC (replaced by latent component)

      - ASD (replaced by autism)

      - TP (replaced by transition probability)

      - MEG (replaced by magneto-encephalogram)

      - MSEL (replaced by Mullen Scale of Early Learnings)

      The remaining abbreviations are:

      HL (high likelihood for autism), LL (low likelihood for autism), EEG (electroencephalogram), PLS-c (partial least square correlation), ERP (event-related potential), RND (random), STR (structured), BSR (bootstrap ratio), AIC (Akaike Information Criterion), PLV (phase locking value), DQ (developmental quotient), APSI (Autism Parent Screen for Infants).

      Moreover, we carefully reviewed how abbreviations were introduced and identified that PLS-c, STR and RND were not defined prior to the Method section. This oversight has been corrected in the reviewed manuscript.

      I would also recommend that the authors be careful with the structuring of the Introduction, particularly with their research questions and hypotheses. The article initially makes clear that the research questions are focused on the developmental trajectory of statistical learning, the levels of word learning that may differentiate high-likelihood versus low-likelihood infants, and the stability of those differences, and associations between statistical learning and various levels of word learning with verbal outcomes. The use of acoustic variability across syllables, while a valuable methodological tool, is somewhat presented as an additional research question, but not clearly stated or tested as such.

      We acknowledge that the introduction (particularly the paragraph from lines 173 to 184) may have implied that speaker variability across syllables was one of our primary research aims. We clarify here that speaker variability was introduced as a mean to increase task difficulty, particularly for high-likelihood (HL) participants, with the aim of amplifying the effect sizes in our analyses.

      To address this, we have removed the theoretical discussion on speaker variability in autism and typical development (lines 173–184) and explicitly stated that speaker variability was not a research question in this study. Crucially, our experimental design did not include a control condition without speaker variability, and thus we could not test its specific effects on statistical learning across age trajectories and groups.

      Page 6, lines 173-176 (pages and lines refer to the reviewed uploaded manuscript): “It is worth noting, however, that our study was not designed to isolate or quantify the specific impact of speaker variability on statistical learning, as the experimental design did not include a baseline control condition omitting this acoustic variation.”

      The authors do a nice job in the Materials & Methods explaining PLS-c and defining the latent components and bootstrapped ratios that will be shared in the Results. An additional brief iteration defining these statistical elements is needed at the beginning of the Results section.

      We thank the reviewer for their appreciation of our Method section. We agree that an additional iteration in the result section would improve readability. We added the following paragraph at the very beginning of the Result section, briefly defining PLS-c and its main statistical output (latent components and bootstrap ratios):

      Pages 6-7, lines 193-202: “Briefly, PLS-c is a data-driven multivariate modelling approach designed to identify significant patterns of electrode clusters (from a brain data matrix containing electrophysiological measures, here PLV) and their associations with “behavioral” variables (from a behavioral design matrix, here age-related parameters). Patterns of brain x behavior associations are called latent components, and their statistical significance is evaluated using permutation testing (n=1000, Bonferroni correction for number of components tested, alpha=.006). Brain and behavioral variables respective contributions to any significant latent component are tested with bootstrapping (500 random samples and replacement), with bootstrap ratios (BSR) greater than 2.3 indicating a stable contribution (for details, see the Materials and Methods section).”

      (1) Page 18 Line 576. The authors need to clarify whether participants were required to be in primarily French-speaking environments and whether there was a minimum amount of French language exposure that participants were required to have if they were exposed to additional languages besides French in their everyday life.

      The reviewer raises a valid concern regarding participants’ language exposure. In this study, all participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare. The parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. However, we did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism.

      The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      To include these considerations, Limitations and Material and methods sections were modified as follows:

      Page 19, lines 604-607: “Second, although all participants were primarily exposed to French, we did not quantify additional language exposure, precluding any analysis of its potential moderator effects on statistical learning in our groups and age-trajectories. However, prior work has reported no effect of bilingualism on auditory triplet segmentation in children (Yim & Rudoy, 2013).”

      Page 20, lines 630-631: “All participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare.”

      References:

      Weiss DJ, Schwob N, Lebkuecher AL. Bilingualism and statistical learning: Lessons from studies using artificial languages. Bilingualism: Language and Cognition. 2020;23(1):92-97. doi:10.1017/S1366728919000579

      Yim D, Rudoy J. Implicit statistical learning and language skills in bilingual children. J Speech Lang Hear Res. 2013 Feb;56(1):310-22. doi: 10.1044/1092-4388(2012/11-0243). Epub 2012 Aug 15. PMID: 22896046.

      (2) Page 18 Line 588. Further, the authors should clarify whether the 7 infants in the HL group, due to early parental concerns were defined by the 18-21-month APSI scores or by parental report prior to study enrollment.

      These 7 infants were recruited based on early parental concerns prior to intake. The APSI score at 18-21 months is only reported to provide an illustration of the amount of early autistic signs that were present in these 7 infants, and to provide an estimation of their probability to develop autism later on based on Sacrey et al., 2018 longitudinal study on the APSI predictive value. We agree with the reviewer that our phrasing suggests that the APSI was used as an inclusion criterion. We rephrased the page 20 lines 642-646 as follows:

      “The 7 other HL infants presented with early parental concerns for autism, based on parental report prior to enrollment. Their Autism Parent Screen for Infants (APSI) total score at their 18-21 months visit was 15.6±6.4, [8-22] range – a score greater than 8 reflecting a 63% positive predictive value for autism in HL populations.”

      (3) Page 20 Line 641. The authors should specify the volume of the stimuli.

      The volume of stimuli was reported in the main text (page 22, lines 695-696) as follows:

      “Stimuli were played on a Bose® Companion 2 Series III at a 50cm distance with an intensity of 75dB.”

      (4) I'd prefer Figure 1 to be reorganized slightly - at present, the placement of the arrows explaining the analysis steps is not intuitive.

      We addressed the reviewer’s comments (4) and (5) together as they both refer to Figure 1B.

      (5) Page 23 lines 718-719. I think it would be helpful to explicitly define each of the interaction variables included in the behavioral design matrix. Further, this matrix should be labeled consistently in both Figure 1B and in the main text.

      We refined figure 1B and its corresponding main text (in Methods section) for clarity. The arrows are now simpler and more parsimonious, labels (e.g., participant i, visit n, behavior design matrix and its parameters) are now standardized between the figure and the main text, and the interaction terms at lines 718-719 are explicitly defined.

      (6) Page 23 lines 726-731: It would be helpful to know whether applying a Bonferroni correction in addition to completing permutation testing is standard when evaluating latent components derived from PLS-c. The authors should also cite justification for a bootstrap ratio cutoff of 2.3 for defining stability.

      In PLS-c analyses, multiple comparisons correction across latent components and bootstrap ratio (BSR) thresholding at 2.3 are commonly adopted practices.

      - Correction for multiple comparisons in PLS-c: PLS-c performs singular decomposition of the data into latent components equal in number to the variables included in the behavior design matrix (7-9 in our study, depending on the inclusion of Verbal outcome as an input variable). Each latent component’s statistical significance is assessed through permutation testing, generating a null distribution for its singular value (Krishnan et al., 2011). Given the multiple tests (one permutation test per latent component), Type I error inflation must be addressed. Recent PLS-c studies commonly applied Bonferroni correction (default procedure in the myPLS toolbox, used by Zoeller et al., 2017, and Delavari et al, 2021), though FDR correction has also been used (Lombardo et al, 2018).

      - Stability threshold for bootstrap and replacement: Within each latent component, saliences’ stability (brain/behavior parameter contributions to each latent component) are evaluated using bootstrapping (Krishnan et al., 2011). The bootstrap ratio (BSR) of each parameter, calculated as the saliency divided by its bootstrap-derived standard error, functions analogously to a z-score under normality assumptions. The BSR can then be used to assess the stability of the saliency (i.e., how stable is its contribution to the latent component). BSR thresholds in the literature typically range from 1.96 to 3.0. Krishnan et al (2011) state that when BSR are “larger than 2 the corresponding saliences are considered significantly stable”. Delavari et al (2021) and our study used a 2.3 thresholding, corresponding to a 99.0% bootstrap confidence interval not crossing the zero line – roughly equivalent to a two-tailed p<.001. Lombardo et al (2018) used a looser threshold of 1.96, corresponding to a 95% confidence interval not crossing the zero line (~two-tailed p<.05), while Zöller et al (2017) used a more stringent 3.0 thresholding (~p<.001, or 99.9% confidence interval not crossing the zero line).

      Thus, our application of Bonferroni correction for multiple comparisons and our 2.3 BSR threshold aligns with established conventions.

      We added following lines in the manuscript:

      Page 25 lines 784-785: “Bonferroni correction was applied to account for multiple comparisons across the 9 tested latent components in the PLS-c, yielding an adjusted alpha of .006 (Zoeller et al, 2017; Delavari et al, 2021).”

      Page 25 lines 789-792: “BSR are analogous to Z-scores and can be used to assess the stability of the saliency. We considered BSR > 2.3 as stable, corresponding to a 99.0% bootstrap confidence interval not crossing zero – roughly equivalent to a two-tailed p<.001 (Delavari et al., 2021; Krishnan et al., 2011).”

      References:

      Delavari F, Sandini C, Zöller D, Mancini V, Bortolin K, Schneider M, Van De Ville D, Eliez S. Dysmaturation Observed as Altered Hippocampal Functional Connectivity at Rest Is Associated With the Emergence of Positive Psychotic Symptoms in Patients With 22q11 Deletion Syndrome. Biol Psychiatry. 2021 Jul 1;90(1):58-68. doi: 10.1016/j.biopsych.2020.12.033. Epub 2021 Jan 18. PMID: 33771350.

      Lombardo, M.V., Pramparo, T., Gazestani, V. et al. Large-scale associations between the leukocyte transcriptome and BOLD responses to speech differ in autism early language outcome subtypes. Nat Neurosci 21, 1680–1688 (2018). https://doi.org/10.1038/s41593-018-0281-3

      Daniela Zöller, Marie Schaer, Elisa Scariati, Maria Carmela Padula, Stephan Eliez, Dimitri Van De Ville. Disentangling resting-state BOLD variability and PCC functional connectivity in 22q11.2 deletion syndrome. NeuroImage, Volume 149, 2017, Pages 85-97, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2017.01.064

      Anjali Krishnan, Lynne J. Williams, Anthony Randal McIntosh, Hervé Abdi, Partial Least Squares (PLS) methods for neuroimaging: A tutorial and review, NeuroImage, Volume 56, Issue 2, 2011, Pages 455-475, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2010.07.034

      (7) I have a few minor grammar/formatting recommendations for the authors as well:

      (a) Should the Geneva Autism Cohort be capitalized? At present, it is not.

      We agree with the reviewer’s suggestion, and we capitalized the Geneva Autism Cohort in the main text (page 18, line 571)

      (b) Page 24, line 750. Do the authors mean that the data was re-referenced to average?

      The preprocessed data is not average-referenced (see section Data pre-processing). Therefore, both for neural entrainment computation and ERPs, the data were average-referenced.

      (c) It would be nice to have a figure of the actual ERP for each condition and age group.

      We agree that PLS-c can be difficult to interpret without the raw actual ERPs on which it was modelled. We direct the reviewer to supplementary figure S6 at page 59, which displays the raw ERPs for each condition (part-word, word, and their subtraction) per age group. Supplementary figures S7-8 at pages 60-61 further illustrate topographical ERPs for each group (high and low likelihood for autism). We deemed these figures too extensive for the main text. Instead, the most relevant ERP topographies are presented in Figures 4-6 to facilitate PLS-c interpretation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and monomethylates PPARγ at K170, and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      We are grateful to the reviewer for the positive feedback and appreciation of our work.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality, and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:

      (a) The co-IP panel in Figure 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.

      We thank the reviewer for this valuable suggestion. The overexpression co-immunoprecipitation experiment referred to by the reviewer has been moved to Supplementary Figure S1 in the revised manuscript. In this experiment, immunoprecipitation was performed using an anti-FLAG antibody to pull down FLAG-tagged PPARγ. In the absence of FLAG-PPARγ, the anti-FLAG immunoprecipitation does not recover a bait protein, and therefore HA-SETD6 is not expected to be specifically immunoprecipitated. Thus, an HA-SETD6-only condition would primarily serve as a negative control for the anti-FLAG pull-down rather than provide additional information regarding the specificity of the interaction.

      Importantly, in the revised manuscript we have substantially strengthened the evidence supporting the SETD6–PPARγ interaction by adding two independent complementary experiments. First, we included a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. Second, we added an independent proximity ligation assay (PLA) (new Figure 2D), which further confirms the interaction between SETD6 and PPARγ in cells. Together with the in vitro binding assay presented in Figure 2A, these orthogonal approaches provide compelling evidence for the specificity of the SETD6–PPARγ interaction. Therefore, we believe that the requested HA-SETD6-only control would not provide additional mechanistic insight beyond the comprehensive validation now included in the revised manuscript.

      (b) In Figure 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors took advantage of a deletion line to validate that the co-IP is specific to the presence of both.

      We thank the reviewer for this helpful suggestion. We have revised the manuscript to strengthen the evidence supporting the endogenous interaction between SETD6 and PPARγ. Specifically, we now include a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous SETD6 co-immunoprecipitates with endogenous PPARγ and, conversely, that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. These reciprocal experiments independently validate the specificity of the interaction.

      In addition, we have revised the figure layout and labeling to more clearly distinguish the immunoprecipitating antibody from the bead control, thereby addressing the reviewer's concern regarding the presentation of the co-immunoprecipitation data.

      Although we did not perform the co-immunoprecipitation in a depletion/knockout background, we believe that the combination of reciprocal endogenous co-immunoprecipitation (New Figure 2B), the independent proximity ligation assay (New Figure 2D), and the direct in vitro binding assay (Figure 2A) provides multiple orthogonal lines of evidence supporting a specific interaction between SETD6 and PPARγ.

      (c) The same IP labeling issue exists for Figure 3B (label is on the same and on top).

      We have revised the labeling in Figure 3B to clearly distinguish the immunoprecipitating antibody from the bead control, thereby improving the clarity of the figure.

      (d) Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      We thank the reviewer for pointing this out. We have now added the missing information regarding the pan-methyl antibody to the Materials and Methods section, including the supplier, catalogue number, and experimental conditions used. Specifically, the pan-methyl antibody used in this study was purchased from Abcam (ab23366) and was used at a 1:500 dilution for western blot analysis and 2 μg per reaction for immunoprecipitation experiments.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA-binding activity of PPARγ. Given that the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediated methylation. Is DNA binding affected/enhanced?

      We thank the reviewer for raising this important point. To investigate whether K170 methylation could directly affect PPARγ binding to DNA, we performed structural modeling based on the available co-crystal structure of PPARγ bound to DNA (PDB: 3DZU). As shown in the new Figure 6F, K170 is positioned near the DNA-binding region; however, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not appear to sterically interfere with the PPARγ–DNA interaction. In addition, modeling of multiple K170me1 rotamers did not suggest any major disruption of the DNA-bound conformation.

      These observations suggest that K170 methylation is unlikely to directly alter the intrinsic DNA-binding affinity of PPARγ. In contrast, our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ occupancy at target promoters in cells. Together, these findings support a model in which K170 methylation promotes PPARγ chromatin association and transcriptional activity through mechanisms other than direct modulation of DNA binding, such as altered cofactor recruitment or protein–protein interactions.

      We agree with the reviewer that future biochemical approaches, including EMSA/gel shift assays or quantitative DNA-binding measurements, will be valuable to directly determine whether K170 methylation affects the intrinsic DNA-binding affinity of PPARγ. We have incorporated this new structural analysis and the corresponding discussion into the revised manuscript.

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or a catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      We thank the reviewer for this insightful suggestion. To address whether SETD6 catalytic activity is required for its association with PPARγ, we performed an additional PLA experiment comparing SETD6 WT and the catalytic mutant SETD6 Y285A. As shown in the revised New Figure 2D, both SETD6 WT and SETD6 Y285A showed comparable proximity to PPARγ in cells, indicating that SETD6 catalytic activity is not required for the physical association between SETD6 and PPARγ.

      These findings support a model in which SETD6 first recognizes and binds PPARγ independently of its catalytic activity. Subsequent methylation of PPARγ at K170 is therefore likely to regulate the downstream functional consequences of this interaction, including enhanced chromatin occupancy and transcriptional activation, rather than the initial SETD6–PPARγ association itself.

      In addition, we generated using AlphaFold a structural model of the SETD6–PPARγ complex as a supportive visualization (New figure S2). Given the limited confidence of the prediction, we interpret this model cautiously and include it in the Supplementary Information rather than the main figures. We agree with the reviewer that future studies examining SETD6 recruitment to PPARγ target promoters and the effect of the PPARγ K170R mutant on SETD6–PPARγ association will further refine the molecular mechanism.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors, and there are some issues with the figure display.

      We thank the reviewer for this comment. We carefully revised the manuscript to correct typographical and grammatical errors throughout the text and also addressed the figure display issues noted by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation. The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism gene promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulates lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing, and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly.

      We thank the reviewer for his/her positive feedback on the manuscript.

      Weaknesses:

      The data presented in the manuscript are largely convincing in support of the authors' conclusions; however, there are some errors in the presentation of the figures and some issues in the text that would benefit from editing. Furthermore, there are some important questions not fully addressed in the results or discussion. 

      It would be great if the authors could speculate more on the diverse roles of SETD6 in methylated transcription factors and/or provide more context regarding the conditions that are likely to support methylation of PPARγ by SETD6. 

      We thank the reviewer for this important suggestion. In the revised Discussion, we expanded the manuscript to better place our findings within the broader context of SETD6-mediated regulation of transcription factors. Previous studies from our group and others demonstrated that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating transcriptional selectivity, chromatin occupancy, and cofactor recruitment. We now discuss the possibility that SETD6 functions as a context-dependent signalling integrator that selectively regulates transcription factor activity through lysine methylation under distinct physiological conditions.

      In addition, we expanded the Discussion regarding potential cellular contexts that may favour PPARγ methylation by SETD6. Because PPARγ activity is strongly induced during lipid overload and fatty acid exposure, conditions associated with steatosis and metabolic stress may enhance the functional importance of SETD6-dependent methylation. We also discuss the possibility that chromatin accessibility, ligand-dependent activation of PPARγ, and metabolic signaling pathways may collectively influence the formation and stability of the SETD6–PPARγ complex.

      Also, while a potential cross-talk between methylation and phosphorylation is described in the discussion, it would be great to provide more structural insight into how this might regulate DNA binding of PPARγ and/or discuss whether there are other possibilities given the location of the target lysine in the DNA binding domain.

      We thank the reviewer for this valuable suggestion. To provide additional structural insight into the potential interplay between methylation and phosphorylation within the PPARγ DNA-binding domain, we performed structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). As shown in the new Figure S6, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not introduce steric clashes with DNA, suggesting that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects. In contrast, phosphorylation of the neighboring residue T166 is predicted to introduce multiple intramolecular steric clashes within the DNA-binding domain. These structural changes could influence the local conformation or dynamics of the DNA-binding domain and thereby indirectly modulate PPARγ DNA binding or transcriptional activity.

      In addition, as suggested by the reviewer, we expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function. While our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ chromatin occupancy at target promoters, the structural modeling suggests that this effect is unlikely to arise from direct steric modulation of the DNA interface. Instead, K170 methylation may influence chromatin occupancy by regulating protein–protein interactions, cofactor recruitment, local conformational dynamics, or other chromatin-associated mechanisms. We have incorporated these new structural analyses and the expanded discussion into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The KEGG panels are too low resolution to read.

      We thank the reviewer for this comment. We improved the resolution of the KEGG pathway enrichment panels and, as suggested by the reviewer (see below), we separated the upregulated and downregulated gene sets to improve clarity and readability. These changes are now reflected in the revised new Figures 5C, 5D, 6B and 6C.

      (2) Figure 3 panel D is hard to visualize; a better version or repeat experiment is needed.

      We thank the reviewer for this comment. To address this concern, we replaced the original Figure 3D with a new independent experiment that more clearly demonstrates the methylation of endogenous PPARγ by SETD6. We believe that the new data provide substantially stronger evidence and improve the clarity of the revised manuscript.

      (3) There is a typo in Figure 2A (line present in "S" of Signal).

      We thank the reviewer for pointing out this typo. The error in Figure 2A has now been corrected in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Overall, the experiments and analyses presented are sufficient to support the conclusions and interpretations of the work. However, there are some issues of presentation and writing that are worth addressing. These comments are listed below:

      (1) My only substantial recommendation is to improve the discussion to provide more context to the findings in terms of both the larger role of SETD6 in methylating transcription factors (some of whom also regulate its expression) and the potential modes through which PPARγ DNA binding activity could be regulated by methylation.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we substantially expanded the Discussion to better place our findings within the broader context of SETD6mediated regulation of transcription factors. We now discuss previous studies demonstrating that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating chromatin occupancy, cofactor recruitment, and transcriptional selectivity. We further propose that SETD6 functions as a context-dependent signaling regulator that integrates distinct cellular pathways through the selective lysine methylation of transcription factors.

      In addition, we expanded the Discussion regarding the potential mechanisms by which PPARγ K170 methylation regulates transcriptional activity. We incorporated new structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU), which suggests that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects, whereas phosphorylation of the neighboring residue T166 may induce intramolecular steric clashes within the DNA-binding domain. We also expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, local conformational dynamics, protein–protein interactions, and recruitment of transcriptional cofactors or chromatin-associated proteins. Finally, we discuss that future biochemical studies will be important to determine whether K170 methylation also influences the intrinsic DNA-binding affinity of PPARγ.

      Can the authors incorporate any other published structural data to speculate on the role of methylation or describe more about how it is expected that the methylation-phosphorylation crosstalk modulates DNA binding? This type of discussion would better highlight the potential importance of this new modification on PPARγ.

      We thank the reviewer for this important suggestion. In the revised manuscript, we expanded the Discussion and incorporated additional structural analyses (new Figure 6F and new figure S6) based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). K170 is positioned within the DNA-binding domain, between the two zinc-finger motifs that mediate DNA recognition and stabilization on PPRE-containing DNA. Our structural modeling predicts that the K170me1 side chain is oriented away from the DNA interface and does not introduce steric clashes with DNA, suggesting that methylation is unlikely to directly alter the DNA-binding interface through steric effects. Nevertheless, lysine methylation can influence protein surface properties, protein– protein interactions, and recognition by regulatory binding partners, raising the possibility that K170 methylation modulates PPARγ chromatin occupancy or promoter selectivity through indirect mechanisms.

      In contrast, structural modeling predicts that phosphorylation of the neighboring residue T166 introduces multiple intramolecular steric clashes within the DNA-binding domain. These clashes could alter the local conformation or dynamics of the DNA-binding domain and thereby indirectly influence PPARγ–DNA interactions and transcriptional activity. Together, these observations suggest that K170 methylation and T166 phosphorylation may represent a regulatory crosstalk that fine-tunes PPARγ function through distinct structural mechanisms.

      We further expanded the Discussion to consider additional, non-mutually exclusive mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, protein–protein interactions, cofactor recruitment, and stabilization of transcriptional complexes at target genes. While these possibilities require further mechanistic investigation, we agree with the reviewer that these structural considerations highlight the potential regulatory importance of this newly identified PPARγ modification.

      (2) In the gene expression experiments presented in Figure 5, it would be useful if the GO terms were described as enriched in either the up- or down-regulated gene sets. From the way it is presented, it is not clear if specific categories of genes are found enriched in those upregulated or downregulated upon KO of SETD6. This is also true for the experiments presented in Figure 6 regarding the PPARγ mutant.

      We thank the reviewer for this important suggestion. In the revised manuscript, we separated the pathway enrichment analyses into upregulated and downregulated gene sets for both the SETD6 knockout RNA-sequencing experiments (Figure 5) and the PPARγ WT versus K170R mutant analysis (Figure 6). This revision provides improved clarity regarding which biological pathways are positively or negatively associated with SETD6 depletion or disruption of PPARγ K170 methylation. The updated KEGG enrichment analyses are now presented in the revised new Figures 5C, 5D, 6B and 6C.

      In addition, if there is a significant overlap in genes misregulated in both mutants, this could be shown in the figure.

      We thank the reviewer for this suggestion. We compared the differentially expressed genes identified in the SETD6 KO cells and the PPARγ K170R mutant cells to evaluate the extent of overlap between the two datasets. However, we did not observe a substantial or statistically significant overlap in misregulated genes under the thresholds used in our analysis. Therefore, we decided not to include this comparison in the revised figure. Nevertheless, both datasets consistently showed enrichment for pathways associated with lipid metabolism and PPAR signaling, supporting a functional connection between SETD6 and PPARγ-mediated transcriptional regulation.

      (3) There are some typos and grammatical errors throughout the work, and it should be carefully edited. One example is the following heading: PPARγ K170 methylation by SETD6 regulates mediates lipid droplets formation.

      We thank the reviewer for this comment. The manuscript was carefully revised to correct typographical and grammatical errors throughout the text. In particular, the heading mentioned by the reviewer was corrected in the revised manuscript.

      Minor errors in the figures:

      (1) Formatting issue in the labeling for the x-axis of Figure 1E.

      We thank the reviewer for pointing out this formatting issue. The labeling of the x-axis in Figure 1E has been corrected in the revised manuscript.

      (2) Figure 2B - HA-SETD6 should be labeled as minus for the first lane.

      We thank the reviewer for pointing this out. We corrected the labeling in Figure 2B (now Figure S1), and the first lane is now properly indicated as negative for HA-SETD6.

      (3) Figure 2D - Labeling needs improvement. Is this FLAG-PPARγ? "NC" was not defined in the legend. If negative control, what type?

      We thank the reviewer for this comment. We improved the labeling and figure legend of Figure 2D for clarity. Specifically, we now clearly indicate that the experiment was performed using Flag-PPARγ, and we defined “NC” in the legend as the negative control condition. In addition, during the revision process we noticed that the original PLA experiment was performed in HeLa cells rather than HepG2 cells, as previously indicated. This has now been corrected throughout the revised manuscript.

      (4) Figure 3 - The CRSIPR control should be described somewhere in the legend or the methods.

      We thank the reviewer for this comment. We clarified the description of the CRISPR control cells in the Materials and Methods section. Specifically, we now explicitly state that the CRISPR control (CT) cells were generated using the empty lentiCRISPR vector without SETD6-targeting sgRNAs.

      (5) Figure 5B and 6A - The legend is not labeled nor defined in the text- fold-change, log2 fold change, z score?

      We thank the reviewer for pointing this out. We revised the figure legends for Figures 5B and 6A to explicitly define the heatmap scale and normalization method. Specifically, we now indicate that the heatmaps represent normalized gene expression values displayed as Z-scores.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Please include an uptake control (early time point) or time‑course to distinguish phagocytosis from intracellular killing.

      We agree that distinguishing uptake from intracellular killing is important for interpreting bactericidal assays. Due to a recent transition, we are unable to conduct additional early‑time‑point assays. To address this transparently, we have revised the manuscript to clarify that our measurements represent overall bacterial load reduction, reflecting the combined effects of uptake and killing.

      The normalization as ‘fold killing’ is non‑standard; please report absolute CFU (log scale).

      We have retrieved the raw data and re‑expressed all bactericidal activity measurements as absolute CFU. All relevant figures, legends, and text have been updated accordingly.

      Authors report quantification of cytokine concentrations, yet no information is provided regarding how these measurements were performed.

      Cytokine concentrations were quantified by ELISA. We have now added details regarding assay kits, sample preparation, and detection parameters to the Methods section.

      While the choice of IL‑1β and IL‑6 is straightforward, the focus on IL‑18 requires explicit justification.

      IL‑18 is a macrophage‑associated pro‑inflammatory cytokine with established links to inflammasome activation and trained immunity pathways, providing a clear justification for its inclusion.

      The methodology used to quantify immune cell populations presented in Figure 2 is not described.

      We have added a detailed description of the flow cytometry methodology, including staining and gating strategies.

      Immune cell quantification would be expected in the context of the challenge experiment as well.

      While we agree that such data would be valuable, additional mouse experiments cannot be performed because the animals used in the challenge model are not currently available during the transition period. If feasible, we are exploring ex vivo flow cytometry data from TWIK2 mutant versus wild‑type macrophages.

      AMs are not considered recruited immune cells; this should be corrected.

      We have corrected this in the figure legend and throughout the manuscript.

      The authors report n = 5 for the survival curves in the figure legend, whereas n = 7 is stated in the Methods section.

      We have corrected the sample size to ensure consistency between the figure legend and Methods.

      ATAC‑seq peaks are referred to as ‘genes’ and ‘differentially expressed genes’.

      We have corrected the terminology in the manuscript. ATAC‑seq identifies differentially accessible chromatin regions, which are then annotated to the nearest downstream gene.

      In Figure 7, trained WT and Nlrp3-/- mice display similar levels of bacterial clearance. How should this result be interpreted?

      A portion of this phenotype is via the normalization of phagocytosis to ‘fold killing’. Presentation of the raw CFU data shows a trend towards reduced bacterial clearance in Nlrp3-/- mice.

      Reviewer #2 (Public review):

      Sample numbers for experiments 1, 2, and 6 are not provided.

      We have added explicit n values for all experiments and verified their accuracy against the original records.

      The Discussion would benefit from a clear summary of study caveats.

      We agree and have added a dedicated paragraph outlining key Caveats as described.

      Specific identities of DEGs are not provided; only pathway enrichment is shown.

      We have now included a supplementary table listing the differentially expressed genes identified in our analysis.

      Controls for subcellular fractionation and dye microscopy should be included.

      Controls for subcellular fractionation have added in figure 3B and controls for dye microscopy have now been added in supplementary figure 1.

      The text states that protease inhibitors diminish ATP‑induced training effects, but the figure does not show significance.

      We have re‑examined the data and updated the figure to include statistical testing where appropriate.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

      We thank the reviewer for the positive assessment of our work and for the constructive suggestions. We agree that certain aspects of the manuscript required clarification, in particular the potential contribution of the RAP80–ABRAXAS pathway, the interpretation of the repair synthesis assay, and the role of RNF168 recruitment. We have addressed these points experimentally where feasible and have revised the manuscript accordingly to avoid overinterpretation.

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) In Figures 4C, 4D, and S4B-D, BRCA1 and RAD51, recruitment to nucleolar caps is partially reduced upon RNF168 depletion. Despite this, the authors broadly conclude that recruitment mainly depends on the MDC1-RNF8-RNF168 pathway. Since the RAP80-Abraxas pathway may also contribute, as briefly mentioned in the Discussion, siRNA knockdown of Abraxas would help clarify the relative roles of these two pathways.

      We thank the reviewer for this important suggestion. To directly address the potential contribution of the RAP80–ABRAXAS pathway, we depleted RAP80 by siRNA and analysed BRCA1 and RAD51 recruitment to nucleolar caps following I-PpoI-induced rDNA damage.

      Strikingly, RAP80 depletion strongly impaired the formation of both BRCA1 and RAD51 nucleolar caps in two independent cell lines (U2OS and RPE1) (new Figure 5). These findings demonstrate that the RAP80–ABRAXAS pathway plays a critical role in BRCA1 recruitment at nucleolar caps.

      Together with our observation that RNF168 depletion only partially reduces BRCA1 and RAD51 recruitment, these results indicate that both RNF168-dependent and RAP80–ABRAXAS-dependent pathways contribute to HDR factor recruitment downstream of RNF8-mediated chromatin ubiquitylation.

      We have revised the model Figure (Figure 9) and the Results and Discussion sections accordingly to reflect this dual-pathway model and to avoid overemphasising the contribution of RNF168.

      (2) In Figure 7C, the EdU-γH2AX PLA assay detects DNA synthesis at rDNA breaks, but it remains unclear whether this signal reflects HDR- or NHEJ-mediated repair. Since MDC1 functions upstream of the DSB repair pathway choice, the observed reduction in PLA signal upon MDC1 depletion does not necessarily reflect impaired HDR alone. Synchronizing cells in G2 or using cell cycle markers would help clarify the repair context and strengthen the interpretation.

      We thank the reviewer for raising this important point. We agree that the EdU–gH2AX PLA assay does not exclusively report on HDR-mediated DNA synthesis and may also capture other forms of repair-associated DNA synthesis.

      In the revised manuscript, we have therefore tempered our interpretation and now describe this assay more cautiously as a readout of DNA synthesis at sites of rDNA damage, rather than as a direct measure of HDR activity.

      Importantly, our conclusion that MDC1 promotes HDR factor recruitment at nucleolar caps is based primarily on the reduced accumulation of BRCA1, PALB2, and RAD51, which are well-established markers of HDR. The PLA assay is now presented as supportive evidence for ongoing DNA synthesis at these sites rather than as a definitive indicator of HDR.

      We agree that further experiments, such as cell cycle synchronization or the use of phase-specific markers, would help to more precisely define the repair context, and we have included this point in the Discussion.

      (3) The authors propose that MDC1 is essential for RNF8-RNF168 recruitment, specifically at nucleolar rDNA breaks. A side-by-side comparison of RNF8 or RNF168 localization in the presence and absence of MDC1, with IR-treated conditions, would provide important validation of this model. Including representative images in Figure S2C would further support the claim.

      We agree with the reviewer that a direct analysis of RNF8 and RNF168 recruitment in the presence and absence of MDC1 would provide valuable mechanistic insight. We therefore attempted to address this experimentally.

      However, despite testing multiple antibodies, we were unable to obtain specific and reproducible signals for RNF8 and RNF168 at nucleolar caps, precluding a reliable analysis of its recruitment under these conditions.

      Given this technical constraint, we have revised the manuscript to avoid overinterpretation regarding direct RNF168 recruitment and instead focus on functional readouts of downstream ubiquitylation-dependent signalling, such as BRCA1 and RAD51 accumulation.

      We note that the requirement for MDC1 in BRCA1 and RAD51 recruitment at nucleolar caps is consistent with a role of MDC1 upstream of RNF8-dependent chromatin ubiquitylation, in line with its established function at IR-induced DSBs.

      Minor comments:

      (1) The legend for Figure 8 should more clearly explain the proposed mechanism and include concise titles or descriptions for each sub-panel.

      We agree with the reviewer that the model should be described in the Figure legend. We have thus updated the model to accommodate the new data and wrote a legend that concisely explains the proposed model. We do not think that titles for each sub-panel are required. Instead, we separately referred to the sub-panels in the legend.

      (2) Typos:

      (a) Page 8: PRE1 MDC1, as "RPE1 MDC1;

      (b) S3 Figure legend: Dhermacon";

      (c) Page 29: "80.103"-please clarify or correct.

      We thank the reviewer for pointing out these errors. These have been corrected in the revised manuscript

      Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

      Weaknesses:

      (1) The recruitment of BRCA2 was not directly demonstrated. This caveat could be recognized, as IF for BRCA2 is challenging.

      (2) PpoI also induces DSBs in the non-rDNA genome. These DSBs would be an ideal control to establish nucleolar specificity of the events described and clarify whether the difference between IR and PpoI is the chemical structure of the DSB or the location of the DSB.

      We thank the reviewer for the positive and insightful evaluation of our work. We appreciate the recognition of the conceptual advance and the robustness of our experimental approaches. We have carefully considered the reviewer’s suggestions and have revised the manuscript to clarify interpretation where appropriate, particularly regarding BRCA2 recruitment and the specificity of I-PpoI-induced DNA damage. Where possible, we have also added new analyses to strengthen the conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The claim that the BRCA1-PALB2-BRCA2 is recruited (abstract, end of results section, discussion page 15) should be qualified as BRCA2 recruitment was not directly demonstrated.

      We thank the reviewer for this important point. We agree that BRCA2 recruitment was not directly demonstrated in our study, as reliable immunofluorescence detection of BRCA2 remains technically challenging.

      We have therefore revised the manuscript throughout (Abstract, Results, and Discussion) to avoid overstatement and now refer more precisely to the recruitment of BRCA1, PALB2, and RAD51, rather than implying direct recruitment of a BRCA1–PALB2–BRCA2 complex.

      We note that BRCA2 function is supported indirectly by the observed RAD51 loading, which depends on BRCA2 activity. However, we have clarified this point to ensure that our conclusions remain fully supported by the presented data.

      (2) The temporal sequence established in Figure 1, 1hr BRCA1 and 2 hrs PALB2, argues against recruitment of a stable BRCA1-PALB2-(BRCA2) complex. This should be acknowledged.

      We thank the reviewer for this insightful observation. We agree that the temporal separation between BRCA1 accumulation (1 h) and PALB2/RAD51 recruitment (2 h) argues against the recruitment of a pre-assembled, stable BRCA1–PALB2–BRCA2 complex.

      We have revised the manuscript to reflect this interpretation and now describe the recruitment of HDR factors as a sequential process rather than as the assembly of a pre-formed complex. This is consistent with current models in which BRCA1 promotes subsequent PALB2 and BRCA2 recruitment, ultimately leading to RAD51 loading.

      (3) The model predicts that MDC1-KO cells are proficient for transcriptional repression after nucleolar DSB induction. Has this been tested?

      We did not specifically test this in the current work, but previous results published by our group revealed that siRNA-mediated depletion of MDC1 in human cells had a minimal effect on rDNA transcriptional inhibition after DNA damage (Larsen et al., 2024).

      (4) The nuclear PpoI DSBs could be analyzed as a specificity control, and clarify whether the difference between IRIF and PpoI DSBs relates to the DSB chemistry or location.

      We thank the reviewer for this important point. We agree that I-PpoI induces DNA breaks both within rDNA repeats and at additional genomic loci.

      To address this, we have now analysed the formation of gH2AX-positive nucleolar caps and non-nucleolar gH2AX foci over time following I-PpoI expression (new Figure 1–figure supplement 2). We find that nucleolar caps form rapidly and are prominent at early time points, whereas gH2AX foci accumulate more gradually.

      These results indicate that nucleolar caps and non-nucleolar DNA damage responses can be distinguished both spatially and temporally, and support the use of nucleolar caps as a specific readout for rDNA damage in our study.

      In addition, we note that RAD51 recruitment to IR-induced foci is not affected by MDC1 loss, suggesting that the requirement for MDC1 in RAD51 loading is specific to nucleolar rDNA breaks rather than reflecting differences in DSB chemistry alone. We have clarified this point in the Discussion.

      Additional points:

      (5) Page 4 top: Shieldin.

      Corrected.

      (6) The general reader will be interested to learn about the connection of the Treacle function with Treacher Collins syndrome. Maybe a paragraph could be added to discuss this?

      We thank the reviewer for this suggestion. We agree that the relationship between Treacle and Treacher Collins syndrome may be of interest to a broad readership. Since the developmental pathology of Treacher Collins syndrome is currently thought to arise primarily from impaired ribosome biogenesis and nucleolar dysfunction rather than defective nucleolar DNA damage signalling, we felt that an extensive discussion would be beyond the scope of the present study. We have, however, added a brief statement introducing Treacle as the product of the TCOF1 gene mutated in Treacher Collins syndrome and noting that whether its DNA damage response function contributes to disease pathology remains an open question.

      (7) Figure 7: A short explanation could be added as to why hypoxia conditions were chosen for the p53-deficient cell lines.

      We thank the reviewer for pointing this out. We have added a brief explanation in the figure legend to clarify that hypoxia conditions were used to stabilise replication stress and enhance detection of DNA repair intermediates in p53-deficient cells.

      (8) A short statement on whether the repair of nuclear DBS is affected by Treacle could be added.

      We thank the reviewer for this interesting point. While our study focuses on nucleolar DNA damage, we did not observe evidence that Treacle is required for the repair of non-nucleolar DSBs. We have added a brief statement in the Discussion to clarify that Treacle appears to function specifically in the nucleolar DNA damage response.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a role of sumoylation at K81 in p66Shc which affects endothelial dysfunction. This explores a new mechanism for understanding the role of PTMs in cellular processes.

      Strengths:

      The experiments are well planned and the results are well represented.

      Vascular tonality experiments were carried out nicely, given the amount of time and effort one needs to put in to get clean results from these experiments.

      Weaknesses:

      (1) The production of ROS has been measured in a very superficial way.

      The term "ROS" confers a plethora of chemical species which exerts different physiological effects on different cells and situations.

      Mitochondria through one of the source, but not the only source of ROS production. Only measuring ROS with mitosox do not reflect the cellular condition of ROS in a specific condition. I would suggest authors consider doing IF of oxidative stress specific markers, carbonyl group and also, maybe, Amplex red for determining average oxidative stress and ros production in the cells.

      As suggested, we employed an additional ROS-sensitive probe, H<sub>2</sub>DCFDA, which revealed an overall increase in intracellular ROS levels upon SUMO2 overexpression; this effect was reversed by knockdown of p66Shc. In addition, we performed the Amplex Red assay on conditioned media to assess extracellular ROS release. SUMO2 overexpression did not alter ROS levels detected in the media, whereas a paradoxical increase was observed following p66Shc knockdown.

      Author response image 1.

      Amplex Red assay performed in HUVECs with and without knockdown of p66Shc expressing SUMO2 (Ad-SUMO2) or a control virus (Ad-LacZ).

      Amplex Red predominantly detects hydrogen peroxide (H<sub>2</sub>O<sub>2</sub>), which may originate from NADPH oxidases or be generated through the dismutation of superoxide. In contrast to superoxide, H₂O₂ is relatively stable, membrane-permeable, and well recognized as a second messenger in endothelial signaling, where it modulates kinase and phosphatase activity and supports physiological vascular functions. Thus, the elevated H<sub>2</sub>O<sub>2</sub> detected following p66Shc knockdown may represent a signaling-competent redox state rather than a pathological increase in oxidative stress. Moreover, the SUMO2-p66Shc-mediated increase in mitochondrial ROS may be efficiently buffered by cellular antioxidant defense mechanisms, thereby limiting detectable extracellular ROS release.

      (2) 8-OHG signal seems very confusing in Figure 7E. 8-ohg is supposed to be mainly in the nucleus and to some extent in mitochondria. The signal is very diffused in the images. I would suggest a higher magnification and better resolution images for 8-ohg. Also, the VWF signal is pretty weak whereas it should be strong given the staining is in aorta. Authors should redo the experiments.

      We have provided a better image for figure 7E. We repeated the staining with another antibody for vWF which showed stronger signal for vWF.

      (3) PCA analysis is quite not clear. Why is there a convergence among the plots? Authors should explain. Also, I would suggest that the authors do the analysis done in Figure 8B again with R based packages. IPA, though being user-friendly, mostly does not yield meaningful results and the statistics carried out is not accurate. Authors should redo the analysis in R or Python whichever is suitable for them.

      We thank the reviewer for their valuable feedback and insightful suggestions.

      Regarding the PCA analysis, the observed convergence among the data points reflects the underlying biological similarity between the samples within each group. Given the relatively low abundance and limited number of quantified peptides in our dataset, the variance captured by PCA is modest, leading to partial overlap between groups. This convergence is therefore likely due to shared biological characteristics and inherent sample variability, rather than technical issues.

      In response to the reviewer’s suggestion on pathway analysis, we have re-performed the analysis using R-based approaches. Specifically, we utilized established pipelines PROGENy-based signaling pathway activity inference. These methods provide statistically robust and reproducible results. The updated analyses and corresponding figures have been included in the revised manuscript, replacing the previous IPA-based results (Fig. 8). We believe these additions strengthen the interpretation of signaling pathway alterations in our dataset.

      (4) The MS analysis part seems pretty vague in methods. Please rewrite.

      We have revised the methodology for MS.

      Reviewer #2 (Public review):

      Summary:

      The article builds on the earlier work that both p66Shc and SUMOylation are essential nitric oxide (NO) based development of endothelial vasculature (PMID: 10580504; 28760777 and 35187108). The current manuscript brings forward a finding of how SUMO2ylation of p66Shc mediated ROS production which is essential for endothelial cells. They further identify that lysine 81 of p66Shc is the residue which is conjugated to SUMO2 and is crucial for mitochondrial localization. They further show that K81 SUMO2ylation is essential for S36 phosphorylation.

      Strengths:

      Convincingly shows that p66Shc is SUMO2ylated on lysine 81 in cells and also shows that the phosphorylation (serine 36) reduces upon loss of this critical SUMOylation site.

      Weaknesses:

      All the experiments performed here are in overexpression background therefore, it would be crucial to show that p66Shc is SUMO2ylated at physiological levels.

      As detecting endogenous SUMO2-p66Shc is technically challenging considering the almost 92% homology between SUMO2 and SUMO3 and the absence of a specific antibody to detect p66Shc, we generated a custom-made antibody which can detect SUMO2-p66Shc (YenZym, CA). Using this antibody, we performed immunoprecipitation which showed endogenous SUMO2-p66Shc at a molecular weight higher than p66Shc suggesting the SUMO2 modification of p66Shc at physiological level.

      Reviewer #3 (Public review):

      Summary:

      The authors set out to determine how SUMO2 impairs endothelial function through direct modification of the protein p66Shc. p66Shc is known to promote reactive oxygen species production, and here the authors demonstrate that SUMO2 modifies p66Shc at lysine-81, resulting in increased phosphorylation, mitochondrial translocation. These are prosed to mediate the detrimental effects of SUMO2 in a mouse model of hyperlipidemia.

      Strengths:

      A major strength of this work is the multi-pronged approach combining biochemical assays, proteomic analyses, and a genetically modified mouse model expressing a SUMOylation resistant mutant of p66Shc. These experiments comprehensively illustrate that lysine-81 SUMOylation of p66Shc is necessary for the observed endothelial dysfunction in hyperlipidemic conditions.

      Weaknesses:

      One notable weakness is that the link between the observed cellular changes and the ultimate in vivo phenotype remains only partially explored. While the authors successfully show that p66ShcK81R knockin mice are protected from endothelial dysfunction in a hyperlipidemic context, additional experiments characterizing the broader tissue-specific roles, or examining further endothelial assays in vivo, would strengthen the mechanistic conclusions. It would also be beneficial to see more direct evaluations of p66Shc subcellular localization in the protective knockin mice to complement the proteomic findings.

      We agree with the reviewer’s suggestion. However, due to limited resources, we could not pursue additional studies.

      Despite these gaps, the data broadly support the authors' main conclusions. The authors lay out a plausible mechanistic pathway for how hyperlipidemia and increased global SUMOylation can converge on the oxidative stress pathway to provoke vascular dysfunction.

      The likely impact of this work on the field is noteworthy. Beyond clarifying how a single post-translational modification event can influence the pathophysiology of endothelial cells, the study provides a model for investigating broader roles of SUMO2 in other cardiovascular conditions and highlights the importance of identifying additional SUMOylation sites and their downstream impact.

      In conclusion, by demonstrating the direct SUMOylation of p66Shc at lysine-81 and linking that modification to endothelial dysfunction in a hyperlipidemic mouse model, this paper offers valuable insights into how broadly acting post-translational modifiers can evoke specific pathological effects.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please rearrange the figures. It is very hard to follow.

      We rearranged some of the figures for a better presentation.

      Reviewer #2 (Recommendations for the authors):

      **Please note that I am not an expert of mouse work therefore I will not be able to comment on the mouse work (Figure 7).

      The following suggested changes major concerns will strengthen the work, and the minor concerns will increase the paper's accessibility to eLife's broad scientific community.

      Major concerns.

      (1) All the work done here is based on overexpression studies, therefore it will be very useful to show that p66Shc gets SUMO2ylated under physiological conditions using SUMO-Trap beads or using tandem SUMO-Interacting Motifs (as shown in Silva-Ferrada et al., 2013: https://doi.org/10.1038/srep01690).

      We have demonstrated endogenous SUMO2 conjugation of p66Shc under physiological conditions by immunoprecipitation using a custom-generated antibody specific for SUMO2-p66Shc. This approach allows detection of SUMO2ylated p66Shc without reliance on overexpression systems, thereby confirming that SUMO2 modification of p66Shc occurs endogenously.

      (2) The authors claim that K81 is the only site of SUMO2ylation and based on the HUVEC experiments (Fig 3C, D) it appears that p66Shc K81R mutant still gets SUMO2ylated indicating that there are other residues which can get SUMO2ylated. Do you see the other sites getting SUMO2ylated in your mass spec data?

      Our mass spectrometry analysis did not identify SUMO2 modification on lysine residues other than K81. However, we acknowledge that SUMOylation detected in vitro on recombinant protein may differ from SUMOylation occurring in a cellular context, where protein conformation, interacting partners, and local enzyme availability can influence modification patterns. These differences may account for the appearance of multiple SUMOylated p66Shc species in cell lysates despite the absence of additional SUMOylation sites in the mass spectrometry dataset. Our emphasis on K81 is based on its localization within the CH2 domain of p66Shc, a region unique to the p66 isoform. Previous studies have demonstrated that post-translational modifications within the CH2 domain, such as phosphorylation, are critical for activation of the oxidative and pro-apoptotic functions of p66Shc. Accordingly, we focused our mechanistic analyses on SUMO2ylation at K81. Nevertheless, we do not exclude the possibility that additional lysine residues within the PTB or CH1 domains of p66Shc may also undergo SUMOylation in cells. Importantly, our functional data support the conclusion that SUMO2ylation at K81 is a key regulatory modification driving the oxidative activity of p66Shc.

      (3) In Figure 4 the authors claim that SUMO2ylation at K81 is a prerequisite for S36 phosphorylation, however the phosphorylation changes are marginal (Figure 4 A, E) and why do they not see the higher molecular weight bands in the western blots.

      We appreciate the critique and agree with the reviewer that our data do not provide direct evidence that K81 SUMO2ylation and S36 phosphorylation occur simultaneously on the same p66Shc molecule. Rather, our findings suggest that SUMO2ylation at K81 may facilitate or enhance S36 phosphorylation, but is not an absolute requirement for this modification. Accordingly, we have revised the wording in the main text to avoid implying a strict prerequisite relationship and to more accurately reflect the magnitude of the observed changes in S36 phosphorylation.

      (4) The authors further claim that the SUMO2ylation at K81 effects S36 phosphorylation which has consequences for mitochondrial localization. However, K81R still translocate to mitochondria which is indicative of SUMO2ylation is important but not necessary (Fig 5A, C).

      We agree with the reviewer’s observation and have revised the main text accordingly, as described in our prior response, to clarify that K81 SUMO2ylation facilitates, but is not essential for, S36 phosphorylation and mitochondrial translocation.

      (5) To know if the K81 SUMO2ylation had any effect on phosphorylation it would be interesting to see how the phosphomimic (S36D) mutant along with/without K18R will behave in mitochondrial localization experiments.

      We agree that examining the double mutant (S36D/K81R) would provide valuable mechanistic insight into the relationship between K81 SUMO2ylation and S36 phosphorylation in regulating mitochondrial localization. Although we are currently unable to perform these experiments due to constraints in manpower and resources, we have acknowledged this as a limitation in the Discussion and plan to pursue these studies in future work to further validate the mechanism.

      Minor Concerns

      (a) In Figure 1A the authors claim that SUMO2ylation of p66Shc promotes ROS production however they do not include any well-known ROS induced proteins in the western blots such as KEAP1 or NRF2 or any other marker.

      We agree that NRF2 is a well-established transcriptional regulator of antioxidant responses. However, the primary aim of Figure 1A was to directly examine the effect of SUMO2ylation on p66Shc-mediated ROS production. While NRF2 and other ROS-responsive proteins reflect downstream adaptive responses, they do not provide a direct measure of ROS generation. Therefore, we focused on measuring SUMO2-induced changes in cellular ROS levels. To further strengthen this conclusion, we have performed additional complementary ROS assays, which are now included in the revised manuscript.

      (b) In Figure 2D, the concentration of Anacardic acid used is low and therefore the SUMO2ylation is not completely inhibited. It would be advisable if the authors go to concentration of complete inhibition.

      We acknowledge that some residual SUMO2ylation is visible in the immunoblots following anacardic acid treatment. However, the assay demonstrates a clear, dose-dependent reduction in SUMO2ylation levels, which is sufficient to interpret the effect of SUMO2 inhibition on p66Shc function. Increasing the concentration further could introduce off-target effects, and the current assay provides a reliable and physiologically relevant readout.

      (c) The figure panels need uniform fonts and also the size/resolution of the western blots needs to be improved

      We thank the reviewer for this suggestion. All figures have been updated to ensure uniform fonts, and the western blot images have been improved for size and resolution in the revised manuscript.

      (d) Supplemental figures are very low resolution.

      High-resolution images have been provided.

      (e) Please introduce abbreviations before using them (like LDLr and ND).

      The abbreviations have been explained.

      (f) Figure 6A seems like data that belongs in the SI rather than the main text.

      We kept it in main figure as the knock-in mouse was generated for this study.

      (a) Figure 7B needs a WT HFD control.

      It is a valid suggestion. However, it is known in the literature (we have prior experience as well) that wild-type mice do not exhibit a drastic increase in serum lipid level with high-fat diet feeding.

      (b) The differences in Figure 7E are not profound and immediately obvious. Is there a better way to show the data, potentially with quantification?

      We agree, the difference is not that huge, we have provided the qualification as Fig. 7F.

      (h) Figure panels 6E-H could use some labels to differentiate the data from WT vs p66ShcK81R mutant mice within the figure panel.

      We mentioned that on the top and have inserted a vertical line to separate the groups.

      (i) Lines 165-168, 182-185, 214-218 and 282-286: These sentences can be rewritten for more clarity.

      We have modified these sentence for clarity.

      (j) Line 512 should read, "Two-way ANOVA, ***P<0.001."

      Thank you, we have corrected the mistake.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely appreciate you and the reviewers for investing time and effort in evaluating our manuscript. After carefully reading the comments and suggestions, we found they are insightful, constructive, and critical for improving the quality of our work. Based on these valuable recommendations, we have substantially revised the manuscript as summarized below.

      Abstract: Inappropriate or ambiguous statements have been revised to improve clarity.

      Introduction: (1) The study purpose and hypotheses have been re-organized in a clearer and more concise way. (2) A mechanistic rationale for sequence-specific degradation has been provided and the use of PMA treatment has been explained. (3) The terminologies related to extracellular DNA and 16S rRNA gene amplicons have been clarified.

      Materials and Methods: 1) More detailed description of the microcosm experiment has been added. 2) The design and rationale of GAPDH F-tagged primers and the use of fusion primers for Illumina library preparation have been clarified. 3) More details about PCR amplification, DNA purification, and pooling strategies have been added. 4) We have corrected and standardized primer naming throughout the manuscript; 5) More details about bioinformatic workflow have been added. 6) We have defined statistical parameters and multiple testing corrections. 7) All abbreviations have been defined and standardized.

      Results: 1) The terminology for PMA-treated DNA has been revised and it has been clarified interpretation as “PMA-treated prokaryotic community” rather than “living community”. 2) The figures and legends have been updated for clarity, and the explicit explanation of “ASV I” and “ASV II” in pairwise comparisons have been added. 3) the figures (e.g., Figs. 2–5, S2–S8) have been reorganized to better reflect results; 4) Inappropriate statements or misleading interpretations have been removed.

      Discussion: A detailed section on technical limitations have been added. The limiatons added mainly include: 1) PCR amplification bias and recommendations for spike-in standards or multi-primer approaches; 2) differential DNA extraction efficiency due to variable cell lysis; and 3) limitations of using 16S rRNA amplicons as proxies for natural extracellular DNA and the limitations of PMA treatment efficiency in soil matrices;

      eLife Assessment

      This valuable study introduces an innovative experimental design to address a crucial and timely issue in microbial ecology: the potential bias in soil microbial community analyses caused by extracellular DNA degradation. While the evidence showing variable degradation rates of extracellular DNA is convincing, additional conceptual, methodological, and statistical clarifications could reinforce the claims and the study's contribution to the field. This research will appeal to microbial ecologists and researchers interested in using molecular techniques to evaluate microbial community structure.

      We sincerely appreciate the editors for the careful assessment of our work and for recognizing the value of addressing extracellular DNA degradation in soil microbial community analyses. We also greatly appreciate the reviewers’ constructive feedbacks concerning the need for additional conceptual, methodological, and statistical clarifications. We agree that further refinement in these areas will strengthen our claims and enhance the study’s contribution to the field. Based on these insightful suggestions, we have carefully revised the manuscript to provide clearer conceptual framework, more detailed methodological descriptions, and more rigorous statistical analyses. We believe these revisions have substantially improved the clarity and robustness of our work. More details about the revisions have been provided in the following responses.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates the degradation dynamics of extracellular DNA in soils and its impact on estimates of microbial abundance and diversity. By combining a broad geographic sampling design with a primer-labeling strategy, qPCR quantification, amplicon sequencing, and PMA treatment, the authors aim to disentangle total versus intracellular DNA signals and explore sequence-specific degradation patterns. The topic is relevant, particularly given the increasing awareness of relic DNA as a confounding factor in microbial ecology. The experimental design is ambitious and potentially impactful. However, several conceptual inconsistencies, methodological ambiguities, and statistical limitations currently weaken the robustness of the conclusions. These issues need to be addressed.

      We sincerely appreciate the reviewer for the constructive assessment of our work. We also appreciate the reviewer’s critical insights regarding the conceptual inconsistencies, methodological ambiguities, and statistical limitations that currently weaken the robustness of the conclusions. We agree with the reviewer that addressing these issues is essential to strengthen our work. Based on these valuable comments, we have carefully revised the manuscript to clarify the conceptual framework. Additionally, we have provided more detailed methodological descriptions, and enhance the statistical rigor of our analyses. We believe these revisions have substantially improved the clarity, consistency, and overall robustness of our conclusions.

      Strengths:

      The manuscript addresses a timely and important question in microbial ecology, particularly given the growing recognition that relic DNA can bias interpretations of community composition derived from amplicon sequencing. The study is ambitious in scope, incorporating a broad geographic sampling design across multiple soil types, which enhances the generalizability of the findings. The use of a controlled microcosm experiment combined with a primer-labeling strategy to track extracellular DNA dynamics is conceptually innovative and provides a structured framework to investigate degradation processes.

      In addition, the integration of multiple approaches, including qPCR for absolute quantification, high-throughput sequencing for community profiling, and PMA treatment to differentiate extracellular from intracellular DNA, represents a comprehensive attempt to disentangle complex sources of bias in soil microbiome analyses. The effort to link degradation dynamics with environmental variables and to explore sequence-level patterns further demonstrates the authors' intent to move beyond descriptive analyses toward a mechanistic understanding.

      We sincerely thank the reviewer for the positive and encouraging comments of our work.

      Weaknesses:

      Several conceptual and methodological issues currently limit confidence in the study's conclusions. Key terms such as "sequence-specific degradation" are not clearly defined or supported by a mechanistic or structural hypothesis, making it difficult to interpret the biological meaning of the results. In addition, the bioinformatic workflow presents inconsistencies, particularly the use of ASVs followed by clustering at 97% similarity, which undermines the resolution required to support sequence-level inferences. Statistical analyses are also insufficiently described, including unclear definitions of "T values," a lack of detail on pairing structure, and no indication of multiple testing correction.

      Furthermore, important methodological details are missing or unclear, including primer design (e.g., GAPDH tag vs ACTF), Illumina library preparation (e.g., adapter and indexing strategy), and validation of PMA treatment efficiency. The interpretation of PMA-treated samples as representing "living communities" is likely overstated, given the known limitations of the method in soil systems. Finally, typographical errors, inconsistent terminology, and unclear phrasing throughout the manuscript reduce readability and further complicate interpretation.

      We sincerely appreciate the reviewer’s thorough and critical evaluation of the manuscript’s weaknesses. The issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, methodological transparency, and the interpretation of PMA treatment have been fully acknowledged. We also recognized that typographical errors, inconsistent terminologies, and unclear phrasing largely reduced readability. In response to these valuable comments, the manuscript has been carefully revised as follows. (1) The clearer definition of the term “sequence-specific degradation” has been provided. (2) The bioinformatic workflow was streamlined to ensure consistency. (3) The descriptions of statistical analyses were substantially expanded, including explicit definitions of “t values,” detailed clarification of the pairing structure, and the application of appropriate multiple testing corrections. (4) Missing details regarding primer design, Illumina library preparation, and PMA treatment validation have been added to the Method section. (5) Interpretations of PMA-treated samples have been revised to more accurately reflect methodological limitations in soil systems. (6) The manuscript has been thoroughly proofread to correct typographical errors, standardize terminology, and enhance overall clarity. These revisions are believed to substantially address the concerns raised and significantly strengthen the manuscript. A point-by-point response to the specific comments is provided below.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes the results of an interesting study examining the rate of degradation of extracellular DNA in soil ecosystems using a clever experimental approach. 16S ribosomal RNA genes were amplified from soil samples, and then purified PCR amplicons, containing a 5' linker sequence on the forward primer, were introduced to soils and monitored over time using real-time quantitative PCR and NGS amplicon sequencing. The study was able to measure rates of overall extracellular DNA degradation, but also sequence-specific degradation rates. I like the idea and execution of the study, and the results are interesting. The manuscript needs some help to improve the overall readability. Please see general and editorial comments below.

      We sincerely thank the reviewer for the positive and encouraging assessment of our study. We have carefully revised the manuscript to enhance clarity, streamline the presentation, and refine the language throughout. We believe these improvements have made the manuscript more readable and easier to follow. We are also grateful for the general and editorial comments provided, which have been addressed as outlined below.

      Strengths:

      Innovative experimental design that is well deployed across a large number of soil types, revealing interesting variability in extracellular DNA degradation.

      We sincerely thank the reviewer for the positive and encouraging assessment of our work.

      Weaknesses:

      (1) The manuscript needs another review to improve the readability of the document.

      We thank the reviewer for this helpful suggestion. We fully agree that improving readability is essential for effectively communicating our findings. Based on the comment, we have carefully revised the manuscript to enhance clarity and readability. We have streamlined sentence structures, standardized terminology, corrected typographical errors, and improved the logical organization of the text. We believe these revisions have substantially improved the overall readability of the manuscript.

      (2) The authors have used 16S genes to look at sequence-specific degradation. But 16S rRNA genes are actually pretty well conserved, and there isn't as much genetic variation across this gene among organisms as there is for other genes. It might be more relevant to look at metagenomic DNA degradation from high AT, high GC organisms, etc. This would be more generalizable than 16S genes.

      We thank the reviewer for this insightful comment. We agree with the reviewer that 16S rRNA genes are relatively conserved compared to functional genes or whole metagenomic DNA, and that studying degradation of more variable sequences (e.g., high‑AT, high‑GC regions, or metagenomic DNA) would provide greater generalizability. However, we would like to clarify the rationale for using 16S rRNA gene amplicons in the present study. First, the 16S rRNA gene remains the most widely used phylogenetic marker in soil microbial ecology (Knight et al., 2018). Demonstrating sequence‑specific degradation with this well‑established marker directly informs a large body of existing research that relies on 16S RNA gene‑based community analyses. Second, despite its conserved nature, the targeted fragment in this study is belong to the highly varied region (V4) of 16S rRNA gene. Accordingly, we indeed observed significant sequence‑specific variation in degradation rates among different 16S rRNA gene amplicon sequence variants (ASVs) (Fig. 2c, 3a). This indicates that even within a conserved marker gene, sequence‑dependent degradation biases exist and can affect diversity estimates. Third, our study was designed as a proof‑of‑concept to establish a methodological framework for quantifying both overall and sequence‑specific degradation rates. Using a single, well‑characterized marker allowed us to develop and validate the primer‑labeling and qPCR/sequencing workflow without the additional complexity of metagenomic DNA (e.g., variable fragment lengths and complex mineral associations). In the revised manuscript, we have added the following sentence to the Discussion section to address the concerns from the reviewer.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015)”

      (3) Consideration of differential cell lysis during soil DNA extraction needs to be considered as well.

      We thank the reviewer for raising this important technical consideration. We agree that differential cell lysis during soil DNA extraction is a well‑recognized source of bias in microbial community analysis. Different microbial taxa (e.g., Gram‑positive vs. Gram‑negative bacteria, spores, or fungi) vary in their cell wall structure and susceptibility to lysis, which can lead to under‑representation of certain groups and over‑representation of others in the extracted DNA. This bias affects both total DNA extracts and PMA‑treated fractions, potentially influencing our estimates of the relative contributions of intact‑cell derived DNA versus extracellular DNA. However, currently, eliminating these biases are still challenging, and thus we have added the following sentence to the Discussion to address this concern.

      L305-311

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (4) It is not clear why the authors didn't put GAPDH linkers on the reverse primer as well. This would have given an easier amplicon to amplify (no degeneracies at all).

      The decision to place the GAPDH linker only on the forward primer (515F) was intentional to balance the need for tracking exogenous extracellular DNA with amplification efficiency, sequencing quality, and cost-effectiveness. Adding linkers to both primers would increase the total amplicon length, potentially reducing amplification efficiency, especially in complex soil samples with degraded or low-quality DNA. More importantly, the reverse primer used in our study is a degenerate primer designed to target the 16S rRNA gene across diverse bacterial taxa, and extending it with an additional GAPDH linker could introduce further complexity, decrease amplification efficiency, and increase primer-dimer formation. Additionally, single-end labeling allows the usage of standard 16S rRNA reverse primers with existing barcodes, whereas dual-end labeling would require synthesis of new barcode-labeled primers, increasing both cost and time. Our preliminary experiments confirmed that single-end labeling produced reproducible amplification curves (~85% efficiency) and high-quality sequencing reads, which were sufficient for quantifying degradation rates. We have added a clarification in the Methods section to explain this rationale.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Inconsistency between ASV inference and 97% sequence recruitment

      The bioinformatic pipeline presents a major conceptual inconsistency. ASVs are inferred using UNOISE3, but reads are subsequently mapped to ASVs at a 97% similarity threshold, effectively reintroducing OTU-level clustering. Given that the manuscript's central claim is sequence-specific degradation, this step undermines the single-nucleotide resolution that ASVs provide and may obscure biologically meaningful differences. The authors should either reanalyze the data using a consistent ASV framework (exact matching), or explicitly treat the analysis as OTU-like and moderate claims of sequence specificity.

      We appreciate the reviewer's critical evaluation of our bioinformatic pipeline. We also apologized for our unclear statements in our original manuscript. We understand the concern that mapping reads to ASVs at 97% similarity might appear to reintroduce OTU-level clustering. Actually, we used the default pipeline provided by the authors of USEARCH with otutab command to generate the ASV table.

      Following the logic and recommendations of the USEARCH/UNOISE3 developer (Robert Edgar), this approach is a standard procedure for robust noise management rather than a conceptual inconsistency. First, in our pipeline, ASVs (ZOTUs) are first inferred using the UNOISE3 algorithm, which effectively identifies “true” biological sequences at single-nucleotide resolution. Secondly, according to the USEARCH manual, while the ASVs themselves represent exact biological sequences, the raw reads inevitably contain stochastic sequencing errors. Using “exact matching” for recruitment would discard a significant portion of the data that originates from a specific ASV but carries minor random errors. Mapping at 97% identity is the recommended method to recruit these noisy reads back to their correct biological origin (the ASV centroid). Meanwhile, during the recruitment process, reads are not randomly assigned to any ASV within the 97% identity radius. Instead, the algorithm follows a “highest similarity first” principle. For instance, if a specific read exhibits 98% similarity to ASV1 and 97% similarity to ASV2, it is strictly assigned to ASV1. A read is only discarded if its highest similarity to any ASV falls below the 97% threshold. Unlike traditional OTU clustering (where sequences are clustered together based on similarity from the start), our approach maintains the ASV as a fixed biological reference. The quantification of degradation rates is performed on these high-resolution centroids. Thus, our claims regarding sequence-specific degradation remain valid, as the underlying biological variation is defined by the ASVs. To avoid any possible confusion, we have rewritten the relevant paragraph in the Methods section (subsection 4.6) as follows.

      L446-457

      “ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework.”

      (2) Undefined "ASV I" and "ASV II" groups

      The manuscript refers to "ASV I" and "ASV II" groups in pairwise comparisons of degradation rates (e.g., Fig. 3), but these groups are not defined anywhere in the text.

      It is unclear whether these represent: predefined biological categories, arbitrary pairwise ASV comparisons, or groupings based on taxonomy, abundance, or degradation rate.

      In addition, the pairing structure underlying these comparisons is not described. While a paired t-test is mentioned, it is unclear how ASVs were paired (e.g., within sites, across samples, or across time points).

      The current terminology ("groups") is potentially misleading and suggests biological structure where none may exist. The authors should explicitly define these terms, clarify the pairing scheme, and revise terminology if these are simply pairwise comparisons.

      We thank the reviewer for this keen observation. We completely agree that the terms "ASV I" and "ASV II" were poorly defined and potentially misleading.

      We would like to clarify that "ASV I" and "ASV II" were not intended to represent predefined biological categories (such as groupings based on taxonomy, abundance, or degradation rates). Instead, they were merely used as a labeling convention to indicate the directionality of pairwise comparisons within the heatmap matrix. Specifically, "ASV I" referred to the ASVs represented in the rows, while "ASV II" referred to those in the columns. In the original Fig. 3, blue indicated that the degradation rate of the row ASV was significantly lower than that of the column ASV, and red indicated the opposite. To avoid any confusion, we have removed the “ASV I/II” terminology throughout the manuscript and figures, replacing them with “Row ASVs” and “Column ASVs”. To address this issue, we have revised the Figure 3 legend to include a more explicit explanation:

      L820-824

      “In the heatmap, each cell represents a pairwise comparison between two ASVs. Blue indicates that the degradation rate of the ASVs listed in the row (row ASVs) is significantly lower than that of the ASVs listed in the column (column ASVs); red indicates that the row ASVs has a significantly higher degradation rate than the column ASV. A positive t value indicates that the row ASVs degrades significantly faster than the column ASVs; a negative t value indicates the opposite.”

      We thank the reviewer for raising the important issue regarding the definition of “t values” in our statistical analysis. We apologize for the lack of clarity in the original manuscript. To clarify, the T values presented in Figure 3a represent the test statistics (t-values) from paired t-tests comparing the degradation rate constants of two ASVs across the 30 study sites. The T-value was obtained from a paired t-test between two ASVs across the same samples. The t-value indicates the magnitude and direction of the difference between the two ASVs’ degradation rates relative to the variability across sites. A positive t-value (colored red in the heatmap) indicates the row ASVs degrades significantly faster than the column ASVs; a negative t-value (colored blue) indicates the opposite.

      L544-548

      “As for the analysis, we performed paired t‑tests across all the study sites. Thus, the degradation rates were essentially compared within each site, with both values originating from a same soil sample under identical incubation conditions. A positive t value indicates that the first ASV has a significantly higher degradation rate than the second one, and a negative t value indicates the opposite. The p values were adjusted for multiple comparisons using the FDR method.”

      (3) Lack of definition and justification of "T values"

      The manuscript reports "T values" for comparisons between ASVs but does not clearly define how these values are calculated. Although a paired t-test is mentioned, it remains unclear how the pairing was constructed, whether assumptions (normality, independence) were evaluated, and whether corrections for multiple comparisons were applied. Given the large number of ASVs, failure to control for multiple testing could inflate false positives. More broadly, the use of a simple paired t-test may not be appropriate given the hierarchical and compositional structure of the data.

      We sincerely thank the reviewer for pointing out the need to clarify the definition and justification of the t-values presented in our manuscript. Each t-value represents the test statistic from a paired t-test comparing the degradation rates of two ASVs across the same set of samples. The paired t-test assumes that the differences between paired observations are approximately normally distributed and that the pairs are independent across columns. We have evaluated the normality of differences using standard diagnostic plots and verified that the assumption is reasonably satisfied given the sample size. We performed a correction for multiple comparisons using the False Discovery Rate (FDR) procedure to control for potential false positives. We have revised the Methods section to clearly define t-values.

      (4) Conceptual validity of "sequence-specific degradation"

      The manuscript repeatedly refers to "sequence-specific degradation" of extracellular DNA; however, this concept is not clearly defined nor supported by a biological or structural hypothesis. It is unclear what "sequence-specific" refers to (e.g., nucleotide composition, GC content, secondary structure, taxonomic identity), whether differences are expected in conserved versus variable regions of the 16S rRNA gene, or what mechanistic basis would explain differential degradation among sequences. Given that the analysis is based on short 16S V4 amplicons, and no structural or biochemical framework is provided, it is difficult to interpret whether the observed differences truly reflect intrinsic sequence-dependent degradation or are instead driven by methodological or statistical artifacts (e.g., abundance effects, amplification bias).

      I believe the authors should explicitly define what is meant by "sequence-specific degradation," provide a biologically grounded hypothesis (e.g., structural accessibility, GC content, stem-loop stability), and align their interpretation with the resolution and limitations of the data.

      We thank the reviewer for this critical conceptual comment. We apologize that “sequence‑specific degradation” was not clearly defined and lacked a biological or structural hypothesis. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework (e.g., GC content, thermodynamic stability, and secondary structures) presented in the paragraph.

      We now define “sequence‑specific degradation” as statistically significant differences in first‑order degradation rate constants among distinct ASVs, mainly arising from intrinsic DNA properties (base composition, secondary structure, and restriction sites) or differential mineral adsorption.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      We also expanded the mechanistic discussion to include GC content and secondary structure.

      L230-235

      “We also examined whether GC content could explain the observed sequence‑specific patterns, but no significant correlation was found (Fig. S4), suggesting that simple base composition is not the primary driver in this study. However, this does not exclude the possibility that higher‑order structural features (e.g., hairpin loops) or sequence‑specific nuclease recognition motifs contribute to differential degradation (Wang et al., 2007). This should be tested in future studies using synthetic DNA constructs with controlled structural elements.”

      We acknowledge that inferring sequence‑specific degradation from combined relative abundance and qPCR data is subject to potential methodological artifacts, including compositional effects, PCR amplification bias, and abundance‑dependent detection limits. However, we have taken several stringent steps to minimize these concerns. Specifically, we restricted our analysis to ASVs that were present in more than 90% of the study sites and for which the degradation curve fits yielded R<sup>2</sup> > 0.5, ensuring that only robustly detected and reliably modeled sequences were retained. Because our analysis tracks the ratio of each ASV across a time series, any sequence-specific PCR amplification bias remains constant for that particular sequence. By focusing on the rate of change rather than absolute read counts, such systematic biases are mathematically canceled out during the calculation of degradation kinetics.

      (5) Conceptual ambiguity in "GAPDH F-labeled 16S rRNA genes"

      The manuscript repeatedly refers to "GAPDH F-labeled 16S rRNA genes," which is confusing and may be misinterpreted as targeting GAPDH rather than 16S. It should be clearly stated that GAPDH refers to glyceraldehyde-3-phosphate dehydrogenase, and a GAPDH-derived sequence is used as a synthetic tag appended to a 16S primer. Additionally, the divergence of this tag from microbial sequences should be justified to ensure specificity. There is also an inconsistency in primer naming (e.g., "GAPDH F" vs "ACTF" in the figures), which should be corrected.

      We sincerely thank the reviewer for this important comment. We agree that the phrase “GAPDH F‑labeled 16S rRNA genes” could be confusing, as it may be misinterpreted as targeting the GAPDH gene rather than the 16S rRNA gene. We have revised the manuscript to avoid this ambiguity and to provide clear justification for the use of the GAPDH tag. GAPDH (glyceraldehyde‑3‑phosphate dehydrogenase) is a human housekeeping gene. Its forward primer sequence (GAPDH F: 5′‑CAT TGG CAA TGA GCG GTT C‑3′) was used as a synthetic tag appended to the 16S primer because (i) no homologous sequences exist in soil DNA (confirmed by PCR), and (ii) its melting temperature is compatible with the reverse primer. This tag allows specific tracking of exogenous DNA without interference from native soil sequences.

      Throughout the manuscript, ambiguous phrases such as “GAPDH F‑labeled 16S rRNA genes” have been replaced with more precise terms, “GAPDH F‑tagged 16S rRNA gene amplicon fragments” clarifying that the tag is an appendage and not the amplification target.

      We have checked the entire manuscript and confirm that “ACTF” does not appear anywhere. To avoid confusion, the primer is now consistently referred to as “GAPDH F” in all figures, legends, and text.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (6) Limitations of using PCR amplicons as proxies for extracellular DNA

      The study uses PCR-generated amplicons to simulate extracellular DNA. While useful for controlled comparisons, these fragments may not reflect the physicochemical diversity of natural extracellular DNA (e.g., adsorption to minerals, fragment size variability, protection within aggregates). This limitation should be explicitly acknowledged, and conclusions should be framed accordingly.

      We appreciate the reviewer’s constructive feedback. We fully acknowledge that using PCR-generated amplicons to simulate extracellular DNA (eDNA) has inherent limitations in capturing the full physicochemical diversity of naturally occurring eDNA in soils. Specifically, we agree that PCR fragments may not replicate features such as highly variable fragment size distributions, associations with complex cellular components (e.g., vesicles or protein complexes), or long-term physical sequestration within soil micro-aggregates. Despite of these limitations, the use of uniform primer-tagged PCR amplicons was a deliberate choice to enable precise tracking of exogenous DNA degradation kinetics while eliminating background interference from endogenous soil eDNA. This design is a prerequisite for the high-resolution kinetic modeling of sequence-specific decay. Furthermore, in our bioinformatic pipeline, the 97% mapping threshold was specifically applied to minimize the influence of stochastic sequencing and PCR errors on abundance quantification. In the revised manuscript, these potential limitations have been addressed.

      L295-299

      “First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020).”

      (7) Interpretation of sequence-specific degradation

      Sequence-specific degradation rates are inferred from combining relative abundance data with total qPCR estimates. This approach is sensitive to compositional effects, amplification biases, and abundance-dependent detection limits. It remains unclear whether observed differences reflect true sequence-specific degradation or methodological artifacts. This limitation should be discussed more explicitly.

      We thank the reviewer for highlighting this critical methodological point. In our study, sequence-specific degradation rates were estimated by combining ASV-relative abundances with total qPCR-derived 16S rRNA gene copy numbers. We acknowledge that this approach may be influenced by compositional effects, PCR amplification biases, and abundance-dependent detection limits. However, the degradation rate constant (k) in our study, represents the rate of change for a specific sequence over time. Since PCR amplification biases are generally sequence-specific and consistent across samples processed under identical conditions, these systematic errors are mathematically canceled out when calculating the relative change (slope) for the same ASV across a time series. Second, all qPCR measurements were performed with three technical triplicates with standard curves to ensure quantitative reliability. Third, relative abundances were converted to absolute abundances using total qPCR estimates, allowing cross-taxa comparisons that reduce compositional bias. This approach is widely recognized in microbial ecology as a robust method. To address this concern, we have added some explanations in the revised manuscript.

      L84-86

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequence variants (ASVs) under identical soil and incubation conditions.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      L311-314

      “While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017)..”

      (8) Overinterpretation of PMA-treated samples as "living communities"

      The manuscript interprets PMA-treated DNA as representing intracellular or "living" microbial communities. While PMA is useful, this interpretation should be treated with caution in soils. PMA efficiency can be affected by soil matrix complexity, DNA adsorption to particles, incomplete light penetration, and permeability of compromised cells. Importantly, no validation of PMA efficiency is presented.

      We thank the reviewer for this important caution. We agree that interpreting PMA‑treated DNA as representing “living” or “intracellular” communities is an overstatement in soil systems. In the revised manuscript, we no longer describe PMA-treated DNA as a direct proxy for the “living community,” but instead refer to it as the “PMA-treated prokaryotic community”.

      Although we did not directly validate PMA efficiency in this study, we used a standardized PMA protocol that has been widely applied in microbial ecology, and our goal was to obtain a comparative estimate of the influence of extracellular DNA on community analysis across soils under a consistent methodological framework. Based on previous studies (Carini et al., 2016; Du et al., 2025), which found that in similar soil types, PMA treatment can significantly reduce the interference of extracellular DNA and alter the community structure, this indirectly proves the effectiveness of this technique.

      Nevertheless, we agree that future studies should include explicit validation controls, such as live/dead cell mixtures, heat-killed controls, or soil-specific PMA efficiency tests, to better quantify method performance across diverse soil matrices. We have added a dedicated paragraph in the "Methodological Considerations and Limitations" section to discuss how soil-specific properties (e.g., turbidity, adsorption capacity) might lead to incomplete exclusion of extracellular DNA, thereby advising a more cautious interpretation of the "viable" community data.

      L496-500

      “To inhibit amplification of eDNA, soils were incubated with propidium monoazide (PMA), as described previously (Carini et al., 2016). Upon photoactivation, eDNA can form covalent bonds through cross-linking, leading to the inhibition of its PCR amplification. In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      Minor Comments

      (1) Line 79: Provide examples of how extracellular DNA contributes to nutrient cycling (e.g., P, N sources) and signal transduction (e.g., horizontal gene transfer).

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have added specific examples to clarify how extracellular DNA contributes to nutrient cycling and signal transduction. Specifically, we now note that extracellular DNA can serve as a source of phosphorus and nitrogen following enzymatic degradation, thereby contributing to soil nutrient turnover. We also clarify that extracellular DNA plays an important role in horizontal gene transfer, acting as a genetic reservoir that can be taken up by competent microorganisms and thereby facilitating the spread of functional traits such as antibiotic resistance. These examples have been added to improve the clarity and biological context of this statement.

      L48-52

      “EDNA serves as a critical vector for horizontal gene transfer (HGT), facilitating the uptake of genetic material by competent microorganisms and promoting the spread of functional traits such as antibiotic resistance (Liu et al., 2024). In addition, eDNA participates in soil biogeochemical cycling because its enzymatic degradation releases bioavailable nutrients, particularly phosphorus and nitrogen, which can be reused by soil microorganisms (Ye et al., 2022).”

      (2) Line 79: Replace "for an extended period of time" with a more precise or referenced timescale.

      We agree with the reviewer. We have replaced the vague phrase with a precise timescale. Extracellular DNA can persist in soils for months to years.

      (3) Line 94: Clarify what is meant by "high-level structure" (e.g., secondary structure, environmental association).

      We thank the reviewer for pointing out this ambiguity. In the original manuscript, the phrase “high-level structure” was not sufficiently precise. In the revised version, we have clarified that this refers primarily to higher-order structural properties of DNA molecules, such as secondary structure, local conformational features, and sequence-dependent interactions with minerals or organic matter in soil. These characteristics may influence the accessibility of extracellular DNA to nucleases and thus affect degradation rates. We have revised the text accordingly to improve clarity and precision.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (4) Line 111: The hypothesis is not clearly linked to the rationale. If sequence-specific degradation is expected, clarify whether it relates to conserved vs variable regions or structural features (e.g., stems vs loops).

      We thank the reviewer for this helpful comment. We agree that the original manuscript did not clearly link the hypothesis regarding sequence-specific degradation to its mechanistic rationale. In the revised manuscript, we have clarified that the expectation of sequence-specific degradation is not simply based on conserved vs variable regions of the 16S rRNA gene, but rather on the potential for sequence differences to influence intrinsic physicochemical properties, including base composition, local conformational features, potential secondary structures, and motif-dependent nuclease susceptibility. These factors may alter DNA accessibility to extracellular nucleases, providing a mechanistic basis for sequence-specific degradation. This clarification is now reflected in the Introduction and linked to the formal hypothesis statement.

      To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses (H1–H3) immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (5) Lines 310-311: Clearly indicate which portion of the primers corresponds to the modified (GAPDH-derived) sequence. Provide full annotated primer sequences.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now clearly indicate which portion of the forward primer corresponds to the GAPDH-derived synthetic tag and which portion corresponds to the 16S rRNA gene primer sequence. We have also provided the full annotated primer sequences in the Methods section to avoid ambiguity.

      Specifically, the modified forward primer is now described as:

      GAPDH-F-515F: 5′-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3′,

      where CAT TGG CAA TGA GCG GTT C is the GAPDH-derived synthetic tag and GTG CCA GCM GCC GCG GTA A is the 16S rRNA gene forward primer sequence (515F).

      The reverse primer is:

      806R: 5′-GGA CTA CHV GGG TWT CTA AT-3′.

      L354-359

      “Briefly, exogenous eDNA was prepared by PCR amplification using a modified forward primer consisting of a GAPDH F tag fused to the 16S rRNA gene primer 515F, together with the reverse primer 806R. The full primer sequences were as follows: GAPDH-F-515F: 5'-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3', in which CAT TGG CAA TGA GCG GTT C represents the GAPDH F tag and GTG CCA GCM GCC GCG GTA A represents the 16S rRNA gene forward primer sequence (515F); and 806R: 5'-GGA CTA CHV GGG TWT CTA AT-3'.”

      (6) Lines 310-311: Explicitly define GAPDH and justify its use as a synthetic tag.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now explicitly define GAPDH as glyceraldehyde-3-phosphate dehydrogenase, a human housekeeping gene. Specifically, the GAPDH-derived sequence was selected for two reasons. First, it is highly divergent from known soil microbial 16S rRNA gene sequences and did not produce detectable amplification when tested with soil DNA using the GAPDH tagged 806R primer pair, indicating that it would not interfere with endogenous soil DNA signals. Second, its melting temperature was compatible with that of the reverse primer, which allowed stable amplification of the tagged 16S amplicons under our PCR conditions.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing”

      (7) Line 346: Start a new paragraph to clearly separate this as a distinct experiment.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have started a new paragraph.

      (8) Line 346: Specify the number of samples analyzed for consistency.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have now explicitly specified the number of samples.

      A total of 120 samples were analyzed in this moisture gradient experiment: 2 ecosystems (Kaiyuan and Dashanbao) × 5 moisture levels (10%, 25%, 50%, 75%, and 100% of water holding capacity) × 6 incubation time points (0, 1, 3, 6, 12, and 24 days) × 2 replicates. Just two technical replicates were performed for this validation experiment, as the primary aim was to assess the trend of moisture effects rather than statistical inference across replicates.

      L408-410

      “This complementary experiment included two sites, five moisture levels, six incubation time points, and two replicates per treatment combination, resulting in a total of 120 soil samples.”

      (9) Lines 351-352: Replace "harvested" with "collected."

      We have replaced “harvested” with “collected” as suggested.

      (10) Line 367: Clarify how Illumina adapters and indices were added (e.g., two-step PCR, fusion primers).

      We thank the reviewer for this helpful suggestion. We have revised the Methods section to clarify how Illumina adapters and indices were incorporated. We used a pooled amplicon library preparation strategy. Individual samples were first amplified with primers containing sample-specific barcode sequences. The barcoded amplicons from multiple samples were then pooled and used for library preparation with the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were first ligated to the pooled amplicons. After bead-based purification, an indexing PCR was performed using the index primer mix, which introduced the complete P5/P7 sequences and a library-level Illumina index into the library molecules. Thus, sample demultiplexing was based on the sample-specific barcodes introduced during amplicon PCR, whereas the Illumina index was used to identify the pooled sequencing library. We have clarified this procedure in the revised manuscript.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).

      (11) Provide more detail on chimera removal, filtering thresholds, and normalization choices.

      We thank the reviewer for this helpful suggestion. The raw paired-end reads were first merged, and primer sequences were removed using the search_pcr2 script in USEARCH. Reads with more than two primer mismatches were discarded. Quality filtering was then performed using fastq_filter, and sequences with quality scores below 20 were removed. Redundant reads were collapsed using fastx_uniques. Amplicon sequence variants (ASVs) were generated using the UNOISE3 denoising algorithm, which also performs built-in chimaera filtering during ASV inference. In addition, ASVs with total sequence counts fewer than 9 were excluded to reduce the influence of low-frequency noise.

      L443-460

      “Briefly, paired-end reads were merged using USEARCH, and primer sequences (GAPDH-F-515F and 806R) were removed using the search_pcr2 script. Reads with more than two primer mismatches were discarded. Quality filtering was performed using the fastq_filter script, and sequences with quality scores below 20 were removed. Redundant sequences were dereplicated using the fastx_uniques script. ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework. Taxonomic annotation of the ASVs was performed in QIIME2 with the Silva v138 database. A total of 89322 prokaryotic ASVs were obtained. To standardize sequencing depth across samples, the read number of each sample was rarefied to 53251 using the rarefy function in the vegan package in R.”

      (12) Line 412: Rephrase to refer to 16S amplicon addition rather than 16S rRNA genes (along the whole text), as only the V4 region is analyzed.

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (13) Ensure consistent primer naming throughout (e.g., GAPDH F vs ACTF).

      We have checked the entire manuscript and confirm that only “GAPDH F” is used as the label primer.

      (14) Finally, the manuscript would benefit from careful language editing. Several typographical errors, grammatical inconsistencies, and unclear phrases are present throughout. Examples include:

      Misspellings such as "diffrence" (e.g., figure legends) and inconsistent capitalization. Inconsistent terminology (e.g., "genes," "amplicons," and "fragments" used interchangeably without clarification). Redundant or awkward phrasing (e.g., repeated use of "extracellular 16S rRNA genes"). Occasional subject-verb agreement issues and missing articles.

      We sincerely apologize for the language issues. The manuscript has now undergone a thorough language editing process by a native English‑speaking colleague.

      Recommendation

      Major revision: The manuscript addresses an important problem and presents a promising approach. However, key issues related to conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA-based results must be resolved. With substantial revision and clarification, the study has the potential to make a meaningful contribution to the field.

      We sincerely thank the reviewer for the thorough, constructive, and critical evaluation of our manuscript. We greatly appreciate the recognition that our study addresses an important problem and presents a promising approach. We also acknowledge the key issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA‑based results. We have taken these comments very seriously and have substantially revised the manuscript accordingly, more details about the revisions are described in the following point-by-point responses.

      Reviewer #2 (Recommendations for the authors):

      Editorial comments:

      (1) Title: I recommend removing "across China" from the title. In many ways, the study has nothing to do specifically with China, and you limit the broad applicability of the study. The same work could have been done with soils from Africa, for example. Also, it might be ok to remove 16S rRNA as well. The 16S rRNA genes are a proxy for rates of extracellular DNA degradation, but the study isn't exactly about 16S either.

      We thank the reviewer for this thoughtful suggestion regarding the title. We have revised the title to “The overall and sequence-specific degradation of soil extracellular DNA fragments: rates and influential factors.”

      (3) L44-45: "...such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis...".

      We thank the reviewer for this suggestion. We have revised the order according to the suggestions of the reviewer.

      L44-45

      “The investigation of soil microbial abundance and diversity heavily relies on DNA-based technologies, such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis.”

      (4) L48: remove "they".

      We agree with the reviewer and have removed the extraneous “they”.

      (5) L51: "noise factor"; "...persistence can lead to...".

      We have revised the sentence as suggested.

      (6) L53: remove theoretical.

      We have removed “theoretical”.

      (7) L58: remove "the".

      We have removed "the".

      (8) L86: Is restriction digestion of DNA a likely extracellular process in soil?

      We thank the reviewer for this thoughtful comment. We agree that the original wording may have overstated the likelihood of classical restriction digestion as a dominant extracellular process in soils. Our intention was not to suggest that intracellular restriction enzyme systems operate directly in the soil matrix in the same manner as they do within living cells. Rather, we aimed to indicate more generally that sequence-dependent nuclease susceptibility could contribute to differential degradation among extracellular DNA fragments.

      L87-95

      “First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (9) L94-99: The authors might also consider the different nitrogen content of different bases; this might also affect sequence-specific selection of DNA for degradation.

      We thank the reviewer for this insightful suggestion. We agree that differences in the elemental composition of DNA bases, including nitrogen content, may provide an additional mechanistic explanation for sequence-dependent degradation. In the revised manuscript, we have incorporated this point into the Introduction.

      L92-95

      “Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (10) L110-112: These are not really written in hypothesis form. Also, what about a hypothesis about degradation rates and soil type/temperature/moisture?

      We thank the reviewer for this constructive critique. We have rewritten the hypotheses. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (11) L116: "GAPDH F-labeled 16S rRNA gene amplicon fragments....".

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (12) L117: "rapidly".

      We agree with the reviewer and have revised.

      (13) L118-120: "After a 48-day incubation period, 0.2 to 3.1% of the initial spike GADPH F-labeled 16S rRNA gene amplicon fragments ...".

      We agree with the reviewer and have revised.

      (14) L125: Spell out SEM in first usage.

      We thank the reviewer for this suggestion. In the revised manuscript, we have spelled out SEM as Structural equation modeling.

      (15) L128: I don't like the idea of putting this Figure in supplemental materials.

      We thank the reviewer for this suggestion. We have moved Figure S2 (moisture gradient microcosm experiment) to the main text as Figure 1f.

      (16) L154: The term "intracellular prokaryotic abundance" is not the right term. This makes one think of an intracellular parasite. I think you want something like: "Approximately 40% of sequences in total soil DNA extraction NGS amplicon libraries were derived from intact cells, while the remaining represented extracellular DNA. Conversely, greater than 80% of observed richness was derived from intact cells." (Please check that I stated this correctly.) I would also suggest some statistics or ranges here.

      We thank the reviewer for this important terminological clarification. We agree that the term “intracellular prokaryotic abundance” is misleading, as it could imply intracellular parasites. In the revised manuscript, we have replaced this with a clearer description and We have also added the across‑site ranges to provide statistical context.

      L163-166

      “The PMA treatment revealed that intact cells accounted for approximately 40% (range: 9–73%) of the total 16S rRNA gene copies. In contrast, over 80% (range: 27–97%) of the observed ASV richness was associated with sequences originating from intact cells (Fig. 4a and b).”

      (17) L168: "...a significant NEGATIVE correlation was observed...".

      We agree with the reviewer and have revised.

      (18) L169: "However, no significant relationship was observed...".

      We agree with the reviewer and have revised.

      (19) L194-195: What about pH and temperature?

      We thank the reviewer for this comment. We agree that pH and temperature are important environmental factors that can influence microbial DNA degradation and community composition. However, our results (Fig. 1c) indicate that soil moisture is the most dominant factor affecting extracellular DNA degradation. Therefore, in the revised manuscript, we have focused the explanation primarily on soil moisture, while acknowledging that pH and temperature may also be important influencing factors.

      L208-211

      “Third, environmental factors, including soil moisture, pH, and temperature, can predominantly govern enzymatic reaction rates (He et al., 2024; Shah et al., 2024). Indeed, strong positive correlations were observed between moisture content and eDNA degradation rates in both the survey and microcosm experiments (Fig. 1d-f).”

      (20) L199: "findings".

      We have revised as suggested.

      (21) L227-229: This sounds more like results.

      We thank the reviewer for this comment. We agree that the original first sentence in L227–229 reads more like results. Our intention was to introduce the discussion by linking extracellular DNA to potential impacts on prokaryotic community analysis, rather than to present specific findings at this point. We have reorganized this section as follows.

      L246-248

      “Accordingly, we further explored how DNA may influence prokaryotic community analyses using PMA treatment, and significant disparities were observed between the profiles of the total and PMA-treated soil prokaryotic communities (Fig. 4).”

      (22) L230: Need to also consider differential cell lysis during DNA extraction.

      We thank the reviewer for this important comment. We agree that differential cell lysis during DNA extraction could influence the observed community profiles, as microbial taxa differ in cell wall composition and resistance to mechanical or chemical lysis. In the revised manuscript, we explicitly acknowledge this limitation in the relevant section. We also clarify that a standardized DNA extraction protocol (DNeasy PowerSoil kit) was used to efficiently lyse a broad range of microbial taxa, but some taxon-specific lysis bias may remain. Future studies could combine multiple lysis methods or spike-in controls to quantify and correct for potential extraction bias.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools (Deshpande and Fahrenfeld, 2023; Wang et al., 2024).”

      L309-311

      “Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (23) L232: Need to also consider that PMA treatment is not perfect and can be affected by substrate, the ability of light to access DNA for crosslinking, etc.

      We thank the reviewer for this important reminder. We agree that PMA treatment is not perfect and that its efficiency can be affected by soil matrix properties (e.g., organic matter, clay minerals) and the ability of light to penetrate the sample for DNA crosslinking. In the revised manuscript, we have explicitly acknowledged these limitations in the discussion.

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (24) L240: Can extracellular DNA have an ecological role?

      We thank the reviewer for this thoughtful question. Yes, extracellular DNA (eDNA) does have important ecological roles beyond being a potential bias in molecular analyses. In the revised manuscript, we have added statements to highlight that extracellular DNA can serve as a nutrient source (e.g., nitrogen and phosphorus) for microbes and may also contribute to horizontal gene transfer. This emphasizes that extracellular DNA may actively influence microbial community structure and function, in addition to its role in potentially inflating observed abundance and richness.

      L267-275

      “We observed a significant correlation between eDNA degradation rates and the overall structure of the prokaryotic community, but this relationship was absent in PMA-treated communities (Fig. 5b). This discrepancy highlights the divergent ecological roles of extracellular and intracellular DNA. Analyses of the total community integrate intracellular DNA from metabolically active cells with eDNA which primarily originates from historical microbial residues (Lennon et al., 2018). EDNA incorporates signals that likely reflect the legacy effects of past environmental conditions (Wang et al., 2021). In contrast, the PMA-treated community reflects transient microbial activity driven by current selective pressures. Additionally, eDNA can serve as a nutrient source and facilitate horizontal gene transfer, which may further shape its interactions with contemporary microbial communities (Levy-Booth et al., 2007).”

      (25) L256: Why would microorganisms selectively degrade one DNA sequence vs another? This seems to be likely to be stochastic in terms of which sequences are taken up by microorganisms. However, different DNA sequences might hydrolyze differently or be otherwise damaged, and that could lead to differential degradation of a viable amplicon. It might be interesting to incorporate long pieces of DNA with different internal primer sites and use quantitative PCR to determine how sequences are degrading.

      We thank the reviewer for this important mechanistic insight. We agree that the observed correlation between degradation rate and sequence abundance does not necessarily imply active microbial preference. It could equally reflect stochastic encounter rates or intrinsic chemical differences (e.g., AT‑rich regions hydrolyzing faster). We have revised the corresponding paragraph in the Discussion.

      L279-292

      “This finding suggests that abundant eDNA degrades at a faster rate compared to rare eDNA. As mentioned earlier, this could be explained by several mechanisms. First, as soil eDNA is subject to enzymatic degradation and microbial recycling, abundant DNA sequences may be more likely to be encountered and degraded by extracellular nucleases simply due to their higher copy numbers (Levy-Booth et al., 2007; Nagler et al., 2018). Similarly, if microbes preferentially take up DNA as a nutrient source, they may degrade abundant sequences more frequently as a stochastic consequence of higher encounter rates (Finkel and Kolter, 2001). However, we also found that the relationships between the sequence-specific degradation rates and the effect sizes of extracellular 16S rRNA gene amplicon fragments varied across the study sites (Fig. S1g). The sequence-specific effect sizes of extracellular 16S rRNA gene amplicon fragments are mainly determined by both their production and degradation rates (Pietramellara et al., 2009; Sirois and Buckley, 2019). These inconsistent correlations emphasize the critical role played by the production rates of extracellular 16S rRNA genes in influencing the analysis of prokaryotic communities. Therefore, future studies should systematically determine both the production and degradation rates of eDNA.”

      (26) L282-283: This belongs in the discussion.

      We agree with the reviewer and have revised accordingly.

      (27) L289: "as well as measurements of total organic carbon".

      We agree with the reviewer and have revised accordingly.

      (28) L338: Any water content for these soils?

      We thank the reviewer for this comment. The water contents of soils from all study sites are reported in Supplementary Table 2.

      (29) L349-350: You mean that you measured the total soil extracted DNA and then added 1% as labeled 16S?

      Yes, for each soil sample, we extracted total soil DNA and quantified its concentration (ng DNA per gram of soil). We then added exogenous GAPDH‑tagged 16S amplicon fragments at an amount equal to 1% of this total DNA concentration. This concentration was chosen to mimic a realistic pulse of extracellular DNA input without overwhelming the endogenous DNA pool. We apologize for any confusion caused by the imprecise wording in the original manuscript.

      L392-398

      “The microcosm experiment was conducted using 30 g of soil for each sample. After pre-incubation at 20℃ for one week, each soil was thoroughly mixed with the GAPDH F‑tagged 16S rRNA gene amplicon fragments and incubated further at 20℃ (Fig. S8). The amount of exogenous GAPDH F‑tagged 16S rRNA gene amplicon fragments added to each soil sample was equivalent to 1% of the total DNA concentration naturally present in that soil, as determined fluorometrically prior to the experiment. This concentration was chosen to approximate natural eDNA fluxes resulting from microbial lysis, ensuring experimental relevance to in situ conditions (Table S2).”

      (30) L354: Remember that soil recovery from intact cells is going to be lower than for extracellular DNA. So, you are probably overestimating the contribution of extracellular DNA to the total DNA in the system.

      We thank the reviewer for this comment. We agree that DNA recovery from intact cells is generally lower than from extracellular DNA due to differential cell lysis efficiencies. Consequently, the contribution of extracellular DNA to total soil DNA may be somewhat overestimated in our study. We have clarified this limitation in the revised manuscript.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools.”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (31) L362: Amplification efficiency is pretty low. I think you would have been better served with GAPDH on both ends, and that would have given you a much higher efficiency qPCR.

      We thank the reviewer for this comment. The actual qPCR amplification efficiency in our assay was approximately 85%, which, although slightly below the ideal range, was still acceptable and produced reproducible amplification curves and reliable quantification for degradation-rate calculations.

      We acknowledge that the amplification efficiency in our qPCR experiments using a GAPDH F-labeled 16S primer on one end was suboptimal. The current design used a single GAPDH tag at the forward primer to avoid potential amplification bias or primer-dimer formation that could arise from extending the degenerate reverse primer. In addition, dual-end labeling would have required synthesis of new barcode-labeled tagged primers, increasing both cost and experimental complexity. Thanks again for the constructive comments, which provided us with the direction for future experiment optimization.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (32) L367: Not enough detail on how barcoded libraries were made. UDIs?

      We thank the reviewer for this helpful comment. We have now clarified the library preparation and indexing strategy in the revised Methods section. This amplicon diversity sequencing used a pooled-library strategy. Individual samples were first distinguished by sample-specific inline barcodes introduced during the amplicon PCR step. After amplification, barcoded PCR products from multiple samples were pooled and subjected to library construction using the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were ligated to the pooled amplicons, followed by an indexing PCR that introduced the complete P5/P7 sequences and a library-level Illumina index. Thus, the Illumina index was used to identify the pooled sequencing library, whereas sample demultiplexing was performed according to the sample-specific inline barcodes. We have revised the Methods section to make this procedure explicit.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (33) L368: Why was the # of cycles increased?

      Thank you for your question. In the original manuscript (L368), we stated that the number of PCR cycles was increased to 35. This was mainly because the exogenously added GAPDH F‑labeled 16S rRNA genes had a relatively low initial abundance in the soil and gradually degraded during the microcosm incubation, with their copy numbers becoming particularly low at the last time points (see Fig. 1a). To ensure sufficient PCR product for high‑throughput sequencing from samples at all time points (especially those with low abundance at later stages), we appropriately increased the cycle number to 35.

      L429-431

      “The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing.”

      (34) L372: Were sequencing adapters ligated onto the pool?

      We thank the reviewer for this question. Yes, in this amplicon diversity sequencing workflow, sequencing adapters were ligated onto the pooled amplicon products. Briefly, individual samples were first amplified with sample-specific barcode sequences, allowing each sample to be distinguished after sequencing. The barcoded PCR products from multiple samples were then pooled for library construction. Universal Illumina-compatible adapters were ligated to this pooled amplicon library using the ALFA-SEQ DNA Library Prep Kit. After adapter ligation and purification, an indexing PCR was performed to introduce the complete P5/P7 sequences and a library-level Illumina index. We have clarified this pooled-library construction workflow in the revised Methods section.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (35) L378: "Amplicon sequence variants".

      We agree with the reviewer and have revised accordingly.

      (36) L380: Why were ASVs with fewer than 9 reads removed?

      We thank the reviewer for this question. The threshold of removing ASVs with fewer than 9 total reads across all samples was applied to reduce noise from sequencing errors and PCR artifacts. Our justification is supported by both the default parameters of the UNOISE3 algorithm and common practice in amplicon sequencing analysis.

      The USEARCH manual specifies that the -minsize parameter in the unoise3 command defaults to 8. This means that unique sequences occurring fewer than 8 times are discarded by the algorithm during ASV inference, as they are unlikely to represent true biological variants. Our threshold of 9 is slightly more conservative than the default (9 > 8), ensuring that only ASVs with a minimal level of abundance are retained. This choice is directly aligned with the algorithm’s intrinsic noise‑filtering logic.

      (37) L402: Please don't forget to discuss that PCR bias can contribute to uncertainty in the abundance of each taxon.

      Thank you for this important reminder. We agree that PCR bias (e.g., primer‑template mismatches, GC content differences, and variable amplification efficiency) can contribute to uncertainty in the abundance estimates of each taxon. Following your suggestion, we have now added a paragraph in the Discussion section to address this issue. We state that sequence‑specific degradation rates and PCR bias may jointly affect the accuracy of taxon abundance estimates, and future studies should incorporate internal standards or multiplex PCR strategies to correct for such biases. Thank you for your careful review.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      (38) L414: Suggest: "To inhibit amplification of extracellular DNA, soils were incubated with propidium monoazide (PMA), as described previously (REF). Briefly, soil (X grams) was mixed with PMA in a total volume of Y (ml).

      We thank the reviewer for this suggestion. We have revised the Methods section to provide a clearer description of PMA treatment, specifying the soil amount (0.50 g) and the total volume (0.5 mL).

      L496-497

      “To inhibit amplification of eDNA, soils were incubated with PMA, as described previously (Carini et al., 2016).”

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS), while the control soil samples were mixed with PBS without PMA.”

      (39) L416: In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.

      We agree and have revised.

      (40) L418-420: wording/sentence is strange and needs work.

      Thank you for pointing this out. We have reviewed the sentence at L418‑420 and agree that the wording is awkward. Moreover, the content only listed the advantages of the PMA method without acknowledging its limitations, making the statement less balanced. Therefore, in the revised manuscript, we have deleted this sentence. The limitations of the PMA method have been addressed in the Discussion section.

      L502-503

      “Currently, PMA treatment is a widely used to suppress PCR amplification of eDNA (Xue et al., 2023; Canini et al., 2024).”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (41) L421-422: PMA treatment is a widely used method for inhibiting the enzymatic processing of extracellular DNA (Xue, Canini).

      We agree and have revised.

      (42) L425: include volume of PBA.

      We thank the reviewer for this comment. We have revised the Methods section to include the volume of PMA used

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS).”

      (43) L429-430: Don't use the word precipitates- use "pellets".

      We agree and have revised.

      (44) L433: "The abundance of 16S rRNA genes was determined using quantitative PCR employing a LightCycler...".

      We agree and have revised.

      (45) L445-: Section 4.9 - needs citations for PERMANOVA, NMDS, SEM, etc.

      Thank you for your suggestion. We have added the necessary citations for PERMANOVA, NMDS, SEM, and other methods in Section 4.9.

      L531-539

      Prokaryotic community structure differences among the study sites and incubation time points were examined through non-metric multidimensional scaling analysis (NMDS), permutation multivariate analysis of variance (PERMANOVA), and Permutational Analysis of Multivariate Dispersion (PERMDISP) (Kruskal, 1964; Anderson, 2001). Random forest modeling was conducted to assess the importance of environmental and soil variables in predicting the overall degradation rates of extracellular 16S rRNA gene amplicon fragments. Structural equation modeling (SEM) was employed to further evaluate the direct and indirect effects of soil moisture, soil pH, MAP, and prokaryotic abundance on the overall degradation rates of extracellular 16S rRNA gene amplicon fragments (Grace, 2006).

      (46) L698: A few comments. It would be nice to know how many different 16S sequences were tracked for differential degradation and shown in the figure.

      We thank the reviewer for this helpful comment. We would like to clarify that Fig. 1A does not track the degradation of individual 16S rRNA gene amplicon sequences, but instead shows the overall degradation dynamics of the total added exogenous DNA pool. The data points are derived from total 16S gene copy numbers measured via qPCR at each incubation time point. Consequently, this quantification inherently includes all sequences present within the added pool. The multiple lines visualized in the figure represent the collective degradation trajectories of the entire DNA pool across different study sites

      To address sequence-level changes, we further analyzed the richness and composition of the GAPDH F-tagged 16S rRNA gene amplicon fragments, which are presented in Fig. 2A and related analyses.

      (47) L699: Better to use "16S rRNA gene amplicon fragment abundance" as the term.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have replaced the original wording with “16S rRNA gene amplicon fragment abundance” where appropriate.

      (48) Y-axis for Figures 1A and 2A should be GAPDH-labeled, not ACTB-labeled.

      We apologize for this mistake. We have corrected this error in the revised manuscript.

      (49) For Figure 1b: Why not use box plots and ANOVA for different soil types?

      Thank you for your valuable suggestion. In the original Figure 1b, we used a bar plot to display the degradation rate constants across the 30 study sites. This choice was intended to emphasize the continuous variation among sites and their gradient relationships with environmental factors (e.g., soil moisture, MAP), which were then used in random forest and structural equation modeling. The bar plot better illustrates the spatial continuum of degradation rates rather than treating ecosystem types as discrete categories.

      Nevertheless, we fully agree that a boxplot grouped by ecosystem type (grassland, forest, cropland, desert) would help readers quickly grasp the overall differences among land‑use types. In the revised manuscript, we have added a boxplot grouped by ecosystem type and performed one‑way ANOVA followed by Tukey HSD post‑hoc tests (Fig. 1.). The results show that degradation rate constants differ significantly among ecosystem types (P < 0.05).

      L124-128

      “The degradation rate constants of the spiked extracellular 16S rRNA gene amplicon fragments displayed considerable variability among the study sites, ranging from 0.05 to 0.16 day<sup>-1</sup> (Fig. 1b). Furthermore, we found that degradation rate constants differed significantly among ecosystem types (Fig. 1c, P < 0.05). Specifically, cropland and forest soils exhibited significantly higher degradation rates than grassland soils (P < 0.05).”

      (50) For Figure 2: Where are PERMANOVA and PERMDISP values for the figure?

      We thank the reviewer for this comment. In the revised manuscript, we have added the PERMANOVA and PERMDISP values corresponding to Figure 2 in the figure legend and Results section (Fig. 2).

      (51) I found Figure 2b to be hard to see. The 48-day circles are almost invisible. Difficult to know what the authors are trying to show here, since there is so much variability associated with soil type.

      We thank the reviewer for this comment. Figure 2b is intended to illustrate the temporal changes in microbial community structure during the incubation. The different colored circles represent samples at different time points (1, 3, 6, 12, 24, and 48 days), showing how communities shift over time. We apologize that in the original Figure 2b, the 48‑day samples were nearly invisible and that the high variability among soil types obscured the intended message. In the revised manuscript, we have added a black border around every data point, which greatly enhances the visibility of the 48‑day samples (and all time points). We now use distinct shapes to represent different ecosystem types (grassland, forest, cropland, desert) in the NMDS ordination, and added PERMANOVA results in both the Results section and the figure legend (Fig. R2b).

      (52) Figure 4A: Y-axis need a label like "16S rRNA gene abundance".

      We agree and have revised.

      (53) Figure 4B: I'd like to see a Shannon index too, not just richness.

      Thank you for your suggestion. We agree that the Shannon index, which integrates both richness and evenness, provides a valuable complement to richness alone. In the revised manuscript, we added an analysis of the Shannon index to compare α‑diversity between total DNA (PMA‑untreated) and intact cell DNA (PMA‑treated) samples (Fig.4).

      (54) Figure 4D: Would be good to have lines linking the intact cell vs total abundance. Also, what about a box plot of Bray-Curtis (or similar) dissimilarity between intact cell and total microbial analysis across the dataset?

      Thank you for your suggestions. Regarding the addition of connecting lines in Figure 4D, after careful consideration we decided not to add them for the following reason: the total and PMA-treated communities from the same site are already coded with the same color (different colors for different sites), which effectively indicates the pairing. Adding lines would greatly reduce readability due to dense overlapping lines, especially given the number of sites. Therefore, we kept the original color‑based pairing design.

      To address your second suggestion, we have added a bar plot showing the distribution of Bray‑Curtis dissimilarities between total (PMA‑untreated) and intact cell (PMA‑treated) communities across all study samples (Fig. R3d).

      (55) Figure 5B: What do correlations with p > 0.05 show? I would remove these from the image.

      We thank the reviewer for this suggestion. We agree that correlations with p > 0.05 do not represent statistically significant relationships and may cause confusion. In the revised manuscript, we have removed these non-significant correlations from Figure 5B.

      (56) Figure 6: "Incubations of 0, 3, 6, 12, 24, and 48 days".

      We agree with the reviewer and have revised as suggested.

      References

      Albertsen, M., Karst, S.M., Ziegler, A.S., Kirkegaard, R.H., Nielsen, P.H., 2015. Back to basics–the influence of DNA extraction and primer choice on phylogenetic analysis of activated sludge communities. PLoS One 10, e0132783.

      Anderson, M.J., 2001. A new method for non‐parametric multivariate analysis of variance. Austral Ecology 26, 32-46.

      Arvizu-Hernandez, E., Ocadiz-Delgado, R., Gariglio, P., 2025. E7HPV16 Oncogene and 17beta-Estradiol Stress Promote Oncogenic microRNA Expression Patterns, Cell Proliferation and Cervical Intraepithelial Neoplasia 1. Cell Biochemistry and Function 43, e70065.

      Buitrago, D., Labrador, M., Arcon, J.P., Lema, R., Flores, O., Esteve-Codina, A., Blanc, J., Villegas, N., Bellido, D., Gut, M., 2021. Impact of DNA methylation on 3D genome structure. Nature Communications 12, 3243.

      Cai, P., Huang, Q., Zhang, X., Chen, H., 2006a. Adsorption of DNA on clay minerals and various colloidal particles from an Alfisol. Soil Biology and Biochemistry 38, 471-476.

      Cai, P., Huang, Q.Y., Zhang, X.W., 2006b. Interactions of DNA with clay minerals and soil colloidal particles and protection against degradation by DNase. Environmental Science & Technology 40, 2971-2976.

      Carini, P., Marsden, P.J., Leff, J.W., Morgan, E.E., Strickland, M.S., Fierer, N., 2016. Relic DNA is abundant in soil and obscures estimates of soil microbial diversity. Nature Microbiology 2, 1-6.

      Deshpande, A.S., Fahrenfeld, N.L., 2023. Influence of DNA from non-viable sources on the riverine water and biofilm microbiome, resistome, mobilome, and resistance gene host assignments. Journal of Hazardous materials 446, 130743.

      Du, Y., Wang, Z., Liu, K., Chai, G., Chi, Y., Li, T., Duan, Y., Xia, T., Liu, D., Che, R., 2025. The performance of different methods in characterizing soil live prokaryotic diversity and abundance is highly variable. iMetaOmics, e70011.

      Emerson, J.B., Adams, R.I., Román, C.M.B., Brooks, B., Coil, D.A., Dahlhausen, K., Ganz, H.H., Hartmann, E.M., Hsu, T., Justice, N.B., 2017. Schrödinger’s microbes: tools for distinguishing the living from the dead in microbial ecosystems. Microbiome 5, 86.

      Finkel, S.E., Kolter, R., 2001. DNA as a nutrient: novel role for bacterial competence gene homologs. Journal of Bacteriology 183, 6288-6293.

      Frostegård, Å., Courtois, S., Ramisse, V., Clerc, S., Bernillon, D., Le Gall, F., Jeannin, P., Nesme, X., Simonet, P., 1999. Quantification of bias related to the extraction of DNA directly from soils. Applied and Environmental Microbiology 65, 5409-5420.

      Grace, J.B., 2006. Structural equation modeling and natural systems. Cambridge University Press.

      He, P., Li, L.-J., Dai, S.-S., Guo, X.-L., Nie, M., Yang, X., Kuzyakov, Y., 2024. Straw addition and low soil moisture decreased temperature sensitivity and activation energy of soil organic matter. Geoderma 442, 116802.

      Heise, J., Nega, M., Alawi, M., Wagner, D., 2016. Propidium monoazide treatment to distinguish between live and dead methanogens in pure cultures and environmental samples. Journal of Microbiological Methods 121, 11-23.

      Huang, C., Xie, D.C., Cui, J.J., Li, Q., Gao, Y., Xie, K.P., 2014. FOXM1c Promotes Pancreatic Cancer Epithelial-to-Mesenchymal Transition and Metastasis via Upregulation of Expression of the Urokinase Plasminogen Activator System. Clinical Cancer Research 20, 1477-1488.

      Knight, R., Vrbanac, A., Taylor, B.C., Aksenov, A., Callewaert, C., Debelius, J., Gonzalez, A., Kosciolek, T., McCall, L.-I., McDonald, D., 2018. Best practices for analysing microbiomes. Nature Reviews Microbiology 16, 410-422.

      Kruskal, J.B., 1964. Nonmetric multidimensional scaling: a numerical method. Psychometrika 29, 115-129.

      Lennon, J.T., Muscarella, M.E., Placella, S.A., Lehmkuhl, B.K., 2018. How, when, and where relic DNA affects microbial diversity. mbio 9, e00637-00618.

      Levy-Booth, D.J., Campbell, R.G., Gulden, R.H., Hart, M.M., Powell, J.R., Klironomos, J.N., Pauls, K.P., Swanton, C.J., Trevors, J.T., Dunfield, K.E., 2007. Cycling of extracellular DNA in the soil environment. Soil Biology and Biochemistry 39, 2977-2991.

      Liu, Q.H., Yuan, L., Li, Z.H., Leung, K.M.Y., Sheng, G.P., 2024. Natural organic matter enhances natural transformation of extracellular antibiotic resistance genes in sunlit water. Environmental Science & Technology 58, 17990-17998.

      Marrone, A., Ballantyne, J., 2008. Sequence Specificity of BAL 31 Nuclease for ssDNA Revealed by Synthetic Oligomer Substrates Containing Homopolymeric Guanine Tracts. PLoS One 3, e3595.

      McKinney, C.W., Dungan, R.S., 2020. Influence of environmental conditions on extracellular and intracellular antibiotic resistance genes in manure-amended soil: A microcosm study. Soil Science Society of America Journal 84, 747-759.

      Morrissey, E.M., McHugh, T.A., Preteska, L., Hayer, M., Dijkstra, P., Hungate, B.A., Schwartz, E., 2015. Dynamics of extracellular DNA decomposition and bacterial community composition in soil. Soil Biology and Biochemistry 86, 42-49.

      Nagler, M., Insam, H., Pietramellara, G., Ascher-Jenull, J., 2018. Extracellular DNA in natural environments: features, relevance and applications. Applied Microbiology and Biotechnology 102, 6343-6356.

      Nocker, A., Sossa-Fernandez, P., Burr, M.D., Camper, A.K., 2007. Use of propidium monoazide for live/dead distinction in microbial ecology. Applied and Environmental Microbiology 73, 5111-5117.

      Pietramellara, G., Ascher, J., Borgogni, F., Ceccherini, M., Guerri, G., Nannipieri, P., 2009. Extracellular DNA in soil and sediment: fate and ecological relevance. Biology and Fertility of Soils 45, 219-235.

      Shah, A., Huang, J., Han, T., Khan, M.N., Tadesse, K.A., Daba, N.A., Khan, S., Ullah, S., Sardar, M.F., Fahad, S., 2024. Impact of soil moisture regimes on greenhouse gas emissions, soil microbial biomass, and enzymatic activity in long-term fertilized paddy soil. Environmental Sciences Europe 36, 120.

      Sirois, S.H., Buckley, D.H., 2019. Factors governing extracellular DNA degradation dynamics in soil. Environmental Microbiology Reports 11, 173-184.

      Vuillemin, A., Horn, F., Alawi, M., Henny, C., Wagner, D., Crowe, S.A., Kallmeyer, J., 2017. Preservation and significance of extracellular DNA in ferruginous sediments from Lake Towuti, Indonesia. Frontiers in Microbiology 8, 1440.

      Wang, X., Ganzert, L., Bartholomaus, A., Amen, R., Yang, S., Guzman, C.M., Matus, F., Albornoz, M.F., Aburto, F., Oses-Pedraza, R., Friedl, T., Wagner, D., 2024. The effects of climate and soil depth on living and dead bacterial communities along a longitudinal gradient in Chile. The Science of the total environment 945, 173846.

      Wang, Y.-T., Yang, W.-J., Li, C.-L., Doudeva, L.G., Yuan, H.S., 2007. Structural basis for sequence-dependent DNA cleavage by nonspecific endonucleases. Nucleic Acids Research 35, 584-594.

      Wang, Y., Yan, Y., Thompson, K.N., Bae, S., Accorsi, E.K., Zhang, Y., Shen, J., Vlamakis, H., Hartmann, E.M., Huttenhower, C., 2021. Whole microbial community viability is not quantitatively reflected by propidium monoazide sequencing approach. Microbiome 9, 1-13.

      Wolpe, J., Guertin, M., 2022. Regional and Single Nucleotide Correction of Sequence Bias in Chromatin Accessibility Data. The FASEB Journal 36.

      Yang, A., Liu, X., Liu, P., Feng, Y.Z., Liu, H.B., Gao, S., Huo, L.M., Han, X.Y., Wang, J.R., Kong, W., 2021. LncRNA UCA1 promotes development of gastric cancer via the miR-145/MYO6 axis. Cellular & Molecular Biology Letters 26, 33.

      Ye, M., Zhang, Z., Sun, M., Shi, Y., 2022. Dynamics, gene transfer, and ecological function of intracellular and extracellular DNA in environmental microbiome. iMeta 1, e34.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback of our work, which led us to think more deeply about the conceptual and methodological aspects of this project and further strengthen the manuscript. Based on the comments, we included more control analyses and revised the manuscript accordingly. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. Two types of learned associations are characterized, one being associations between stimulus features and control states (SC), and the other being stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates, which are used to test properties of their dynamics.

      The conclusion is reached that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing SC and SR correlates are: (1) not entirely overlapping in cross-decoding; (2) simultaneously observed on average over trials in overlapping time bins; (3) independently correlate with RT; and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      Weaknesses:

      I still have my concern from the first round that the decoders are overfit to temporally structured noise. As I wrote before, the SC and SR classes are highly confounded with phase (chunk of session). I do not see how the control analyses conducted in the revision adequately deal with this issue.

      In the figures, there are several hints that these decoders are biased. Unfortunately, the figures are also constructed in such a way that hides or diminishes the salience of the clues of bias. This bias and lack of transparency discourage trust in the methods and results.

      I have two main suggestions:

      (1) Run a new experiment with a design that properly supports this question.

      I don't make this suggestion lightly, and I understand that it may not be feasible to implement given constraints; but I feel that this suggestion is warranted. The desired inferences rely on successful identification of SC and SR representations. Solidly identifying SC and SR representations necessitates an experimental design wherein these variables are sufficiently orthogonalized, within-subject, from temporally structured noise. The experimental design reported in this paper unfortunately does not meet this bar, in my opinion (and the opinion of a colleague I solicited).

      An adequate design would have enough phases to properly support "cross-phase" cross-validation. Deconfounding temporal noise is a basic requirement for decoding analyses of EEG and fMRI data (see e.g., leave-one-run-out CV that is effectively necessary in fMRI; in my experience, EEG is not much different, when the decoded classes are blocked in time, as here). In a journal with a typical acceptance-based review process, this would be grounds for rejection.

      Please note that this issue of decoder bias would seem to weaken the rest of the downstream analyses that are based on the decoded values. For instance, if the decoders are biased, in the within-trial correlation analysis, how can we be sure that co-fluctuations along certain dimensions within their projected values are driven by signal or noise? A similar issue clouds the LMM decoding-RT correlations.

      We appreciate the reviewer’s concern with the potential confound of temporally structured noise (TSN) in the EEG data. As we understand it, TSN refers to a process that the noise structure drifts over time. It follows that noise structure should be more similar for temporally closer trials and that the TSN’s bias on decoding accuracy is stronger for test trials that are closer to the training data. In the previous round of revision, we conducted a control analysis that reduced the influence of TSN by maximizing the temporal distance between training and test data (the distance between the centers of the training and test data of the same SC/SR manipulation is about 400 trials given the experimental design) and showed comparable decoding accuracy with the main results. As the reviewer finds this analysis unconvincing, we reason that the reviewer believes that the TSN has a long-term effect, such that it remains relatively stable over time and can be picked up by trials temporally distant from the training data. With this assumption and the assumption that this effect may not be linear, constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment. Instead, we focused on testing whether and how much TSN systematically biased the reported decoding accuracy.

      Please note that the existence of TSN in the EEG data is not sufficient to rule that the decoding results are biased. As TSN is stronger for trials closer to each other, the idea that auto-correlation biases decoding results would predict a distance effect, such that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. To test this predicted negative correlation, in each fold and each repetition of the cross-validation reported in the SC-SC and SR-SR decoders in Fig. 4, we calculated the distance (mean=5.84 trials, SD=2.05, 5th percentile =2.87, 95th percentile=9.45, one trial = 2.4-2.6s) between each test trial and its closest training trial of the same trial type. This distance was used as the predictor to predict decoding accuracy in a linear regression. Note that even if the relation between distance and decoding accuracy is non-linear, the linear relation will be negative because the relation is monotonic (similarity in noise structure decreases monotonically with temporal distance between trials). Similarly, because the effect is monotonic, if a long-range effect exists, it should also exist in short-range and be picked up by the distance range in this analysis. The regression coefficient is averaged across cross-validation folds and repetitions for each subject to match how the decoding accuracy was reported in the main text. Finally, the averaged regression coefficient was tested against 0 using a one-sample t-test. This analysis was conducted at each time point (from -250ms to 1500ms) separately. As shown in the figure below, no time point exhibited the negative correlation as predicted by the auto-correlation account. An alternative explanation is that this result indicates that TSN remains stable over time. If this is the case, TSN will be shared by all trials and will be unable to bias decoding results. Together with the control analysis introduced previously, this new control analysis supports the notion that the decoding results are not inflated by TSN in the EEG data. We included all the control analyses in the revised manuscript (page 13-14). Please note that this analysis is specific for the present dataset and we strongly agree with the reviewer that TSN is a key confounding factor in EEG analysis in general and should be carefully addressed.

      Lastly, we understand the concern with the early onset of above-chance decoding accuracy. Here, we provide an explanation: because of the blocked design (i.e., participants performed hundreds of trials with the same SC/SR associations), it is possible the participants learned the associations and used them to guide proactive cognitive control. As proactive cognitive control is anticipatory and sustained (Braver, 2012; Khan et al., 2025), it may be able to be decoded early on a trial, or even before trial onset. In the revised manuscript, we discussed this account along with the TSN issue as a limitation of the current project and directions for future research (page 24).

      (2) Increase transparency in the reporting of results throughout main text.

      Please do not truncate stimulus-aligned timecourses at time=0. Displaying the baseline period is very useful to identify bias, that is, to verify that stimulus-dependent conditions cannot be decoded pre-stimulus. Bias is most expected to be revealed in the baseline interval when the data are NOT baseline-corrected, which is why I previously asked to see the results omitting baseline correction. (But also note that if the decoders are biased, baseline-correcting would not remove this bias; instead, it would spread it across the rest of the epoch, while the baseline interval would, on average, be centered at zero.)

      Please use a more standard p-value correction threshold, rather than Bonferroni-corrected p<0.001. This threshold is unusually conservative for this type of study. And yet, despite this conservativeness, stimulus-evoked information can be decoded from nearly every time bin, including at t=0. This does not encourage trust in the accuracy of these p-values. Instead, I suggest using permutation-based cluster correction, with corrected p<0.05. This is much more standard and would therefore allow for better comparison to many other studies.

      I don't think these things should be done as control analyses, tucked away in the supplemental materials, but instead should be done as a part of the figures in the main text -- including decoding, RSA, cross-trial correlations, and RT correlations.

      Thank you for your suggestions. we have added the baseline period from 200 to 0 ms prior to the stimulus onset in all the stimulus-locked analyses and tested the significance with cluster-based permutation test (cluster-forming threshold p < 0.001, cluster-level p < 0.05, (Collins & Frank, 2018)) in all the analyses including decoding, RSA, cross-trial correlations and RT correlations. The results showed similar patterns, and they are all reported in the main text (please see all the figures and page 30-32 in the main text).

      Other issues:

      Regarding the analysis of the within-trial correlation of RSA betas, and "Cai 2019" bias:<br /> The correction that authors perform in the revision -- estimating the correlation within the baseline time interval and subtracting this estimate from subsequent timepoints -- assumes that the "Cai 2019" bias is stationary. This is a fairly strong assumption, however, as this bias depends not only on the design matrix, but also on the structure of the noise (see the Cai paper), which can be non-stationary. No data were provided in support of stationarity. It seems safer and potentially more realistic to assume non-stationarity.

      This analysis was included in the supplemental material. However, given that the correlation analysis presented in the Results is subject to the "Cai 2019" bias, it would seem to be more appropriate to replace that analysis, rather than supplement it.

      Regardless, this seems to be a moot issue, given that the underlying decoders seem to be overfit to temporally structured noise (see point above regarding weakening of downstream analyses based on decoder bias).

      Thank you for this important point. We now replaced the previous control analysis with a new one that does not assume stationary noise structure (page 19 in the revised manuscript). In Cai et al (2019), the source of confound is the covariance between observations. Specifically, as the observations in fMRI data are the BOLD signal at different time points, TSN can introduce covariance between nearby observations, which further biases the observed correlation between experimental conditions/trial types. In our case, the observations are decoding accuracy for different trial types. Thus, bias in the correlation may come from covariance between trial types. In this study, potential covariance between trial types includes the constrain that the decoding accuracy of all trial types adds up to 1 for a given trial (although we transformed the accuracy into logits prior to RSA, so the constrain may not hold), and the blocked design (as discussed above). Thus, to establish a baseline level of correlation between SC and SR representation strength, we took a similar shuffling approach as in Cai et al (2019) and randomly shuffled the trial types within each block. The reason to shuffle within each block is to preserve the covariance structure in the blocked design. We then repeated the same analysis using the shuffled data. The results of 10 shuffled analysis were averaged to form a baseline. Please note that (1) this control analysis was performed separately at each time point, hence removing the assumption of stationary noise structure, (2) this analysis also included as noise any covariance introduced by the proactive cognitive control guided by the learned SC and SR associations (see response to comment 1), thus it is more stringent than intended and (3) this control analysis started from decoding and was intended to provide a baseline for all downstream analysis. As shown in figures 2A, 3A, 7C and 8C, the reviewer was correct that the bias was not stationary, as the baseline of correlation coefficient varies over time. Additionally, the SC-SR representation strength correlation remained significantly above baseline between ~100 and ~ 450 ms following stimulus onset and between -180 and + 50 ms relative to response, suggesting that the noise structure (even when including potential proactive cognitive control) cannot fully explain the observed the SC-SR representation strength correlation. Considering the fact that this control analysis treated proactive control as a source of confound, this result does not necessarily contradict the absence of distance effect reported above.

      Outliers and t-values:

      More outliers with beta coefficients could be because the original SD estimates from the t-values are influenced more by extreme values. When you use a threshold on the median absolute deviation instead of mean +/-SD, do you still get more outliers with beta coefficients vs t-values?

      Thank you for your suggestion. We calculated the proportion of outliers with a threshold of median absolute deviation (defined as values beyond median ± 5 median absolute deviation) for each subject. The outliers remained less frequent for t-values than for beta coefficients (t-values: mean = 1.08%, SD = 0.12%; beta-values: mean = 4.45%, SD = 0.28%). Based on these results and to maintain consistent with previous studies employing the methods (Cellier et al., 2022; Kikumoto & Mayr, 2020; Kikumoto et al., 2022a; Kikumoto et al., 2022b; Rangel et al., 2023), we still decided to stay with t-values.

      Random slopes:

      Were random slopes (by subject) for all within-subject variables included in the LMMs? If not, please include them, and report this in the Methods.

      Thank you for your suggestion. The model failed to converge with random slopes of all variables. Thus, we chose not to add random slopes in the LMM. But we have added the random effects structure in the methods (see page 34).

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a sold design and the analyses are appropriate for the research questions.

      Strength:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed all your concerns raised below point by point.

      Weakness:

      (1a) The distinction between Phase 2 and Phase 1&3 behavioral results, specifically the opposite effect of MC/MI in congruent trials raises some concerns with regard to the effectiveness of the ISPC manipulation. Why do RTs and error rates under MC congruent condition in Phase 2 seem to be worse than MI congruent?

      Thank you for raising these issues. In Phase 1, one color set (red and blue) was assigned to the MC condition, whereas another color set (yellow and green) was assigned to the MI condition. In Phase 2, these assignments were flipped, and they were flipped back again in Phase 3. Thus, the MC condition consisted of red and blue in Phases 1 and 3 but yellow and green in Phase 2, whereas the MI condition consisted of yellow and green in Phases 1 and 3 but red and blue in Phase 2 (Fig. 1b in the manuscript). This manipulation leads to seemingly opposite patterns between Phases 1 & 3 and Phase 2.

      However, when considering specific colors, the pattern is consistent across phases. In Phase 2, RTs and error rates for yellow and green (MC congruent) were worse than those for red and blue (MI congruent), which mirrors the pattern observed in Phases 1 and 3, where RTs and error rates for yellow and green (MI congruent) were worse than those for red and blue (MC congruent)

      We interpreted the results in Phase 2 as reflecting a typical ISPC effect, which is defined as a smaller conflict effect in the MI condition (MI incongruent – MI congruent) compared with the MC condition (MC incongruent – MC congruent). To our knowledge, the ISPC paradigm does not impose a specific prediction regarding the relative difference between MC-congruent and MI-congruent conditions.

      (1b) Could there be other factors at play here, e.g. order effect?

      We agree that order effect could play a role, such that memory from Phase 1 may influence the pattern in phase 2. For example, in phase 1, yellow and green were assigned to the MI condition, and participants therefore have associated these colors with a high control state (SC) and incongruent responses (SR). These prior associations could interfere with the newly learned mappings in Phase 2, where yellow and green were reassigned to the MC congruent condition (i.e., low control state and congruent responses). As a result, memory from Phase 1 may have weakened the expected MC in phase 2. A similar effect could also apply to the MI condition. Consequently, the same condition does not show parallel performance between phase 1 and phase 2, which may lead to different patterns in the difference between MC congruent and MI congruent conditions in phase 2.

      (1c) How does this potentially affect the neural analyses where trials from different phases were combined?

      Thank you for the question. As we mentioned above, the order effect could slow down the newly learned associations. However, we still found the ISPC effect in each phase, suggesting that all kinds of both SC and SR associations were formed and could be applied to the decoding and the following analyses cross phases. Relatedly, there might be confounded with temporal structured noise (TSN) when the neural analyses on decoding were combined the trials from different phases. However, we have performed the control decoding analyses and distance effect tests and confirmed that our decoding results were not driven by TSN (Please see comment #1 of R1).

      (1d) the manuscript does not mention whether there is counterbalancing for the color groups across participants, so far as I can tell.

      Thank you for the reminder. We have balanced the color groups by randomly dividing the participants into two groups and assigning different color sets to each group. The related interpretations have been included in task overview of the revised manuscript (page 6), which reads:

      “The color groups were counterbalanced across participants by red and blue as the color set of MC in one group while as the color set of MI in another group in the phase 1.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I commend the authors for addressing and clarifying my previous questions. One new comment regarding the newly added Figure 9: the response-locked behavioral correlation is much weaker compared to the stimulus-locked one, even never reaching the significance level. I think this difference should be discussed instead of simply glossing over it.

      Thank you for your suggestion. We discussed the difference in the discussion of revised manuscript (page 25), which reads:

      “Note that we found the negative prediction of the strength of SC and SR to RTs did not reach statistical significance with response-locked analysis as stimulus-locked analysis. It is possible that SC and SR representations have occurred before the stage of response processing, which is usually aligned with stimulus onset (Jiang et al., 2020a; Kang & Yu-Chin, 2024; Khan et al., 2025)”

      References

      Braver, T. S. (2012). The variable nature of cognitive control: a dual mechanisms framework. Trends Cogn Sci, 16(2), 106-113. doi:10.1016/j.tics.2011.12.010

      Cellier, D., Petersen, I. T., & Hwang, K. (2022). Dynamics of Hierarchical Task Representations. J Neurosci, 42(38), 7276-7284. doi:10.1523/JNEUROSCI.0233-22.2022

      Collins, A. G., & Frank, M. J. (2018). Within- and across-trial dynamics of human EEG reveal cooperative interplay between reinforcement learning and working memory. Proceedings of the National Academy of Sciences, 115(10), 2502-2507. doi:10.1073/pnas.1720963115

      Khan, A. U., Hoy, C. W., Anderson, K. L., Piai, V., King-Stephens, D., Laxer, K. D., . . . Bentley, J. N. (2025). Neural dynamics of proactive and reactive cognitive control in medial and lateral prefrontal cortex. iScience, 28(9), 113375. doi:10.1016/j.isci.2025.113375

      Kikumoto, A., & Mayr, U. (2020). Conjunctive representations that integrate stimuli, responses, and rules are critical for action selection. Proc Natl Acad Sci 117(19), 10603-10608. doi:10.1073/pnas.1922166117

      Kikumoto, A., Mayr, U., & Badre, D. (2022a). The role of conjunctive representations in prioritizing and selecting planned actions. Elife, 11. doi:10.7554/eLife.80153

      Kikumoto, A., Sameshima, T., & Mayr, U. (2022b). The Role of Conjunctive Representations in Stopping Actions. Psychol Sci, 33(2), 325-338. doi:10.1177/09567976211034505

      Rangel, B. O., Hazeltine, E., & Wessel, J. R. (2023). Lingering Neural Representations of Past Task Features Adversely Affect Future Behavior. J Neurosci, 43(2), 282-292. doi:10.1523/JNEUROSCI.0464-22.2022

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study of Drosophila mating behaviors has offered a powerful entry point for understanding how complex innate behaviors are instantiated in the brain. The effectiveness of this behavioral model stems from how readily quantifiable many components of the courtship ritual are, facilitating the fine-scale correlations between the behaviors and the circuits that underpin their implementation. Detailed quantification, however, can be both time-consuming and error-prone, particularly when scored manually. Song et al. have sought to address this challenge by developing DrosoMating, software that facilitates the automated and high-throughput quantification of 6 common metrics of courtship and mating behaviors. Compared to a human observer, DrosoMating matches courtship scoring with high fidelity. Further, the authors demonstrate that the software effectively detects previously described variations in courtship resulting from genetic background or social conditioning. Finally, they validate its utility in assaying the consequences of neural manipulations by silencing Kenyon cells involved in memory formation in the context of courtship conditioning.

      Strengths:

      (1) The authors demonstrate that for three key courtship/mating metrics, DrosoMating performs virtually indistinguishably from a human observer, with differences consistently within 10 seconds and no statistically significant differences detected. This demonstrates the software's usefulness as a tool for reducing bias and scoring time for analyses involving these metrics.

      (2) The authors validate the tool across multiple genetic backgrounds and experimental manipulations to confirm its ability to detect known influences on male mating behavior.

      (3) The authors present a simple, modular chamber design that is integrated with DrosoMating and allows for high-throughput experimentation, capable of simultaneously analyzing up to 144 fly pairs across all chambers.

      Weaknesses:

      (1) DrosoMating appears to be an effective tool for the high-throughput quantification of key courtship and mating metrics, but a number of similar tools for automated analysis already exist. FlyTracker (CalTech), for instance, is a widely used software that offers a similar machine vision approach to quantifying a variety of courtship metrics. It would be valuable to understand how DrosoMating compares to such approaches and what specific advantages it might offer in terms of accuracy, ease of use, and sensitivity to experimental conditions.

      (2) The courtship behaviors of Drosophila males represent a series of complex behaviors that unfold dynamically in response to female signals (Coen et al., 2014; Ning et al., 2022; Roemschied et al., 2023). While metrics like courtship latency, courtship index, and copulation duration are useful summary statistics, they compress the complexity of actions that occur throughout the mating ritual. The manuscript would be strengthened by a discussion of the potential for DrosoMating to capture more of the moment-to-moment behaviors that constitute courtship. Even without modifying the software, it would be useful to see how the data can be used in combination with machine learning classifiers like JAABA to better segment the behavioral composition of courtship and mating across genotypes and experimental manipulations. Such integration could substantially expand the utility of this tool for the broader Drosophila neuroscience community.

      (3) While testing the software's capacity to function across strains is useful, it does not address the "universality" of this method. Cross-species studies of mating behavior diversity are becoming increasingly common, and it would be beneficial to know if this tool can maintain its accuracy in Drosophila species with a greater range of morphological and behavioral variation. Demonstrating the software's performance across species would strengthen claims about its broader applicability.

      Reviewer #2 (Public review):

      This paper introduces "DrosoMating," an integrated hardware and software solution for automating the analysis of male Drosophila courtship. The authors aim to provide a low-cost, accessible alternative to expensive ethological rigs by utilizing a custom acrylic chamber and smartphone-based recording. The system focuses on quantifying key temporal metrics-Courtship Index (CI), Copulation Latency (CL), and Mating Duration (MD)-and is applied to behavioral paradigms involving memory mutants (orb2, rut).

      The development of open-source behavioral tools is a significant contribution to neuroethology, and the authors successfully demonstrate a system that simplifies the setup for large-scale screens. A major strength of the work is the specific focus on automating Copulation Latency and Mating Duration, metrics that are often labor-intensive to score manually.

      However, there are several limitations in the current analysis and validation that affect the strength of the conclusions:

      First, the statistical rigor requires substantial improvement. The analysis of multi-group experiments (e.g., comparing four distinct strains or factorial designs with genotype and training) currently relies on multiple independent Student's t-tests. This approach is statistically invalid for these experimental designs as it inflates the family-wise Type I error rate. To support the claims of strain-specific differences or learning deficits, the data must be analyzed using Analysis of Variance (ANOVA) to properly account for multiple comparisons and to explicitly test for interaction effects between genotype and training conditions.

      Second, the biological validation using $w^{1118}$ and $y^1$ mutants entails a potential confound. The authors attribute the low Courtship Index in these strains to courtship-specific deficits. However, both strains are known to exhibit general locomotor sluggishness (due to visual or pigmentation/behavioral defects). Since "following" behavior is likely a component of the Courtship Index, a reduction in this metric could reflect a general motor deficit rather than a specific lack of reproductive motivation. Without controlling for general locomotion, the interpretation of these behavioral phenotypes remains ambiguous.

      Third, the benchmarking of the system is currently limited to comparisons against manual scoring. Given that the field has largely adopted sophisticated open-source tracking tools (e.g., Ctrax, FlyTracker, JAABA), the utility of DrosoMating would be better contextualized by comparing its performance - in terms of accuracy, speed, or identity maintenance - against these existing automated standards, rather than solely against human observation.

      Finally, the visual presentation of the data hinders the assessment of the system's temporal precision. While the system is designed to capture time-resolved metrics, the results are presented primarily as aggregate bar plots. The absence of behavioral ethograms or raster plots makes it difficult to verify the software's ability to accurately detect specific transitions, such as the exact onset of copulation.

      We sincerely thank the reviewers for their constructive and detailed feedback, which has substantially improved the clarity and rigor of our work. Below is a summary of the major revisions.

      (1) Comparison with existing tools. We added a new main figure (Figure 5) and Table 1 systematically benchmarking DrosoMating against Ctrax and FlyTracker on identical low-quality, high-throughput mating videos. Both conventional tools failed under our recording conditions: Ctrax showed severe segmentation instability and fragmented trajectories, while FlyTracker frequently crashed during feature computation. These results demonstrate that DrosoMating's state-detection approach bypasses the pose-tracking limitations that impair established pipelines. We also revised the Discussion to clarify that DrosoMating is specialized for robust extraction of mating timing metrics from low-quality videos, not a replacement for general-purpose tracking tools.

      (2) Statistical analysis. Following Reviewer #2's recommendation, we re-analyzed Figure 3 using one-way ANOVA with Tukey's HSD for strain comparisons, and Figure 4 using two-way ANOVA with Sidak's post-hoc test for learning assays, including the critical Genotype by Training interaction term.

      (3) New supplementary data. We added Figure S4 showing basal locomotor velocity of single-housed males across all four strains to decouple motor defects from courtship deficits in w1118 and y1 mutants. We also added Figure S3 with behavioral ethograms for individual flies to visually demonstrate the system's temporal resolution and detection accuracy.

      (4) Additional improvements. We standardized statistical reporting across all figure legends, explicitly defined the segmentation threshold parameter s in the Methods, consistently used "Mating Duration (MD)" throughout, standardized MB247-GAL4 labeling, clarified that occluded frames are retained for mating-state detection via merged-contour analysis, and corrected minor errors including duplicate references and unnecessary quotation marks.

      We believe these revisions substantially strengthen the manuscript and clearly position DrosoMating within the existing ecosystem of behavioral analysis tools.

      We remain grateful for the valuable feedback from both reviewers and the editorial team, and we hope the revised version meets the standards for publication in eLife.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It's difficult to assess the utility of this tool in relation to the variety of alternative methods available for the automated scoring of courtship behavior, some of which offer more granular behavioral data than what DrosoMating has been presented to produce. A direct comparison to other approaches (e.g., FlyTracker, DeepLabCut, SLEAP) would be highly valuable for understanding what use cases DrosoMating is best suited for. Ideally, this analysis would include (i) a comparison of the accuracy of scoring courtship metrics and constituent behaviors, (ii) an assessment of the sensitivity of the various software to different experimental conditions, such as lighting and alternative chamber designs, and (iii) testing across Drosophila species, which could help highlight the particular strengths of DrosoMating over more established methods. At a minimum, a table comparing key features, requirements, and capabilities would help position DrosoMating within the existing ecosystem of tools.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure(Figure 5) and comparative table (Table1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      "DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines.

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Evaluation of accuracy, robustness, and cross-species potential

      We addressed the three key points requested by the reviewer:

      (i) Accuracy: DrosoMating achieved 98–99% agreement with manual scoring. Under our experimental conditions, neither Ctrax nor FlyTracker completed an end-to-end workflow capable of reliably extracting copulation latency and mating duration. Ctrax produced fragmented trajectories with unstable target counts, and FlyTracker terminated with runtime errors during feature computation.

      (ii) Robustness: DrosoMating is highly robust to low-contrast lighting and common behavioral chamber setups, whereas conventional tools require high-quality videos and fail under fly occlusion.

      (iii) Cross-species testing: We did not perform cross-species validation in this revision. DrosoMating relies on mating state detection rather than species-specific morphology, suggesting potential transferability, although this remains to be experimentally validated. We have therefore revised the text to avoid claiming demonstrated cross-species performance and now describe this as a potential future application.

      The revised manuscript text is as follows (line 572-576):

      "Cross-species testing was not performed in this study. Although DrosoMating's state-detection approach is morphology-agnostic and therefore theoretically applicable across Drosophila species, we have revised the text to avoid claiming demonstrated cross-species performance. Formal validation across diverse species remains a promising future direction."

      We added Table1 to compare key features across DrosoMating, Ctrax, and FlyTracker, including low-quality video compatibility, multi-chamber support, robustness to fly overlap, direct output of mating timing metrics, accuracy.

      (3) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioral features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have also added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (2) The coarse nature of the metrics captured by DrosoMating may limit the usefulness of this tool for many researchers. Consider integrating DrosoMating with one or more behavioral classifiers (e.g., JAABA) and validating its performance to increase its utility across a wider range of uses. Demonstrating the feasibility of this would substantially increase the tool's appeal to those who need access to more discrete behavioral elements.

      We thank the reviewer for this constructive suggestion. We agree that DrosoMating is currently optimized for rapid, high-throughput quantification of core mating timing metrics (CL, CI, and MD) rather than discrete behavioral classification (e.g., wing extension, circling, or aggressive postures). This reflects a deliberate design trade-off: by prioritizing robust state detection over detailed pose tracking, DrosoMating achieves reliable performance on low-quality videos where conventional identity-based pipelines fail.

      We appreciate the reviewer’s vision for expanding the tool’s utility. While DrosoMating does not presently generate the per-frame kinematic features (position, orientation, wing angles, etc.) required as input for JAABA classifiers, its modular architecture and underlying video-processing framework provide a foundation for future integration. In the revised Discussion, we have clarified that extending DrosoMating to export trajectory-derived features compatible with downstream classifiers such as JAABA represents a promising future direction—one that would broaden its applicability to laboratories requiring granular behavioral elements, without necessitating a complete overhaul of the existing pipeline.

      We hope this clarification addresses the reviewer’s concern and accurately reflects both the current capabilities and future potential of the tool.

      The revised manuscript text is as follows (line 614-620):

      "While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays."

      (3) Please clarify how DrosoMating handles fly identification during mating? I would think mounted flies would be occluded, and if these frames are excluded from analysis, it would be expected to skew mating duration scores. It would be helpful if the authors could discuss these details.

      Occluded frames are excluded from CI calculation because individual courtship actions cannot be reliably assigned during prolonged overlap. However, these frames are not discarded from mating-duration analysis. Instead, prolonged merged contours are used as evidence for copulation-state detection, and the MD timer continues until physical separation:

      (1) When two flies are separate, the system detects two contours; when they overlap during mounting, it detects one merged contour. Our code processes both cases, so occluded frames are retained, not skipped.

      (2) To distinguish a single fly from two overlapping flies, we analyze the shape of the merged contour. Two overlapping flies produce a characteristically different aspect ratio (elongated shape) compared to one fly. This geometric cue is fed into the state classifier to label the frame as "mating."

      (3) Because these frames are classified as mating rather than excluded, the mating duration timer runs continuously through the occlusion period. Thus, mating duration is not artificially shortened.

      The high agreement between DrosoMating and manual scoring for mating duration (within 10 s, no significant difference) confirms that this approach does not introduce bias.

      We revised the manuscript as follow (line 191-193):

      "Occluded frames are excluded from CI calculation but retained for mating-state detection via merged-contour analysis, ensuring continuous MD measurement through copulation."

      (4) Please indicate the statistical tests used in Figures 3, 4, and S1. What methods were used to address multiple hypothesis testing?

      We thank the reviewer for this important comment. For comparisons between two groups, two-sided Student’s t-tests were used. For comparisons among three or more groups, one-way ANOVA followed by post-hoc tests were applied. For two-factor experimental designs, two-way ANOVA was used. We have clearly stated these statistical tests in the Statistical Analysis section and have added this information to the figure legends for Figures 3, 4, and S1 in the revised manuscript.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (5) Line 43: "High resolution video tracking enables...".

      We thank the reviewer for noting this imprecise expression. We have revised the statement to accurately describe that our method uses image-based video analysis to identify courtship and copulation events and quantify their temporal parameters. The description has been corrected in the revised manuscript line 44-46:

      "Our image-based video analysis enables precise identification of courtship and copulation events, as well as quantification of their timing and duration under controlled conditions. "

      (6) Line 51: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 51 has been removed in the revised manuscript.

      (7) Line 107: Chen et al, 2024 is listed twice in the references.

      We thank the reviewer for the careful correction. The duplicate reference of Chen et al., 2024 has been removed from the reference list in the revised manuscript.

      (8) Line 242: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 242 has been removed in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Please re-analyze the data in Figures 3 and 4 using ANOVA followed by appropriate post-hoc tests (e.g., Tukey's HSD). Specifically, use a One-way ANOVA for strain comparisons in Figure 3 and a Two-way ANOVA for the learning assays in Figure 4. The interaction term (Genotype $\times$ Training) is critical for demonstrating specific learning deficits. Update the "Statistical Analysis" section and figure legends accordingly.

      We appreciate the reviewer’s recommendation. We have re-analyzed the data in Figures 3 and 4 using the suggested ANOVA approaches: one-way ANOVA with Tukey’s HSD for strain comparisons (Figure 3), and two-way ANOVA (including the Genotype × Training interaction term) with Sidak's post-hoc test for learning assays (Figure 4). The updated statistical methods are now described in the Statistical Analysis section and figure legends.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (2) To decouple motor defects from courtship deficits in $w^{1118}$ and $y^1$ mutants, please use your tracking data to calculate and present a "General Locomotion" metric (e.g., average velocity or total distance traveled in the absence of a female).

      We thank the reviewer for this suggestion. We have now added Figure S4 showing basal locomotor velocity of single-housed males across all four strains. As expected, w^1118 and y^1 mutants move more slowly than wild-type controls.

      These data reveal that both motor and sensory factors contribute to the observed courtship phenotypes. Reduced basal locomotion likely limits the males' ability to approach and follow females. However, this generalized hypoactivity is compounded by strain-specific sensory deficits: w<sup>1118</sup> males suffer visual impairment that compromises female detection (Krstic et al., 2013), while y<sup>1</sup> males display altered cuticular hydrocarbons that disrupt pheromonal communication (Drapeau et al., 2006). These sensory defects impair courtship initiation and female recognition independent of locomotor capacity. Thus, the reduced CI, CL, and MD in these mutants reflect the combined effects of slower movement and courtship-specific sensory impairments, rather than motor defects alone.

      We have revised the manuscript to incorporate the velocity data and clarify this interpretation (line 219-226):

      "Notably, reduced basal locomotor activity in w<sup>1118</sup> and y<sup>1</sup> mutants has been well documented in previous studies, independent of courtship behavior (Drapeau et al., 2006; Krstic et al., 2013). Consistent with these reports, our tracking data show that single-housed w<sup>1118</sup> and y<sup>1</sup> males exhibit lower average velocity than Canton-S and Oregon-R controls (Fig. S4A). These general locomotor differences are insufficient to fully explain the observed courtship and mating timing phenotypes, indicating that additional courtship‑related processes contribute to the observed behavioral differences."

      (3) Please expand the discussion or provide a small comparative dataset contrasting DrosoMating with established tools like JAABA. Explain the specific advantages of your pipeline (e.g., cost, simplicity, focus on CL/MD) to justify its adoption over these alternatives.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure (Figure 5) and comparative table (Table 1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      " DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioural features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (4) Complement the aggregate bar plots with behavioral ethograms or raster plots for representative individual flies. Color-code these plots for specific states (resting, following, courting, copulating) to visually demonstrate the system's temporal resolution and detection accuracy.

      We thank the reviewer for this valuable suggestion. To address this comment, we have added a new supplementary figure (Figure S3) that presents behavioral ethograms for individual flies, directly complementing the aggregate bar plots in the main text. The figure displays the full temporal progression of courtship and copulation behaviors for all wells that exhibited successful mating. As requested, behaviors are color-coded (orange: courting, red: copulating) to clearly delineate different states. These ethograms visually demonstrate the system’s ability to resolve behavioral transitions with high temporal precision, confirming the accuracy of our automated detection of courtship initiation, duration, and copulation events. This addition provides critical individual-level validation that supports the aggregate statistical results presented in the main text.

      (5) Standardize the use of estimation statistics (DBMs). If used in Figure 3, they should also be applied to Figure 4, with appropriate statistical comparisons between groups.

      We thank the reviewer for this suggestion. We have standardized the statistical reporting format across all figure legends, with each legend explicitly stating the test used, the post hoc method (where applicable), and significance thresholds.

      Our approach is as follows:

      Fig. 2 and Fig.3D-I (two-group comparison, manual vs. automated scoring): DBM + Student's t-test.

      Fig. 3A-C and Fig. S1B-C, E-F (multi-group comparison, 4 strains or 2 rearing conditions): One-way ANOVA + Tukey's HSD.

      Fig. 4B-D and F (two-factor design, genotype × training): Two-way ANOVA + Sidak's post hoc.

      We have also added a sentence to the Methods clarifying that estimation statistics (DBM) are used for single two-group contrasts, while ANOVA-based approaches are used for multi-group or multi-factor designs. The statistical methods are now consistently documented in both the Methods section and the corresponding figure legends. We hope this clarification addresses the reviewer's concern.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (6) Fix inconsistent labeling (e.g., "MB-247" vs. "MB247") and redundant axis labels.

      Thank you for pointing this out. We have now standardized all labels to "MB247-GAL4" throughout the text and figures.

      We have also removed redundant axis labels from multi-panel figures. And redundant DBM and metric labels have been streamlined in Figures 2 and 4. All statistical tests and metric definitions are now fully described in the corresponding figure legends rather than being repeated on each sub-panel.

      (7) Significantly increase the size of data panels to make individual data points and error bars legible.

      Thank you for this suggestion. We have standardized the figure formatting and simplified the panels by removing redundant axis labels and consolidating descriptive details into the figure legends. We believe the current panel sizes, combined with these clarifications, provide sufficient legibility for both data points and error bars in the final high-resolution PDF.

      Minor Corrections:

      (1) Abstract: Rephrase "lack of certain timing-related behavioral repertoires" (Line 108) to "lack of precise quantification for temporal parameters of post-copulatory behavior."

      Thank you for the suggested rephrasing. We have updated the sentence accordingly.

      (2) Define the physical parameter "s" (e.g., is it a pixel threshold?) to ensure reproducibility.

      Thank you for this important suggestion. We have now explicitly defined the parameter s in the Methods section. Briefly, s is the grayscale intensity threshold (0–255, 8-bit) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value of three user-selected reference flies plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.

      The revised manuscript text is as follows (line 499-503):

      “Note on the s value: The parameter s represents the grayscale intensity threshold (range: 0–255 for 8-bit images) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value at the three reference fly positions plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.”

      (3) Figure 1 Legend: Clarify the "clockwise selection" of pillars and their relation to well numbering.

      Thank you for this suggestion. We have revised the Figure 1 legend to clarify that:

      (1) The four pillars are selected in clockwise order starting from the top-left corner to define the chamber corners for perspective transformation (homography), which corrects for camera angle and standardizes the field of view.

      (2) Well numbering is independent of pillar selection order. After automated perspective correction, wells are numbered sequentially in a left-to-right, top-to-bottom order (1–36) based on the standardized chamber layout.

      The updated Figure 1 legend now reads:

      "Columns (pillars) should be selected in clockwise order starting from the top-left corner to define the chamber boundaries for perspective transformation (Fig. 1A, lower). Well numbering (1–36) follows a left-to-right, top-to-bottom sequence after perspective correction and is independent of pillar selection order."

      (4) Add the citation for Eastwood and Burnet (1977) regarding Courtship Latency.

      Thank you for pointing this out. We have added the citation Eastwood and Burnet (1977) to the Introduction where Courtship Latency is first defined

      (5) Select one term ("Mating Duration" or "Copulation Duration") and use it consistently throughout the text and figures.

      Thank you for this suggestion. We have now standardized the terminology throughout the manuscript and figures. "Mating Duration" (MD) is used consistently in all instances where "Copulation Duration" previously appeared. The abbreviation MD has been retained for consistency with existing figure labels

      References

      Branson, K., Robie, A.A., Bender, J., Perona, P., & Dickinson, M.H. (2009). High-throughput ethomics in large groups of Drosophila. Nature Methods, 6(6), 451–458.

      Claridge-Chang, A., & Assam, P.N. (2016). Estimation statistics should replace significance testing. Nature Methods, 13(2), 108–109.

      Drapeau, M.D., Cyran, S.A., Viering, M.M., Geyer, P.K., & Long, A.D. (2006). A cis-regulatory sequence within the yellow locus of Drosophila melanogaster required for normal male mating success. Genetics, 172(2), 1009–1030.

      Eastwood, L., & Burnet, B. (1977). Courtship latency in male Drosophila melanogaster. Behavior Genetics, 7(3), 359–372.

      Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X.P., Hoopfer, E.D., Schor, J., Anderson, D.J., & Perona, P. (2014). Detecting social actions of fruit flies. Lecture Notes in Computer Science, 8692, 772–787.

      Gil-Martí, B., Barredo, C.G., Pina-Flores, S., Poza-Rodriguez, A., Treves, G., Rodriguez-Navas, C., Camacho, L., Pérez-Serna, A., Jimenez, I., Brazales, L., Fernandez, J., & Martin, F.A. (2023). A simplified courtship conditioning protocol to test learning and memory in Drosophila. STAR Protocols, 4(1), 101572.

      Greenspan, R.J., & Ferveur, J.F. (2000). Courtship in Drosophila. Annual Review of Genetics, 34, 205–232.

      Hall, J.C. (1994). The mating of a fly. Science, 264(5163), 1702–1714.

      Kabra, M., Robie, A.A., Rivera-Alba, M., Branson, S., & Branson, K. (2013). JAABA: interactive machine learning for automatic annotation of animal behavior. Nature Methods, 10(1), 64–67.

      Kitamoto, T. (2001). Conditional modification of behavior in Drosophila by targeted expression of a temperature-sensitive shibire allele in defined neurons. Journal of Neurobiology, 47(2), 81–92.

      Krstic, D., Boll, W., & Noll, M. (2013). Influence of the White locus on the courtship behavior of Drosophila males. PLoS ONE, 8(9), e77904.

      Levin, L.R., Han, P.L., Hwang, P.M., Feinstein, P.G., Davis, R.L., & Reed, R.R. (1992). The Drosophila learning and memory gene rutabaga encodes a Ca2+/calmodulin-responsive adenylyl cyclase. Cell, 68(3), 479–489.

      Pavlou, H.J., & Goodwin, S.F. (2013). Courtship behavior in Drosophila melanogaster: towards a ‘courtship connectome’. Current Opinion in Neurobiology, 23(1), 76–83.

      Yapici, N., Kim, Y.J., Ribeiro, C., & Dickson, B.J. (2008). A receptor that mediates the post-mating switch in Drosophila reproductive behaviour. Nature, 451(7176), 33–37.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1 (Public review):

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% <sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph).

      On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2 (Public review):

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev.

      Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews [refs 25-27]. The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"."

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer #3 (Public review):

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (2) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A crossmutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted. 

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

      Recommendations for the authors:

      Reviewing Editor Comments:

      (1) The membrane composition used in this study is highly specific. The manuscript would benefit from a clearer justification of the lipid composition and discussing the transferability of the approaches to other relevant membrane systems.

      We now acknowledge the limitation of our work to a single composition (p. 19, 3rd paragraph). This composition was chosen because it is widely used in both computational and experimental studies (e.g., PMID: 21144818; 21344950; 29845130; 29995324; 34813727; 37406927). However, as we now point out, future developments should account for the effects of lipid composition.

      (2) This work heavily relies on computational modeling. Thus, it remains unclear to what extent the model captures the thermodynamics of membrane insertion, rather than reproducing the behavior of the PPM framework. Further comparisons with available experimental results will make this work more impactful.

      We note that the manuscript already has substantial experimental support. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) Overall, the manuscript is lengthy. Shortening will improve the readability, clarity, and accessibility to a broad audience.

      We have shortened some text, as explained below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript uses a single membrane composition… The high PIP<sub>2</sub> content (5%) in the simulations may overemphasize electrostatic contributions from basic residues. Please discuss how different membrane compositions (e.g., lower <sub>2</sub>, presence of cholesterol, cardiolipin in mitochondrial membranes) might alter the q parameters and whether AroMIP predictions would change qualitatively or quantitatively.

      We now discuss how different membrane compositions may alter q parameters (p. 19, 3rd paragraph). We believe that these alterations will change our predictions quantitatively but not qualitatively, given that our validation is against experimental results acquired on a variety of membrane compositions.

      (2) The limitations of the current framework should be discussed more explicitly. For example, the applicability of AroMIP beyond isolated 9-residue motifs remains unclear. In full-length IDPs, membrane insertion may be modulated by longer-range sequence interactions, transient secondary structure formation, multivalent interactions, or post-translational modifications.

      We now discuss the limitation of the 9-residue motif, and note these neglected factors for future developments (paragraph running from p. 19-20).

      (3) The AroMIP web server is user-friendly… The utility of the server could be enhanced by enabling visualization of full-length disordered proteins. For example, if users input a >300 residue IDP, the server could output residue-wise or sliding-window membrane insertion propensity profiles.

      We have revised the web server. We now display insertion propensity profiles as a plot and have added a link for users to download.

      (4) The manuscript is quite long and in several sections overly descriptive, e.g., in the first MD Results section and portions of the "Additional test cases" section.

      We have shortened the first subsection in Results and placed the expanded presentation in Supporting Information. However, we have kept the “Additional test cases” subsection, because Reviewer 2 appears to have overlooked the experimental validation of our method in these additional test cases.

      (5) The authors used the Berendsen barostat… known not to reproduce correct volume fluctuations and is generally considered less rigorous for equilibrium simulations, although this does not affect the main conclusions of the work. For future studies, the authors may consider using more modern barostats.

      Thank you for the suggestion! We will definitely be using the more modern barostats in future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) The idea of the article seems very interesting. The problem of membrane association mediated by aromatic residues is definitely worth studying. Aromatic residues, especially Tryptophan (W), but also, albeit to a lesser extent, Phenylalanine (F), and Tyrosine (Y) are well known to partition preferentially to the headgroup region of the lipid bilayer. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, however, none of these articles is cited. Have the authors read them?

      We now cite the relevant references as explained above.

      (2) The authors propose to decipher the sequence code for insertion of sequences containing aromatic residues in the membrane employing three types of calculation methods with decreasing order of detail and complexity, but increasing order of efficiency. First, all-atom MD simulations; second, the PPM method (protein positioning in membranes) from Lomize et al (2006), Protein Sci 15, 1318; and third, AroMIP, a mathematical model developed by the authors. Incidentally, I don't see anywhere in the text what AroMIP stands for. Is it Aromatic Membrane Insertion Prediction, or something like that? Please define. In any case, the proposed endeavor is commendable.

      We now spell out the acronym at its first occurrence in the Abstract and Introduction (Aromatic Membrane Insertion Predictor).

      (3) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"." Thus, there must be a comparison with experiment, as elaborated in the next point.

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM.

      (4) I understand that the authors are computational or theoretical physical chemists and would not expect them to perform the experiments. However, there are plenty of data in the literature that can be used to corroborate the calculations. First and foremost, though, the authors must calculate the binding constants for their set of peptides and then compare them with experiment, or use a different set of peptides for which the experimental results are available, and test how PPM and AroMIP perform on those peptides. There are two possible approaches to calculate the binding constants. First, calculate the Gibbs energy of binding from simulation or calculation, and, from that, calculate the binding constant via the Boltzmann factor. Second, and better, is to calculate the probability of binding in the simulations or calculations from the fraction of time that the peptide spends bound to the membrane or in water. According to the ergodic principle, the ratio of the two times is the binding constant. I understand that most amphipathic peptide sequences whose binding constants to membranes have been determined by experiment are long. But some are not. For example, Mastoparan X is a 14-residue antimicrobial peptide, and its dissociation constant from POPC vesicles is known to be about 300 μM.

      We would like to make it clear that the aim of our study is to predict membrane insertion propensities of aromatic-centred motifs, not the membrane binding affinity of peptides. These two properties are related but require different ways of validating predictions. For membrane insertion, validation requires that not only the motifs are bound to membranes but also the aromatic side chains are placed in the acyl chain region. Our validation of the 12 IDRs listed in Table S2 targeted these requirements.

      That said, we note that free energy of binding and probability of binding calculations, mentioned by the reviewer, have been reported previously, including the free-energy cost of Ala substitutions of aromatic residues located at various depths reported by Waheed et al. (ref 35) and residue-specific insertion depths reported by Wang et al. (ref 13). Both of these studies highlighted the propensities of aromatic side chains in inserting into the acyl chain region.

      (5) Furthermore, the entire Wimley-White interfacial hydrophobicity scale was determined using pentapeptides, measuring the equilibrium binding constant to POPC membranes experimentally and then using those data to build a residue-based Gibbs energy of binding (Wimley & White, Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nature Struct. Biol. 1996, 3, 842-848; White & Wimley Membrane protein folding and stability: Physical principles. Annu. Rev. Biophys. Biomol. Struct. 1999, 28, 319-365.) This allows for the calculation of the binding affinity for any sequence. The calculations have been compared to experiment and shown to have a high accuracy. A note of caution: Be careful, though, because Wimley and White used a mole fraction concentration scale, which makes the Gibbs energy of binding more favorable than the value calculated using the more common molar concentration scale by -2.4 kcal/mol. Calculating the binding constant for the peptides used by the authors from the WW interfacial scale and comparing the results with the authors' results would be a good place to start.

      This is an excellent suggestion! We now compare our insertion scores with the binding free energies calculated from the WW interfacial scale (new Figure S10; p. 15, 2nd paragraph). Interestingly, we found moderately higher correlations with binding free energies calculated from the octanol scale, which we suggest is more in line with our insertion scores since octanol mimics the hydrophobic region as suggested by White and Wimley in their 1999 Annu Rev paper.

      (6) When we speak of insertion in a bilayer, we normally mean insertion in the nonpolar core, not in the interfacial region, which is what the authors mean in this paper. This is extremely misleading and should be changed, namely in the title, but also throughout the paper.

      Actually, by insertion we precisely mean into the nonpolar core, NOT the interfacial region. Throughout the Introduction, when we used the word “insert”, we added “into the acyl chain region” (p. 3, line 6 from bottom; p. 5, lines 6-7 from top and line 7 from bottom; p. 6, line 4 from top). To avoid any confusion, we now also explicitly add “into the membrane hydrophobic core” when the word “insertion” first occurs in the Abstract and in the opening paragraph of Introduction (replacing the previous “deep insertion”).

      (7) What is q? It appears in equation (1) on page 14, and is referred to several times afterwards, but it is never defined.

      We now elaborate, in the text above and below equation (1), on the meaning of the q parameters: they represent the contributions of flanking residues to the insertion score of a central aromatic residue.

      (8) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is wrong even in a qualitative sense.

      We have revised these figures.

      (9) In Figure 1, the location of the bilayer midplane should be indicated, for example, with a line. Currently, there is a red line on the figures, but that is not the bilayer midplane. A reader may easily - and is likely to - misinterpret the figure.

      As explained in the Figure 1B, C caption, the red line is drawn at Z = -3.1 Å, which is the mean position of glycerol C2 carbon atoms (setting Z = 0 for the phosphate plane) and thus the start of the acyl chain region.

      Reviewer #3 (Recommendations for the authors):

      (1) Please refrain from using the term "transmembrane" when describing Drp1-membrane insertion, as Drp1 is a soluble, peripheral membrane-binding protein that reversibly associates with membranes.

      We have removed the sentence mentioning “transmembrane orientation”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Naim et al. use genetically engineered mouse models and tissue culture cell lines to investigate the role of the SLAP adaptor protein in colonic epithelium and colon tumour formation. The SLAP adaptor protein is known to be a negative regulator of tyrosine kinase signaling in hematopoietic cells, but its role outside the immune system is less well defined. Here, the authors use genetically engineered SLAP-deficient mice, tissue-specific SLAP KO, and colonic organoids to demonstrate that SLAP is expressed in cells of the colonic epithelium, where it acts as a cell-autonomous regulator of proliferation and differentiation. In addition, they provide biochemical evidence that loss of SLAP expression in cultured colonic organoids results in increased Src family kinase activity and global tyrosine phosphorylation, consistent with its known role as a suppressor of tyrosine kinase activity in immune cells. Consistently, treatment with an SRC kinase inhibitor inhibited the growth of SLAP-deficient organoids. These data provide solid evidence of a cell-autonomous role of SLAP in the colonic epithelium.

      This work would be improved by further description and interpretation of the SLAP expression pattern shown in the constitutive and tissue-specific KO to further support the conclusions made. In Supplementary Figure 1, magnification of the colon epithelium areas with SLAP expression shown by b-gal and anti-SLAP staining, highlighting regions of interest, would better support the conclusions regarding SLAP expression in specific regions of the colon epithelium. In Supplementary Figure 1B, the authors should indicate that the SLAP staining referred to is epithelial and in resident immune cells, as is mentioned in the text. Also, magnification of the boxed area of LRG5 staining in Figure 1 would improve this figure.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We have revised Fig 1 and S1 to better highlight SLAP expression patterns. Specifically, we have included higher-magnification images of the colonic epithelial regions, with clearly indicated regions of interest (new Figure S1). We have also clarified in the legend that SLAP staining is observed in both epithelial and resident immune cells, as described in the text. Additionally, we have provided a magnified view of the boxed area showing LGR5 staining in Figure 1 to improve clarity.

      Using a chemically induced model of colitis-associated cancer, the authors demonstrate that inactivation of SLAP shows a trend toward increased tumor formation (though this did not reach significance) as well as increased Src family kinase activity within tumors. Tumor spheres from SLAP-deficient animals showed enhanced growth that was suppressed by treatment with a Src family kinase inhibitor. Of note, the latter effect was specific to SLAP-deficient tumor spheres. These observations are convincing and support the authors' conclusion that SLAP has a tumor suppressor role in CRC through inhibition of SFK signaling.

      Mechanistically, elevated expression of the RTK, EphB2, was detected in immunoblots of SLAP KO colonic crypts, while overexpression of SLAP in CRC cell lines downregulated EphB2 protein levels. Using an EPHB2 inhibitor, the role of EPHB2 in the growth of SLAP-deficient colonic organoids was demonstrated. While these data generally support the authors' conclusion that SLAP limits colonic organoid growth by downregulating RTKS such as EphB2 and downstream Src family kinase activity, they do not show which cell types/regions in the colonic epithelium have increased EPHB2 protein and how this relates to SLAP and phospho-SRC expression, as shown in Figure 1 and Figure S1 immunocytochemistry. The expression of EphB2 and its role in colonic tumorsphere growth were not investigated.

      Overall, this work provides evidence of SLAP adaptor function in restricting tyrosine kinase signaling in the colonic epithelium, and suggests that loss of SLAP expression could promote tumorigenesis in this context.

      We thank the reviewer for their positive assessment of our tumour studies and for recognizing the evidence supporting a tumor suppressor role for SLAP through inhibition of SFK signaling.

      To address the reviewer’s mechanistic concerns, we performed additional experiments that are now included in the revised manuscript. We confirmed that loss of Slap is associated with increased EPHB2 expression in colonic crypts by IHC (new Figure 4B) and directly tested the role of EPHB2 in the Slap-deficient phenotype: EPH inhibition reduced both pTyr levels and SRC activation in Slap-deficient organoids (new Figure S2), demonstrating that SFK hyperactivation depends on upstream EPHB2 signaling. Consistent with this mechanism, we also observed increased EPHB2 tyrosine phosphorylation and active SRC (pSRC) association in isolated colonic epithelial cells following Slap deletion (new Figure 4A).

      To extend these findings to the tumour context, we examined the effect of EPH inhibition in human CRC tumoroids (new Figure 5). Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      Reviewer #2 (Public review):

      Summary:

      Protein tyrosine kinases are subject to diverse regulatory mechanisms controlling their activity in normal situations. The authors previously identified SLAP (Src-like adaptor protein), a negative regulator of receptor tyrosine kinase (RTK) signaling, as a key suppressor of the cytoplasmic tyrosine kinase SRC in the normal colon and demonstrated that SLAP is downregulated in a majority of colorectal cancers (CRCs).

      In this study, the authors further explored SLAP functions in mouse models using constitutive and inducible epithelial-specific Slap deletion (villin-CreERT2 model). They found that loss of SLAP augments colonic epithelial cell proliferation and that induction of tumorigenesis by the AOM/DSS protocol mimicking CRC leads to more aggressive tumors in the absence of SLAP. This effect is apparently cell-autonomous as growth of normal and tumoral colonic organoids is SLAP-dependent in in vitro settings. Finally, the authors define that, in colon, SLAP represses EphB2, an RTK lying upstream of SRC, and show that inhibitors of EphB2 can partially limit tumorigenic development in vitro.

      Strengths:

      The manuscript is clearly and concisely written, making it easy to follow. The data obtained in the mouse models are very convincing.

      Weaknesses:

      Direct evidence that EphB2 is activated/phosphorylated in the absence of SLAP is lacking, as conclusions are only based on results obtained with inhibitors. Some other issues have to be addressed before acceptance, in particular, the relevance of the findings in CRC patients.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We agree that direct evidence linking SLAP loss to activation of the EPHB2-SRC pathway would strengthen the study. To address this point, we performed additional experiments that are now included in the revised manuscript. In addition to demonstrating increased EPHB2 expression upon Slap deletion, we found that loss of Slap enhances the EPHB2 tyrosine phosphorylation (an index of EPHB2 activity) and association between EPHB2 and pSRC in isolated colonic epithelial cells, supporting increased signaling through this pathway (new Figure 4A). Furthermore, pharmacological inhibition of EPHB2 reduced both SRC activation and the hyperproliferative phenotype observed in Slap-deficient organoids (new Figure S2). Together, these findings provide functional and biochemical evidence that deregulated EPHB2 signaling contributes to SRC activation in the absence of SLAP. We also examined the effect of EPHB2 inhibition in human CRC tumoroids (new Figure 5).

      Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      To address the relevance of our findings in CRC patients, we also extended our analyses to human datasets (new Figure 5D). We observed a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. In addition, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients, whereas no such association was observed in microsatellite instability (MSI) tumours. These findings support the clinical relevance of the SLAP-EPHB2 signaling axis and are consistent with a role for SLAP in restraining EPHB2-dependent CSC signaling in CRC.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers have reported that the study of the EPHB2-SLAP-Src axis is not very developed, as most results are derived from using inhibitors. A key question is whether EPHB2 is activated by SLAP depletion and whether it is critical to SRC activation. Addressing these questions would greatly improve the paper.

      We thank the Reviewing Editor for this important suggestion. We would be happy for the editors to assess the revised version without involving the reviewers again. In the revised manuscript, we have substantially strengthened the mechanistic link between SLAP loss, EPHB2 activation, and SRC signaling. We show that Slap deletion increases both EPHB2 expression and tyrosine phosphorylation, enhances EPHB2-pSRC association in colonic epithelial cells. Importantly, pharmacological inhibition of EPHB2 suppresses SRC activation and rescues the hyperproliferative phenotype of Slap-deficient organoids. We further demonstrate that EPHB2 inhibition selectively impairs growth of CRC tumoroids with low SLAP expression, whereas this effect is largely abolished upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new biochemical and functional data establish EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and significantly reinforce the central conclusions of the study.

      Reviewer #1 (Recommendations for the authors):

      (1) Evidence of SLAP expression in the colon is an important basis for these studies and could be moved to the main Figure 1 rather than being in the supplementary material.

      We thank the reviewer for this suggestion. While we agree that documenting SLAP expression in the colon is important, we have retained these data in Figure S1 to maintain a concise main figure set, consistent with the recommended format for Short Reports.

      (2) Define AOM/DSS and briefly describe the model at first mention. In addition, the model in 3A includes TAM treatment at 45 days, but this is not mentioned in the text. Why is this done?

      We have better defined the AOM/DSS protocol at first mention in the revised manuscript and specified the rationale for tamoxifen administration at day 45, which is required to maintain efficient SLAP deletion throughout the duration of the experiment (90 days).

      (3) Evidence that the SRC inhibitor decreased phospho-tyrosine levels in addition to inhibiting the growth of organoids should be included.

      We included data showing the inhibitory effect of the used SRC inhibitor on global phospho-tyrosine levels in organoids in the revised manuscript.

      (4) Further experiments investigating the involvement of EphB2 in colonic tumor formation are of interest and would increase the significance of this work.

      We included data showing that SLAP modulation affects the response of tumoroids derived from cell lines to EphB2 inhibition, providing complementary mechanistic insights.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should confront their findings with data obtained in normal and pathological tissues: are SLAP, SRC, and EphB2 co-expressed, at the single cell level, in normal colon and CRC? In which cell populations? Is loss of SLAP associated with poor prognosis in CRC patients?

      We thank the reviewer for this important suggestion. We agree that assessing the relevance of the SLAP-EPHB2-SRC axis in human CRC is important. However, transcriptomic datasets have inherent limitations in this context, as SLAP primarily regulates signaling at the post-transcriptional level and SRC activity cannot be reliably inferred from mRNA expression.

      To address the clinical relevance of our findings, we performed additional analyses of CRC patient datasets. We found a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. Furthermore, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients. These findings are consistent with a role for SLAP in restraining EPHB2-dependent signaling in CRC. Finally, while our data identify EPHB2 as a critical upstream regulator of SRC signaling controlled by SLAP, we do not exclude the possibility that additional receptor tyrosine kinases contribute to the effects of SLAP loss during colorectal tumorigenesis.

      (2) In Figure 4A, total EphB2 levels are increased in the absence of SLAP in colonic crypts. However, the level of EphB2 phosphorylation is not shown. This is an important point to address. Which ligand(s) may activate EphB2?

      We agree that assessing EPHB2 activation is important. To address this point, we have included new data in the revised manuscript showing that Slap deletion increases EPHB2 tyrosine phosphorylation in isolated colonic epithelial cells, providing direct evidence that EPHB2 signaling is enhanced in the absence of SLAP. In addition, we show that loss of Slap increases EPHB2-pSRC association and that pharmacological inhibition of EPHB2 reduces SRC activation and suppresses the hyperproliferative phenotype of Slap-deficient organoids. Together, these findings establish EPHB2 as a critical upstream regulator of SRC signaling following SLAP loss.

      Regarding EPHB2 activation, previous studies have shown that EPHB2 is primarily activated by ephrin-B ligands expressed within the intestinal crypt compartment (Batlle et al., 2002). EPHB2 signaling may also be reinforced through cooperation with other Eph receptors, particularly EPHB3, which is highly expressed in intestinal stem and progenitor cells (Holmberg et al., 2006; Genander et al., 2009). In addition, SRC has been reported to phosphorylate EPH receptors, raising the possibility of bidirectional signaling that could further amplify EPHB2-SRC pathway activity (Leroy et al., 2009; Hochgräfe et al., 2010).

      (3) In the absence of SLAP, inhibitors of EphB2 should also decrease SRC activity as EphB2 lies upstream of SRC (Figure S3C). Does this occur in organoids?

      We now show that EphB2 inhibition reduces SRC activity in SLAP-deficient organoids.

      (4) What is the status of EphB2 and SRC (total, phosphorylated) in SW620 and HT29 CRC cells in the absence of SLAP?

      SW620 and HT29 cells are SLAP-low CRC models. Given that SLAP expression is already minimal in these cells, further depletion is unlikely to provide meaningful additional insight.

      (5) Expression of SLAP is associated with a decrease in the stem cell compartment in CRC cell lines (Figure S2). Is there a stem cell signature associated with low SLAP levels in CRC?

      We analyzed TCGA colorectal cancer datasets using a published colorectal cancer stem cell (CSC) signature. We found a low but significant inverse correlation between SLAP expression and the CSC-like activity score, supporting our experimental observations that SLAP restrains stem cell properties in CRC cells and organoids.

      (6) Does overexpression of the mutant form of SLAP (SLAPmut) limit SLAP effects in SW620 and HT29 CRC cells in Figure S2?

      We have now performed the requested experiments and found that, unlike wild-type SLAP, SLAPmut failed to inhibit tumoroid growth in CRC cells. These results are consistent with our previous findings showing that SLAPmut lacks tumour suppressor activity in CRC cells (Naudin et al., Nat Commun, 2014) and further support the requirement of SLAP signaling functions for the regulation of CRC stem-like properties.

      (7) Total SRC is missing in Figure 2B

      Total SRC levels are now included in the revised figure.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      The manuscript by Kostanjevec et al. investigates the mechanism behind spiral pattern formation in the cornea. The authors demonstrate that the spiral motion pattern on the mammalian corneal surface emerges from the interaction between the limbus position, cell division, extrusion, and collective cell migration. Using LacZ mosaic murine corneas, they reveal a tightening spiral flow pattern and show that their cell-based, in silico model accurately reproduces these patterns without global guidance cues. Additionally, they present a continuum model that extends the XYZ hypothesis to describe cell flux on the cornea, offering a quantitative explanation for tissue-scale processes on curved surfaces.

      Strengths

      The manuscript is well-written, with a systematic approach that clearly explains experimental setups, model construction, assumptions, parameter selection, and predictions. The discussion also provides insightful perspectives on the broader implications of the results for both physics and biology.

      We thank the reviewer for their positive assessment of the manuscript. We are pleased that the reviewer found the work well written and systematic, and that the experimental design, model construction, assumptions, parameter selection, predictions and broader discussion were clearly presented. We have aimed to preserve these strengths in the revised manuscript while substantially expanding the discussion and analysis in response to the reviewer’s concerns.

      Weaknesses

      The central premise of the manuscript, that the spiral patterning of epithelial corneal cells occurs without guidance cues, is not fully supported. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by pre-existing neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; rather, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea.

      Although our work shows that a swirling prepattern of nerves is not required to obtain radial patterns of epithelial cell migration, we now address why nerve swirling occurs and how it affects interpretation of our data. Previous literature indicates that nerve swirling is visible from approximately 3 weeks, before epithelial striping patterns become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks. However, both experimental observations and our modelling indicate that spiral epithelial cell migration proceeds for many days before reporter stripe patterns become visible. Thus, epithelial migration may already be underway before epithelial striping is detectable. If axons are following epithelial migration, they would therefore be expected to become radially aligned before the reporter epithelial stripes are evident.

      We also considered the alternative possibility that epithelial cells follow an axonal prepattern. However, several experimental observations argue against the hypothesis that epithelial migration is primarily guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, corneal axon extension. In chimeric Pax6+/− LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6+/− epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns. These findings suggest that axonal organization is at least partly dependent on epithelial behaviour rather than the reverse.

      Importantly, we do not exclude the possibility that additional cues, including axonal contact, neurotrophins or other environmental signals, may modulate or refine epithelial migration. We have therefore expanded the revised manuscript to discuss these possibilities and the relevant literature more fully.

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      What role do boundary conditions play?

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      We thank the reviewer for identifying these points, which we have now addressed in the revised manuscript.

      First, we have examined the role of geometry more systematically. The spiral pattern also appears in a disk geometry, and the same qualitative migration pattern is observed across a broader range of simulated geometries, including different curvatures, cap angles, a prolate ellipsoid, an oblate ellipsoid and a disk. The precise shape of the spiral depends on geometric features such as curvature and cap-angle opening, but spiral formation is robust across convex cornea-like geometries. These results are now discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in the revised Fig. 10.

      Second, we have clarified the role of boundary conditions, particularly the role of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters unchanged, the cornea fails to produce the radial striping pattern. At the alignment strength where robust spiral formation normally appears, the simulated tissue flow is disordered and resembles our in vitro calibration simulations. At higher alignment strengths, spiral formation is still not recovered; instead, defects become anchored to the boundary. These findings show that ordered influx from the limbus is important for promoting the spiral state. They are now discussed in the same new section and shown in revised Fig. 9.

      Third, we have revised the quantitative comparison between model and experiment. We agree that the original comparison was limited. We have therefore reanalysed both experimental and simulation data, focusing on the time-averaged hydrodynamic velocity field rather than short-range fluctuations amplified by divisions in the numerical model. We also replaced the Fourier-space velocity correlation functions with spatial velocity correlation functions, which are more directly interpretable. The revised analysis shows that the experiments have mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and about one hour in time, and that these are well captured by the simulations for both plastic and explant substrates. We have also added representative experimental and simulation snapshots in revised Fig. 4g-j.

      For the full cornea, we acknowledge that direct quantitative comparison remains limited by the available experimental data. We can compare with inferred migration direction fields and resurfacing timescales, but we do not yet have direct live measurements of the full corneal velocity field.

      The authors emphasize polar alignment as a key feature of the spiral pattern based on simulation results. However, they do not provide experimental evidence for this polar alignment. The manuscript includes discussions of polar and nematic symmetries that, without supporting data, feel somewhat distracting. If direct experimental evidence for polar alignment is not available, the authors could instead quantify nematic alignment as the spiral forms. This would also allow them to explore potential crosstalk between nematic cell orientation and the polar alignment of self-propulsion, especially considering recent studies showing alternative mechanisms for vortex formation in similar systems.

      We thank the reviewer for pointing out that the discussion of polar and nematic alignment was confusing. We have substantially revised this part of the manuscript.

      We agree that we do not have direct experimental evidence for polar alignment. However, several observations support the interpretation that the system is dominated by substrate-based polar motility with weak polar alignment. We have now quantified nematic alignment of cell orientations and find no evidence of significant elongation or local nematic order in corneal epithelial cells, as shown in the new Fig. 3d.

      Our computational model begins from uncorrelated, substrate-based polar active cell migration. We tested whether polar alignment is necessary and found that the in vitro data are inconsistent with the complete absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than expected from persistent driving alone.

      We also discuss that cell-cell stress patterns in the epithelium may remain nematic, but any such effect must be sufficiently weak not to dominate the observed axon motion. We have further revised the discussion to include recent related work showing spiral formation in substrate-based cell migration models with polar dynamics, as well as recent work indicating that nematic-like phenomenology can arise from minimal ingredients such as uncorrelated polar activity and cell deformability. This revision is intended to make the interpretation clearer and to avoid overemphasising unsupported claims.

      Reviewer #2 (Public review):

      In K. Kostanjevec et al, the authors study a possible mechanism for the formation of spiral patterns in the cornea. First the authors analyze an inferred velocity field, which is deduced from images of fixed corneas, and then determine the position-dependent spiral angle of this velocity fields. Next, the authors analysed two possible markers of cell polarity: the direction of the centrosome-nuclei and the axis of mitosis. Then the authors introduce a stochastic agent-based model of self-propelled particles with over-damped dynamics and with aligning interactions to the orientation of the nearest neighbors and to the particle's velocity. The authors claim to be able to reproduce the equal-time autocorrelation function and the velocity Fourier spectrum. Then the authors introduce the geometry of the cornea by constraining the dynamics on a spherical cap and show that their model can reproduce a typical trajectory in experiments. Finally, the authors produce a phase diagram of the states at a fixed time point as a function of the spherical cap radius and the strength of the coupling aligning constant. Finally, the authors propose an interpretation of the cell fluxes based on the equation of mass conservation.

      We thank the reviewer for their careful assessment of the manuscript and for recognising the work as a solid theoretical study. We have revised the manuscript substantially in response to the reviewer’s major concerns, particularly regarding the terminology of topological defects and stagnation points, the comparison with experiments, and the role of corneal geometry.

      Regarding the terminology of topological defects, we have clarified the distinction between a velocity-field stagnation point and the topological classification of the velocity direction field. Stagnation points can be assigned a topological index when one considers the direction of the vector field away from the core, while ignoring the magnitude. We agree that the physical origin of interactions in a velocity field differs from that in a director field, and we have revised the text to avoid confusion. The revised manuscript now includes Box 1, which summarises the relevant topological concepts and caveats.

      Regarding the comparison with experiments, we have expanded and clarified the validation of the inferred velocity field. The LacZ reporter system was designed for lineage tracing and therefore reports coarse-grained cell motion and growth patterns. Direct live imaging of the full cornea remains technically difficult because of the macroscopic size, curved surface and long resurfacing time. However, live fluorescent reporter systems are consistent with our in vivo model and support the interpretation that inferred velocity fields recapitulate epithelial migration in vivo. We have also clarified the interpolation procedure using the XY model: approximately 30% of the corneal surface is covered by directly estimated velocity directions from stripe edges, rising to more than 50% near the central spiral. These measured directions provide sufficient boundary conditions for the annealing procedure to converge to a slowly varying field consistent with the observed stripe geometry.

      We have also revised the comparison between simulation and experiment. For in vitro data, we now compare spatial velocity correlation functions of the time-averaged hydrodynamic velocity field, rather than relying on Fourier-space correlations. The revised comparison shows good agreement between experiments and simulations for both plastic and explant substrates. For the full cornea, we acknowledge that the available quantitative data are limited to inferred migration direction fields and resurfacing timescales.

      Regarding the role of geometry, we have now simulated a wider range of substrate geometries, including spherical caps with different curvatures and cap angles, prolate and oblate ellipsoids, and a flat disk. The spiral migration pattern is robust across these convex geometries, although the detailed spiral shape depends on geometric features. We have also clarified that the boundary condition of inward limbal influx is crucial: without limbal stem cells, the radial striping pattern does not form, and defects may instead anchor to the boundary. These results are now shown in revised Figs. 9 and 10 and discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells.”

      Overall, these additions clarify that the proposed mechanism relies on the combination of polar motility, weak alignment, limbal influx, cell division and extrusion, and cornea-like confinement, while also acknowledging the current experimental limitations.

      Recommendations for the authors:

      Reviewing Editor:

      The authors could strongly improve the manuscript by following the recommendations given below.

      Reviewer #1 (Recommendations for the authors):

      There are, however, substantial shortcomings that authors need to address to make their claims supported by enough evidence:

      Major concerns:

      (1) Neglect of Potential Axon Guidance

      As it stands the premise of the paper is unfortunately faulty. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by preexisting neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      See for example:

      from: https://iovs.arvojournals.org/article.aspx?articleid=2126612:

      It is therefore not clear why the authors clearly show the axon vortex in Figure 1, then talk about prepatterning - and never mention the axon vortex again.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; however, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea. Although our work shows that a swirling ‘prepattern’ of nerves is not required to obtain radial patterns of epithelial cell migration, we address below why it occurs and how it affects our data.

      An intuitive interpretation of corneal anatomy would be that sensory axons follow the path of least resistance between migrating epithelial cells. We believe this is the most likely explanation in the normal wild-type cornea; however, the reviewer correctly notes that previous literature indicates that axon patterning precedes epithelial cell patterning. Nerve swirling is observed from approximately 3 weeks, earlier than the epithelial striping patterns that become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks (e.g. Collinson et al., 2002 [PMID: 12203735]; Iannaccone et al., 2012 [PMID: 22347498]; McKenna and Lwigale 2011 [PMID: 20811061]). We now address this in the revised manuscript and below.

      The early appearance of radial axonal projections can be readily explained even if axons are following the epithelial cells. Experimental observations (Collinson et al., 2002 [PMID: 12203735]) and our modelling (Fig. 6a,b) both indicate that spiral epithelial cell migration proceeds for many days before stripe patterns become visible in reporter mosaics. Thus, epithelial migration is already underway before the striping pattern becomes detectable. If axons are following the epithelial migration, they would be expected to become radially aligned before the reporter epithelial stripes were evident. The apparent precedence of radial axonal projections over epithelial striping is therefore fully consistent with the biological scenario in which axons follow the migrating epithelial cells. Corneal epithelial basal cells have been shown to wrap around individual and grouped subbasal axons, acting as surrogate glia (Stepp et al., 2016, Investigative Ophthalmology & Visual Science 57, 1292), which represents a plausible mechanism to allow migrating epithelial cells to shepherd axons in their direction of movement.

      We considered that epithelial cells may follow an axonal prepattern, but several experimental observations argue against the hypothesis that epithelial migration is guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, the extension of corneal axons (Leiper et al., 2009 [PMID: 19029029]; Song et al., 2004 [PMID: 14744881]). Furthermore, in chimeric Pax6<sup>+/−</sup> LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6<sup>+/−</sup> epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns (Leiper et al., 2009 [PMID: 19029029]). These results suggest that axonal organization is at least partially dependent on epithelial behaviour rather than the reverse.

      Importantly, none of this evidence excludes the possibility that additional cues (including axonal contact or other environmental signals) may modulate or refine epithelial migration. For example, Walczysko et al. 2016 [PMID: 27563231] showed that isolated epithelial cells cultured on de-epithelialised and de-nervated corneal stroma still migrate with a small but significant radial bias, indicating that epithelial cells can respond to physical features of their environment. Our model likewise includes a component of alignment with environmental structure. Since the central radial striping in many of our simulations is somewhat less ordered than in vivo, it is plausible that additional biological guidance cues or axonal neurotrophins help refine the pattern.

      We now discuss these issues and the relevant experimental evidence more fully in the revised manuscript.

      (2) Model Validation and Complexity

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      To address the Reviewers’ concerns about the role of geometry, we first note that the spiral pattern also appears in a disk geometry. In fact, the geometric parameters of the cornea vary across mammals, including humans, yet the spiral pattern is preserved (Dua, et al., 1993 [PMID: 8325424]; Zander and Weddell, 1951 [PMID: 14814019]). Similar variations arise in disease states, e.g. in keratoconus the cornea becomes elongated.

      On a spherical cap, and all shapes with the same topology, the boundary winding number fixes the interior index, so ongoing limbal influx maintains a total index of 1. To explore the role of geometry more systematically, we simulated a broader range of geometries, including different curvatures, cap angles, a prolate and an oblate ellipsoid, and finally a disk, and compared the resulting patterns with published data across mammals and with disease states.

      For all of these shapes, we find the same qualitative migration pattern – a central spiral, although its precise shape depends on geometric features such as curvature and cap-angle opening. These results are discussed in a new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in Fig. 10 of the revised manuscript.

      What role do boundary conditions play?

      In the revised manuscript, we address the role of boundary conditions, in particular the presence of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters kept unchanged, the cornea fails to make the radial striping pattern. We observe two distinct changes: First, at J=0.1, the amount of alignment at which robust spiral formation appears normally, the simulated tissue flow is instead disordered, in fact very similar to our in vitro calibration simulations (revised Figure 4). This shows that the boundary cue of an ordered influx from the limbus promotes flocking when it otherwise would not (yet) appear. Second, when we increase the alignment to J=0.15 or J=0.2, we still do not observe spiral formation. Instead of the expected central vortex shape, we have anchoring of defects to the boundary, facilitated by the effective compressibility of the tissue because of the density feedback in the division/extrusion rates.

      These findings are discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and in Fig. 9 in the revised manuscript.

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      Regarding the in vitro cell data and matching simulations: We are aware that we have limited data to work with, and our match is intended as a rough estimate of physical parameters.

      For the revision, we have reanalysed both experimental and simulation data carefully. We realised that certain details of our numerical model amplify short-range fluctuations, namely the way cell divisions induce stress dipoles, and the fact that we did not include cell-cell friction forces. This is not merely a guess but emerged from related theoretical work by some of us (Keta and Henkes [PMID: 40556485]; Kammeraat et al, arXiv:2508.01046 (2025)).

      We therefore compared the time-averaged velocity field excluding divisions to the experiment, the same quantity that we already introduced as hydrodynamic velocity for the corneal surface. We also carried out further simulations that included cell-cell friction forces for comparison. They led to very similar results, albeit with a transition to flocking at somewhat lower alignment strength J.

      Instead of the Fourier-space velocity correlation functions that are hard to interpret, we switched to spatial velocity correlation functions. Our experiments have ‘swirly’ velocity fields with mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and one hour in time. As can be seen in the revised panels 4c,d for the spatial correlations of the hydrodynamic velocity, the match between experiment and simulations is in fact good for both plastic and explant substrates. There are also systematic changes in length and time scales between the two experiments that emerge without fine-tuning from simulation. We furthermore have included snapshots of both experimental in vitro conditions and matching simulations as new panels 4g-j, showing that the mesoscale correlations appear in the hydrodynamic velocity.

      Ultimately the parameter values and length and time scales that we infer for our systems (see Table 1) are quantitatively consistent with three other estimates of in vitro epithelia (Henkes et al. 2020 [PMID: 32179745]; Saraswathibhatla et al, Extreme Mechanics Letters 48, 101438 (2021), Kammeraat et al, arXiv:2508.01046 (2025)).

      For cell flows on the full cornea, regrettably we lack further quantitative data to compare to beyond the migration direction fields inferred in Figure 2, and the time scales of corneal resurfacing (Figure 6).

      (3) Importance of polar alignment.

      Polar alignment is put forward as one main feature for the observed patterns based on the simulation results. However, no experimental confirmation for such polar alignment is presented. The authors instead present various scattered discussions about polar versus nematic symmetry, which at times reads rather unnecessary and distracting from their main message.

      If they are not providing direct experimental evidence on the polar alignment, at least they could quantify nematic alignment of cell orientation as the spiral forms. One possibility is that nematic alignment of cell orientations has a crosstalk with the polar alignment associated with the self-propulsion of the cells. This is important because, as acknowledged by authors, several recent works in the context of in vitro epithelial under disk confinement have revealed alternative mechanisms for spiral vortex formation.

      We thank the Reviewer for pointing out the confusion, and we have therefore completely revised our discussion of alignment. While the evidence remains indirect, the following observations lead us to conclude that our system is dominated by substrate-based polar motility and weak polar alignment:

      We have investigated nematic alignment of cell orientations. As can been seen in new Figure 3d, corneal epithelial cells do not show evidence of significant elongation in any direction, i.e. we find no local nematic order.

      Our computational model starts from uncorrelated, substrate-based polar active cell migration. We have carefully investigated if polar alignment is a necessary ingredient and found that the in vitro data are inconsistent with the absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than those expected from persistent driving only.

      Still, cell-cell stress patterns in the epithelium may remain nematic. While one cannot exclude anything, the effect must be sufficiently weak to not affect axon motion. Their naturally long, thin shapes would strongly react to nematic stresses, but we do not see ±1/2 defects in their growth patterns, only polar ±1 defects.

      We were recently made aware the work of Lång et al. (2024) [PMID: 38630812], in which the authors also observe spiral formation in cell migration on a substrate, with +1 topological defects. They explain their observations using a polar model, and our results are consistent with their observations and model.

      In the active matter community, the ‘active nematic cell sheet’ paradigm is currently undergoing a revision. Notably, nematic-like phenomenology, in particular ±1/2 defects, can also arise from the minimal ingredients of uncorrelated polar activity and cell deformability (Chiang et al. 2024 [PMID: 39302997]).

      Minor comments:

      - It would help the reader to see some representative images of the cells, explants, and the simulation (maybe with vectors overlaid), as it could be hard to imagine what exactly is happening and the other relevant properties to compare.

      Please see new panels Fig. 4g-j, and new supplementary videos S3 and S4.

      - Add a colorbar that represents direction to Fig. 1a.

      We assume the Reviewer meant Fig. 2a. We have added a circular orientation colour chart to the image.

      - When do additional +1,-1 defects appear? The authors just say it is unlikely, but it is not clear when they are observed. Disease state?

      The reviewer is correct about pathology. We now cite (in conclusion) Collinson et al. 2004 [PMID: 15037575], which shows disruption and discontinuity in Pax6-mutant corneas with chronic corneal degeneration. We also have a publication in prep that shows extra +1 and -1 discontinuities occur during wound healing. We don’t want to include these data in this manuscript, but cite Sagga et al. (in prep).

      Reviewer #2 (Recommendations for the authors):

      Overall, the manuscript presents a solid theoretical work. However, I have a few major concerns on this manuscript. (1) The authors use the concept of topological defect to refer to stagnation points of the velocity field (2) The comparison to experiments remains qualitative and it is unclear whether the proposed mechanism is at work, (3) The role of geometry remains unclear.

      Major points:

      (1) On page 3 the authors claim that topological defects exist in velocity fields, however this statement is incorrect and can confuse readers. Unlike the director field of an ordered phase, their velocity field is a vector in R^2 with a norm that is not fixed, and therefore the velocity field has no topological defects. In fluid dynamics, these special points in the velocity field are called stagnation points, fluid sinks or sources. This distinction is important because the physical nature of the interaction forces between two "topological defects" in a velocity field is fundamentally different to the interaction forces between two topological defects in a director field. For these reasons, it can confuse readers to mix the two concepts (stagnation points vs topological defects). Note that if the authors address this concern, many parts of the main manuscript should be rewritten.

      Stagnation points are topological defects since the topological classification ignores the magnitude and looks only at the direction of the vector field away from the core. There are some caveats related to the role of boundary conditions, since in the case of a fluid, if boundary conditions are not fixed, one can eliminate the defect by setting the flow field to zero everywhere. Mathematically, a pair of point sources or sinks in an incompressible potential flow interacts through the same Green’s function that gives the elastic interaction between two-point disclinations in a 2D nematic director field, so their long-range pair potentials are formally identical (~ln(r) in 2D). The physical mechanism behind these forces is, as correctly pointed out by the reviewer, quite different. In addition, our coarse-grained velocity field is effectively compressible, which has consequences for the hydrodynamic equations we (can) write, see below.

      In the revised version, we clarify these points by adding Box 1, which summarizes the idea of topological defects.

      (2.1) How did the authors check that the velocity field that is inferred from the stripe edges matched the coarse-grained cell velocity field in a life sample? How did the authors validated the interpolation of the inferred velocity field using an XY model? What is the scale of the inferred velocity field? Can the authors clarify also this point?

      The murine cornea LacZ reporter system was specifically designed to allow for lineage tracing, i.e. following coarse-grained cell motion and growth patterns. The combination of macroscopic size (3.6 mm diameter), two-week resurfacing time and the curved surface however stymied our early attempts to directly measure the cell velocity field on the cornea using confocal time-lapse microscopy. However live sample fluorescent reporter systems are fully consistent with our in vivo model and shown conclusively that inferred velocity fields are recapitulated in vivo (Park et al., 2019 [PMID: 31843909]).

      For the XY model inference: We first note that approximately 30% of the corneal surface is covered by directly estimated velocity directions from the stripe edges (red arrows, Fig. 11f). This rises to more than 50% near the central spiral (red arrows, Fig. 11h) due to the way the stripes narrow near the centre due to cell extrusion. Therefore, the inferred areas are only slightly more than half of the cornea, and in regions where we expect the flow field to be largely uniform with no defects. The red arrows provide sufficient boundary conditions that a simple annealing simulation (or equivalently an energy minimisation) of the XY model rapidly converges to a slowly varying field consistent with those boundary conditions.

      Furthermore, in simulation we observe that stripe edges and the macroscopic velocity field correlate strongly with each other once the spiral has fully formed (Fig. 6b-c).

      The scale of the inferred velocity field is the distance between arrows in our digital version of the cornea, approximately 40 concentric rings over a 70° cone angle for a R = 1800 μm micron cornea, resulting in a spacing of 55 μm between velocity arrows. This is about at the scale of the in vitro velocity correlations. We are not able to obtain velocity magnitudes using this procedure.

      Furthermore, the spiral angle profile reported in Fig. 7b-c appear to be different to that found in experiments Fig. 2b. Can the authors clarify if their theoretical framework reproduces the spiral angle profiles?

      First, we note that empirically, simulations with the largest two radii (R = 1000 μm, R = 1500 μm), approaching the full experimental size, match the observed angle profile best. They both consist of a radially inward profile α(0) = 0° at the edges, only increasing, corresponding to a tightening spiral, below about θ = 20°. Note that very near the corneal centre, few cells contribute to data, and additionally the central defect position fluctuates somewhat. Therefore, the angle profile not reaching α(0) = 90° in the simulations is due to fluctuations and lack of statistics. As these simulations were run with our best fit experimentally matched parameters, this convergence is meaningful and cannot be scaled out.

      Second, our partial model (eq. 4, reproduced below) links the spiral profile angle α(θ) with the macroscopic velocity magnitude v(θ) and the net cell loss rate A(θ), and it depends explicitly on the corneal radius R.

      That is one equation for three radial fields. If we make the reasonable assumption that simulations at fixed 𝐽 that differ only in R have the same constitutive law A(ρ), and would follow the same continuum velocity equation that ultimately sets v(ρ), the radius R still appears explicitly in the equation. This is consistent with the different radial profiles for different R that we observe in Fig. 7c. Furthermore, in Fig. 7b we observe that above the flocking threshold, different J lead to very similar profiles. This indicates that the system enters a fully polar phase.

      The same is true in experiment: We expect that cell mechanics and planar cell polarisation coordination are local effects that will set J, A(ρ), and v(ρ). Thus, we do predict that the radius R will explicitly affect the spiral profile consistent with the equation above, but we would need more direct measurements of J, A(ρ), and v(ρ) to go any further.

      Properly answering this question would first require simulations of a non-dimensionalised model with different radii, boundary influx, alignment strengths and constitutive laws for the division / extrusion dynamics. Then one would want to construct a matching equation for the polarisation and / or velocity field, going beyond the Malthusian flock approximations of constant magnitude v(θ) and net cell loss rate A(θ) = 0. This is well beyond the scope of this publication, and there is certainly no experimental data to compare to.

      (3.1) It appears that the formation of the spiral velocity pattern results from a combination of a radial flow of cells due to cell division at the outer boundary and apoptosis at the geometrical center and azimuthal flow of cells due to flocking in confined geometries. Is this correct? Can the authors explain how does the 3d geometry of the spherical cap modify each of these flow fields?

      The Reviewer is broadly correct; however, the true picture is more subtle. Almost all of the divisions and extrusions occur in TA cells, neither at the limbus or at the corneal centre (please see the A(θ) profiles in Figure 8). Flow is then a spiral flock that is partially radial and azimuthal, and with a variable velocity magnitude. It ultimately all has to follow the flux equation 4, in steady state.

      For the modification of the flow field due to 3d geometry: Broadly, they simply change the amount of corneal surface available at different angles from the limbus when switching to isomorphic surfaces like, e.g. the disk. Thus, we still observe spirals, but of somewhat modified shapes. Please see the reply to Reviewer 1 above, and new Figure 10 for simulations of different geometries.

      A full theory of radially symmetric corneal shapes would again need additional velocity and constitutive equations and is beyond the scope of this publication.

      (3.2) In their model, it appears that the geometry is introduced by constraining the dynamics of agents, and it has not direct influence on the alignment of agents. Can the authors explain how the geometry of the spherical cap influences the emergent spiral states? Can the authors show whether their results are robust to changes in the substrate geometry? For example, by changing the substrate geometry from a spherical cap to another convex shape. Can the authors identify differences between the spiral patterns on a spherical cap vs that on a flat disk?

      Please see the response to Reviewer 1, above, and new Figure 10. Briefly, the results are robust to changes in the corneal geometry as long as they are other convex shapes. There are differences in details of the spiral shape.

      Minor comments:

      - On page 3, the authors claim that the Euler characteristic of a spherical cap or a disk is 1. Unfortunately, this statement is incorrect. The Euler characteristic of a spherical cap or a disk is determined by the winding of the vector field around the open boundary. In their case the Euler characteristic is 1 because the velocity field is oriented towards the top of the spherical cap, which give a winding of +1. The authors explain this correctly on page 13.

      The Euler characteristic is a topological invariant of the surface and does not depend on the tangent vector field. A disc or spherical cap has Euler characteristic 1. What depends on the boundary winding is the index formula for a vector field on a surface with boundary: the winding of the field along the boundary determines the corresponding boundary contribution, and hence the sum of interior indices. Thus, while the reviewer is correct about the role of boundary winding in computing the field’s index, this does not alter the Euler characteristic of the surface itself.

      We note that the orientation of the velocity field does indeed matter in our simulations: In new Figure 10, we show that in the absence of the boundary condition of inward flux at the limbus, we do not observe a spiral robustly, and we do see anchoring of defects to the boundary, with winding numbers that are now different.

      - On page 5, the authors claim the clockwise and counterclockwise -oriented spirals are equally likely, however no quantification is provided to support this claim. Can the authors clarify this point?

      The approximately equal likelihood is as described in Collinson et al. (2002) [PMID: 12203735].

      - On page 8 the authors state "Dipolar active force cannot cause a single cell to migrate, and the flow is an emergent collective phenomenon", here I was confused, because a bacteria swimming in a Newtonian fluid can self-propel by exerting a dipolar active force on the fluid. See for instance work by E. Lauga or I. Aronson. Can the authors clarify this point?

      The Reviewer is correct about how bacteria can swim using dipolar active forces. However, and unfortunately, that language was straight ported to very different conditions, that of cells migrating on a frictional substrate with no induced flow, with the equation of motion . Unless there is a net force arising from the stress profile, the cell cannot move, and with microscopic models where the active stress is a single value per cell, that statement is always true. Of course, more detailed cell models with active stresses exist, but their motion is still due to the net force arising from them. We have clarified the statement in the revised manuscript.

      - On page 16, the expression for the velocity seems to miss a parenthesis.

      Fixed. Thanks!

      - On page 16, the authors claim that the flux profiles in the simulations and in the XYZ model are in good agreement. However, there are clear differences for theta> 50 deg. Can the authors discuss the possible explanation for these differences? Note that one may expect a between agreement between two theoretical approaches.

      These disagreements are due to imperfections in the observed spiral patterns in simulations. They are not perfectly radially symmetric, and the defect is not always in the dead centre of the cornea. Thus, the theoretical predictions are not quite accurate. Numerically, what happens for θ > 50° is that the radial bins sometimes include part of the limbal zone with strong proliferation, and the corneal edge itself. Both are not described by the flux prediction. We decided not to crop out this region and rather explain where it comes from in the text.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Wang, Po-Kai et al., utilized the de novo polarization of MDCK cells cultured in Matrigel to assess the interdependence between polarity protein localization, centrosome positioning and apical membrane formation. They show that the inhibition of Plk4 with Centrinone does not prevent apical membrane formation, but does result in its delay, a phenotype the authors attribute to the loss of centrosomes due to the inhibition of centriole duplication. However, the targeted mutagenesis of specific centrosome proteins implicated in the positioning of centrosomes in other cell types (CEP164, ODF2, PCNT and CEP120), as well as the use of dominant negative constructs to inhibit centrosomal microtubule nucleation did not affect centrosome positioning in 3D cultured MDCK cells. A screen of proteins previously implicated in MDCK polarization revealed that the polarity protein Par-3 was upstream of centrosome positioning, similar to other cell types.

      Strengths:

      The investigation into the temporal requirement and interdependence of previously proposed regulators of cell polarization and lumen formation is valuable. The authors have provided a detailed analysis of many of these components at defined stages of polarity establishment, and well demonstrate that centrosomes are not necessary for apical polarity formation, but are involved in the efficient establishment of the apical membrane.

      Weaknesses:

      Key questions remain regarding the structure of the intracellular cytoskeleton following depletion of centrosomes, centrosome proteins, or abrogation of centrosome microtubule nucleation. The authors strengthen their model that centrosomes are positioned independently of microtubule nucleation using dominant negative Cdk5RAP2 and NEDD-1 constructs, however, the structure of the intracellular microtubule network remains unresolved and will be an important avenue for future investigation.

      We thank the reviewer for raising this important point. We agree that understanding the organization of the intracellular microtubule network following centrosome depletion, disruption of centrosomal proteins, or inhibition of centrosomal microtubule nucleation will be important for further mechanistic insight. However, a detailed analysis of cytoskeletal architecture under 3D culture conditions would require super-resolution or other advanced microscopy techniques and therefore falls beyond the scope of the current study. Nevertheless, several previous studies conducted under conventional 2D culture conditions provide relevant mechanistic context for interpreting our findings.

      (1) Centrosome depletion by centrinone treatment

      Previous studies have shown that upon centrosome loss induced by the Plk4 inhibitor centrinone, alternative microtubule-organizing centers (MTOCs), particularly the Golgi apparatus, can compensate by nucleating non-centrosomal microtubules (Chen et al., 2022; Gavilan et al., 2018; Martin, Veloso, Wu, Katrukha, & Akhmanova, 2018; Wu et al., 2016). Consistent with these findings, our microtubule regrowth assays in 2D MDCK cells (see Author response image 1) demonstrated that centriole depletion markedly altered microtubule organization. Control cells displayed the typical radial microtubule array emanating from a centralized centrosome, whereas centrinone-treated cells exhibited a dispersed microtubule network with enhanced microtubule growth from Golgi-associated sites.

      Author response image 1.

      Staining of control or centrinone (CN) treated MDCK cells for α-tubulin (magenta), Golgi GM130 green) and DNA (DAPI, blue) 1 min after nocodazole washout. Z-maximum projections of confocal images. The boxed cells in the overview images are magnified, and the microtubule regrowth regions are further enlarged. Scale bars: 50 μm (overview images), 10 μm (magnified cells), and 5 μm (enlarged microtubule regrowth regions).

      (2) Disruption of centriole or centrosomal proteins

      Previous studies have reported that depletion of the subdistal appendage protein ODF2 reduces centrosome–microtubule interactions and destabilizes centrosomal microtubules (Hung, Hehnly, & Doxsey, 2016; Ibi et al., 2011; Tateishi et al., 2013). In addition, in cells lacking the PCM protein pericentrin (PCNT), AKAP450 has been shown to partially compensate for centrosomal microtubule nucleation activity, although at a somewhat reduced level compared with wild-type cells (Figure 4—figure supplement 3A, B) (Gavilan et al., 2018).

      (3) Inhibition of centrosomal microtubule nucleation

      For dominant-negative Cdk5RAP2 and NEDD1 constructs, a previous study demonstrated that these constructs displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this report, our MDCK cells expressing dominant-negative Cdk5RAP2 or NEDD1 also exhibited reduced γ-tubulin localization at centrioles (Figure 4—figure supplement 3C, D, and G).

      Together, these findings support the interpretation that centrosomal microtubules are not strictly required for polarized vesicle trafficking, centrosome migration, or epithelial polarization, but instead enhance the efficiency and robustness of these processes. We agree with the reviewer that future studies using advanced imaging approaches under 3D culture conditions will be important to resolve the spatial organization and dynamics of the intracellular cytoskeleton during epithelial polarization.

      Reviewer #3 (Public review):

      Here the Wang et al resubmit their manuscript describing the events in the establishment of polarity in MDCK cells cultured in vitro. As with the original version, the description is throughout and is important to the field to report as it establishes a hierarchy of events in polarization, placing Par3 upstream of centrosome positioning and apical membrane component trafficking. Unfortunately, in the revised version, the authors addressed almost none of my points. They did a cursory job of responding in the rebuttal letter but made little attempt to actually address what was being asked or to incorporate any of my suggestions into the manuscript. The particularly egregious examples are cited below:

      Comments on revisions:

      (1) My original main experimental concern was not addressed: I had originally asked what role microtubules play in the process of polarization (either centrosomal or non-centrosomal). An obvious model is that Gp135, Rab11, etc. are delivered to the AMIS on centrosomal microtubules. Centrosomes might also be pulled to the AMIS via cortically derived microtubules as is the case in the C. elegans intestine where the centrosome moves apically on apical microtubules via dynein directed transport to the cortically anchored minus ends. The authors do not explore the role of microtubules in the revision, citing that it was not possible to observe the microtubules directly or to perform nocodazole experiments during polarization.

      Instead, the authors use a relatively new genetic tool to disrupt centrosomal microtubules. They appear to succeed in displacing centrosomal g-tubulin using this tool, but without being able to observe microtubules, a remaining caveat of this experiment is that it is still unclear whether the authors have removed centrosomal microtubules. Compounding this issue is that this tool has never been used in MDCK cells. The authors conclude "we found that cells lacking centrosomal microtubules were still able to polarize and position the centrioles apically.", but they have not shown this, instead the data suggest this conclusion and the authors should acknowledge the caveat that they have no idea whether centrosomal microtubules are abolished.

      We appreciate the reviewer’s important comments regarding the role of microtubules during epithelial polarization. We agree that determining how centrosomal and/or non-centrosomal microtubules contribute to apical trafficking and centrosome positioning represents an important mechanistic question.

      We previously attempted to directly visualize microtubules during live imaging using SPY-tubulin labeling (see Author response image 2). However, under 3D Matrigel culture conditions, MDCK cells rapidly become rounded and densely packed, substantially reducing image contrast and making the majority of intracellular microtubule networks difficult to resolve, except for spindle microtubules and the cytokinetic bridge. In addition, nocodazole treatment caused mitotic arrest under our experimental conditions, thereby preventing de novo polarization from proceeding and precluding interpretation of polarity establishment.

      Author response image 2.

      Time-lapse maximum-intensity z-projections of MDCK cells expressing EGFP-PACT (yellow; centrosome marker) and H2B-mCherry (magenta; nuclei) embedded in Matrigel. Microtubules were labeled with SiR-tubulin (cyan) before live-cell imaging. Images show a representative dividing cell. Time stamps indicate hours and minutes relative to anaphase onset (0:00). Scale bar, 10 μm.

      To partially address the role of centrosomal microtubules, we used dominant-negative Cdk5RAP2 and NEDD1 constructs that have previously been shown to displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this study, we observed substantial loss of γ-tubulin from centrioles in MDCK cells expressing these constructs (Figure 4— figure supplement 3C, D, and G). However, as the reviewer correctly points out, we were unable to directly visualize centrosomal microtubules under our 3D imaging conditions. Therefore, we cannot definitively conclude that centrosomal microtubules were completely abolished. We have revised the manuscript to clarify this limitation and to more cautiously state that our data suggest centrosomal microtubules may not be strictly required, for apical polarization and centrosome positioning under these conditions.

      Similarly, the authors also state: "Additionally, although PCNT knockout cells show reduced microtubule nucleation ability, they still recruit a small amount of γ-tubulin". Where are the data that show that microtubule nucleation is reduced in these PCNT knock out cells?

      We thank the reviewer for this comment and apologize for not sufficiently presenting these data in the previous revision. To directly assess microtubule nucleation activity in PCNT-KO cells, we performed microtubule regrowth assays in MDCK cells following nocodazole washout (Figure 4— figure supplement 3A, B). Compared with wild-type cells, PCNT-KO cells showed reduced centrosomal microtubule regrowth, indicating impaired microtubule nucleation capacity.

      Importantly, microtubule nucleation was not completely abolished in PCNT-KO cells, consistent with previous reports showing that AKAP450 can partially compensate for the loss of pericentrin and maintain residual centrosomal microtubule nucleation activity (Gavilan et al., 2018).

      (2) Many of my comments were addressed in the rebuttal, but not in the text.

      We sincerely thank the reviewer for the valuable suggestions. We have carefully considered all comments and incorporated many of the recommended revisions into the revised manuscript. However, we found that including every additional analysis, experiment, and discussion in the main manuscript would substantially reduce its coherence and readability. Therefore, while not all new analyses and experimental results are included in the revised manuscript, we have addressed every comment comprehensively in this response letter. Where necessary, we performed additional experiments and analyses to obtain the requested data, and the corresponding results and explanations are provided in our responses. We hope the reviewer will understand our effort to thoroughly address all comments while preserving the clarity and overall flow of the manuscript.

      The non-centrosomal GP135 in Figure 2 is not acknowledged or explained.

      We apologize for not sufficiently addressing the non-centrosomal Gp135 signal in Figure 2. We have now revised the manuscript to explicitly describe and discuss this point in the text (Page 5, Paragraph 4).

      That the polarity index does not actually measure polarity, but nuclear-centrosome distance is not acknowledged or explained in the paper.

      We have revised the manuscript to explicitly state that the “polarity index” represents the distance between the nucleus and the centrosome, which we use as a quantitative indicator of the degree of cell polarity (Page 5, Paragraph 1).

      I still don't believe that the quantification in Figure 3D matches the images I am being shown in Figure 3A. In the centrinone treatment condition, there is certainly an enrichment of GP135 at the AMIS that is not detected in the quantification. The method described in the rebuttal might miss this enrichment if it is offset from line drawn between the centroid of the two nuclei.

      We thank the reviewer for this comment. To better address this concern, we have now included a 3D view of the corresponding image data (see Author response image 3 and Author response image 4). This analysis clarifies that, in the centrinone-treated condition, the Gp135 signal is not localized at the geometric center of the cell doublet, but is instead offset from the axis used in our line-scan quantification. As a result, the enrichment visible in the projection image was not fully captured by the original quantification method. SiR-DNA Gp135

      Author response image 3.

      Author response image 4.

      3D reconstructions of p53-KO control and centrinone-treated MDCK cell doublets expressing EGFP-Gp135. Images are shown after 90° rotations about the x-axis (or y-axis) to visualize the spatial distribution of Gp135. Fluorescence intensity profiles of EGFP-Gp135 were measured along the line connecting the two nuclei. White arrows indicate the central fluorescence intensity value used to quantify Gp135 accumulation at the apical membrane initiation site (AMIS) (a.u., arbitrary units). Time stamps indicate hours and minutes.

      Cell height changes in the centrosome depleted cysts are still referenced in the text ("the cell heights of the centrosome-depleted cysts are less uniform"), but no specific data or image is called out. Currently, Figure 3G is referenced, but that is a graph of GP135 intensity

      We have revised the manuscript to indicate representative images, ensuring that the text and figures are consistent (Page 7, Paragraph 3).

      In my original review, I called on the authors to comment on the striking similarity of the mechanisms they documented in MDCK cells to what has been shown in in vivo systems. The authors did not do this, instead restating in the rebuttal some features of what they found. But, the mechanisms shown here are remarkably similar to the polarization of primordia that generate tubular organs in vivo. Perhaps most striking is the similarity to the C. elegans intestine where Par3 localizes to the cortex at the site of an apical MTOC that pulls the centrosome to the apical surface via dynein (Feldman and Priess, 2012). Instead of discussing this similarity, the authors state: "Par3 is likely to regulate centrosome positioning through some intermediate molecules or mechanisms, but its specific mechanism is still unclear and requires further investigation." Given the acetylated tubulin signal emanating from the Par3 positive patch in Figure 5E and F, I suspect similar mechanisms to the C. elegans intestine are at play here. Such a parallel should be noted in the Discussion.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we have expanded the Discussion section to compare our findings with epithelial polarization mechanisms described in in vivo systems, including the C. elegans intestine and other tubular epithelial tissues (Page 13, Paragraph 2).

      We agree that the hierarchical relationship we observe between Par3 localization, centrosome positioning, and apical membrane formation bears important conceptual similarities to mechanisms reported in the C. elegans intestine (Feldman & Priess, 2012). We also considered the possibility that Par3 may regulate centrosome positioning through dynein-dependent mechanisms.

      To examine this possibility, we performed immunofluorescence staining for the dynein cofactor dynactin subunit p150<sup>Glued</sup>. However, we did not observe enrichment of p150<sup>Glued</sup> at the center of cell doublets during the cytokinetic pre-abscission stage (see Author response image 5), suggesting that dynein is not strongly concentrated together with Par3 near the AMIS under our conditions.

      In addition, pharmacological inhibition of dynein resulted in cytokinesis failure and the formation of binucleated cells, preventing reliable assessment of centrosome migration and polarity establishment.

      We would also like to clarify that the acetylated tubulin signal observed in Figure 5E and F does not emanate from the Par3-positive patch. Rather, this signal corresponds to the cytokinetic bridge, adjacent to which Par3 accumulates during cytokinesis. Consistent with this interpretation, γ-tubulin was not detected at the Par3-positive region (Figure 1A).

      Author response image 5.

      Single MDCK cells after 12 h of culture in Matrigel. Immunostaining signals of the indicated markers are shown: centrosome marker PACT-mKO1, dynactin subunit p150Glued, Gp135, and DAPI. Single confocal sections through the middle of a cyst are shown. The order of polarization is arranged from single cell (1-cell), metaphase (Meta), telophase (Telo), cytokinetic pre-abscission (Pre-Abs), post-cytokinesis (Post-CK), to lumen open (LO). Scale bar: 5 μm.

      I had originally commented that "I find the results in Figure 6G puzzling. Why is ECM signaling required for Gp135 recruitment to the centrosome. Could the authors discuss what this means?" The authors responded that "The data in Figure 6G do not indicate that ECM signaling is required for the recruitment of Gp135 to the centrosome". In Figure 6G, the localization of GP135 to the centrosome appears significantly delayed compared to its localization to the centrosome in images where cells were cultured in Matrigel.

      Indeed, the authors argue that the centrosomal localization precedes and contributes to its localization to the AMIS. In the absence of ECM, GP135 localizes to the membrane before it localizes to the centrosome and its localization to the centrosome appears significantly reduced. Thus, my original and current interpretation is that ECM signaling is somehow required for the centrosomal targeting of GP135. One could make a competition argument, i.e. that the cortex in the absence of ECM is somehow a more desirable place to localize than the centrosome, but this experiment also argues that the centrosome does not need to be a source of this material in order for it to end up on the cortex.

      We agree that the absence of ECM substantially alters the trafficking behavior of Gp135.

      Our interpretation is that ECM primarily promotes the endocytosis and internal trafficking of Gp135, thereby enabling its redistribution to membrane domains lacking ECM contact and facilitating AMIS formation (Buckley & St Johnston, 2022; O'Brien et al., 2001; Yu et al., 2005).

      Under ECM-free conditions, a larger fraction of Gp135 remains associated with the plasma membrane, resulting in reduced internalized Gp135 available for centrosome-associated trafficking. During anaphase to telophase (Figure 6G, 0:05–0:20), Gp135 predominantly redistributes along the plasma membrane toward the cleavage furrow. Only after cytokinesis initiation (Figure 6G, 0:30) do we observe a small amount of internalized Gp135 associated with centrosomes near the center of the cell doublet.

      Importantly, we agree with the reviewer that these findings suggest centrosomal trafficking is not absolutely required for Gp135 to localize to the plasma membrane. Rather, our data support a model in which centrosome-associated trafficking contributes specifically to the efficient and spatially restricted delivery of Gp135 to the AMIS during epithelial polarization.

      We have revised the manuscript to clarify this interpretation and to avoid overstating the role of centrosome-associated Gp135 trafficking.

      (3) There needs to be precision in the language used in many places:

      I don't understand this line in the abstract: "When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with mitosis." If a cell has divided, it is no longer a single cell.

      We have revised the sentence to: “When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with cytokinesis (Page 1, Paragraph 1).” This indicates that polarization happens as the cell divides. We thank the reviewer for this helpful suggestion, which has improved the readability of the sentence.

      The authors state in the Introduction "Because of its strong ability to nucleate microtubules, the centrosome functions as the primary microtubule organizing center", but then state ""In polarized epithelial cells, the centrosome is localized at the apical region during interphase, which contributes to the construction of an asymmetric microtubule network conducive to polarized vesicle trafficking". In the latter statement, I assume the authors are describing the well-characterized apical microtubule network in epithelial cells that is non-centrosomal. Thus, the latter sentence is at odds with the former.

      We did not intend to refer to the apical non-centrosomal microtubule network present in mature, fully polarized epithelial cells. Rather, we were referring to the off-center centrosome functions as an off-center MTOC, creating an asymmetric microtubule network during the early stages of epithelial polarization, as mentioned in previous review papers (Meiring, Shneyer, & Akhmanova, 2020).

      The apical non-centrosomal microtubule network is a feature of mature, fully polarized epithelial cells. In fact, its formation is also driven by the release of microtubule minus-ends from the off-centre centrosome, which are then transferred to the apical membrane (Goldspink et al., 2017; Moss et al., 2007; Sanchez & Feldman, 2017).

      The authors continually refer to Par3 as a tight junction protein. "Par3, which controls tight junction assembly to partition the apical surface from the basolateral surface". To my knowledge, PARD3 is an apical protein with similar localization to C. elegans PAR-3 and Drosophila Bazooka. PARD3B is a junctional protein. I assume that the antibody that the authors are using is to PARD3 and not PARD3B? Can the authors please clarify this in the text?

      The antibody used for PARD3 staining was the Merck Millipore rabbit polyclonal antibody (Cat. No. 07-330), generated against a GST-tagged recombinant fragment corresponding to 288 amino acids from the internal region of mouse PAR-3.

      Canine PARD3 and PARD3B are encoded by distinct genes located on chromosomes 2 and 37, respectively. In our study, we used two independent shRNA constructs specifically targeting canine PARD3, both of which reduced the immunoblot signal detected by this antibody, supporting the conclusion that the antibody primarily recognizes PARD3 rather than PARD3B.

      However, in immunofluorescence staining, the signal was mainly localized at tight junctions, and we did not observe significant signal at the apical membrane. We will further clarify in the revised manuscript that this antibody targets PARD3 rather than PARD3B (Page 10, Paragraph 1).

      Reference

      Buckley, C. E., & St Johnston, D. (2022). Apical-basal polarity and the control of epithelial form and function. Nat Rev Mol Cell Biol, 23(8), 559–577. doi:10.1038/s41580-022-00465-y

      Chen, F., Wu, J., Iwanski, M. K., Jurriens, D., Sandron, A., Pasolli, M., . . . Akhmanova, A. (2022). Self-assembly of pericentriolar material in interphase cells lacking centrioles. Elife, 11. doi:10.7554/eLife.77892

      Feldman, J. L., & Priess, J. R. (2012). A role for the centrosome and PAR-3 in the hand-off of MTOC function during epithelial polarization. Curr Biol, 22(7), 575–582. doi:10.1016/j.cub.2012.02.044

      Gavilan, M. P., Gandolfo, P., Balestra, F. R., Arias, F., Bornens, M., & Rios, R. M. (2018). The dual role of the centrosome in organizing the microtubule network in interphase. EMBO Rep, 19(11). doi:10.15252/embr.201845942

      Goldspink, D. A., Rookyard, C., Tyrrell, B. J., Gadsby, J., Perkins, J., Lund, E. K., . . . Mogensen, M. M. (2017). Ninein is essential for apico-basal microtubule formation and CLIP-170 facilitates its redeployment to noncentrosomal microtubule organizing centres. Open Biol, 7(2). doi:10.1098/rsob.160274

      Hung, H. F., Hehnly, H., & Doxsey, S. (2016). The Mother Centriole Appendage Protein Cenexin Modulates Lumen Formation through Spindle Orientation. Curr Biol, 26(6), 793–801. doi:10.1016/j.cub.2016.01.025

      Ibi, M., Zou, P., Inoko, A., Shiromizu, T., Matsuyama, M., Hayashi, Y., . . . Inagaki, M. (2011). Trichoplein controls microtubule anchoring at the centrosome by binding to Odf2 and ninein. J Cell Sci, 124(Pt 6), 857–864. doi:10.1242/jcs.075705

      Martin, M., Veloso, A., Wu, J., Katrukha, E. A., & Akhmanova, A. (2018). Control of endothelial cell polarity and sprouting angiogenesis by non-centrosomal microtubules. Elife, 7. doi:10.7554/eLife.33864

      Meiring, J. C. M., Shneyer, B. I., & Akhmanova, A. (2020). Generation and regulation of microtubule network asymmetry to drive cell polarity. Curr Opin Cell Biol, 62, 86–95. doi:10.1016/j.ceb.2019.10.004

      Moss, D. K., Bellett, G., Carter, J. M., Liovic, M., Keynton, J., Prescott, A. R., . . . Mogensen, M. M. (2007). Ninein is released from the centrosome and moves bi-directionally along microtubules. J Cell Sci, 120(Pt 17), 3064–3074. doi:10.1242/jcs.010322

      O'Brien, L. E., Jou, T. S., Pollack, A. L., Zhang, Q., Hansen, S. H., Yurchenco, P., & Mostov, K. E. (2001). Rac1 orientates epithelial apical polarity through effects on basolateral laminin assembly. Nat Cell Biol, 3(9), 831–838. doi:10.1038/ncb0901-831

      Sanchez, A. D., & Feldman, J. L. (2017). Microtubule-organizing centers: from the centrosome to noncentrosomal sites. Curr Opin Cell Biol, 44, 93–101. doi:10.1016/j.ceb.2016.09.003

      Tateishi, K., Yamazaki, Y., Nishida, T., Watanabe, S., Kunimoto, K., Ishikawa, H., & Tsukita, S. (2013). Two appendages homologous between basal bodies and centrioles are formed using distinct Odf2 domains. J Cell Biol, 203(3), 417–425. doi:10.1083/jcb.201303071

      Vinopal, S., Dupraz, S., Alfadil, E., Pietralla, T., Bendre, S., Stiess, M., . . . Bradke, F. (2023). Centrosomal microtubule nucleation regulates radial migration of projection neurons independently of polarization in the developing brain. Neuron, 111(8), 1241–1263 e1216. doi:10.1016/j.neuron.2023.01.020

      Wu, J., de Heus, C., Liu, Q., Bouchet, B. P., Noordstra, I., Jiang, K., . . . Akhmanova, A. (2016). Molecular Pathway of Microtubule Organization at the Golgi Apparatus. Dev Cell, 39(1), 44–60. doi:10.1016/j.devcel.2016.08.009

      Yu, W., Datta, A., Leroy, P., O'Brien, L. E., Mak, G., Jou, T. S., . . . Zegers, M. M. (2005). Beta1-integrin orients epithelial polarity via Rac1 and laminin. Mol Biol Cell, 16(2), 433–445. doi:10.1091/mbc.e04-05-0435

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a semi-naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade-off between reward seeking and threat avoidance. By combining behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethological setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit, specifically in dominant mice. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. However, it is worth noting that the authors use an extremely low contrast for the low threat condition (20%), which may to some extent be insufficient to reliably trigger escape responses. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      We agree that the low-contrast condition (20%) represents a relatively weak threat signal by design. This level was intentionally chosen to fall within a regime that biases behavior toward freezing and escape after assessment rather than direct escape, allowing us to examine graded decision-making across threat intensities. The higher escape probability observed under low-contrast conditions relative to previous studies (Evans et al., 2018; Fratzl et al., 2021) is likely attributable to the linear arena used here, which generally promoted escape behavior over freezing.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships sustain elevated vigilance should be interpreted carefully, as it does not fully align with broader views of social stability as protective against anxiety and stress and generally beneficial for mental health and resilience. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for this important comment and agree that findings obtained in male mice may not necessarily generalize to female mice. We have acknowledged the limitation in the Discussion and note that future studies will be required to determine the extent to which the present findings extend to female mice.

      We would also like to clarify our interpretation regarding vigilance and social buffering. Vigilance should not be conflated with stress or anxiety, and the slower habituation observed in group-housed mice reflects sustained sensitivity to repeated threat exposure rather than elevated anxiety. Thus, our data do not contradict the concept of social buffering; rather, they are consistent with it, as pair-housed mice exhibited reduced defensive responses compared with individually housed animals (Lenz et al., 2022). Furthermore, in an independent study (Li, Gao, and Li, 2026; eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and found clear evidence of social buffering. These findings suggest that social interactions can attenuate defensive responses while prolonging vigilance during repeated threat exposure. We have revised the Discussion accordingly.

      Another important limitation is that the neural mechanisms underlying these effects remain highly speculative. Although the manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, these interpretations go far beyond the data presented in the study and are not directly supported by experimental evidence within the paper itself. The discussion gives substantial weight to potential circuit mechanisms based primarily on previous literature rather than on findings from the current study. Given the complexity and distributed nature of the circuits likely involved in integrating vigilance, reward, social context, and defensive behavior, the present work is better viewed as providing a strong behavioral framework rather than direct mechanistic insight into the underlying neural substrates. In this context, some references discussing how animals learn to suppress defensive responses to repeated looming threats and the neural mechanisms supporting this process could further strengthen the discussion (Salay et al 2021; Fratzl et al. 2021; Conway et al. 2025; Mederos et al. 2025).

      We agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation and have included a discussion of learning-dependent suppression of defensive responses to repeated looming threats and the underlying circuit mechanisms.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions. At the same time, some statements in the discussion slightly overstate the novelty of the methodological approach. For example, the claim that the study differs from earlier work by using machine learning rather than manual annotation overlooks that several previous studies have already implemented automated or semi-automated strategies to classify looming evoked defensive behaviors beyond manual scoring alone.

      We have revised the manuscript to moderate our claims regarding the novelty of the machine-learning–based classification approach.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to minimize differences in internal state across conditions, whereas parts of the methods describe experiments in which animals were water deprived. This distinction is not always clearly explained across the different experimental sections, despite internal state being central to the interpretation of the behavioral findings. A clearer separation and description of these conditions would further strengthen confidence in the work. In addition, it was somewhat surprising that the low contrast (20%) looming condition was still sufficient to trigger robust escape responses, and additional clarification or discussion regarding stimulus saliency at this contrast level could help readers better contextualize these findings.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in semi-naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      - Clarify which experiments involved water deprivation and which did not in methods section.

      We have revised the Methods section to clearly specify which experiments involved water deprivation and which were conducted without water deprivation.

      - Add discussion on why the 20% low contrast stimulus is still sufficient to trigger escape responses.

      We have expanded the Discussion regarding the 20% contrast stimulus to clarify why it was sufficient to elicit escape responses.

      - Tone down statements regarding the novelty of the machine learning based behavioral classification.

      We have revised the manuscript to moderate claims regarding the novelty of the machine-learning–based classification approach.

      - Include references related to learning-dependent suppression of looming responses and circuit mechanisms

      We have incorporated the suggested references and expanded the Discussion of learning-dependent suppression of looming responses and the underlying circuit mechanisms.

      - Moderate the discussion of candidate neural circuits, particularly the SC related interpretations, as these are not directly tested in the study.

      We have revised the discussion of candidate circuit mechanisms.

      - Clarify that the interpretation linking stable social relationships to elevated vigilance may be specific to this ethological context.

      We have revised the Discussion to clarify this interpretation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Garcia-Alcala, Kratz and Cluzel investigate to what extent our understanding of bacterial physiology in bulk experiments can be applied to single-cell observations. They find that intrinsic noise may be powerful enough to even inverse the trends found in the bulk. The authors hypothesize that the asymmetric distribution of ribosomes to daughter cells during cell division plays the dominant role in the intrinsic noise and is able to generate the observed phenomenon. They do not show it directly, but the data and its agreement with the model are sufficient to support this claim.

      Strengths:

      The experimental part is convincing: the positive correlation between the elongation rate and promoter activity of unnecessary protein is clear, as well as the negative correlation between the mean values while changing the promoter strength. This was demonstrated in both rich and poor media. The causality between the growth rate and the promoter activity was shown using the negative lag time of the cross-correlation function. A simple, reasonable model accounts well for the data. This paper demonstrates an interesting phenomenon and provides a plausible theory for it, advancing our understanding of bacterial physiology on the single-cell level.

      Weaknesses:

      (1) Mean-reversion timescales were assumed to be longer than the simulation time and much longer than the cell cycle time. It is not clear whether the results are robust in case mean reversion timescales become of the order of the cell-cycle or smaller. If not, is there an argument for such practically infinite reversion timescales?

      Due to an error in the simulation code, the incorrect mean-reversion timescales were reported in the manuscript and have now been corrected. Instead of 1000 h and 100 h for 𝜏<sub>R</sub> and 𝜏<sub>U</sub>, respectively, they are 6.25 h and 0.0625 h, which are both much shorter than the total simulation time (75 h). The error had no effect on simulation results and was purely a timescale conversion mistake. We appreciate the referee’s careful review of the manuscript which allowed us to catch this error.

      Given the correct values, 𝜏<sub>U</sub> is well within the cell-cycle time, as is expected since we assume the source of noise in unnecessary protein expression is from stochastic gene expression. In contrast, 𝜏<sub>U</sub> is still significantly longer than the cell-cycle time and needs to be in order to generate the experimentally observed behavior (i.e., the increase in growth rate with an increase in unnecessary protein expression at the single-cell level). This is consistent with our biological hypothesis that some cells inherit ribosomal surpluses from their mothers which enable bursts in protein production. 𝜏<sub>U</sub> less than the cell-cycle time would correspond to a case where ribosomal composition quickly decays to the population average, meaning that daughter cells would never have time to capitalize on the benefit of receiving a ribosome surplus. Furthermore, 𝜏<sub>U</sub> being longer than the cell-cycle is biologically justifiable as proteins such as ribosomes are passed from mother to daughter at division, thus allowing for memory to persist over longer timescales than a single generation.

      (2) It is not easy to understand the simulation part unless one reads Ref [14]. k(t) is assumed Equation (1) from Reference [14]? Is it crucial that the ribosome noise appears only at the division? The ribosome noise strength σ<sub>R</sub> =0.06 - is it lower or higher than the naively expected binomial division? Also, a more intuitive explanation of the Simpson paradox would help the reader.

      𝜅(𝑡) is indeed Eq. (1) from Ref. [14]. To make the computational results clearer, the methods section has been updated to include the full set of equations used to perform the simulations, and the code used to produce the computational figures is now on GitHub. The relative standard deviation expected by modeling ribosome distribution at division by a binomial distribution with equal probability of being inherited by either daughter cell is in the range of 1-3% (0.01-0.03), assuming N~10<sup>3</sup>-10<sup>4</sup> ribosomes. Thus, 0.06 is reasonable as it is in the same order of magnitude as what is predicted by binomial division. Furthermore, if clustering is present as already demonstrated in [39, 40, 43], we would expect the relative standard deviation to increase as clustering reduces the effective number of proteins which are distributed between daughters.

      (3) It would be useful for the reader to see the raw data and not only the filtered one to appreciate the measurement noise level.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the direct measurements and the corresponding smoothed traces for the flagellar reporter strains.

      (4) Negative lag time of the cross-correlation function is visible, but consider adding a statistical test for it.

      We have included a statistical analysis of the lags of maximum correlation of elongation rate and activity for the flagellar reporter strains in Fig. S10. For each strain, we now show an overlay of the cross-correlation functions for all lineages and the distribution of the lag corresponding to the maximum correlation. In addition, we have added a section in Materials and Methods describing in detail how the cross-correlation and the strain-averaged lag were computed.

      (5) Can you make similar cross-correlation plots using the model? Can you infer by using it, whether the data agrees better with the assumption that ribosomal noise appears only at division or continuous fluctuations during the cell cycle?

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one process of protein production and thus lacks any delay or memory mechanisms which would make a significant positive or negative correlation.

      Both continuous fluctuations and noise from division could in principle contribute to our observation of Simpson’s paradox. We are unable to use the model in its present form to dissect the contributions from both mechanisms. However, our model simulations clearly show that noise at division by itself is sufficient to explain the observed effect with parameter values that are biologically plausible. As Chao et al., [40] has previously shown experimentally that a significant source of ribosomal noise comes from unequal distribution at division, which is the hypothesis we retained for the model.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Garcia-Alcala et al. reports an interesting paradox: the cost of gene expression slows the population-average growth rate, whereas at the single-cell level, expression levels from these genes positively correlate with the growth rate. The effect is observed in the expression of flagellar genes and a gene under a synthetic promoter in E. coli. The findings are explained by the inheritance of growth factors, including ribosomes, during asymmetric division.

      Strengths:

      (1) The manuscript adds strength to an emerging body of literature showing that the population-level bacterial growth laws do not match correlations based on single-cell data. The evidence presented here is more striking than in previous works (such as Pavlou et al., Nat. Commun. 2025), as the trends in population-level data and single-cell data are reversed.

      (2) A relatively simple model correctly explains the trends in the data.

      Weaknesses:

      (1) It is not clear whether flagellar proteins are expressed proportionally to the reporter signal. Furthermore, it is questionable if E. coli bacteria in the mother machine channels are flagellated. If they are, they could potentially swim out of the channels, which is not the case when they do not carry the MotA E98K mutation. The authors should provide some evidence that E. coli expresses the actual filament proteins in the channels.

      We agree that it is important to demonstrate that our reporter reflects the production of functional flagellar structures under our experimental conditions. To this end, we first tested the swimming capabilities of our strains before introducing the MotA E98K mutation, using a standard soft‑agar motility assay. For strains in which only the Class‑1 promoter was modified (Pro2, Pro4, Pro5), after 12 h of inoculation, the diameters of the swarming rings followed the expected order based on promoter strength: Pro2 showed the smallest ring, followed by Pro4, WT, and Pro5. In contrast, control strains carrying Pro4 together with either MotA E98K (non‑rotating motors) or ΔfliC (no flagellin filament) did not form rings, consistent with their inability to swim. We also included an MG1655 strain carrying an IS5 insertion upstream of the Class‑1 promoter, which is known to enhance flagellar expression [56]; this strain showed a larger ring, as expected. A qualitative summary of ring diameters for all strains is provided in Author response table 1. We have included the figure and section “Experimental validation of functional flagella expression in reporter strains” on Supplementary Material.

      Author response table 1.

      We also tested whether cells assemble functional flagella on the mother‑machine. We compared MG1655 WT and MG1655 carrying the MotA E98K mutation under identical microfluidic growth conditions. Many WT cells left the channels during the experiment (∼35% of lineages over ~20 h after the onset of exponential growth inside the device), consistent with active swimming, whereas the non‑motile MotA E98K strain did not leave the channels. Because the only difference between these two strains is the MotA E98K mutation, which disables motor rotation but not flagellar assembly, this result indicates that (i) cells do express functional filaments and motors in the mother‑machine environment, and (ii) the strain used for our main experiments is immobilized by MotA E98K.

      Both experiments are included in the Supplementary Information section “Experimental validation of functional flagella expression in reporter strains” and Fig. S3.

      (2) It is unclear what fraction of the total proteome mVenus represents in different measurements. Some quantification is needed (for example, using the Coomassie staining). Using f_U as high as 14.4% in simulations is questionable.

      We agree that we do not currently know the exact fraction of the proteome occupied by mVenus in our experiments, as we only quantified fluorescence and did not perform Coomassie staining or proteomics. The primary goal of our simulations is not to reproduce exact experimental conditions, but to illustrate how unequal ribosome partitioning affects daughter cells across a range of protein synthesis burdens. Consequently, our use of values up to F<sub>U</sub> = 14.4% in the simulations was intended as an exploratory upper range rather than as a direct estimate of the experimental condition.

      To put our experimental burden in context, we compare our growth-rate reduction to a well– characterized high-burden case in the literature. In the study by T. Hwa’s group [6], overexpression of β–galactosidase such that it constituted approximately 27% of the total proteome led to a 67% reduction in growth rate for E. coli growing in a medium similar to ours (differing only in the carbon source: glucose in their case, glycerol in ours). In our system, overexpression of mVenus from a plasmid result in a substantially smaller growth–rate reduction of 9% relative to the non–expressing control, and this is observed on the poorer carbon source (glycerol), under which burden effects are typically less pronounced [6].

      Although we cannot convert our fluorescence measurements into an exact proteome fraction, the much smaller growth defect compared to the 27% β‑galactosidase case strongly suggests that mVenus does not approach such an extreme fraction of the proteome under our conditions. Under the simplifying assumption that the qualitative relationship between unnecessary‑protein fraction and growth‑rate reduction is similar in the two systems, our data are therefore consistent with a modest fraction of unnecessary protein and make it unlikely that mVenus reaches very high fractions such as 27%. This gives us confidence that exploring F<sub>U</sub> values up to 14.4% in the simulations represents a conservative upper range relative to our experimental burden, rather than an underestimate.

      (3) The data from the MC4100 strain does not directly match the trends of MG1655. The justification for filtering out the low-frequency components of MC4100 is not particularly convincing. It appears unlikely that ribosomes or other growth factors partition significantly differently in the MC4100 strain than in the MG1655 strain. Further discussion and a plot similar to Figure 1 (Left) for this strain are warranted.

      We thank the reviewer for this helpful suggestion. We agree that the filtering analysis from the initial manuscript was not clear enough to be used as robust supplementary information, and we therefore removed it. Instead, we carried out with the ‘unfiltered’ MC4100 strain, the same single-cell analyses as with MG1655, including the binning analysis of instantaneous elongation rate versus Class-2 promoter activity, and a cross-correlation analysis between such measurements.

      Both analyses did not reveal a significant correlation between growth and flagellar promoter activity in MC4100. We now present these results in the revised manuscript (Fig. S15), where we explicitly show the lack of association in a plot directly comparable to Fig. 1.

      While ribosomes partition mechanism is likely to be the same between these two strains, MC4100 is known to exhibit very long oscillations of growth rate over 10 generations, which are absent in MG1655.

      We now present MC4100 as an explicit counterexample to highlight that the positive correlation between short timescale growth fluctuations and flagellar expression observed in MG1655 is not universal across all E. coli strains, especially in strains like MC4100 whose growth rate fluctuations are dominated by long timescales, much longer than the division time. In MC4100 these slow modes are largely decoupled from flagellar gene regulation. We have also revised the text to clarify this point.

      (4) The model needs to be described in more detail. A closed set of equations that have been simulated must be presented, along with all values of the model parameters and their sources. The authors should consider depositing their code on GitHub or another publicly accessible repository.

      The methods section has now been updated to include the full set of equations used to perform the simulations along with all parameter values and their sources. Additionally, the code used to produce the computational figures is now on GitHub. For a full derivation and biological justification of each model component, we still refer readers to ref [14] where this model was first published.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The first paragraph of Results still belongs to the Introduction Section, I feel.

      Thank you for the recommendation, we agreed with the change.

      (2) Figure 1, left, is after some filter (Savitzky-Golay)? It might be useful to see the raw points.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the raw (direct) measurements and the corresponding smoothed traces for the flagellar reporter strains.

      Reviewer #2 (Recommendations for the authors):

      (1) Is there a statistically significant positive slope in single-cell data of the elongation rate as a function of Class-2 activity? There should be an analysis of statistical significance for the slopes in Figure 1 (Right) and Figure 4A, including both population-average and single-cell data.

      Thank you for your recommendation. We have now added statistical analyses of the correlations between activity and elongation rate at both the single-cell and population levels. The results for the flagellar strains are presented in Fig. S5, S6, and S7. Also, the corresponding analysis for constitutive Venus expression is shown in Fig. S16.

      (2) How does FlhC-YFP data compare with the CFP signal in the first set of measurements? What would Figure 1, Left look like for this signal?

      At the single‑cell level, the relationship between Class‑1 promoter activity and elongation rate within each strain is still positive, but clearly weaker than for Class‑2, as shown in Fig. S8. When we bin the data, we can observe that the binned averages don’t display a clear positive trend, even though the Pearson correlation for each flagellar reporter strain is still positive. We think this is because Class‑1 controls a much smaller part of the proteome (it only encodes the two subunits of FlhDC) while Class‑2 promoters drive many structural and export proteins. So, changes in Class‑2 activity more directly reflect shifts in global translational capacity and are more tightly linked to growth, whereas Class‑1 activity adds only a small translational load.

      (3) The information from the ER-Activity cross-correlation functions is interesting but has not been interpreted or compared with the model. Which signal precedes the other? What can explain the observed lag time on the order of Tdiv? Why do some cross-correlations show negative values in Fig. S9 while others are positive (as they should)? Can the model explain the experimentally observed cross-correlation function?

      In our cross‑correlation analysis, a negative lag at the maximum means that fluctuations in elongation rate (ER) precede fluctuations in promoter activity (A) by that lag. We now describe this explicitly in Materials and Methods and quantify lag distributions for all reporter strains in Fig. S10.

      In Fig. 13 (previously Fig. 9), we mainly observe positive correlations between ER and PA with negative or near zero peak lags, i.e., ER tends to lead PA by a lag by about a division time. The reason governing the negative lags is not immediately clear, but it is consistent with a resource driven mechanism: cells that inherit more growth factors (e.g. ribosomes) at division, use them to prioritize housekeeping processes and thus increase growth rate first, and only subsequently use these extra resources to increase flagellar promoter activity.

      As for Class–1 activity, the correlation with growth is much weaker as it was already observed in Kim et al. Averaging the cross–correlations across all lineages yields a modest peak (mean correlation 0.10, SD 0.12; see Author response image 1) at a small positive lag of 0.5 h (while average division time is T<sub>div</sub> ~ 1.7h). This indicates a weak but real positive correlation between Class 1 activity and ER at short lags. However, the lags of the individual maxima are widely distributed (SD ≈ 9 h; mean −1.8 h, median −0.28 h, mode ≈ 0), with only ~53% of the cells showing a negative time lag. Thus, delays are roughly symmetrically spread around zero with only a slight negative bias.

      Author response image 1.

      Left: cross−correlation between elongation rate and Class−1 activity for each lineage (N = 96, colored lines), and their average as a function of lag (black line). Right: distribution of the lags at which each lineage’s cross−correlation attains its maximum. The dashed line indicates zero lag, and the full line marks the lag of the peak of the mean cross correlation.

      An in-depth analysis about the sign and magnitude of the lag would require measurements from a broader set of promoters. The lag is likely to depend on what genes the promoter controls (e.g. stress response, housekeeping, or large structural modules), on its strength and regulation, and on growth conditions.

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one single process of protein production and thus lacks any delay or memory mechanism that could generate a phase shift.

      Nonetheless, this memory-free formulation shows that stochastic redistribution of growth factors at division is by itself sufficient to generate the Simpson’s paradox behavior. Capturing the experimentally observed lag time would require including explicit transcription/translation delays and additional regulatory dynamics between housekeeping and flagellar genes whose information we do not have.

      (4) Figure 3, Left - it is not clear what this plot shows. Red and blue are scattered over the whole plot. How have daughter 1 and daughter 2 been assigned? Perhaps choosing one of the daughters with a higher growth rate and then plotting the data could reveal some trends.

      We agree that the original left panel of Fig. 3 was difficult to follow. In the original version, the daughter labels were assigned as “Daughter–1” for the cell at the closed end of the channel and “Daughter–2” for the cell closest to the open side. After discussing this with Dr Camilla Ulla Rang (whom we now acknowledge in the manuscript), we relabeled the daughters as “new–pole daughter” and “old–pole daughter,” following the convention used in studies of aging and ribosome distribution in E. coli.

      We now show the class–2 promoter activity comparison between new pole vs old pole, and in a separate panel, we plot the mean ratios of elongation rate, and class–1/class–2 promoter activities between the new–pole and old–pole daughters. These ratios show that the new–pole daughter tends to have both greater growth rate and flagellar gene activity than the old–pole daughter. Importantly, this new plot is in line with Rang and colleagues who demonstrated that new–pole daughters have greater ribosome density and faster growth rates than their old–pole sisters. Together, these results further support our hypothesis that excesses of ribosomes inherited at division underlies the observed growth boosts.

      (5) Flagellar activity -> activity of flagellar gene synthesis (presumably no flagellar activity in these cells).

      We corrected the terms used to reference the flagellar gene activity.

      (6) Page 7: "The strain MC4100, known to exhibit slow, long period oscillations in growth ..." - some reference is- needed here.

      We placed the reference some lines after, as such paper also includes the information of MG1655 short-term oscillations. Now the reference is [44]: Tanouchi, Y., et al., A noisy linear map underlies oscillations in cell size and gene expression in bacteria. Nature, 2015

      (7) Page 7: Figure S11B - is Figure S11C perhaps meant?

      The correct panel indeed was Fig. S11C, and we have now corrected and updated the figure label and text accordingly.

      (8) Page 11: Savitsky-Golay filter.

      We have corrected it, thank you.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study utilizes polarized second-harmonic generation (pSHG) microscopy to investigate myosin conformation in the relaxed state, distinguishing between the disordered, actin-accessible ON state and the ordered, energy-conserving OFF state. By pharmacologically modulating the ON/OFF equilibrium with a myosin activator (2deoxyATP) and inhibitor (Mavacamten), the authors demonstrate that pSHG can sensitively quantify the ON/OFF ratio in both skeletal and cardiac muscle. Validation with X-ray diffraction supports the accuracy of the method. Applying this approach to a hypertrophic cardiomyopathy model, the study shows that R403Q/MYH7-mutated minipigs exhibit an increased ON state fraction relative to controls. This difference is eliminated under saturating concentrations of myosin modulators, indicating that the ON/OFF balance can be pharmacologically shifted to its extremes. Additionally, ATPase assays reveal elevated resting ATPase activity in R403Q samples, which persists even when the ON state is saturated, suggesting that increased energy consumption in this mutation is driven by both a shift toward the ON state and inherently higher myosin ATPase activity.

      Strengths:

      This is a well-written and well-conducted study that clearly reveals the power of SHG microscopy. The study clearly establishes the great utility of SHG to study thick filament regulation.

      Weaknesses:

      (1) Several studies have shown that the ON state of the thick filament is sensitive to both temperature and filament lattice spacing, with a common recommendation to conduct skinned fiber experiments at temperatures above 27 °C and in the presence of dextran to better preserve physiological conditions. The authors should clarify the experimental temperature used in their skinned fiber studies, indicate whether dextran was included, and discuss whether adherence to these recommended conditions would have impacted their results.

      Additional experiments were performed on skinned psoas muscle strips to assess the influence of temperature and lattice spacing on the measured γ values. In particular, psoas strips were tested under relaxing conditions at three different temperatures (10 °C, 15 °C, and 20 °C). In a separate set of experiments, psoas strips were examined under relaxing conditions in the absence and in the presence of 5% dextran.

      These measurements confirmed previous reports indicating that the ON state of the thick filament is sensitive to both temperature and lattice compression. Specifically, increasing the temperature from 10 °C to 20 °C resulted in a shift toward lower γ values, consistent with a greater radial displacement of myosin heads. A similar trend, although less pronounced, was observed in the presence of dextran, which also produced a decrease in γ compared with control conditions.

      Figure 1 and text have been modified including these new findings (revised manuscript, pages 11-12).

      (2) On page 13, the authors report the proportion of disordered heads as approximately 30% in wild-type and 65% in R403Q fibers. They should clarify whether these values represent the percentage of total myosin heads, or rather the percentage of heads that are responsive to Mavacamten and dATP.

      The proportions reported in the original version referred to the fraction of myosin heads assumed to be responsive to Mavacamten and dATP. These estimates were obtained using a simplified model in which dATP was assumed to produce a complete (100%) shift of myosin heads toward the ON state, whereas Mavacamten was assumed to cause a complete depletion (0%).

      Since these assumptions are not fully supported by structural data, we decided to remove these values from the Results section and instead discuss these considerations in more general terms in the Discussion (page 19). Figure 3 and text have been modified accordingly (pages 14-15).

      (3) In Figure 5, regarding ATPase measurements, the content of contractile material per unit volume of muscle preparation will influence the results. Did the authors account for this variable, and if not, how might it have affected the conclusions?

      We thank the reviewers for raising this concern and agree that the content of contractile material per unit volume can influence the absolute values obtained in ATPase measurements. In our experiments, however, we primarily assessed the effects of compounds using paired measurements (e.g., the same strip measured before and after Mavacamten treatment).

      In addition, based on our previous structural data obtained from similar preparations of human myocardium (https://doi.org/10.1161/CIRCRESAHA.122.321956), the ratio of contractile tissue volume to total muscle volume is highly preserved. This supports comparisons between muscle strips from different experimental groups (e.g., WT and R403Q).

      (4) For readers primarily interested in assessing the ON/OFF state of thick filaments, could the authors list the specific advantages of polarized second harmonic generation (pSHG) microscopy compared to X-ray diffraction?

      In the present work, we deliberately chose to frame pSHG microscopy and X-ray diffraction as complementary approaches, rather than emphasizing a direct one-to-one comparison of their respective advantages and disadvantages.

      That said, we agree that for readers specifically interested in assessing the ON/OFF state of thick filaments, a clearer delineation of the respective strengths and limitations is useful. We have therefore slightly revised the Discussion (page 21) to better clarify the advantages and constraints of both approaches.

      (5) Given that many data points were derived from the same fiber or myocyte, how did the authors address the risk of type I errors due to non-independence of measurements? Was a nested or hierarchical statistical approach used?

      We thank the reviewer for raising this important point and agree that the original statistical analysis did not fully account for the non-independence of measurements derived from the same fibre.

      To address this issue, we have now reanalyzed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes.

      Importantly, this revised analysis substantially confirms the original results, with no major changes in the observed trends or in the statistical significance of the findings. All figures have been modified accordingly, and the statistical methods used are now described in the Methods section (page 8).

      Reviewer #2 (Public review):

      Summary:

      In striated muscle, myosin motors can dynamically switch between an energyconserving OFF state and an activated ON state. This switching is important for meeting the body's needs under different physiological conditions, and previous studies have shown that disease-causing mutations associated with cardiomyopathies can affect the population of these states, leading to aberrant contractility. Studying these structural states in muscle has previously only been possible via X-ray diffraction, which requires access to a beam line. Here, Arecchi et al. demonstrate that polarized second-harmonic generation microscopy (pSGH), a technique that is more accessible, can be used to probe the ON/OFF states of myosin in both permeabilized and intact muscle.

      Strengths:

      (1) There is an outstanding need in the field to better understand the regulation of the ON/OFF states of myosin. Currently, this is studied using X-ray diffraction, meaning that it is accessible to only a few labs. The authors demonstrate that pSGH can be used to probe the ON/OFF states of myosin both in intact and permeabilized muscle. This is a significant advance, since it makes it possible to study these states in a standard research laboratory.

      (2) The authors demonstrate that this approach can be employed in both skeletal and cardiac muscle. Importantly, it works with both porcine and mouse cardiac muscle, which are two of the most important animal models for preclinical studies.

      (3) The authors manipulate the ON/OFF equilibrium using both drugs and a genetic model of hypertrophic cardiomyopathy that has been shown to modulate the ON/OFF equilibrium. Their results generally agree with previous studies conducted using X-ray diffraction as well as biochemical measurements of myosin autoinhibition.

      Weaknesses:

      (1) While the application of pSGH to the ON/OFF equilibrium is an important advance, there are limited new biological insights since the perturbations used here have been extensively characterized in previous studies.

      We acknowledge that the biological insights provided in this study largely confirm previous findings reported in the literature. However, the primary aim of our work was to demonstrate the applicability and robustness of pSHG in probing the ON/OFF equilibrium of thick filaments. In this context, obtaining results that are consistent with established knowledge represents an important validation of the technique.

      We therefore believe that the combination of these findings with the advantages of pSHG makes our work noteworthy and supports the broader utility of this approach in future biological and physiological studies.

      (2) SGH has previously been applied to study the nucleotide-dependent orientation of myosin motors in the sarcomere (PMID: 20385845). The authors have previously interpreted the value of gamma as being a readout of lever arm position, but here, it is interpreted as a measure of ON/OFF equilibrium. When this technique is applied to intact muscle, it is not clear how to deconvolve the contributions of lever arm angle from the ON/OFF population (especially where there is a mix of states that give rise to the gamma value). This is an important limitation that is not discussed in the manuscript.

      We thank the reviewer for this insightful and important observation. The SHG signal arises from the coherent contribution of peptide bonds within the myosin molecule and therefore reflects an average angular distribution rather than a single structural state. In the study cited (PMID: 20385845), by reconstructing each contributing second-harmonic emitter within the actomyosin motor array at the atomic scale, we were able to establish a direct relationship between myosin conformation and the γ value. In particular, this analysis revealed an overall trend: γ increases as the average angle of myosin relative to the thick filament backbone becomes larger, and decreases as this angle is reduced. Based on these observations, we proposed that γ could serve as a proxy sensitive to changes in the ON/OFF equilibrium.

      However, we fully acknowledge that, especially in intact muscle, it is not possible to disentangle the respective contributions of lever arm orientation and population shifts between different structural states. The reviewer’s point is therefore well taken.

      To address this, we have revised the Discussion (pages 18-19) to more clearly acknowledge this limitation and to better articulate the interpretative framework underlying our analysis.

      Finally, we would like to emphasize that, in the present study, we aimed to minimize the contribution of actomyosin-bound states in intact trabeculae by performing experiments under conditions expected to strongly favor relaxation (low stimulation frequency, 0.1 Hz, and low temperature, 21 °C). Under these conditions, the number of strongly bound actomyosin cross-bridges during the diastolic phase is expected to be minimal, and variations in γ are therefore primarily associated with changes in the relaxed myosin population.

      (3) The R403Q mutation has previously been shown to cause an increase in ATP usage. Here, the authors measure an elevated basal ATPase rate under relaxing conditions, and they interpret this as showing increased myosin ATPase activity intrinsic to the motors; however, care should be used in interpreting these results. Work from the Spudich lab has shown that the R403Q mutation can appear as increasing motor function in some assays but depressing motor function in others (see PMID: 32284968, 26601291). Moreover, the actin-activated ATPase rate is an order of magnitude higher than the basal ATPase rate, and thus, small changes in the basal ATPase rate are unlikely to be important for physiology.

      We thank the reviewer for this thoughtful comment, and for pointing out the complexity of interpreting the functional consequences of the R403Q mutation. We agree that the functional effects of the R403Q mutation on myosin motor activity have been reported to vary depending on the experimental system and assay used, with studies showing both reduced and enhanced motor performance (PMID: 32284968, 26601291, 20560002).

      Consistent with this complexity, previous biochemical and physiological studies have reported altered energetic properties associated with the R403Q mutation in cardiac muscle. In a previous study, the relationship between cross-bridge kinetics and energetics was investigated in single cardiac myofibrils and multicellular cardiac muscle strips from human HCM samples with and without the R403Q mutation. In those experiments, cross-bridge relaxation was faster in R403Q samples and correlated with an increased energetic cost of tension generation. Basal ATPase activity measured in human samples was also elevated (4.4 ± 0.5 vs 6.6 ± 1.2 μmol L<sup>-1</sup> s<sup>-1</sup> in HCMsn and R403Q, respectively), supporting the idea that the mutation affects energetic balance in the relaxed state (PMID: 24928957).

      We also agree that actin-activated ATPase activity is substantially higher than basal ATPase activity. However, cardiac muscle spends a large fraction of the cardiac cycle in the relaxed (diastolic) state, during which myosin heads are predominantly detached from actin. Even relatively small changes in ATP turnover during this phase could therefore influence the overall myocardial energetics, and potentially contribute to the activation of signaling pathways involved in pathological remodelling.

      To address the reviewer’s concern, we have now expanded the Discussion (pages 19-20) to clarify the heterogeneous results reported in the literature for the R403Q mutation, and to more cautiously interpret the physiological implications of the observed resting ATPase activity.

      (4) The authors interpret some of their data based on the assumption that the high concentrations of drugs cause the myosin to either adopt 100% OFF or ON states. This assumption is not validated, limiting the ability to interpret the fraction of myosins in the ON/OFF states.

      We fully agree: these assumptions are not fully supported by structural data. We consistently decided to remove these analysis from the Results section and instead discuss these considerations in more general terms in the Discussion (page 18). Figure 3 and text have been modified accordingly (pages 14-15).

      (5) The ATPase measurements are innovative but hard to interpret. dATP and ATP do not have identical ATPase kinetics, meaning that it is hard to deconvolve whether the elevated ATPase rate with dATP is due to changes in the ON/OFF population and/or intrinsic ATPase activity. Similarly, mavacamten reduces the rate of phosphate release from myosin, and this effect is not strictly coupled to the formation of the OFF state (e.g., see PMID: 40118457). As such, it is difficult to deconvolve drug-based changes in the inherent ATPase kinetics of the myosin from changes in the OFF-state population.

      We thank the reviewer for this important comment. We recognize that dATP and ATP do not exhibit identical ATPase kinetics, and that Mavacamten slows steps in myosin nucleotide release that are well documented for ATP and not for dATP (for both S1; PMID: 28808052 and, more markedly, HMM; PMID: 30018063). At the same time, both compounds have been shown to perturb the regulatory state of myosin heads along the thick filament. These effects are, in turn, mediated by shifts in the distribution of ATPase intermediate states, which alter the likelihood of myosin adopting autoinhibited conformations and thereby biasing the system toward OFF or ON states (see, e.g., PMID: 39444161; PMID: 30018063). This configures a dual mechanism of action for small molecules, which we now describe more clearly in the Discussion (page 20). As suggested by the referee, we also highlight the intrinsic difficulty of disentangling compound-induced changes in the intrinsic ATPase kinetics of myosin from shifts in the population of the OFF state. Consequently, we have now better focused results and discussion section considering mainly differential effects of the drugs on SHG and ATPase measurements.

      Moreover, under the strongly relaxing conditions used in our experiments (pCa 10), the population of actomyosin-bound cross-bridges is expected to be extremely small (<0.1%), thereby minimizing any dATP-mediated activation of myosin arising from enhanced electrostatic interactions with actin (see, e.g., PMID: 31110001). This is supported by the absence of a significant effect of the nucleotide on the resting tension of myofibrils (new Fig. S1).

      We have now better clarified these points in both result and discussion section of the revised manuscript.

      Reviewer #3 (Public review):

      This is a very interesting paper extending the use of SHG to the study of relaxed muscle and its use to assess the order-disorder (and on /off) states of myosin heads in the thick filament. The work convincingly shows that SHG and the parameter gamma provide a reliable measure of the state of the myosin heads in a range of different relaxed muscle fibres, both intact and skinned, and in myofibrils. In mini pig cardiac fibres, the use of dATP and mavacamten increased or decreased the number of heads in the disordered state, respectively. On the assumption that these treatments push myosins fully into the disordered or ordered state, then this allows the fraction of ordered heads to be assessed under a wide variety of conditions. It is unfortunate that dATP treatment was not used (as mavacmten was) on rabbit psoas and mouse samples to further test this hypothesis.

      The results with the myosin mutant R403Q support the idea that this mutation reduces the fraction of myosin heads in the ordered state and that mavacamten can recover the WT situation.

      The results from SHG were compared with parallel studies using X-rays to validate the conclusions. Independent fibre ATPase data further support the conclusions.

      The work is solid and provides a novel approach to assessing the activity state of muscle thick filaments. The authors point out some of the potential uses of this approach in the future, including time-resolved SHG measurements. Indeed, jumps in mavacamten or dATP concentration with time-resolved SHG could measure the rates of entry and exit from the ordered, off state of the filament. A measurement is urgently needed in the field.

      Strengths:

      (1) The SHG signal is convincingly shown to assess the fraction of ordered/disordered myosin heads in the thick filament of a variety of muscle fibres.

      (2) The results are similar for rabbit psoas, mouse, and minipig cardiac fibres. Skinning the fibres and production of myofibrils do not change the SHG signal.

      (3) Use of myosin R403Q mutant in mini pig confirms a loss of ordered myosin heads, and the ordered heads can be recovered by mavacamten.

      (4) Parallel X-ray scattering and ATPase data support the conclusions.

      (5) Assuming that dATP and mavacamten generate 100% disordered vs ordered myosin heads respectively, then the percentage of ordered heads can be calculated for a variety of conditions.

      Weaknesses:

      (1) Issues like the effect of fibre disarray and lattice spacing on the SHG signal are not well defined.

      We thank the reviewer for raising this important point. Regarding fibre disarray, we took advantage of the spatial resolution of SHG imaging in thick samples to selectively analyse regions of the preparation in which myofibrillar organization was preserved. This approach proved particularly useful in samples such as those from HCM, where a pronounced global disarray is present but locally well-organized regions can still be identified and reliably analysed.

      Concerning lattice spacing, we have performed additional experiments on psoas muscle in which lattice spacing was modulated using dextran. This allowed us to directly assess the sensitivity of the technique to changes in inter-filament spacing.

      Figure 1 and the corresponding text have been revised accordingly to incorporate and clarify this point (page 11-12).

      (2) The, now well-defined heterogeneity of thick filament structure is not acknowledged.

      We agree that the heterogeneity of thick filament structure is an important aspect that should be acknowledged. In the revised manuscript, we have now explicitly addressed this point in the Discussion (page 21). In particular, we highlight that the capability of pSHG to probe the ON/OFF state with sub-sarcomere spatial resolution offers the future opportunity to investigate spatial heterogeneity in thick filament organization. This aspect is now clearly acknowledged and discussed in the context of the potential applications of the technique.

      (3) dATP was only used on minipig cardiac fibres. The effect of dATP on rabbit psoas and mouse cardiac fibres would be a useful comparison and would help validate the calculation of % ordered heads.

      We agree that, in the original version of the manuscript, there was a methodological imbalance in the use of dATP across preparations. To address this point, we have performed additional experiments in which the effect of dATP was also evaluated in rabbit psoas and mouse cardiac fibres.

      Importantly, the inclusion of these data has also proven useful in the discussion of the ON/OFF equilibrium across different muscle types and species. The corresponding results and discussion have been added to the revised manuscript (page 18-19).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the points in the Public Review, please also address the following points.

      (1) There are some issues with the calculated ionic strengths of the solution (or some details are missing). For example, on p. 4, it is stated that there is a 200 mM ionic strength solution that contains both 100 mM KCl and 2 mM MgCl2, meaning that the ionic strength is higher than 200 mM. There are other similar issues in other places.

      We thank the reviewer for pointing this out. We agree that there was an error in the reported ionic strength calculations and that some details were not sufficiently clear.

      We have now corrected the ionic strength values throughout the manuscript and revised the Methods section to provide a clearer and more consistent description of the solution composition (pages 6-7).

      (2) Please report standard deviations rather than standard errors.

      We thank the reviewer for this suggestion. Following the recommendations of Reviewer 1, we have reanalyzed the entire dataset using an updated statistical framework. As part of this process, we carefully evaluated the most appropriate measure of variability and have now consistently reported the corresponding error estimator throughout the manuscript.

      (3) For the statistical testing, please mention what tests were done for normalcy. It appears that some data is not normally distributed, in which case nonparametric tests should be used.

      We have now reanalysed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes. As part of this updated statistical framework, normality tests were performed for all datasets. When the assumption of normality was not met, appropriate nonparametric or model-based approaches were used.

      (4) Please discuss sex as a biological variable.

      Sex as a biological variable was not specifically investigated in the present study, and we agree that this represents a limitation. This point has now been explicitly acknowledged in the revised manuscript (page 5).

      (5) Please add an explicit section on limitations.

      We thank the reviewer for this suggestion. In the revised manuscript, the Discussion has been expanded to more clearly highlight the limitations of the technique and to better contextualize them in comparison with other approaches (page 21). We believe that integrating these aspects within the Discussion provides a more coherent and balanced presentation, and we have therefore chosen not to include a separate, dedicated limitations section.

      (6) Please discuss the limitations of using saturating concentrations of the drug. For example, the sensitivity of this method to detect changes in ON/OFF equilibrium at physiological concentrations is likely lower.

      The Discussion has been implemented to highlight that the sensitivity of this technique to the ON/OFF ratio could be further explored across species by employing a range of concentrations of Mavacamten and dATP, thereby better capturing physiologically relevant conditions (pages 18-19).

      (7) It is stated on p. 13 that R403Q has a higher sensitivity to mava versus dATP. I'm not sure this is supported by the data. The R403Q starts at a higher percentage of ON, and thus the effect size will be larger with mava, but this doesn't imply anything about sensitivity (which implies concentration dependence).

      Our statement was not intended to imply a difference in sensitivity in terms of concentration dependence, but was instead based on the statistical outcome of our measurements. Specifically, while a significant response was observed in the presence of mavacamten, no statistically significant response was detected upon dATP application in the R403Q condition.

      We acknowledge that this does not constitute evidence of differential sensitivity per se, and we have revised the text accordingly by removing the concept of sensitivity to avoid potential misinterpretation.

      (8) Please add a discussion of the potential contributions of RLC phosphorylation to ON/OFF regulation.

      We agree with the reviewer that the potential involvement of RLC phosphorylation in ON/OFF regulation is an important aspect to consider in future work, and we have now acknowledged this point in the revised Discussion.

      Reviewer #3 (Recommendations for the authors):

      Some things that need to be clarified:

      (1) The comparison of skinned and intact muscle. These appear to give an unaltered SHG gamma signal (Figure 2), but a change in lattice spacing is expected between skinned and intact fibres. There is no mention of the lattice spacing of the samples or whether this was controlled. The implications are that SHG is independent of lattice spacing and/or the fraction of ordered myosin heads is independent of lattice spacing - each of which would be a useful result.

      We thank the reviewer for this important observation and agree that the role of lattice spacing is a relevant factor in the interpretation of the SHG signal. To address this point, we have performed an additional series of experiments on rabbit psoas muscle in which lattice spacing was modulated using dextran. These measurements allowed us to directly assess the sensitivity of the SHG signal to changes in interfilament spacing. We observed relatively small effects, but in the expected direction.

      The lack of appreciable differences between skinned and intact preparations may therefore be explained by the limited sensitivity of the technique to detect the relatively small variations in lattice spacing associated with these conditions.

      (2) In comparing the WT and mutant mini pig cardiac data, the authors note that the mutant fibre has more disarray. A comment on the effect of disarray on the SHG signal would be helpful. How much of the difference between WT and mutant could be due to this disarray? What happens if more vs less disordered areas of the fibre are compared?

      Myofibrillar disarray is indeed a characteristic feature of HCM tissue and is more evident in the R403Q minipig samples compared to WT. To minimize potential polarization artefacts related to structural disorganization, the entire field of view was first examined to identify regions of interest (ROIs) where sarcomeres showed minimal local disarray. Data acquisition and analysis were restricted to these locally well-aligned regions, as described in the Methods section. Therefore, the comparison between WT and mutant samples was performed on locally well-oriented regions rather than on highly disorganized areas.

      We have now clarified this point in the revised manuscript to better explain how ROI selection, restricted to locally aligned regions, minimizes the potential contribution of myofibrillar disarray to the pSHG measurements (page 14).

      (3) I note in Figure S3 that there is a bigger dispersion in the minipig data than mouse or rabbit. The minipig may also show a non-normal distribution. Has this been considered?

      Following the reviewer’s suggestion to include dATP measurements in additional species, we have extended the interspecies analysis and revised Figure S3 to provide a direct comparison across rabbit psoas, mouse cardiac, and minipig cardiac samples under the investigated conditions.

      From this more comprehensive dataset, no substantial differences in data dispersion are apparent among the different muscle types. Moreover, as part of the updated statistical framework, normality tests were performed for all datasets

      (4) A note about the heterogeneity in the regulation of thick filaments along their length should be added. There is significant evidence for a difference in regulation between the MyBP-C regions and the rest from single molecule studies (Kad lab) and interference X-ray signals (London Kings group). This is probably beyond the current resolution of the SHG, but the complexity should be acknowledged. Heterogeneity in the thick filament is also apparent from cryo-EM images of relaxed muscle and thick filaments.

      This important point is now acknowledged the revised Discussion (page 21).

      (5) Interpretation of the effects of mavacamten on the ATPase data ae complicated by the observation that mava is an inhibitor of myosin independent of the effect on the order- disorder of thick filaments. I.e. mavacamten will inhibit myosin S1.

      This point has now been addressed (see response to Reviewer #2, point 5 weaknesses).

      Minor issues:

      (1) P7 line 2: where/were purchased from Sigma.

      (2) P7 last but one line: R403Q/R4303Q.

      (3) P8 last line: 2-deaoxyATP 2-deoxyATP.

      (4) Results, p9: in the section title and first line, replace psoas with rabbit psoas.

      (5) Line 5: ROI is not defined, only in the Figure legend.

      (6) Mavacamten: 50 uM used in psoas and 10 uM in cardiac fibres. A note on why would help those not familiar with this literature. Similarly, in Figure S3, presumably the mava concentration was saturating in each case.

      (7) P14, last paragraph, line 3: Figure 4 should be Figure 5.

      (8) P17, 5 lines from the end: as effective at 2-deoxyATP at/as.

      (9) P19, lt line: fibber/fiber.

      All these minor points have been fully addressed. We thank the reviewer for noting them.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback on our work. We also thank the editor for communicating with reviewer #1 regarding our thoughts on their comments. We hence revised the manuscript based the new feedback from reviewer #1. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      In my view, the ability to address this issue with additional analyses is relatively limited. I cannot think of a solid way to escape this issue in the present design. Adding a nuisance regressor to their RSA regression, which I suggested in my previous response, may reduce the bias, but the efficacy of this would be limited to controlling only a certain kind of phase-dependent confound (a 'main effect' component of study phase, i.e., one that is constant across other conditions; see next message for discussion).

      Rather than new analyses, I think a more straightforward revision might be to modify the conclusions advanced in the paper, so that they are more solidly supported by the design and evidence. In my opinion, this design is ill-posed to solidly identify SC and SR representations. As a result, I think that any framing that alleviates pressure on this design to yield 'solid' evidence for identification of SC and SR representations, and instead emphasizes stronger areas of this work, would be constructive.

      For example, one potential framing is to advance the idea of SC and SR representations, and discuss an idealized design that could identify them by orthogonalizing stimulus features from ISPC, which are genuinely novel and useful ideas. The current design could then be presented as an opportunistic or initial case study of testing this question, while acknowledging its limitation upfront. The goal here would be to frame the study in a way that allows for the results to be presented with an appropriate grain of salt, while also illustrating the authors thoughtfulness and creativity in devising analyses to test for latent associative representations. In this case the 'solid' label would reference the authors' reasoning and analyses rather than design and conclusions.

      I only intend this example as an illustration; there may be several ways of framing this paper so that it is on more 'solid' ground, and I don't want to dictate how exactly authors should write their paper.

      Nonetheless I think this issue is important, not only to avoid faulty inference, but also to avoid establishing counterproductive precedents in this field. For example, if students read a paper whose conclusions are labeled "solid" but that nevertheless has critical flaws in its design, then those students may be misled in their own work. But if the flaws were discussed transparently and critically, and the strength of the conclusions were de-emphasized relative to other aspects of the paper, students may not only be inspired by the ideas developed in the paper, but also come away with knowledge about the issues of experimental design.

      Discussion of the weaknesses in the conclusions and an additional potential analysis:

      The key goal of this study is to identify SC/SR representations, which requires decoupling stimulus features from item-specific proportion congruency (ISPC), but the study-phase confound contaminates this decoupling. I think this impacts both their cross-phase decoding and RSA analyses.

      In their results (lines 139-144):

      "This standard ISPC manipulation can test whether neural representations of controlled and non-controlled information are wrapped on the same trial by combing with the following EEG analysis (See Methods). However, it potentially mixes color identity with SC, word identity with SR, and the ISPC between SC and SR. To deconfound these factors when estimating SC and SR association representations on each trial, we modified this paradigm by flipping the ISPC contingencies across different phases of the task."

      SC/SR representations are higher-order conjunctive classes, formed by the interaction of lower order variables (Color, Word, and ISPC). This means that successfully identifying these representations relies on demonstrating that each class can be reliably individuated from every other class in a manner that cannot be explained by (1) representation of shared lower-order features, such as stimulus color or response, and (2) trivial nuisance factors such as study phase. However, in this design, the trivial factor of study phase is strongly confounded with the ISPC contingency flip.

      Regarding RSA: If study phase indeed leads to trivial separability, it seems that similarity among conditions within the same phase would be inflated because the study phase does not appear in the RSA model (Figure 12). In which case, the SC/SR coefficients would be inflated, as the SC and SR models are entirely within-phase.

      Nevertheless, an additional control analysis may be possible here. Looking at this regressor set, it seems possible to me to fit a model where the predominant phase is entered as an additional covariate. (If I am reading this correctly, phase would appear as a 2x2 block-diagonal matrix). I suggested this in my last review, but in their recent letter, authors refused. I do not understand why, as I think this regressor set should be identifiable, but perhaps I am wrong here.

      That being said, I do not think adding this nuisance covariate would fully solve the issue. The lower-order regressors (Color, Word, ISPC) are defined as the EEG responses shared across phases of the study. To interpret the higher-order SR/SC coefficients, the lower-order components must be fully partialled out. But the putative impact of study phase contaminates this interpretation, as study phase could trivially decrease the similarity of lower-order terms (e.g., decreasing Color similarity between phase 2 vs 3). In which case, the partialling would be expected to be incomplete.

      This incomplete partialling is the RSA analogue of the issue in interpreting cross-phase decoding discussed above. As there, so too here I do not see a solid way around it in the present design. This is why in my previous review I referred to the addition of a phase covariate in the RSA regression as a "band-aid": it controls for some problems (main effect of phase), but not all (interactions of phase and lower-order terms).

      To summarize, I think that RSA would offer an additional opportunity to control for this potential confound, albeit in a limited sense (study phase effects that are consistent across conditions). But in a more general and rigorous sense, to me, the SC and SR terms in the RSA regression also seem susceptible to the same weakness as the cross-phase decoding analysis.

      We thank the reviewer for taking the additional time and effort to provide the new comments. Following the reviewer’s suggestion, we revised the language regarding the conclusions of this project (page 2,5,27-28) and explicitly discussed the weaknesses of the design on decoding and outline a possible solution for future studies in the Discussion section:

      “A limitation of the current design is that in theory temporally structured noise (e.g., autocorrelation in EEG data) may bias the decoding accuracy due to the blocked design. Although the present data provided no evidence that the decoding results in this study were biased by temporally structured noise, future studies should aim to develop experimental designs that eliminate this potential confound at the source. One potential solution would be to introduce additional phases flipping ISPC manipulations. At the same time, enough trials must be included in each phase to ensure the strength of the ISPC effect within each phase. A careful balance between session number and length will be helpful to optimize the duration of such a design.”

      As eLife also publishes review report, below we also summarize the three control analyses we ran and our reasoning of how they (would) address the issue of temporally structured noise for interested readers to assess:

      We acknowledge the theoretical issue of temporally structured noise (TSN) in our design when classes were decoded across different phases. However, the key question for the current data is whether there is empirical evidence that the decoding results were actually driven by TSN. To clarify this issue, we summarize several lines of evidence suggesting that the decoding results were not attributable to TSN:

      (1) Split-half cross-validation. We split the EEG data from Phase 2 and the combined Phases 1 and 3 into chronological first and second halves. Phases 1 and 3 were combined because they shared the same MC and MI assignments. This resulted in four possible combinations, each consisting of eight classes drawn from different phases: combination 1 included the first half of Phase 2 and the first half of Phase 3; combination 2 included the first half of Phase 2 and the second half of Phase 3; combination 3 included the second half of Phase 2 and the first half of Phase 3; and combination 4 included the second half of Phase 2 and the second half of Phase 3. We trained the decoders on one combination and tested them on another and then averaged the decoding results across all possible training-test assignments. The similar decoding patterns observed across these analyses (Fig. 6a,b) further confirmed that the decoding results were not driven by TSN.

      This analysis is conceptually similar to the “cross-phase” decoding analysis suggested by the reviewer in the first round of review. We also performed an additional distance-based control analysis (see below) to further test whether the decoding results could be explained by TSN, without imposing the constraint used in the split-half cross-validation that trials from the two phases had to fall within a 400-trial window.

      (2) Distance analysis. We predicted that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. However, we did not observe such a negative pattern (Fig. 6c). Note that this distance was defined with respect to trials of the same type, rather than absolute chronological time.

      (3) Shuffled analysis. If the decoding results were primarily driven by TSN, either at a short-term or long-term timescale, then shuffling the condition labels within each mini block should preserve the TSN structure present in the real data. In that case, the decoding results from shuffled data should not differ from those observed from real data. However, we found the significant difference between real data and shuffled data as shown in Author response images.

      Author response image 1.

      Shuffling analyses with stimulus-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 1b.

      Author response image 2.

      Shuffling analyses with response-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 2b.

      We thank the reviewer for the suggestion on RSA with phase. There are some concerns for this analysis:

      First, we think that the suggested analysis may be difficult to interpret. Because the SC and SR conditions differ across phases. Regressing out phase in RSA could also remove SC and SR information.

      Second, based on the reviewer’s comment, we understand that the suggested analysis may still not provide a clear falsifiable criterion for determining whether the results could be driven by the theoretical TSN issue inherent in the design.

      Third, the three control analyses we have performed examine this issue from different perspectives and collectively provide no evidence that the results were driven by TSN.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Thank you for your suggestion regarding the design. It is possible that the (re-)learning of ISPC can be fast. That said, enough trials are required to obtain a robust ISPC effect for each phase after the ISPC flips. Given that the EEG scanning (not including capping) in current design was about 1.5 hours, it is impractical to have both more sessions for a more orthogonal design and long sessions for robust within-session ISPC effects. We chose to maximize the latter because flipped behavioral ISPC effect in each session is the basis for the following EEG analysis. We have included the reviewer’s suggestion as a potential design solution for future studies in the Discussion section mentioned above on page 27.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Thank you for raising this important point. Although ISPC effects are considered reactive, in our design the long sessions may create a temporal context for the participants to differentiate the current control demand linked to each color. The maintenance of such contextual information needs to span across trials, leading to pre-stimulus coding that proactively guides the control demand for each color. This claim is not central to this manuscript, which investigates whether SC and SR representations simultaneously guide behavior. Additionally, we do not think the current design is well-equipped to test this hypothesis because the pre-stimulus onset is the only supporting evidence. In the revised manuscript, we discussed this as a future research direction and proposed a design that aims at better isolating proactive control signal on page 25.

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

      We indeed tried both the full model of random effects (i.e., considering covariance between all slopes and intercept) and a reduced model without any covariance (i.e., listing each random slope separately without intercept in lme4). However, neither model converged at all time points.

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a solid design and the analyses are appropriate for the research questions.

      Strengths:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed your concerns raised below.

      Weaknesses:

      I still have some doubts on the effectiveness of the experimental manipulation on Phase 2: although the ISPC effect is still present, it is much weaker in comparison, suggesting the participants did not learn the contingency statistics in Phase 2 as well as they did in the other phases, due to either the lingering effect of the previous phase or an inherent bias towards one color pairs. Perhaps by separately plotting the earlier and later blocks of Phase 2 any difference can be revealed if it exists. This behavioral difference could result in unequal levels of SC/SR representation across phases, which may raise problems when data were combined for analyses that assume the neural effects are equivalent.

      Thank you for your concern about this important issue. We agree with the reviewer that the true SR/SC levels may not be equivalent between Phase 2 and Phase 1/3. Nevertheless, because the manipulation of ISPC is binary, the decoders were trained to test whether the neural signals represent the two levels of ISPC (i.e., a higher vs. a lower level) differ systematically. The decoding analysis does not require that the neural effects of SC and SR must be numerically equivalent between phases (i.e., it is not necessary that the two levels are equidistant from the center point of SC/SR. Indeed, the decoding analysis only requires that the two levels are different). For example, if the ISPC level ranges from -1 to 1 and the EEG signals can reliably decode ISPC levels of -0.5 and 0.7 (i.e., two unequal levels), it can still be treated as supporting evidence that ISPC levels are encoded in the EEG signals. The same logic applies to the representational subspace analysis. As to the RSA, as can be seen in Fig. 12, the regressors are also binary, encoding whether two experimental conditions share the same SC/SR level without assuming equivalent neural effects. Thus, we argue that the reported decoding and RSA can still test the encoding of SR and SC. We discussed this issue on page 24.

      Following the reviewer’s comment, we plotted the ISPC effects in first and second half of Phase 2 separately (the figure below). We also tested whether the ISPC effects differ qualitatively between the two halves using a 3-way ANOVAs separately on RT and Error rate. The results showed that the time (the first half vs. the second half) × Congruency × ISPC interaction was not significant for either RT data (F<sub>(1,39)</sub> = 3.40, p > 0.05) or error rate (F<sub>(1,39)</sub> = 2.49, p > 0.05), suggesting that ISPC effect did not systematically change over time in Phase 2 See Supplementary Figure 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have expanded the real-data analyses, clarified the relationship between CMP and prior modulated Poisson models, and added discussion of model limitations and future extensions. In summary, the major changes include:

      (1) We revised the Introduction, Results, Methods, and Goris-model appendix to clarify the relationship between CMP and prior modulated Poisson models. In particular, we now emphasize that the key distinction is CMP’s continuous-time stochastic gain process.

      (2) We moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Figure 3 in the revised manuscript), making the validation of the inference procedure more visible to readers.

      (3) We added new analyses of the inferred gain process. Specifically, we now show the cross-trial gain mean and cross-trial gain variance in Figure 4A to assess whether gain captures stimulus-locked structure, and we added an analysis of pre- versus post-stimulus cross-trial gain variability in Figure 4C to test for gain-variability quenching during stimulus presentation.

      (4) We clarified the definitions and implementation of the Baseline Poisson, Poisson-GP, and Goris-style comparison models, including the role of the smoothness prior on the stimulus drive.

      (5) We expanded the Discussion to describe future extensions to population recordings, including a GPFA-inspired extension with low-dimensional shared gain activity across neurons.

      (6) We added a Discussion paragraph clarifying that the current CMP model captures Poisson and super-Poisson variability, but not sub-Poisson variability, and outlined possible extensions using spike-history terms, renewal-process likelihoods, or alternative count distributions.

      eLife Assessment

      This work of fundamental significance introduces a novel statistical model of spiking activity that incorporates continuous−time gain modulation. The authors provide exceptional evidence that the model outperforms earlier approaches and alternative candidates in capturing spiking responses across multiple visual areas in the macaque. Beyond its methodological contribution, the study offers new insights into how stimulus−driven variability and internally generated gain fluctuations evolve over time and between brain areas. The framework is likely to find broad application beyond the datasets examined here.

      We sincerely thank the Senior Editor, Reviewing Editor, and both reviewers for their careful evaluation and constructive feedback. We are encouraged by the positive assessment of the work and by the recognition of its methodological and conceptual contributions. We especially appreciate the acknowledgement that the continuous-time formulation provides a useful framework for modeling gain modulation in spiking activity, improves upon earlier approaches in capturing responses across multiple visual areas, and offers new insights into how stimulus-driven variability and internally generated gain fluctuations evolve over time and across brain regions.

      In the revised manuscript, we have addressed the reviewers’ comments by clarifying the relationship between CMP and prior modulated Poisson models, strengthening the presentation of the simulation-based recoverability analysis, adding new validation analyses of the inferred gain process, and expanding the Discussion of model scope, limitations, and future directions. In particular, we now more clearly distinguish the continuous-time gain process in CMP from Goris-style models with constant or piecewise-constant gain, move the simulation recoverability analysis into the main Results, examine trial-averaged inferred gain and gain-variability quenching, clarify the definitions of the baseline and comparison models, and discuss extensions to population recordings and sub-Poisson variability.

      We believe these revisions improve the clarity, rigour, and scope of the manuscript. Below, we address each reviewer comment in turn and describe the corresponding changes made in the revised manuscript.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, Rupasinghe and co−authors introduce a new statistical model for spiking neurons. Building on earlier work, they propose to model spikes as arising from a Poisson process whereby the firing rate is the product of stimulus drive and astimulus−independent gain signal. The critical innovation of this work is that the gain signal is modeled in continuous time. Earlier explorations of this statistical construction treated the gain−signal as constant within a trial. This innovation is elegant and important. It makes the model richer, more plausible, and more broadly applicable. The authors show that the model parameters are recoverable from realistic amounts of data and then apply the framework to previously studied datasets. They show that the new model outperforms earlier models and alternative candidates in capturing spiking data across four visual areas of the macaque monkey. Analysis of the model parameters replicates some earlier findings and uncovers several new insights. The model and fitting methods can be broadly applied to partition different types of signals and noise from spiking data and are likely to be widely adopted in the systems neuroscience community.

      Strengths:

      (1) Through clever use of advanced statistical techniques, the authors manage to infer critical information from single−trial single−cell data.

      (2) The question of which aspect of a spike train is signal and which is noise is omnipresent in neuroscience. By improving our ability to characterize the distinct factors that shape spiking activity, this work makes a fundamental contribution to the literature.

      We sincerely thank the reviewer for the thoughtful and detailed evaluation of our manuscript. We are pleased that the continuous-time formulation and its methodological contributions were viewed as elegant, important, and broadly applicable. We also appreciate the reviewer’s recognition that the framework provides a useful way to separate stimulus-driven and modulatory components of neural variability from single-trial, single-cell data. The reviewer’s comments helped us improve the precision of our framing, clarify the relationship between CMP and prior modulated Poisson models, and strengthen the validation of the inferred gain process. Below, we respond to each point in turn and describe the revisions made in the manuscript.

      Weaknesses:

      Overall, I find the work impressive and important. I have a couple of questions and suggestions.

      (1) The work is entirely focused on single−cell data. While this is a great starting point, expanding the approach to spiking activity in neural populations is an importantfuture goal.

      We thank the reviewer for this important suggestion. We agree that extending the CMP framework to population recordings is a natural and important direction for future work. In the present study, we focus on single-neuron responses to establish the continuous-time model, validate the inference, and characterize how stimulus-driven activity and stochastic gain fluctuations can be separated at the level of individual cells. However, the same modeling principles could be extended to simultaneously recorded neural populations by introducing shared latent structure across neurons. For example, one natural direction would be to combine CMP with ideas from Gaussian Process Factor Analysis [Keeley et al., 2020], using low-dimensional shared gain activity to capture population-wide fluctuations, while retaining neuron-specific stimulus-driven components. Such an extension would allow the model to capture correlated variability and shared modulatory dynamics across neural ensembles. In the revised manuscript, we have expanded the Discussion to describe this possible future extension to population recordings.

      To address this comment, we expanded the Discussion (Page 14: lines 473-478) to describe a possible GPFA-inspired extension of CMP to population recordings.

      (2) Line 49−53: These statements seem incorrect to me. The modulated Poisson model , as introduced in Goris et al (2014), is a process model that can perfectly be used to generate spike trains (within a trial, spiking emerges from a Poisson process, which canbe homogeneous or inhomogeneous). Moreover, the model contains a parameter thatrepresents the duration of the counting window (delta t). The dependency of over− dispersion on the size of the time bins for real neurons is shown in Figure 1b (inset plot) of that paper (and shown to resemble the model prediction). This time− dependency was further explored by the same authors in Goris et al (2018 − Journal ofVision) and also in Henaff et al (2020 − Nature Communications). I suggest that the authors rephrase this argument (here and at some later points in the paper). They could just say that the Goris model makes the simplistic and implausible assumption that, within a given trial, gain does not fluctuate. This is clearly an important limitation and the key difference with the continuous model introduced here.

      We sincerely thank the reviewer for identifying this lack of clarity in our original description. We agree that our original description was not sufficiently precise. The modulated Poisson model introduced by Goris et al. (2014) is indeed a generative process model and can be used to generate spike trains, with spiking arising from a Poisson process that may be homogeneous or inhomogeneous within a trial. We apologize for implying otherwise.

      Our intended point was that, in the original formulation, the modulatory gain is represented as a scalar random variable associated with a counting window or trial, and therefore does not explicitly model gain as a continuously time-varying process within a trial. Thus, the key limitation addressed by CMP is not the use of a Poisson process, but the assumption that gain is constant or piecewise constant over the relevant interval.

      In the revised manuscript, we have rephrased the Introduction to clarify this distinction. We now describe the Goris model more accurately as a modulated Poisson framework in which gain is constant over the counting window, and we emphasize that CMP extends this framework by replacing this assumption with a continuous-time stochastic gain process. We have also added discussion of related time-dependent analyses and extensions [Goris et al., 2018, H´enaff et al., 2020], as thoughtfully suggested by the reviewer.

      In addition, we revised the Results and Methods to clarify how the Goris-style baselines were implemented in our comparisons. Specifically, all Goris-style results reported in the main model comparisons use versions with a smoothness prior on the stimulus drive, where the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the Poisson-GP model. This ensures that the comparisons focus on different assumptions about the temporal structure of the gain process, rather than differences in stimulus-drive estimation. We also clarified the comparison to Goris-style variants without this smoothness prior, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood (Figure 5 - figure supplement 2). These results show that the smoothness prior on the stimulus drive substantially improves model performance. Finally, we revised the Figure 1 caption and the Goris-model appendix to make these distinctions explicit.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21), Figure 1 caption, and Goris-model appendix to clarify that CMP extends the Goris framework by modeling gain as a continuously time-varying process within trials.

      (3) Line 54−55: I think the first part of the claim is a bit misleading. There is nothing in the Goris model that would inherently limit it to homogeneous Poisson processes, as seems to be implied by this description. The model is built on the assumption thatspike generation within a trial arises from a Poisson process. This may very well be an inhomogeneous Poisson process (i.e., a stimulus−dependent time−varying firing rate). Homogeneous and inhomogeneous Poisson processes both give rise to Poisson distributed spike counts (and thus a mixture of Poisson distributions across trials in the Goris model). I suggest the authors clarify this description a bit. Note that the two model variants illustrated in Figure 1b and c were also explored in Henaff et al (2020 − Nature Communications).

      We thank the reviewer for this helpful clarification. We agree that the Goris model is not limited to homogeneous Poisson spiking and can incorporate a stimulus-dependent, time-varying firing rate within trials. We did not intend to imply otherwise, and we have revised the relevant text to avoid this misunderstanding.

      Our intended point was that, in formulating continuous-time extensions of the modulated Poisson framework, we explicitly model the time-varying stimulus drive using a smoothness prior, as in the CMP framework, and then consider different assumptions about the temporal structure of the gain process, including constant gain and independently resampled gain across time bins. This highlights the distinction between piecewise-constant gain assumptions and the fully continuous gain process introduced in CMP.

      In the revised manuscript, we have clarified this distinction in the Introduction, Results, and Methods. We now state that the Goris-style variants use stimulus-dependent, time-varying Poisson firing rates, and that the main difference between these variants and CMP lies in the temporal structure assumed for the gain process. We have also acknowledged related variants explored in Goris et al. [2018] and H´enaff et al. [2020], and clarified that our continuous-time formulations of the Goris model differs by imposing a smoothness prior on the stimulus drive. This allows us to estimate a regularized time-varying stimulus component while comparing different assumptions about gain dynamics, ensuring that the comparison focuses on the temporal structure of the gain process rather than differences in stimulus-drive estimation. We also highlight in Figure 5 - figure supplement 2 that even for the Goris-style models, versions that use a smoothness prior on the stimulus drive outperform versions that do not, which are closer to the original modulated Poisson formulation.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21) to clarify that the Goris-style variants allow stimulus-dependent time-varying firing rates and differ from CMP primarily in their assumptions about gain dynamics. We also added citations to related time-dependent extensions of the modulated Poisson framework.

      (4) The extension to the continuous case is very elegant!

      We thank the reviewer for the positive comment and are pleased that the continuous-time formulation was viewed as elegant.

      (5) I find the result shown in Appendix 3 critically important. The recoverability of the model for realistic amounts of data is foundational for the rest of the paper. I wouldconsider including this analysis in the main results section. Not all readers may check Appendix 3, but they should know about this result.

      We thank the reviewer for emphasizing the importance of this result. We agree that demonstrating parameter recoverability is foundational to the paper and should be visible to readers in the main Results section. In the revised manuscript, we have moved the simulation-based validation from Appendix 3 into the main Results. This section now describes the synthetic CMP dataset, the inference procedure used to estimate the latent stimulus-drive and gain processes, and the comparison between true and inferred GP hyperparameters. These results show that the proposed inference framework can accurately recover the ground-truth stimulus drives, gain processes, and hyperparameters from realistic amounts of simulated data.

      To address this comment, we moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Page 6: Lines 200-211 and Figure 3).

      (6) Figure 3: I am wondering whether the inferred gain is capturing some response fluctuations that originate from the cell’s phase−selectivity. Could the authors compute the trial−averaged inferred gain (ideally, aligned to stimulus−phase at the start of the trial if this experimental parameter varied across repeats)? If they have successfully partitioned the response variance, the trial−averaged gain should have no systematic temporal structure. If it has a sinusoidal modulation, it may partially capture stimulus−drive. This could be an interesting test to run on all model fits to further validate that the partitioning into a signal and noise component succeeded as intended.

      We thank the reviewer for this insightful suggestion. We agree that verifying that the inferred gain does not capture stimulus-driven structure is an important validation of the model. In the revised manuscript, we have added the trial-averaged inferred gain to Figure 4A for the example neuron. This analysis shows that the trial-averaged inferred gain is relatively flat and neither resembles the inferred stimulus drive nor exhibits clear stimulus-locked temporal structure. This suggests that trial-specific gain fluctuations largely average out across repeats, consistent with the interpretation that the gain process captures random trial-to-trial variability rather than stimulus-driven activity.

      We also note that a direct comparison of this inferred gain trace across methods is not possible for the Goris-style baselines, because these models do not infer a continuous trial-specific gain process. Instead, they marginalize over scalar or time-bin-independent gain variables when computing likelihoods and Fano factor curves. Thus, the trial-averaged gain diagnostic is specific to the CMP model, where the posterior over the continuous-time gain process is explicitly inferred.

      To address this comment, we added the trial-averaged inferred gain to Figure 4A and clarified that it does not show a clear stimulus-locked temporal structure (Page 7: Lines 229-236).

      (7) One common observation that is currently not explored is the quenching of neuronal response variability following stimulus onset (Churchland et al 2010 − NatureNeuroscience), which was suggested to reflect a quenching of gain variability in Goris et al (2024 − Nature Reviews Neuroscience). Building on the previous suggestion, the authors could compute the temporal evolution of cross−trial gain variability from the inferred gain traces. Do they recognize a reduction in gain variability following stimulus onset? If so, it would be worthwhile to show this.

      We sincerely thank the reviewer for this valuable suggestion. We agree that examining whether gain variability decreases following stimulus onset provides an important test of the inferred gain process. In the revised manuscript, we have added an analysis of the temporal evolution of cross-trial gain variability before and after stimulus onset.

      First, in Figure 4A, we now show the cross-trial variance of the inferred gain for the example neuron. This trace shows larger gain variability during the stimulus-off period and a reduction following stimulus onset, suggesting that the inferred gain captures a stimulus-related quenching of trial-to-trial variability. To quantify this effect across the population, we also added a pre- versus post-stimulus comparison in Figure 4C. Following the approach of Churchland et al. [2010], we compared gain variability in two matched 400-ms windows: a pre-stimulus window ending at stimulus onset and a stimulus-period window beginning 100 ms after stimulus onset. For each neuron and stimulus condition, we computed the cross-trial variance of the inferred gain at each time bin, averaged this quantity within each window, and then compared the pre- and post-stimulus values across neuron-stimulus pairs.

      This analysis revealed a significant reduction in inferred gain variability following stimulus onset (one-sided paired Wilcoxon signed-rank test, p≤ 10<sup>−15</sup>), consistent with gain variability quenching [Churchland et al., 2010, Goris et al., 2024]. We now report this result in the main text and illustrate it in Figure 4A and Figure 4C. This provides additional evidence that the inferred CMP gain captures meaningful trial-to-trial variability and its temporal modulation around stimulus presentation.

      To address this comment, we added the cross-trial gain variance trace to Figure 4A and a population-level pre- versus post-stimulus gain-variability quenching analysis (Page 7 and 8: Lines 239-248) in Figure 4C.

      (8) Line 543−565: I want to make sure I understand the Baseline Poisson model and Poisson−GP correctly. For the baseline model, I had imagined that the authors would simply use the stimulus−conditioned PSTH as an estimate of the time−dependent firing rate, coupled with an inhomogeneous Poisson process assumption. But they additionally assume a Gamma prior on the firing rate to compensate for the sparsenessof the data (sometimes only 5 repeats per condition). The Poisson−GP includesexactly the same model components, but now the time−dependent firing rate is modeled by a Gaussian process. Doing this massively improves the goodness−of−fit (Fig 4A). Do I understand this correctly?

      We thank the reviewer for this careful reading. Yes, this understanding is broadly correct, and we have revised the manuscript to clarify the relationships among the Baseline Poisson, Poisson-GP, and Goris-style models. The Baseline Poisson model estimates a stimulus- and time-dependent firing rate independently for each stimulus condition and time bin, using a Gamma-Poisson formulation to regularize the estimate when the number of repeats is limited. The Poisson-GP model uses the same conditionally Poisson observation model, but replaces these independent time-bin-wise rate estimates with a smooth stimulus-specific Gaussian process model for the log firing rate.

      We have also clarified how the Goris-style models were implemented. All Goris-style results reported in the main model comparisons use versions with a GP prior on the stimulus drive. In these versions, the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the PoissonGP model, and the gain parameters are then fit under either the independent-gain or constant-gain assumptions. We used these GP-smoothed versions as stronger baselines. In Figure 5, Figure Supplement 2, we additionally compare these models to Goris-style variants without the GP prior on the stimulus drive, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood. This comparison shows that adding a GP smoothness prior to the stimulus drive substantially improves held-out model fit. Together, these analyses clarify that the GP-smoothed stimulus drive improves the Goris-style baselines, while the continuous-time gain process in CMP provides an additional improvement by capturing temporally structured trial-to-trial variability.

      To address this comment, we clarified the definitions of the Baseline Poisson, Poisson-GP, and Goris-style models (Pages 8-9: Lines 255-259, 263-266, 273-277), and revised the text (Page 11: Lines 319-326) describing Figure 4 - figure Supplement 2 to make explicit how this existing comparison isolates the effect of the GP prior on the stimulus drive.

      Reviewer #2 (Public Review):

      Summary:

      Neurons have varied responses to external stimuli that cannot be explained by naive Poisson models. Previous work has quantified and partitioned higher−than−Poisson variability in the brain into different components. The authors improve on these methods to infer how both the stimulus drive and internal gain dynamics impact neuronal variability continuously in time. The clean and well−reasoned model is rigorously developed and then applied to neural data across the visual hierarchy. This lends new insights into how variability is partitioned, agreeing with and extending previous work on how that variability changes from early visual areas (LGN, V1) through to higher, motion−sensitive areas (area MT). Another key contribution is that this partitioning can be fully addressed as a continuous−time process, which allows for the dissection of how the timescale of fluctuations in these two components changesacross the brain’s processing arc.

      Strengths:

      (1) The model is cleanly derived and thoroughly documented, including usable code shared in a GitHub repo. This makes the method immediately portable to other neural systems.

      (2) This is a clear and well−presented piece of work. The figures and writing are clear and understandable, and all pieces of the derivations are included in the main text and supplementary information.

      (3) Comparisons to other models, particularly the one from Goris et al., 2014 shows how this Continuous Modulated Poisson (CMP) model outperforms previous work.

      (4) New insights about how variability partitioning changes across the visual stream from LGN to MT are revealed, including how the gain fluctuates on longer timescales in higher visual areas. Another key result about the anticorrelation between the variance in stimulus drive and gain fluctuations comports with theories about how neurons maintain efficient, reliable encoding.

      (5) In addition to the results reported here, this work will serve as an excellent tutorial for students and postdocs first delving into the sources of variability in the brain.

      We sincerely thank the reviewer for the thoughtful and positive assessment of our work. We are pleased that the model development, empirical analyses, and presentation were viewed as clear, rigorous, and useful for the broader neuroscience community. We also appreciate the reviewer’s recognition that the continuous-time formulation meaningfully extends prior variability-partitioning approaches by allowing stimulus drive and internal gain dynamics to be characterized across temporal scales. The reviewer’s comments helped us further clarify the positioning of the work, expand the Discussion of model scope and limitations, and better articulate future extensions. Below, we address the specific suggestions raised by the reviewer and describe the revisions made in the manuscript.

      Weaknesses:

      The work is somewhat incremental, building on previous studies of the partitioning of variability in the brain, but it provides important new extensions, as noted above.

      Regarding the comment on incremental contribution, we agree that our framework builds directly on previous variability-partitioning approaches, especially the modulated Poisson framework of Goris et al. However, the main goal of this work is to move this class of models from a count-based formulation to a continuous-time spike-train framework. This extension is important because it allows us to model gain as a temporally structured latent process, characterize how variability depends on the timescale over which spikes are counted, and infer the temporal covariance structure of stimulus-independent fluctuations. In addition, the CMP framework provides analytic expressions for the Fano factor as a function of bin size, introduces the EPL covariance function for slowly decaying gain dynamics, and enables direct comparisons of gain amplitude and timescale across visual areas. In the revised manuscript, we have clarified this positioning and emphasized how CMP extends prior variability-partitioning models while preserving their interpretability.

      To address this comment, we revised the Introduction (Pages 3-4: Lines 108-111 and Lines 118-121) and Discussion (Page 13: Lines 425-429) to clarify better how CMP builds on prior variability-partitioning models while extending them to continuous-time spike-train data.

      The only major gap I would suggest addressing in the Discussion is the observation of sub−Poisson variability in the brain. It seems clear that this model can extend to sub− Poisson variability and its partitioning and perhaps even show how that varies in real time, with an animal’s attentional state. That is, of course, beyond the scope of the current work, but could be mentioned in the Discussion.

      We thank the reviewer for this suggestion. We agree that sub-Poisson variability is an important phenomenon observed in neural data. Because the CMP model uses a conditionally Poisson observation model with stochastic gain modulation, it naturally captures Poisson and super-Poisson variability but does not generate sub-Poisson spike count statistics in its current form. In the revised manuscript, we have clarified this limitation in the Discussion and outlined possible extensions that could address sub-Poisson variability, including spike-history terms, renewal-process likelihoods, and alternative count distributions [Truccolo et al., 2005, Paninski et al., 2007, Aghamohammadi et al., 2024]. We also note that such extensions could allow future models to examine how sub-Poisson and super-Poisson components vary with behavioral state, attention, or arousal.

      To address this comment, we added a Discussion paragraph describing the current model’s limitation for sub-Poisson variability and possible extensions to capture it (Page 14: Lines 460-471).

      References

      Stephen Keeley, Mikio Aoi, Yiyi Yu, Spencer Smith, and Jonathan W Pillow. Identifying signal and noise structure in neural population activity with gaussian process factor models. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13795–13805. Curran Associates, Inc., 2020.

      Robbe L. T. Goris, Corey M. Ziemba, J. Anthony Movshon, and Eero P. Simoncelli. Slow gain fluctuations limit benefits of temporal integration in visual cortex. Journal of Vision, 18(8):8–8, 08 2018. ISSN 1534-7362. doi: 10.1167/18.8.8. URL https://doi.org/10.1167/18.8.8.

      Olivier J H´enaff, Zoe M Boundy-Singer, Kristof Meding, Corey M Ziemba, and Robbe L T Goris. Representation of visual uncertainty through neural gain variability. Nat. Commun., 11(1):2513, May 2020.

      Mark M Churchland, Byron M Yu, John P Cunningham, Leo P Sugrue, Marlene R Cohen, Greg S Corrado, William T Newsome, Andrew M Clark, Paymon Hosseini, Benjamin B Scott, David C Bradley, Matthew A Smith, Adam Kohn, J Anthony Movshon, Katherine M Armstrong, Tirin Moore, Steve W Chang, Lawrence H Snyder, Stephen G Lisberger, Nicholas J Priebe, Ian M Finn, David Ferster, Stephen I Ryu, Gopal Santhanam, Maneesh Sahani, and Krishna V Shenoy. Stimulus onset quenches neural variability: a widespread cortical phenomenon. Nat. Neurosci., 13(3):369–378, March 2010.

      Robbe L T Goris, Ruben Coen-Cagli, Kenneth D Miller, Nicholas J Priebe, and M´at´e Lengyel. Response sub-additivity and variability quenching in visual cortex. Nat. Rev. Neurosci., 25(4):237–252, April 2024.

      Wilson Truccolo, Uri T. Eden, Matthew R. Fellows, John P. Donoghue, and Emery N. Brown. A point process framework for relating neural spiking activity to spiking history, neural ensemble, and extrinsic covariate effects. Journal of Neurophysiology, 93(2):1074–1089, 2005. doi: 10.1152/jn.00697.2004. URL https://doi.org/10.1152/jn.00697.2004. PMID: 15356183.

      Liam Paninski, Jonathan Pillow, and Jeremy Lewi. Statistical models for neural encoding, decoding, and optimal stimulus design. In Paul Cisek, Trevor Drew, and John F. Kalaska, editors, Computational Neuroscience: Theoretical Insights into Brain Function, volume 165 of Progress in Brain Research, pages 493–507. Elsevier, 2007. doi: https://doi.org/10.1016/S0079-6123(06)65031-0. URL https://www.sciencedirect.com/science/article/pii/S0079612306650310.

      Cina Aghamohammadi, Chandramouli Chandrasekaran, and Tatiana A. Engel. A doubly stochastic renewal framework for partitioning spiking variability. bioRxiv, 2024.

    1. Author response:

      Reviewer #1:

      We thank the reviewer for their comments. They raised an issue with the correlational nature of the analysis, suggesting alternative explanations, including stable individual differences in response styles, latent third variables, or measurement properties of the confidence scales. We agree with the reviewer’s comments that a latent third variable may confound our findings and higher-order metacognitive beliefs (which we did not assess) are a possible candidate. We will broaden our discussion to include other plausible confounds that may jointly relate to depression and confidence dynamics. Regarding stable individual differences in response styles, our analysis decomposed confidence into within and between-person components and standardised the within-person component by each participant’s own variability. This allowed the interaction and moderated mediation analyses to separate within-person associations from stable between-person differences (e.g., range of confidence scale used; Epskamp et al., 2018). With regards to measurement properties, we will conduct additional analyses investigating whether trait depression is associated with altered response patterns (e.g., non-linear or more variable mapping of latent evidence into confidence reports).

      The reviewer also noted that the temporal resolution we selected may not be optimal to measure co-fluctuation of depression and confidence. We agree this remains a challenge for the depression research (Jamalabadi et al., 2025; Tamm et al., 2024), and metacognitive confidence research (da Fonseca et al., 2023; Wright et al., 2024). We chose the interval as it allowed us to balance retention and data quality across an 8-week study (Eisele et al., 2022) and address a gap in the literature – frequent but shorter (~1-week) EMA-based assessments of existing metacognitive confidence studies (da Fonseca et al., 2023; Wright et al., 2024) precludes insights into the dynamics of mood and metacognitive confidence across longer timescales. We have explored one-day and half-day lagged associations but did not find significant effects across most symptoms for both local and global confidence. This was omitted from the original manuscript as the analyses could not symmetrically control for local/global confidence (as these were only assessed bi-daily) but will be included in the revised manuscript’s supplement.

      The reviewer noted that the temporal ordering of the moderated mediation analysis cannot be conclusively established. We agree that there is a lack of research investigating alternative orderings. We selected the local-to-global ordering because prior work has similarly modelled and demonstrated global confidence estimates as a cumulative integration of local confidence estimates across the block (Cavalan et al., 2023; Katyal et al., 2025; Lee et al., 2021; Rouault et al., 2019, 2022). Nevertheless, we modelled an alternative process (i.e. whether global confidence interacted with trait-level depression in predicting subsequent local confidence) but did not find a significant frequentist interaction effect. We will elaborate on our rationale for the current ordering and include analyses of this alternative process in our revised manuscript and supplement.

      Finally, we agree with the reviewer’s comments that the sample constraints generalisability. The revised manuscript will more clearly discuss the generalisability of our findings to clinical populations and the implications of self-selection and limited within-person variability in depressive symptoms for interpretation and future work. We will also tone down the causal language in the manuscript.

      Reviewer #2:

      We thank the reviewer for their comments. We agree with the reviewer that our strongest effects live at the trait level, whereas our cross-lagged findings provided insufficient evidence for depression driving underconfidence or vice versa (at least with a two-day lag). However, our moderated mediation finding centres on a different angle - trait level depression moderating how within-person confidence is integrated into more global beliefs. Nevertheless, we agree that the findings provide a plausible account for why underconfidence (and possibly metacognitive beliefs) remain persistent in depression, but not how either underconfidence or depression arose in the first place. We will make this clearer in our revised manuscript.

      Consistent with reviewer #1’s comments, we agree that the non-clinical nature of sample does constrain generalisability and inference. We will conduct a sensitivity analysis with a larger sample (more lenient inclusion criteria) and discuss its implications in more detail in our revised manuscript. We also agree with the reviewer’s comments that our findings do not preclude the possibility of temporal precedence and will further clarify in our revised manuscript with reference to shorter time lags. The reviewer noted that our use of terminologies (e.g., “depression”, “depressive symptoms”, “depressive mood”) was not clearly defined. Our revised manuscript will make this point clearer whilst also making the usage of terminologies consistent.

      References

      Cavalan, Q., Vergnaud, J.-C., & de Gardelle, V. (2023). From local to global estimations of confidence in perceptual decisions. Journal of Experimental Psychology: General, 152(9), 2544–2558. https://doi.org/10.1037/xge0001411

      da Fonseca, M., Maffei, G., Moreno-Bote, R., & Hyafil, A. (2023). Mood and implicit confidence independently fluctuate at different time scales. Cognitive, Affective, & Behavioral Neuroscience, 23(1), 142–161. https://doi.org/10.3758/s13415-022-01038-4

      Eisele, G., Vachon, H., Lafit, G., Kuppens, P., Houben, M., Myin-Germeys, I., & Viechtbauer, W. (2022). The Effects of Sampling Frequency and Questionnaire Length on Perceived Burden, Compliance, and Careless Responding in Experience Sampling Data in a Student Population. Assessment, 29(2), 136–151. https://doi.org/10.1177/1073191120957102

      Epskamp, S., Waldorp, L. J., Mõttus, R., & Borsboom, D. (2018). The Gaussian Graphical Model in Cross-Sectional and Time-Series Data. Multivariate Behavioral Research, 53(4), 453–480. https://doi.org/10.1080/00273171.2018.1454823

      Jamalabadi, H., Koosha, T. A., Stocker, E., Jansen, A., Ebner-Priemer, U. W., Proppert, R. K. K., Rieble, C. L., Tutunji, R., & Fried, E. I. (2025). Optimizing the frequency of ecological momentary assessments using signal processing. Psychological Medicine, 55, e358. https://doi.org/10.1017/S003329172510264X

      Katyal, S., Huys, Q. J., Dolan, R. J., & Fleming, S. M. (2025). Distorted learning from local metacognition supports transdiagnostic underconfidence. Nature Communications, 16(1), 1854. https://doi.org/10.1038/s41467-025-57040-0

      Lee, A. L. F., de Gardelle, V., & Mamassian, P. (2021). Global visual confidence. Psychonomic Bulletin & Review, 28(4), 1233–1242. https://doi.org/10.3758/s13423-020-01869-7

      Rouault, M., Dayan, P., & Fleming, S. M. (2019). Forming global estimates of self-performance from local confidence. Nature Communications, 10(1), 1141. https://doi.org/10.1038/s41467-019-09075-3

      Rouault, M., Will, G.-J., Fleming, S. M., & Dolan, R. J. (2022). Low self-esteem and the formation of global self-performance estimates in emerging adulthood. Translational Psychiatry, 12(1), 1–10. https://doi.org/10.1038/s41398-022-02031-8

      Tamm, J., Takano, K., Just, L., Ehring, T., Rosenkranz, T., & Kopf-Beck, J. (2024). Ecological Momentary Assessment versus Weekly Questionnaire Assessment of Change in Depression. Depression and Anxiety, 2024, 9191823. https://doi.org/10.1155/2024/9191823

      Wright, A. C., Palmer-Cooper, E., Cella, M., McGuire, N., Montagnese, M., Dlugunovych, V., Liu, C.-W. J., Wykes, T., & Cather, C. (2024). Experiencing hallucinations in daily life: The role of metacognition. Schizophrenia Research, Hallucinations: Neurobiology and Patient Experience, 265, 74–82. https://doi.org/10.1016/j.schres.2022.12.023

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting and well-written manuscript in which the authors set out to answer a simple, old question with a modern toolkit: where in crab evolution did sideways walking arise, how often has it been lost or regained, and is it plausibly linked to the ecological and taxonomic success of true crabs. To do this, they record locomotion from 50 live species, convert each species' movements into a quantitative index that compares forward versus sideways bouts, and then map the resulting states onto a recent crab phylogeny to infer the most likely evolutionary history of locomotor direction.

      We thank the reviewer for this positive summary of the study and for recognizing the value of our comparative behavioral dataset and phylogenetic approach.

      Strengths:

      The strongest part of the study is the dataset itself. Comparable behavioral measurements across dozens of crab species are rare. The authors have done the field and husbandry work needed to make this possible. The overall pattern they recover, that most true crabs are strongly biased toward sideways movement (while a smaller set of lineages move predominantly forward), is interesting and likely to be useful to others. The phylogenetic mapping is also a reasonable way to address the "how many times" question (although this is peripheral to my expertise). The manuscript makes a convincing case that sideways locomotion is not simply a trivial byproduct of a crab-like body plan.

      We appreciate the reviewer’s recognition of the dataset and the overall value of the study. We have revised the manuscript to make the conclusions more robust and better aligned with the strength of the evidence.

      (1) Where I am less convinced is in how strongly the authors describe the discreteness of the behavioral categories and the absence of intermediates. The manuscript states that the Forward-Sideways Index shows a clear separation between two locomotor types with little evidence for intermediates, and it cites a statistical test rejecting a single peak in the distribution. However, the histogram in Figure 3 appears structured within each labeled category, with subclusters inside both the forward and sideways groups rather than a single tight peak per group. This matters because the index is built by first placing each movement bout into "forward" versus "sideways" bins using a fixed angle boundary and then collapsing the result into a single ratio. That approach is simple and transparent enough, but it can also hide mixed strategies. For example, a species that produces substantial amounts of both forward and sideways walking can still end up with a strongly positive or negative index, and therefore be classified as a pure "type," even though the underlying behavior is mixed. In that context, rejecting a single peak in the across-species distribution does not, by itself, justify the stronger claim that intermediates are rare or absent.

      Related to this, a key methodological choice is the use of 60 degrees as the cutoff between forward and sideways bouts. This boundary may be reasonable as a convention, but the paper does not explain why it is the right place to draw the line, and there is a plausible biological concern that a fixed angular cutoff does not mean the same thing across taxa.

      Crabs vary in body shape and in how the legs are arranged around the body. In my own comparative work, for example, some species show an elliptical stance pattern elongated along the preferred direction of travel, while others show a more circular leg arrangement, and the latter can express more mixed forward and sideways behavior. When limb arrangement and body geometry differ across species, the same measured angle can correspond to different underlying mechanics and different functional "degree of sidewaysness." The practical implication is that the reported binary separation may partly reflect the imposed classification rule, rather than a sharp biological divide.

      We thank the reviewer for this important point. We agree that the across-species distribution of FSI values alone does not justify a strong statement that intermediate or mixed locomotor tendencies are absent. We also agree that reducing continuous bout-angle distributions to a single index could potentially obscure mixed directional strategies. We have therefore revised the manuscript to avoid implying a strict absence of intermediates and have added an additional analysis of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2).

      Specifically, we fitted one- and two-component mixture models to the continuous bout-angle distributions of each taxon and examined the supported number of components, peak locations, and mixture weights (Results, lines 196-204; Table S2). This analysis showed that 14 taxa were best described by a one-component model, whereas 36 taxa were best described by a two-component model. Importantly, among the 36 taxa best described by a two-component model, 33 had a dominant component explaining at least 70% of the distribution, whereas only three taxa showed relatively balanced two-component distributions. Thus, although some taxa do show mixed directional tendencies, most taxa are dominated by a primary directional component rather than showing an even mixture of forward and sideways locomotion.

      We also clarified the rationale for the 60° threshold used in the FSI calculation (Methods, lines 125-130). This threshold was not intended to represent a taxon-specific biological boundary between forward and sideways locomotion. Rather, it was used to divide the 360° space into three equal directional sectors: forward, sideways, and backward. This equal partitioning provides a consistent reference under a null expectation of uniformly distributed movement directions.

      To further assess whether our classification depended on the original FSI-based classification, we performed an additional data-driven check based on continuous bout-angle distributions (Results, lines 204-208; Fig. S2). We extracted the dominant peak location from each taxon’s continuous bout-angle distribution (Table S2) and estimated a boundary from the distribution of these dominant peak locations using a Gaussian mixture model. This yielded a data-informed cutoff of approximately 49.4°. The resulting peak-based classification was identical to the original FSI-based classification, with 15 forward-moving and 35 sideways-moving taxa. Thus, no taxon changed category under this independent classification approach.

      (2) Another limitation that affects interpretation is the decision to use one individual per species. I understand the logistics, and for some questions, a single representative individual can be a reasonable first pass. But it is not strong support for negative claims about intermediates, especially in a group where individuals can change substantially with growth and allometry. Crabs can grow dramatically, often with pronounced allometric shifts in limb proportions that can alter the center of mass location. Size alone can alter the kinematics and choice of locomotor behaviors in crustaceans. In species where appendage proportions change with size, or where certain legs become disproportionately large (or calcified), it is plausible that locomotor direction and the distribution of movement angles shift across ontogeny. That makes it hard to treat a single individual as a complete description of a species-level strategy, particularly for species that fall closer to the boundary between categories.

      We thank the reviewer for raising this important limitation. We agree that using one representative individual per species cannot capture the full range of within-species variation, including ontogenetic, size-dependent, or allometric changes in locomotor behavior. We also agree that this limitation is particularly relevant to strong claims about the absence of intermediates.

      As described in our response to Comment #1, we have therefore toned down statements implying a strict absence of intermediates and added analyses of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2). These additional analyses showed that some taxa do exhibit mixed directional tendencies, although most taxa were dominated by a primary directional component.

      We have also revised the manuscript to clarify the scope of our conclusions. Specifically, we now state that our single-individual sampling design does not capture possible ontogenetic, size-dependent, or allometric variation within species (Methods, lines 105-107). We also clarify that our conclusions are intended to identify broad interspecific patterns in the predominant direction of locomotion across major brachyuran lineages, rather than to describe the full range of locomotor variation within each species (Methods, lines 107-108). Thus, we no longer treat a single individual as providing a complete description of species-level behavioral variation, but instead use it as a standardized representative observation for broad comparative and phylogenetic analyses.

      In sum, this is a valuable and useful behavioral comparative study with a dataset that many in the field will appreciate. The main conclusions about the likely evolutionary placement of sideways walking are plausible, but several of the stronger claims about discrete locomotor types, the absence of intermediates, and the relationship to diversification would be more convincing if the analysis were less dependent on a fixed angular cutoff and on single individuals per species, or if the manuscript framed those points more cautiously so the conclusions track the strength of the evidence.

      We thank the reviewer for this constructive summary and for recognizing the value of our behavioral comparative dataset. We have addressed these concerns in detail in our responses above and revised the manuscript to make the main claims better aligned with the strength of the evidence.

      Reviewer #2 (Public review):

      Summary:

      The current work investigates the evolution of sideward locomotion in Brachyura in light of a single evolutionary origin. To this end, the authors first analysed the mode of locomotion in 50 crab species and observed mutually exclusive presence of sideways vs. forward movement. The phylogenetic analysis confirmed that there is indeed a single evolutionary origin for sideways movement, which was sometimes followed by several reversions to forward locomotion. This way, authors demonstrate how locomotor movement modes shape evolutionary diversification in animals by showing that species richness is much higher in side-ways-moving crabs than in the nearest groups. This is an interesting work that integrates behavioural analysis and phylogenetic relations, capitalising largely on crabs. I have a few suggestions and questions.

      We thank the reviewer for the positive assessment of the study and for recognizing the value of integrating behavioral analysis with phylogenetic relationships. We address the specific suggestions and questions below.

      (1) Firstly, I think the paper spends too much time on a straightforward analysis of the mode of locomotion.

      We agree that the final classification of taxa into predominantly forward- and sideways-moving groups is conceptually simple. However, because our study compares locomotor behavior across a broad range of crab taxa and then uses these behavioral data for phylogenetic reconstruction, we considered it important to describe the behavioral quantification in a transparent and reproducible way. The purpose of this section is therefore not to make a simple endpoint unnecessarily complex, but to show how discrete locomotor states were derived from raw trajectory data using a standardized procedure. For this reason, we retained the current analytical description.

      (2) I was also wondering whether the phylogenetic analysis could be simply achieved by maximising an objective function in which the modes of movement are inversely coded for two putative groups, with all values calculated at all possible nodes.

      The proposed objective-function approach may be useful for identifying a node that best separates two predefined locomotor groups. However, in the present study, we aimed not only to locate a possible boundary between forward- and sideways-moving lineages, but also to reconstruct the evolutionary history of locomotor transitions under an explicit phylogenetic model.

      For this reason, we used standard ancestral state reconstruction and stochastic character mapping rather than maximizing an ad hoc objective function across possible nodes. This approach allowed us to compare alternative transition-rate models (ER and ARD), estimate uncertainty in ancestral states at internal nodes, and quantify the posterior distribution of gains and reversals. We therefore retained the current phylogenetic framework, as it provides a model-based and more informative reconstruction of locomotor evolution across true crabs.

      (3) Unfortunately, I find that the authors did not sufficiently discuss differences in the ecological niches of species with forward vs. sideways locomotion modes (including challenges of locomotion and substrate).

      Likewise, what are the anatomic correlates of forward vs. sideways locomotion? For instance, how are the advantages assumed for sideways movement associated with a flattened body? Is it possible that the mode of motion is secondary to flattened/narrow body structure, which basically limits the distance between legs and thus makes the forward movement difficult - under this logic, the mode of movement would be a secondary phenomenon to body shape traits. How can one differentiate between this alternative and the one that puts the mode of movement in the centre of the story? On a related note, how do different modes of movement relate to the ability to fit into tight spaces - how does it relate to differences in leg joints?

      Is it possible that the sideways movement maximises the scanned visual field per unit time/displacement, which may be beneficial for mostly forward-moving predators?

      We thank the reviewer for this helpful comment. We agree that the previous version did not sufficiently address the possible relationship between locomotor mode and body shape, especially the alternative explanation that sideways locomotion may be secondary to carapace flattening. In response, we added a new morphological analysis using two carapace shape indices: relative carapace length (CL/CW) and relative carapace depth (CD/CS) (Methods, lines 178–185). In the revised Results, we report that relative carapace length differed significantly between forward- and sideways-moving taxa (phylogenetically informed ANOVA: F = 26.90, p < 0.001), whereas relative carapace depth did not differ significantly between the two groups (F = 1.18, p = 0.403) (Results, lines 209–214; Fig. S3). We also added this interpretation to the Discussion, noting that locomotor mode is associated with some aspects of carapace shape but is not explained by simple carapace flattening alone (Discussion, lines 318–325).

      We also revised the Discussion to address the reviewer’s suggestions about possible functional advantages of sideways locomotion beyond rapid bidirectional escape. Specifically, we now mention that other possible advantages may include movement through confined spaces and visual-field sampling during locomotion (Discussion, lines 310-318).

      Finally, we retained the existing discussion of ecological specializations in forward-moving lineages, including coordinated collective movement in soldier crabs, decoration and concealment in majoid crabs, and life inside confined host spaces in pea crabs. This discussion supports the broader point that the adaptive value of sideways locomotion may depend on ecological context.

      (4) It is really difficult to decipher the information contained in the nodes (circles) in the printed black-and-white version of the manuscript.

      We have changed the color scheme and strengthened the outlines of the node pie charts so that the ancestral-state probabilities can be more easily distinguished (Fig. 5). We also applied the same revised color scheme to Figure 3 and Figure S4 for consistency across the manuscript.

      (5) Briefly, although I find the study interesting, the presented complexity may not be necessary given the endpoints; it can be achieved much more simply. Furthermore, the degree to which the conceptual analysis of different modes of locomotion was exercised was limited. The general approach may serve as a good model for the evolutionary analysis of other traits. The demonstration of traceability of the relations in question is a major contribution of the work.

      We thank the reviewer for this constructive summary and for recognizing the broader value of our approach. We have addressed the methodological and conceptual points raised here in our responses to the specific comments above.

      Strengths:

      The research question and the novel combination of different data types.

      We thank the reviewer for highlighting the research question and the novel combination of different data types as strengths of the study. We have revised the manuscript to further strengthen this integrative framework.

      Weaknesses:

      The complexity of the methods used, along with a limited discussion of the potential dynamics that may underlie the evolution of the sideways movement mode.

      We have addressed these concerns in our responses to the specific comments above, particularly by clarifying the rationale for the behavioral quantification and expanding the discussion of morphology, ecological context, and functional hypotheses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Unbiased analysis of angle distribution. The authors already extract continuous bout angles prior to binning. I recommend using these distributions directly to assess modality at the species level (e.g., unimodal vs bimodal, peak locations, mixture weights) before collapsing behavior into the Forward-Sideways Index. Even a simple circular density estimate or mixture model would clarify whether species classified as "forward" or "sideways" are behaviorally pure or mixed, and would provide a quantitative basis for claims about intermediacy.

      We added mixture-model analyses of the continuous bout-angle distributions, including modality, peak locations, and mixture weights (Results, lines 196-204; Table S2).

      (2) Justify or stress-test the 60° cutoff. The manuscript should either provide a clear biological or data-driven justification for using 60° as the boundary between forward and sideways bouts, or demonstrate that the main conclusions are robust to reasonable alternative cutoffs (e.g., 45°, 75°). A brief sensitivity analysis in the supplement would be sufficient and would greatly strengthen confidence in the classification. Alternatively (my preference) would be to let the data inform the cutoff.

      We clarified the rationale for the 60° sector definition used to calculate FSI (Methods, lines 125-130; Fig. 2) and added an independent data-driven boundary analysis based on dominant peak locations, which yielded the same forward/sideways classification (Results, lines 204-208; Fig. S2).

      (3) Sampling justification. I recommend explicitly acknowledging that sampling a single individual per species limits the ability to detect ontogenetic, size-dependent, or allometric variation in locomotor strategy. If feasible, adding even limited replication across size classes or individuals for a small subset of taxa (particularly those near the classification boundary) would substantially strengthen the conclusions; otherwise, the manuscript should more clearly delimit which claims do and do not rely on the assumption of within-species invariance.

      We clarified that our conclusions concern broad interspecific patterns of predominant locomotor direction, rather than the full range of within-species variation (Methods, lines 102-108).

      (4) I would suggest toning down or reframing statements about "no intermediates". If additional analyses are not added, I recommend revising statements that imply a strict absence of intermediates to language that reflects what is directly shown (e.g., bimodality in an index derived from binned data). This would better align the claims with the current evidence.

      We revised the manuscript to avoid implying a strict absence of intermediates and now acknowledge that some taxa show mixed directional tendencies (Abstract, lines 27-29; Results, lines 191-208).

      (5) Framing and claims about diversification. The discussion of sideways locomotion as a key innovation would benefit from clearer separation between observed correlations and causal inference. If trait-dependent diversification analyses are not added, I suggest consistently framing this section as a hypothesis supported by comparative patterns rather than a demonstrated mechanism.

      We revised the Discussion to more clearly frame sideways locomotion as a possible key innovation associated with diversification, rather than as a demonstrated causal mechanism (Abstract, lines 31-34; Discussion, lines 294-309).

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors perform an analysis of the relationship between the size of an LMM and the predictive performance of an ECoG encoding model made using the representations from that LMM. They find a logarithmic relationship between model size and prediction performance, consistent with previous findings in fMRI. They additionally observe that as the model size increases, the location of the "peak" encoding performance typically moves further back into the model in terms of percent layer depth, an interesting result worthy of further analysis into these representations.

      Strengths:

      The evidence is quite convincing, consistent across model families, and complementary to other work in this field. This sort of analysis for ECoG is needed and supports the decade-long enduring trend of the "virtuous cycle" between neuroscience and AI research, where more powerful AI models have consistently yielded more effective predictions of responses in the brain. The lag analysis showing that optimal lags do not change with model size is a nice result using the higher temporal resolution of ECoG compared to other methods like fMRI.

      We thank the reviewer for their thoughtful assessment! We agree that the “virtuous cycle” between neuroscience and AI research has been, and will continue to be, a driving force in advancing our understanding of brain function through more powerful predictive models. We are especially pleased that the reviewer appreciated the lag analysis, as we view this as a valuable complement to the existing fMRI work.

      Weaknesses:

      I would have liked to have seen the data scaling trends explored a bit too, as this is somewhat analogous to the main scaling results. While better performance with more data might be unsurprising, showing good data scaling would be a strong and useful justification for additional data collection in the field, especially given the extremely limited amount of existing language ECoG data. I realize that the data here is somewhat limited (only 30 minutes per subject), but authors could still in principle train models on subsets of this data.

      We thank the reviewer for their valuable suggestion. For the revised manuscript, we performed a new analysis where we trained encoding models using subsets of the data (randomly sampling contiguous chunks of 50%, 25%, and 10% of all words in each of the training folds) and tested these models on all words in the test fold. As expected, we found that encoding performance increases as the training dataset size increases, suggesting that model performance scales with data quantity even within the constraints of our relatively small dataset. This result reinforces the importance of collecting dense ECoG data. We have added the following text to our Results section: “We also built encoding models using subsets of the data and found that encoding performance increases as the volume of training data increases (Fig. S6)” and included the results as a supplementary figure 6 in the revised manuscript.

      Separately, it would be nice to have better justification of some of these trends, in particular the peak layerwise encoding performance trend and the overall upside-down U-trend of encoding performance across layers more generally. There is clearly something very fundamental going on here, about the nature of abstraction patterns in LLMs and in the brain, and this result points to that. I don't see the lack of justification here as a critical issue, but the paper would certainly be better with some theoretical explanation for why this might be the case.

      We thank the reviewer for this insightful comment. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Lastly, I would have wanted to see a similar analysis here done for audio encoding models using Whisper or WavLM as this is the modality where you might see real differences between ECoG and other slower scanning approaches. Again, I do not see this omission as a fundamental issue, but it does seem like the sort of analysis for which the higher temporal resolution of ECoG might grant some deeper insight.

      We appreciate this suggestion. In a separate project, we focused on multimodal audio-to-speech-to-language large language models (LLMs), building encoding models using Whisper embeddings (from both the encoder and decoder stacks) to predict electrocorticographic (ECoG) signals during naturalistic conversations (Goldstein, Wang, et al., 2025). The higher temporal resolution of ECoG enables us to trace the temporal flow of information from the superior temporal gyrus (STG) and somatomotor areas (SM) to the inferior frontal gyrus (IFG) during speech comprehension. Conversely, during speech production, encoding in IFG peaked significantly earlier than in the STG and SM. We agree that scaling encoding models using multimodal approaches and our ECoG conversation datasets could yield valuable insights, and we look forward to exploring this in future work. However, we feel that the added complexity of multimodal encoding models falls beyond the scope of this paper.

      We have modified the following text to our Discussion section:

      “Since we exclusively employ textual LLMs, which lack inherent temporal information due to their discrete token-based nature, future studies utilizing multimodal LLMs integrating continuous audio or video streams, like Whisper or WavLM may better unravel the relationship between model size and temporal dynamic representations in LLMs (Goldstein, Wang, et al., 2025; Millet et al., 2023; Vaidya et al., 2022).”

      Reviewer #2 (Public review):

      Summary:

      This paper investigates whether large language models (LLMs) of increasing size more accurately align with brain activity during naturalistic language comprehension. The authors extracted word embeddings from LLMs for each word in a 30-minute story and regressed them against electrocorticography (ECoG) activity time-locked to each word as participants listened to the story. The findings reveal that larger LLMs more effectively predict ECoG activity, reflecting the scaling laws observed in other natural language processing tasks.

      Strengths:

      (1) The study compared model activity with ECoG recordings, which offer much better temporal resolution than other neuroimaging methods, allowing for the examination of model encoding performance across various lags relative to word onset.

      (2) The range of LLMs tested is comprehensive, spanning from 82 million to 70 billion parameters. This serves as a valuable reference for researchers selecting LLMs for brain encoding and decoding studies.

      (3) The regression methods used are well-established in prior research, and the results demonstrate a convincing scaling law for the brain encoding ability of LLMs. The consistency of these results after PCA dimensionality reduction further supports the claim.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) Some claims of the paper are less convincing. The authors suggested that "scaling could be a property that the human brain, similar to LLMs, can utilize to enhance performance", however, many other animals have brains with more neurons than the human brain, making it unlikely that simple scaling alone leads to better language performance.

      We thank the reviewer for this insightful comment. We agree that simply having more neurons does not automatically confer more complex or human-like cognitive or linguistic capabilities. This suggestion deserves a more nuanced treatment than we had included in the original manuscript.

      Research in comparative neuroscience has argued that human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012). However, the critical aspect is not merely the number of neurons, but how these neurons contribute to computational power within a specific evolutionary and cultural context. The uniqueness of human cognition has been argued to result from a global adaptation for increased information processing capacity (Cantlon & Piantadosi, 2024). Moreover, the language network in humans is likely grounded in the evolution of particular structural networks in the primate brain (Friederici & Becker, 2025). This suggests that the way brain regions are connected and the expansion of certain pathways are critical, not just the overall scale. Furthermore, the specialized structure must be tuned by its learning environment and training data. For example, both humans and LLMs learn from language data generated by other humans, which reflects world knowledge that has accumulated over many generations.

      We have modified the following text in the Introduction:

      “Research in comparative neuroscience has suggested that uniquely human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012).”

      We also added a caveat to the Discussion on this point:

      “As in the human brain, while scaling alone may yield emergent cognitive abilities (Cantlon & Piantadosi, 2024; Herculano-Houzel, 2012), specialized architectural features likely also play a critical role (Friederici & Becker, 2025).”

      Additionally, the authors claim that their results show 'larger models better predict the structure of natural language.' However, it remains unclear to what extent the embeddings of LLMs capture the "structure" of language better than the lexical semantics of language.

      We appreciate the reviewer's point about how well LLM embeddings capture the "structure" of language versus just lexical semantics. It's true that distinguishing these aspects is complex. From our perspective, a model's ability to predict/produce natural language entails that the model captures various levels of linguistic structure, including morphology, syntax, semantics, and contextual dependencies. We use "structure" inclusively in this sense. A model cannot achieve high predictive accuracy without representing, to some extent, all of these structural elements (Linzen & Baroni, 2021; Manning et al., 2020; Pavlick, 2022). There is a very active field of research into understanding exactly how these models represent these different structures of language (Ameisen et al., 2025; Chemla et al., 2024; Elhage et al., 2021, 2022; Hewitt & Manning, 2019). Our results confirm the core trend that larger models tend to better reproduce the various structures of language (i.e., yield lower perplexity; Fig. 2A).

      In previous work, we have shown that LLM embeddings better predict neural activity during natural language processing than lexical embeddings (e.g., GloVe) that do not contain other elements of linguistic structure (Goldstein et al., 2022; Kumar et al., 2024; Zada et al., 2024). In response to the following comment, we also compare LLMs to simpler models capturing specific speech and language features (see next comment). To clarify our intended use of the word “structure”, we’ve added a brief explanation in the Methods section:

      “In this study, we use the term “structure” to refer to a variety of linguistic patterns (e.g., morphology, syntax, semantics, context) that LLMs encode in order to better predict natural language.”

      (2) The study lacks control LLMs with randomly initialized weights and control regressors, such as word frequency and phonetic features of speech, making it unclear what the baseline is for the model-brain correlation.

      We’ve added several supplementary analyses to the revised manuscript to address these concerns. To establish a baseline, we extracted embeddings from each layer of the SMALL model with randomly initialized weights and constructed encoding models. The encoding performance is significantly higher for pretrained SMALL than for untrained SMALL for every layer (Fig. S4). For the untrained model, the performance is the highest for the 0th layer and decreases in subsequent layers. This is because at the 0th layer, every instance of the same word receives an identical, albeit random, embedding (See Supplementary Figure 4).

      We also compared the encoding performance of LLMs with more classical speech/language features and static GloVe embeddings (Goldstein, Wang, et al., 2025; Kumar et al., 2024). First, we extracted features capturing lower-level speech features. Using the stimulus transcript as input, we created one-hot vectors for phonetic and articulatory features. Phoneme classes (39 total classes) were obtained from the Carnegie Mellon Pronouncing Dictionary (The CMU Pronouncing Dictionary, n.d.). We further classified the phonemes based on their place of articulation (9 classes), manner of articulation (9 classes), and voiced or voiceless status (3 classes), according to the general American English consonants of the International Phonetic Alphabet. Given that each word consists of multiple phonemes, we averaged the one-hot vectors for all phonetic and articulatory features for each word.

      Second, we extracted linguistic features using spaCy (Honnibal et al., 2020), including part of speech (17 classes), tag (50 classes), function or content word (3 classes), dependency (45 classes), whether the word is an alpha character (binary), and whether the word is a stop word (binary). We also extracted prefix (30 classes) and suffix (44 classes) information using the Cambridge Dictionary. We constructed one-hot vectors for each multi-class feature and one-dimensional vectors for each binary feature.

      Third, for each word, we obtained word frequency from the Google Web Trillion Word Corpus (Brants & Franz, 2006) and from our own dataset.

      Fourth, we generated static word embeddings of dimension 50 using GloVe (Pennington et al., 2014).

      We then built encoding models in the same way as the contextual embeddings for each of the three categories of speech features, all speech features concatenated, and the GloVe embeddings. To control for the different dimensions of the embeddings, we also standardized all embeddings to the same size (50 dimensions) using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression. For both ridge and OLS encoding, our contextual embeddings from LLMs showed significantly better performance than the classic speech features and GloVe embeddings.

      We have added the following text to our manuscript and updated our Figures S4, S5, Table S1, and the methods section:

      “To establish a general baseline for encoding performance, we built encoding models using embeddings from the SMALL model with randomly initialized weights. The trained SMALL model exhibits significantly higher encoding performance across all layers compared to the untrained SMALL model (Fig. S4). We also assessed the encoding performance of contextual embeddings from LLMs against classic speech features and static GloVe embeddings (Table S1). The SMALL and XL embeddings achieved markedly higher encoding correlations than the speech features and GloVe embeddings (Fig. S5).”

      (3) The finding that peak encoding performance tends to occur in relatively earlier layers in larger models is somewhat surprising and requires further explanation. Since more layers mean more parameters, if the later layers diverge from language processing in the brain, it raises the question of what aspects of the larger models make them more brain-like.

      We thank the reviewer for this insightful comment; this point was also highlighted by Reviewer 1. We agree that this result is somewhat surprising, and we aim to provide a more detailed explanation in the revised manuscript. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models.

      Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Reviewer #3 (Public review):

      This manuscript studies the connection between neural activity collected through electrocorticography and hidden vector representations from autoregressive language models, with the specific aim of studying the influence of language model size on this connection. Neural activity was measured from subjects who listened to a segment from a podcast, and the representations from language models were calculated using the written transcription as the input text. The ability of vector representations to predict neural activity was evaluated using 10-fold cross-validation with ridge regression models.

      The main results are that (as well summarized in section headings):

      (1) Larger models predict neural activity better.

      (2) The ability of language model representations to predict neural activity differs across electrodes and brain regions.

      (3) The layer that best predicts neural activity differs according to model size, with the "SMALL" model showing a correspondence between layer number and the language processing hierarchy.

      (4) There seems to be a similar relationship between the time lag and the ability of language model representations to predict neural activity across models.

      Strengths:

      (1) The experimental and modeling protocols generally seem solid, which yielded results that answer the authors' primary research question.

      (2) Electrocorticography data is especially hard to collect, so these results make a nice addition to recent functional magnetic resonance imaging studies.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) The interpretation of some results seems unjustified, although this may just be a presentational issue.

      (a) Figure 2B: The authors interpret the results as "a plateau in the maximal encoding performance," when some readers might interpret this rather as a decline after 13 billion parameters. Can this be further supported by a significance test like that shown in Figure 4B?

      We agree that this could be a subjective interpretation, so we conducted an additional analysis. We performed paired two-sided t-tests between best layer encoding performances averaged across electrodes (df = 159 electrodes), comparing all models with larger models. We found that after 13 billion parameters, only the encoding performance for OPT-66B, the largest model in the OPT family, is significantly worse than the encoding performance of some other smaller models, supporting the claim that the maximal encoding performance declines after 13 billion parameters. However, we did not find conclusive statistical evidence of a decline in encoding performance for other model families.

      We have added the statistical results as Supplementary Figure 1.

      We have also modified the following text in the manuscript:

      “We also observed a plateau in the maximal encoding performance, occurring around 7 billion parameters (Fig. 2B), with a decline in performance for the OPT-66B model (Fig. S1).”

      (b) Figure S1A: It looks like the drop in PCA max correlation is larger for larger models, which may suggest to some readers that the same trend observed for ridge max correlation may not hold, contra the authors' claim that all results replicate. Why not include a similar figure as Figure 2B as part of Figure S1?

      PCA is an unsupervised dimensionality reduction technique and may discard model features with small eigenvalues that nonetheless contribute to encoding performance. Ridge regression, a supervised method, can capitalize on these features. We suspect that this is why there appears to be a larger drop in model performance for larger models with PCA than with ridge regression. We replicated the logarithmic relationship between model size and encoding performance using PCA and ordinary least-squares (OLS) regression encoding models. We have updated Supplementary Figure 2.

      (2) Discussion of what might be driving the main result about the influence of model size appears to be missing (cf. the authors aim to provide an explanation of what seems to drive the influence of the layer location in Paragraph 3 of the Discussion section). What explanations have been proposed in the previous functional magnetic resonance imaging studies? Do those explanations also hold in the context of this study?

      We suspect that the increased expressivity of larger models - that is, their improved sensitivity to nuanced structure in natural language - yields improved alignment to brain activity (given large enough samples of brain activity) (Antonello et al., 2023). This effect persists even when dimensionality is tightly controlled in our PCA-based analysis, indicating that the improved alignment with the brain is not a modeling artifact of dimensionality alone, but results from the structural representations learned by these larger models.

      We have added the following text to our Discussion section:

      “We suspect that the improved alignment with brain activity in larger models is driven by their increased expressivity and sensitivity to nuanced linguistic structure present in large-scale naturalistic datasets (Antonello et al., 2023).”

      (3) The GloVe-based selection of language-sensitive electrodes (at least to me) isn't explained/motivated clearly enough (I think a more detailed explanation should be included in the Materials and Methods section). If the electrodes are selected based on GloVe embeddings, then isn't the main experiment just showing that representations from larger language models track more closely with GloVe embeddings? What justifies this methodology?

      We selected electrodes based on previously established methods (Goldstein et al., 2022). Our use of GloVe embeddings for electrode selection does not imply that larger language model representations are simply more closely aligned with GloVe embeddings. On the contrary, contextual embeddings from LLMs, which incorporate the word’s previous context, consistently outperform static embeddings like GloVe or word2vec (Fig. S3). Selecting electrodes using LLM embeddings would likely result in a slightly different, potentially larger set of electrodes (Goldstein et al., 2022), but would be more circular (Kriegeskorte et al., 2009). The GloVe-based electrode selection represents a more conservative approach by identifying words encoding linguistic content without biasing the selection directly toward any LLMs.

      We have added the following text to our Method section:

      “We used GloVe embeddings for electrode selection to avoid biasing our main results toward a particular LLM.”

      (4) (Minor weakness) The main experiments are largely replications of previous functional magnetic resonance imaging studies, with the exception of the one lag-based analysis. Is there anything else that the electrocorticography data can reveal that functional magnetic resonance imaging data can't?

      We thank the reviewer for this thoughtful question. While we agree that a key contribution of our work corroborates previous fMRI findings, we would argue that using ECoG is not merely a replication but a crucial validation and extension of that work. It is important to validate these effects across distinct measurement modalities. In our work, we further observed a novel trend where the peak encoding performance tends to occur in relatively earlier layers for larger models. This is supported by recent studies suggesting that later layers of large LLMs may not significantly contribute to benchmark performance (Csordás et al., 2025). While scaling has been an effective method to improve LLM performance, including in encoding models, future research should explore the potential underutilization of the later layers as models scale.

      Furthermore, ECoG data offers temporal resolution on the order of milliseconds, far superior to fMRI’s. Although we did not observe a relationship between model size and temporal lags in this study, future work should investigate the temporal dynamics of encoding that are accessible with ECoG (Goldstein, Ham, et al., 2025; Goldstein, Wang, et al., 2025).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you to the authors for the fun and personally useful read.

      I see in Supplementary Figure 1 the authors show a comparison of the performance between OLS vs. Ridge regression. Is the OLS model the only one that is working over PC features, or are both models using PC features? The current text is a bit unclear. My current understanding is that the comparison is between (OLS + PCA) and (Ridge with no PCA), but I am not sure.

      The OLS model is the only one that works over PC features, following previous methods (Goldstein et al., 2022).

      We have added the following text to our Results and Methods section for clarity:

      “To control for the different embedding dimensionality across models, we standardized all embeddings to the same size using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression, replicating the logarithmic relationship but with significantly lower encoding performance overall (Fig. S2). The PC features are used by the OLS models only.”

      Clarification in the text would be appropriate. If this is the correct understanding, the authors should note in the main text that the ridge approach is more effective than the PCA approach, which is still the dominant approach to building linear encoding models in the field for some unjustifiable reason.

      We thank the reviewer for pointing out the confusion. We have updated Supplementary Figure 2.

      How were the alpha values for ridge regression determined? Do you use the same ridge parameter for all electrodes or fit a different parameter for each electrode? This is not mentioned anywhere.

      The alpha values are determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022). Specifically, we perform a grid search over cross-validation folds in the training data to find the best-performing alpha. The alpha parameter is specific to each ridge regression model, meaning each fold, lag, and electrode has a different alpha parameter.

      We have added the following text to our manuscript:

      “For each ridge regression model (for each fold, lag, and electrode), the alpha parameter is determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022).”

      It's not entirely clear to me how the authors handle tokens that do not terminate in words (such as the "there" + "'s" example in the text). My current reading of the text is that authors essentially ignore these half-word embeddings, doing one forward pass per word, rather than per token, but the current text is somewhat ambiguous.

      If a word is tokenized into several tokens, like “there” and “‘s”, we average the token embeddings to get a word embedding.

      We have added the following text to our Method section:

      “To facilitate a fair comparison of the encoding effect across different models, we aligned all tokens in the story across all models. We averaged the token embeddings if a word is split into multiple tokens, resulting in one embedding per word for each model.”

      The authors describe the scaling relationship they find as a "log-linear" relationship. I believe this is a misnomer derived from the original paper describing this relationship in fMRI as log-linear (Antonello et al.) The correct term is simply "logarithmic", and for what it's worth, the authors of the original fMRI work have made this correction as well.

      Thank you! We have made this correction.

      Is the data publicly available? If not, there should be some basic justification as to why (consent reasons, etc.).

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      The authors assert that ECoG has "superior spatiotemporal resolution". While this is unquestionably true for temporal resolution, the story is a bit more complicated for spatial resolution, where ECoG has far less cortical coverage than fMRI. Perhaps this sentence should be revised.

      Thank you for pointing out the typo! We have changed it to “superior temporal resolution”.

      Minor Points:

      The bolded title of Figure 3 probably shouldn't be bolded, as this is just actually the title of Figure 3A.

      Fixed.

      Figure 4d is has a typo: "Best Encoidng Layer".

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      The authors could consider adding control regressors such as word rate, word frequency, phonetic features, and syntactic features like node counts, as well as control LLMs of comparable size to serve as baselines. The authors could also include correlation analyses of the embeddings from different layers of the same LLM to further illustrate how distinct the layers are within the models.

      We have added untrained LLM embeddings as a baseline and included a comparison of encoding models between LLM contextual embeddings and classical speech features. We have also performed some preliminary correlation analyses of embeddings. In some models, we found evidence of the “two-phase abstraction process” (Cheng & Antonello, 2024). However, the result is inconclusive across different LLM families. Since each LLM layer accesses and modifies the residual stream (Elhage et al., 2021), the embeddings across layers are inherently correlated. Future work could instead explore the isolated transformations within each layer to illustrate the distinct information across layers (Kumar et al., 2024).

      The analysis codes and data should be made available.

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      Reviewer #3 (Recommendations for the authors):

      Most of my concrete recommendations are in the public review. Below are some additional minor ones:

      (1) Introduction: "Remarkably, these models learn from much the same shared space as humans: from real-world language generated by humans."

      I think this is an extremely strong claim due to e.g. the different nature of child-directed speech vs. written text corpora, the lack of multimodality and grounding in language models, etc. I might suggest re-wording this sentence or removing it entirely.

      We thank the reviewer for their suggestion! We have removed the sentence from the manuscript.

      (2) Introduction: "EleutherAI, n.d." reference for GPT-Neo

      GPT-NeoX-20B has an associated paper, which the authors might cite instead: https://aclanthology.org/2022.bigscience-1.9

      Thank you! We have added the reference for GPT-NeoX-20B (Black et al., 2022).

      (3) Figure 4D: Encoidng -> Encoding

      Fixed.

      (4) Materials and Methods, Contextual embeddings: "except for GPT-Neox-20b, which assigns additional tokens to whitespace characters."

      What do the authors mean by "additional tokens to whitespace characters?" The tokenizer for GPT-NeoX-20B works in much the same way as that of GPT-Neo, just with a different vocabulary set.

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The quick brown fox jumps over the lazy dog.").input_ids) ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      >>> t2.convert_ids_to_tokens(t2("The quick brown fox jumps over the lazy dog.").input_ids)

      ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      If the authors are referring to Ġ as the "additional token to whitespace characters," then these are in all other tokenizers as well (not only that for GPT-Neo, but also those for GPT-2 and OPT).

      We agree that “additional tokens to whitespace characters” is an oversimplification. The GPT-Neo model family, which includes the 125M, 1.3B, and 2.7B models, utilizes the same Byte Pair Encoding (BPE) tokenizer as GPT-2. This common tokenizer has a vocabulary size of 50,257 tokens, providing compatibility and seamless integration across the models.

      The GPT-NeoX-20B model introduces a modified tokenizer to address limitations observed in the GPT-2 tokenizer (Black et al., 2022). As detailed in Section 3.2, this new tokenizer incorporates a few key improvements:

      (1) New BPE tokenizer: A more general-purpose BPE tokenizer was trained using the Pile dataset.

      (2) Space Delimitation: Unlike the GPT-2 tokenizer, which treats tokenization at the start of a string as a non-space-delimited token, the GPT-NeoX-20B tokenizer applies consistent space delimitation regardless. This change resolves inconsistencies related to the presence of prefix spaces in the tokenization input.

      (3) Whitespace Handling: The tokenizer includes tokens for repeated space characters (up to 24 consecutive spaces), enhancing efficiency in tokenizing text with substantial whitespace, such as program source code or LaTeX documents.

      These modifications result in the GPT-NeoX-20B tokenizer representing the Pile validation set with approximately 10% fewer tokens than the GPT-2 tokenizer. This efficiency gain is particularly beneficial for processing texts with extensive whitespace.

      In our analysis, we extracted embeddings by setting `add_prefix_space = True` to all tokenizers, so space delimitation does not result in tokenizer differences. We highlight here examples of the other two tokenizer differences using the Huggingface `AutoTokenizer`:

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The Downing Street.").input_ids) ['The', 'ĠDowning', 'ĠStreet']

      >>> t2.convert_ids_to_tokens(t2("The Downing Street.").input_ids)

      ['The', 'ĠDown', 'ing', 'ĠStreet']

      >>> t1.convert_ids_to_tokens(t1("Hello !").input_ids)

      ['Hello', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ!']

      >>> t2.convert_ids_to_tokens(t2("Hello !").input_ids)

      ['Hello', ' ', '!']

      More examples showing the differences between the GPT-2 tokenizer and the GPT-NeoX-20B tokenizer can be found in Appendix F: Tokenizer Analysis (Black et al., 2022).

      We have added the following text to our manuscript for simplicity:

      “All models within the same model family adhere to the same tokenizer convention, except for GPT-Neox-20B, which utilizes a different tokenizer (Black et al., 2022).”

      References

      Ameisen, E., Lindsey, J., Pearce, A., Gurnee, W., Turner, N. L., Chen, B., Citro, C., Abrahams, D.,  Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C.,  Cunningham, H., Henighan, T., Jermyn, A., Jones, A., … Batson, J. (2025). Circuit Tracing:  Revealing Computational Graphs in Language Models. Transformer Circuits Thread. https://transformer-circuits.pub/2025/attribution-graphs/methods.html

      Antonello, R., & Huth, A. (2024). Predictive coding or just feature discovery? An alternative account of why language models fit brain data. Neurobiology of Language (Cambridge, Mass.), 5(1), 64–79.

      Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 2023. https://doi.org/10.48550/ARXIV.2305.11863

      Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell,  K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., & Weinbach, S. (2022). GPT-NeoX-20B: An Open-Source Autoregressive Language Model.  Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models. Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, virtual+Dublin. https://doi.org/10.18653/v1/2022.bigscience-1.9

      Brants, T., & Franz, A. (2006). Web 1T 5-gram Version 1 [Dataset]. Linguistic Data Consortium. https://doi.org/10.35111/CQPA-A498

      Cantlon, J. F., & Piantadosi, S. T. (2024). Uniquely human intelligence arose from expanded information capacity. Nature Reviews Psychology, 3(4), 275–293.

      Chemla, E., D’Ascoli, S., Diego-Simón, P., King, J.-R., & Lakretz, Y. (2024). A Polar coordinate system represents syntax in large language models. Advances in Neural Information Processing Systems 37, 105375–105396.

      Cheng, E., & Antonello, R. J. (2024). Evidence from fMRI supports a two-phase abstraction process in language models. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2409.05771

      Csordás, R., Manning, C. D., & Potts, C. (2025). Do language models use their depth efficiently?  In arXiv [cs.LG]. https://doi.org/10.48550/ARXIV.2505.13898

      Dupré la Tour, T., Eickenberg, M., Nunez-Elizalde, A. O., & Gallant, J. L. (2022). Feature-space selection with banded ridge regression. In bioRxiv. https://doi.org/10.1101/2022.05.05.490831

      Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., & Olah, C. (2022). Toy Models of Superposition. Transformer Circuits Thread.

      Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A.,  Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A.,  Kernion, J., Lovitt, L., Ndousse, K., … Olah, C. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread.

      Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., & Wang, Z. (2024). Not all Layers of LLMs are Necessary during Inference. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.02181

      Friederici, A. D., & Becker, Y. (2025). The core language network separated from other networks during primate evolution. Nature Reviews. Neuroscience, 26(2), 131–132.

      Goldstein, A., Ham, E., Schain, M., Nastase, S. A., Aubrey, B., Zada, Z., Grinstein-Dabush, A.,  Gazula, H., Feder, A., Doyle, W., Devore, S., Dugan, P., Friedman, D., Brenner, M., Hassidim, A., Matias, Y., Devinsky, O., Siegelman, N., Flinker, A., … Hasson, U. (2025). Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications, 16(1), 10529.

      Goldstein, A., Wang, H., Niekerken, L., Schain, M., Zada, Z., Aubrey, B., Sheffer, T., Nastase, S. A., Gazula, H., Singh, A., Rao, A., Choe, G., Kim, C., Doyle, W., Friedman, D., Devore, S., Dugan, P., Hassidim, A., Brenner, M., … Hasson, U. (2025). A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02105-9

      Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A.,  Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369–380.

      Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., & Roberts, D. A. (2024). The Unreasonable Ineffectiveness of the Deeper Layers. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.17887

      Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences of the United States of America, 109 Suppl 1(supplement_1), 10661–10668.

      Hewitt, J., & Manning, C. D. (2019). A Structural Probe for Finding Syntax in Word Representations. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North (pp. 4129–4138). Association for Computational Linguistics.

      Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength Natural Language Processing in Python.

      Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F., & Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience, 12(5),  535–540.

      Kumar, S., Sumers, T. R., Yamakoshi, T., Goldstein, A., Hasson, U., Norman, K. A., Griffiths, T. L., Hawkins, R. D., & Nastase, S. A. (2024). Shared functional specialization in transformer-based language models and the human brain. Nature Communications, 15(1), 5523.

      Linzen, T., & Baroni, M. (2021). Syntactic Structure from Deep Learning. Annual Review of Linguistics, 7(1), 195–212.

      Manning, C. D., Clark, K., Hewitt, J., Khandelwal, U., & Levy, O. (2020). Emergent linguistic structure in artificial neural networks trained by self-supervision. Proceedings of the National Academy of Sciences of the United States of America, 117(48), 30046–30054.

      Millet, J., Caucheteux, C., Orhan, P., Boubenec, Y., Gramfort, A., Dunbar, E., Pallier, C., & King, J.-R. (2023). Toward a realistic model of speech processing in the brain with self-supervised learning. NeurIPS 2022. https://doi.org/10.48550/ARXIV.2206.01685

      Pavlick, E. (2022). Semantic structure in deep learning. Annual Review of Linguistics, 8(1),  447–471.

      Pennington, J., Socher, R., & Manning, C. (2014). Glove: Global vectors for word representation.  Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar. https://doi.org/10.3115/v1/d14-1162

      Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the United States of America, 118(45), e2105646118.

      The CMU Pronouncing Dictionary. (n.d.). Retrieved May 27, 2025, from http://www.speech.cs.cmu.edu/cgi-bin/cmudict

      Vaidya, A. R., Jain, S., & Huth, A. G. (2022). Self-supervised models of audio effectively explain human cortical responses to speech. ICML 2022. https://doi.org/10.48550/ARXIV.2205.14252

      Zada, Z., Goldstein, A., Michelmann, S., Simony, E., Price, A., Hasenfratz, L., Barham, E., Zadbood,  A., Doyle, W., Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Nastase, S. A., & Hasson, U. (2024). A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations. Neuron, S0896627324004604. Zada, Z., Nastase, S. A., Aubrey, B., Jalon, I., Michelmann, S., Wang, H., Hasenfratz, L., Doyle, W.,  Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Goldstein, A., & Hasson, U. (2025). The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension. Scientific Data, 12(1), 1135.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We are grateful to you and the reviewers for your careful and positive assessment.  In response to the comment below from reviewer #2, we have expanded Figure 1- figure supplement 1 to show the results (images plus quantification) of ERG and PU.1 immunostaining together with DAPI staining that underly the calculation of the percent on non-endothelial cells that are recombined by Cdh5-CreER.

      However, I found their explanation as to how they arrived at only 18% of the fibroblasts being recombined confusing. Mostly because it looks like there are many GFP+/ERG- cells in panels A & D, these would be the recombined fibroblasts and AB cells. Considering that Cdh5 gene is pretty broadly expressed across LPM, the prediction is the recombination rate would be higher (recognizing mice strains vary and this is inducible Cre).

      Ideally, they would perform recombination analysis with a TF expressed by all fibroblasts in combination with Erg (like Foxc1). However, a potentially simpler approach could be to quantify this using DAPI and Erg in existing images, of the total DAPI+, what are GFP+/DAPI+/Erg- (fibroblasts) vs GFP+/ERG+/DAPI+ (endothelial). The figure would be improved by adding DAPI to one set of panels with GFP/ERG, this would show a lot of DAPI+/GFP- cells, encompassing CD206+ cells and non-recombined fibroblasts.

      Thank you for overseeing this manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This well-conceived manuscript investigates the mechanisms that shape the chromatin landscape following fertilization, using the Drosophila embryo as a model system. Importantly, the authors revisit conflicting data using new approaches and analysis to show that the silent H3K27me3 mark deposited by PRC2 is established de novo in the embryo in coordination with the slowing of the nuclear division cycle and activation of zygotic transcription. Unexpectedly, they demonstrate that the transcription factor GAF is not required for the deposition of this mark, but that the well-studied pioneer factor Zelda, which is required for widespread gene expression, is required for H3K27me3 deposition at a subset of regions. The experiments are rigorously performed, and interpretations are clear. Strengths of this manuscript include the rigor of the experimental design, careful analysis, and well-supported conclusions. Some additional citations, analysis, and broadening of the Discussion section to include additional models and data would further strengthen this manuscript.

      We appreciate the reviewer’s positive assessment and in revision we have revised the Discussion to clarify some of the mechanistic insights of the work as well as including references to related studies in the zebrafish model system.

      Reviewer #2 (Public review):

      Summary:

      Epigenetic silencing of target genes by the Polycomb pathway is central to maintenance of cell fates during development and depends on repressive chromatin states involving Polycomb complexes and histone modifications. However, the mechanisms by which these chromatin states are built at the earliest stages of development are unclear. Here, Gonzaga-Saavedra and colleagues use the premier experimental system for studying Polycomb gene regulation, Drosophila development, to investigate when Polycomb domains emerge and how they are assembled. Using a combination of CRISPR gene editing, imaging, and genomic profiling, they determine that while H3K27me3 is initially present in the first nuclear cycles, it quickly dissipates and does not re-emerge until mid-nuclear cycle 14, during the major wave of zygotic genome activation (ZGA). This finding helps resolve current discrepancies in the field, informs potential mechanisms of transgenerational inheritance, and indicates that repressive Polycomb domains are built de novo on target genes in embryogenesis. The authors then set out to examine how Polycomb domains are built. Through live imaging and immunofluorescence, they determine that the histone H3K27 methyltransferase, E(z), is present in nuclei at high levels throughout cleavage and blastoderm stages. By contrast, they determine that several Polycomb proteins that bind PREs (cis elements that demarcate Polycomb targets in the genome) are absent from early cleavage nuclei and progressively increase following nuclear cycle 10. These findings suggest that the absence of H3K27me3 in early embryos may be due to failure to assemble functional Polycomb complexes at target genes. Lastly, the authors test the requirement of two transcription factors with important roles in ZGA, GAF, and ZLD. Despite binding to many PREs and regulating chromatin accessibility in early embryos, they find that GAF is largely dispensable for the emergence of H3K27me3 domains. On the other hand, they find that the pioneer factor ZLD is required for proper H3K27me3 emergence; in its absence, some Polycomb domains accumulate greater levels of H3K27me3, whereas other Polycomb domains accumulate less H3K27me3.

      Strengths:

      The strengths of this study are manifold. It studies an important topic with broad interest to the chromatin and epigenetics fields. It is well-written with detailed method descriptions. In addition, the experimental design and rigor of execution are exceptional despite working with very small amounts of biological material. Example strengths include that the Polycomb proteins studied were tagged with the same epitope, permitting direct quantitative comparisons in imaging and in genomics experiments. Microscopy studies are quantified and performed both via live imaging and via immunofluorescence. The microscopy studies reinforce and extend conclusions made via ChIP. Sophisticated loss-of-function analyses allow for direct mechanistic tests of Polycomb domain emergence.

      Weaknesses:

      Overall, the study is quite strong already, but it can be further strengthened in several ways. First, several conclusions should be refined based on the data presented. Second, the extent to which ZLD is important for initiating Polycomb domain formation should be made clearer. Third, additional genomic profiling experiments are needed to provide insight into models explaining why H3K27me3 is absent prior to NC14.

      We are grateful for the reviewer’s thorough and supportive comments. We have revised certain assertions and conclusions for objectivity. For the point about providing “insight into models explaining why H3K27me3 is absent prior to NC14,” we have a separate study that addresses this issue directly (Degen, Gonzaga-Saavedra, and Blythe, bioRxiv 2025, in press). In summary, we find evidence that a maternal PcG imprint is indeed maintained through cleavage divisions, albeit through lower-order methylation states (maximally, H3K27me2). We chose not to include these additional results in this manuscript to maintain the focus of this study on ZGA. Our revision of the manuscript includes a reference to this associated study in the Discussion.

      Reviewer #3 (Public review):

      Gonzaga-Saavedra et al report an analysis on genomic binding of Polycomb group proteins, and of H2Aub1 and H3K27me3 domain formation in the early Drosophila embryo. Using carefully staged embryos during the nuclear cycles (NC) leading up to the cellular blastoderm stage, the authors provide compelling evidence that H3K27me3 domains at PcG target genes are only established during NC14 and do not exist in NC13. In contrast, H2Aub1 domains already start to appear during NC13. The authors show that E(z), the catalytic subunit of the H3K27 histone methyltransferase PRC2, is readily detected in interphase nuclei during the rapid nuclear divisions in pre-blastoderm embryos. In contrast, the DNA-binding proteins Pho, Cg, and GAF that are known (Pho) or have been postulated (Cg, GAF) to anchor PRC2 and PRC1 to Polycomb Response Elements (PREs) in Polycomb target genes only start to show nuclear localization from NC10 onwards with gradually increasing nuclear concentrations, reaching a maximum during NC14. These data strongly corroborate the simple, straightforward view that targeting of PRC2 and PRC1 to PREs by sequence-specific DNA-binding proteins is a prerequisite for the formation of H3K27me3 and H2Aub1 domains at Polycomb target genes.

      The authors then explore the potential role of GAF/Trl in this process. They find that in embryos depleted of GAF/Trl, H3K27me3 domain formation is largely unperturbed.

      The authors also depleted the pioneer factor Zelda (Zld) and found that removal of Zld results in a more complex outcome. Zelda appears to counteract the accumulation of H3K27me3 at the Polycomb targets eve and zen, but also appears to be required for effective H3K27me3 domain formation at Polycomb targets such as amos or atonal.

      This is a very thorough study that reports data of superior technical quality that are highly relevant for the field. The study by Gonzaga-Saavedra et al extends and strengthens previous work from the labs of Eisen (Li et al, eLife 2014) and Zeitlinger (Chen et al, eLife 2013) to convincingly demonstrate that Polycomb domain formation in the early embryo occurs during ZGA but that such domains do not exist prior to ZGA. This should now finally put to rest earlier claims by the Iovino lab (Zenk et al, Science 2017) that H3K27me3 domains present in the zygote nucleus would be propagated and partially maintained during the rapid nuclear cleavage cycles and serve as seeds for H3K27me3 domain formation during ZGA.

      The experiments analyzing H3K27me3 domain formation in embryos depleted of GAF/Trl or Zelda will be of great interest to the field.

      We thank the reviewer for recognizing the strength of our data and conclusions, and we agree that our results help settle conflicting claims in the field. We have emphasized Zelda’s context-dependent effects more clearly in the revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      It would strengthen the manuscript to more fully acknowledge and discuss related work in the field. Addressing the caveats in the functional analyses, either through editorial clarification or additional experiments, would also improve the study.

      For recommendations to the authors, comments from each reviewer are listed below.

      Reviewer #1 (Recommendations for the authors):

      Some additional analysis would clarify the relationship between pPREs and transcription factors.

      Because of the limitations of depleting Pho, the model that nuclear levels of Pho, Cg, and GAF regulated E(z) activity is purely correlative, as GAF depletion did not change H3K27me3 distribution. As it stands, it is possible that this NC14 nuclear enrichment of these factors is not relevant. The statement on lines 295-297 regarding the correlation between re-establishment of the modification state and nuclear localization of Pho, Cg, and GAF is true, but a bit misleading since there is no evidence to support the necessity of these factors for H3K27me3 establishment. Other models remain possible, and the discussion should be toned down to account for this. Furthermore, the only factor that is shown to influence H3K27me3 is Zelda, which does not show an increase in nuclear localization at NC14.

      We thank the reviewer for highlighting this issue. As the reviewer indicates, we lack definitive mechanistic evidence that limited nuclear localization of any recruitment/nucleating factor is limiting for H3K27me3 deposition. However, our data are, we feel, definitive in terms of demonstrating that nucleation from PREs and PRE-like regions arises at mid-NC14 for H3K27me3, and in late cleavages for H2Aub. Ideally, we would test each known nucleating factor for a necessary role in mediating this activity. We have chosen not to include such measurements in this manuscript because a proper mechanistic treatment of each factor (Pho, Cg, and others) would need to be extensive, and we feel better suited for an independent study. We have added language to the “Limitations of the Study” section to reflect the remaining need to identify the key nucleating factors responsible for establishment of the zygotic PcG landscape.

      We re-read the lines the reviewer suggested were misleading and we respectfully disagree. The lines (“Taken together, these observations are consistent with a model where…re-establishment of this modification state is restricted to late cleavage divisions by limiting nuclear localization of nucleating factors such as Pho, Cg, and GAF.”) are expressing a hypothesis/model, and in our opinion these lines are suitably framed as to not be misleading.

      Additional analysis/discussion regarding the relationship between Zelda, H3K27me3, CBP, and paused polymerase would provide further clarity into how Zelda might promote this methylation. It would be useful to discuss how the various classes defined correlate with enhancers versus promoters, and also the gene expression of the underlying gene. It would be clarifying to discuss how the H3K27me3 at these Zelda-dependent regions relates to gene silencing since many of the genes depend on Zelda for expression. Are these genes expressed prior to NC14 and then silenced at this time point? How do H3K27ac levels, which Zelda promotes through recruitment of CBP, relate to the various pPRE classes? The authors should also consider the report from the Mannervik lab that CBP is instrumental in promoting H3K27me3 (Hunt, Boija, Mannervik et al. Mol Cell 2022 82:3580-3597) and the relationship with paused RNA Pol II. Given this possible connection, it could be useful to consider that paused polymerase is also first evident at NC13/14. Overlaying the pPREs with paused polymerase from Chen et al. 2013 eLife (current citation 30) could be informative. At a minimum, a discussion of these additional mechanisms and the implications of the data in Hunt et al. would strengthen the Discussion section.

      We thank the reviewer for this comment. We agree that additional analysis of the relationship between Zelda, H3K27me3 and CBP would be essential for providing mechanistic clarity on the role of Zelda for putting these loci into play for apparently either positive or negative regulation. At present, we feel that extensive additional analysis, including work with CBP, PRC1, and nucleating factors such as Pho would be better suited for a future study.

      H2Aub is clear earlier in development than H3K27me3, and in mice and zebrafish, it promotes PRC2-mediated H3K27me3 (Hickey et al. eLife 2022 doi: 10.7554/eLife.67738, citations 86, 89). As such, it remains possible that this mark is instructive for the H3K27me3 deposition observed. As such, a bit more analysis of where H2Aub is deposited and how it overlaps with pPREs might help determine whether similar mechanisms could be important in Drosophila.

      We suspect that similar mechanisms are important in Drosophila as well. The omission of the Hickey…Cairns reference was an oversight in the original document. We have revised this sentence to refer to both mouse and zebrafish and have added the citation.

      Prior work has noted the sudden increase in GAF nuclear concentration. Please cite Dima and Reeves. Development 2025 152:dev204460 in support of the observations shown in Figure 3D.

      Thank you. Yes, this article was published shortly after we submitted this manuscript for review and we have now added it to reflect its independent replication of the GAF nuclear concentration result.

      It is not clear that Figure 4 warrants an entirely new figure, since the conclusions drawn are similar to/the same as Figure 3. Perhaps change to a supporting figure?

      We agree that this figure was repetitive. We have now moved it to a figure supplement of Figure 3.

      The authors have developed a powerful modified ChIP protocol that enables them to perform the experiment on the equivalent of 10 embryos! This is not highlighted in the manuscript, despite the vast improvement this provides. The authors should highlight this in the manuscript, unless a separate manuscript describing this technique is being written/published. Regardless, this is a very exciting protocol.

      Thank you. A methods paper describing this approach is in preparation.

      Minor:

      (1) Line 49-53: clarify that, as opposed to the mechanisms described earlier in the paragraph, these mechanisms are specific to Drosophila.

      Done.

      (2) Line 61: cite Sun et al. and Schulz et al. (current citations 60 and 61) since these papers demonstrated the pioneering function of Zelda.

      Done.

      (3) Line 95: a word seems to be missing. Perhaps "approach (STAN) to identify a set of PcG "domains" from our"?

      We have made this revision.

      (4) Line 217: Calling an embryo a "specimen" is odd. Can you just say all embryos?

      Ok.

      (5) Line 316: GAF is encoded by Trithorax-like, not Trithorax-related.

      Revised.

      (6) Figure 5: The arrowheads and asterisk are so small that they are nearly impossible to see when the figure is printed. Please make it larger and perhaps use a color that stands out more.

      We have made the arrowheads and asterisk larger and changed the color to yellow to improve visibility.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding weakness 1, the title states "nucleation-dependent propagation". This is an overstatement of the study's conclusions because these features are inferred and not directly tested here. This is an easy fix with text revisions.

      We acknowledge that we have not directly tested the role of specific nucleating factors for the process of nucleation-dependent propagation. This is reflected in the Discussion text. We have added to the “Limitations of the Study” section the following text: “Finally, we acknowledge that further loss-of-function analysis will be necessary to determine the key nucleating factors responsible for the initial establishment of the zygotic H2Aub and H3K27me3 landscape.” However, we feel that we have demonstrated definitively that, by mid-NC14, nucleation of H3K27me3 sites is first detected genome-wide. As such, we have left the title as-is.

      (2) Also, regarding weakness 1, the abstract makes additional overstatements. These are easy fixes with text revisions.

      (a) "A large subset of targets requires ZLD..." (line 24). This phrase makes it sound like the majority of Polycomb domains depend on ZLD, which I am not sure is accurate.

      Thank you for pointing this out. We agree with this assessment and have revised the abstract to remove the word “large” so that now the sentence reads “; a subset of targets requires Zelda…”

      (b) "...requires ZLD not for PcG factor recruitment" (same sentence). This sentence should state "E(z)" instead of "PcG factor" because E(z) was the only one tested in ZLD mutants.

      We agree as well and have made the requested change to exchange “PcG factor” to “E(z)”.

      (c) "to license a loaded PRE" (same sentence). Whether PREs are fully loaded in ZLD mutants was not tested. In addition, see comments below for feedback on the use of the term "license."

      We have revised this to read “to license an E(z)-loaded PRE,” to keep consistent with the above suggested change. We also removed the mention of “H2Aub” because now the PRE-loading acknowledges we have only measured E(z), although we see effects on both modifications. We acknowledge, in light of Reviewer 3’s comment, that we have not directly measured whether PRC1 still can bind to certain PREs in the absence of Zelda like we see with E(z)/PRC2.

      (3) Also regarding weakness 1, the authors set up a dichotomy for ZLD's role in initiating Polycomb domain formation, either acting as a pioneer or as a licensing factor. However, based on the data presented (e.g. the atonal browser shot), it appears that chromatin accessibility is lost in both sub-classes of H3K27me3 domains that depend on ZLD, meaning that ZLD's role as a pioneer would explain both sub-classes, and that a licensing role is not supported by the data. To assess whether there are pioneer-independent roles of ZLD in the emergence of Polycomb domains, it may help to examine the role of chromatin accessibility changes directly (e.g., is there a substantial fraction of changes in H3K27me3 domains or E(z) peaks that cannot be attributed to changes in chromatin accessibility in ZLD mutants?). The authors should refine their language or provide additional support for the non-pioneering role of ZLD.

      We acknowledge that the pioneer/licensing distinction is still not clear from a mechanistic perspective and that future work will be needed to address this issue. We do not mean to imply that the licensing is necessarily independent of pioneering, rather that at sites like atonal, clearly chromatin accessibility is not the only job that Zelda performs. There are at least two possible mechanisms: 1) it is all about pioneering, and in the absence of accessible chromatin, some other factor required for stimulating H3K27me3 deposition is unable to bind; or 2) it is all about Zelda, and in mutants Zelda both does not confer accessible chromatin, and also does not stimulate H3K27me3 deposition through whatever means. For either mechanism, the remarkable feature is that E(z) still localizes to its genomic target and requires additional information to deposit H3K27me3. This is distinct from the other class (represented by amos in Figure 6 in the final manuscript) where loss of accessibility correlates with loss of E(z) –and presumably PRC2– binding. To explain this activity, we have invoked the term “licensing” which we feel captures the effect of Zelda (whether it be direct or indirect). The possibility of indirectness is a significant caveat, so we have added clarifying text to indicate this possibility. We have added to paragraph 2 of the Discussion the sentence, “As such, we emphasize that Zelda-dependent licensing could stem either directly or indirectly from Zelda function.”

      (4) Regarding weakness 2, as currently written, it seems like a small fraction of H3K27me3 domains depend on ZLD. Is this accurate? Some of my confusion may stem from alternating use of bins and runs. Can the fraction of domains that depend on ZLD be made more explicit through text revisions and additional bioinformatics? More comprehensive bioinformatics analyses can be performed with the ZLD mutant datasets by incorporating ATAC and E(z) peaks. For instance, what fraction of E(z) peaks inside and outside of Polycomb domains are affected in ZLD mutants, and do these E(z) changes correlate with H3K27me3 changes? Similarly, what fraction of PREs change in accessibility in ZLD mutants, and are these accessibility changes correlated with H3K27me3 changes? When do PREs become accessible during wild-type embryogenesis? Are there unique features of ZLD-dependent H3K27me3 domains or E(z) peaks? The authors' perspective on the extent to which ZLD is required for Polycomb domain initiation should also be added to the Discussion.

      We find that 38 PcG domains are sensitive to Zelda (9 have increases, 29 have decreases in H3K27me3), and this is stated in the Results. This accounts for 16% of domains, which is a small fraction of the total. We have added a sentence to the results reporting this fraction. An earlier draft of this manuscript included an analysis of accessibility: both in terms of the timing of when PREs gain accessibility and their dependency on Zelda function. This section was omitted from the submitted manuscript because it did not add clarity the distinction between classes that we report. As for the final request that we add to the Discussion our perspective on the extent to which Zelda is required for PcG domain initiation, we now address this in the second paragraph of the Discussion.

      (5) Regarding weakness 3 (why is H3K27me3 missing pre NC14?), it is suggested that the absence of H3K27me3 is due to the short duration of nuclear cycles relative to the rate of me2->me3 catalysis. And although the authors suggest that Polycomb complex assembly occurs at pPREs without H3K27me3, their microscopy studies indicate that binding of Polycomb proteins to PREs may be regulated via nuclear accumulation. Therefore, it remains unclear whether (and when) the lack of H3K27me3 is due to incomplete Polycomb complex assembly. Additional genomic profiling experiments are needed to directly test when Polycomb complexes assemble on chromatin relative to the emergence of H3K27me3 domains. These experiments would also support claims of nucleation.

      (a) ChIP of E(z) and a PRE binding protein (Pho, Gc, GAF) should be performed at NC13 to test whether Polycomb proteins (especially PRC2) are bound at PREs prior to H3K27me3 emergence.

      (b) ChIP of E(z) should also be performed at NC10 when the PRE binding proteins appear to be absent but when E(z) is hyperabundant in the nucleus. Can E(z) bind PREs in the absence of these "nucleators"? Or does E(z) promiscuously interact with chromatin, helping to explain the broad, low-level H3K27me1 enrichment profile?

      We thank the reviewer for this comment. We have addressed this issue in a separate study (preprinted and accepted for publication as of this writing). Although H3K27me3 is not detectable on cleavage-stage chromatin, lower-order H3K27me2 is maintained on chromatin and detected throughout cleavages by immunostaining and by ChIP. The maintenance of cleavage-stage H3K27me2 depends on both E(z) and Esc. Notably, this lower-order state reflects maintenance of a maternally supplied H3K27 methylation state: H3K27me2 is only detected on maternal (not paternal) chromatin during the period (prior to NC10) when Pho/Cg/GAF do not localize to nuclei. Overall, these observations underscore the limiting nature of early cleavages to support de novo establishment of H3K27 methyl states, but also demonstrate the competency of the PcG system during this time to engage in some degree of H3K27 maintenance.

      Minor Comments:

      (1) It is interesting that the great majority of E(z) peaks (72%) are outside of H3K27me3 domains at NC14. What are these sites? Do these sites correspond to Polycomb domains later in embryogenesis (ie, do they become marked by H3K27me3 later)? Or, do these E(z) peaks disappear at later stages of embryogenesis?

      We agree that this observation is interesting. Figure 2 shows that the majority of these sites are concurrent with transcription start sites. Visual comparison between ChIP datasets generated here and any of the various publicly available datasets generated at later stages (e.g., stage 16 embryos) reveals that the set of PcG domains observed at ZGA is fairly consistent with the set of PcG domains observed later on. Therefore, on a bulk level, these extra-domain E(z) sites observed at ZGA do not represent later-onset PcG domains. However, at this level of resolution, we cannot rule out that in some cell type these sites are converted to a cell-type-specific PcG domain. We have not addressed the perdurance of these E(z) peaks at later stages. The function and the fate of these extra-domain binding sites remains an open question for future investigation.

      (2) Figure 3A live imaging indicates that E(z) nuclear signal intensity diminishes over subsequent nuclear cycles. Can the authors expand on whether photobleaching may contribute to this decrease? The E(z) signal in IF experiments should be quantified to test whether it also diminishes over successive nuclear cycles.

      The imaging conditions for this experiment were controlled to minimize photobleaching in the EGFP channel. While we cannot rule out some small contribution of photobleaching to the overall signal intensity, the overall trend reported in the quantification of Figure 3A reflects, to the best of our abilities to measure, a biological effect. This effect is also evident in the immunofluorescence imaging without need for additional quantification: From NC10 to NC14, when nuclei are presented on the embryo surface, a clear decrease in staining intensity is observed (Figure 3-figure supplement 1A in the final manuscript). Our observations indicate that E(z) decreases in concentration with increasing nuclear content during the cleavage divisions.

      (3) Can the authors provide further interpretation of the H3K27me1 signal profile? There is very little difference in signal amplitude inside relative to outside domains in Figure 1C, and across a 20kb window surrounding E(z) peaks in Figure 1B. Is all this chromatin considered to be H3K27me1-enriched, or is there a high level of noise? The anticorrelation with H3K27me3 at NC14 is compelling and seems to argue that the broad H3K27me1 signal is real.

      The H3K27me1 signal profile reflects a likely broad distribution of H3K27me1 that includes not only canonical “PcG domains” (i.e., regions that ultimately gain high-level H3K27me3) but also non-canonical domains. Consistent with observations in other systems (e.g., PMID: 24289921), H3K27me1 is broadly distributed across the Drosophila genome. We have added the following sentence to the Results section to contextualize our description of the H3K27me1 domains: “The broad genome-wide distribution of H3K27me1, including in regions outside of canonical PcG domains, resembles profiles previously measured in mammalian tissue culture.” (citing the Ferrari et al study).

      Reviewer #3 (Recommendations for the authors):

      My suggestions for changes are mainly of an editorial nature.

      My comments below are not in order of priority, but grouped into different types of suggestions for changes.

      Comments on the presentation of results:

      (1) The relationship between the genomic regions shown in the heat map in Figure 1A and Figure 2A is not clear. In Figure 2A, 1264 regions with E(z) peaks in PcG/H3K27me3 domains are shown. In Figure 1A, the text (line 97/98) states that 237 PcG/H3K27me3 domains contain 1 or more E(z) peaks, and on the left of the H3K27me3 heat map panel, it says "PcG domains". Are these then the 237 regions that are shown in Figure 1A? Only the top ones seem to have high levels of H3K27me3. Please clarify.

      We could have been more clear about this. As the reviewer points out, there are different numbers of peaks shown in the heatmaps in Figures 1A and 2A. Figures 1A and B plot one E(z) peak per domain, determined by finding the one maximal E(z) peak per and plotting it. This is stated in the figure legend: “One representative E(z) peak per domain was selected for plotting.” Figure 2A plots all E(z) peaks within Domains (n = 1264). For the heatmaps and the average plots in Figures 1A and 1B, we found that selecting the maximal E(z) peak yielded a more accurate representation of the accumulation of H3K27me1/3, compared with plotting this for all E(z) peaks within domains, presumably because some called peaks (e.g., many of the minor E(z) peaks shown in Figure 1E) do not appear to be the primary sites of nucleation within the domain. For Figure 2, we wished to compare all E(z) peaks, inside and outside of domains, and for this we plotted heatmaps over all of the peaks.

      (2) In general, in the text, the figure panels should be discussed in order of appearance. For example, it is not ideal that after describing the data in Figure 1A, the text jumps to discuss results shown in Figure 2A and only then goes back to discuss Figure 1B and C. This should be easy to resolve by reorganizing the text or perhaps re-arranging figure panels (e.g., perhaps moving elements to additional supplemental figures?).

      We acknowledge that this is inconvenient and we apologize for this. However, we feel the flow of the manuscript and the figures themselves benefit from this somewhat awkward order of discussion and we have chosen to leave them as-is.

      (3) In general, I felt that the text could be improved by putting some of the more detailed technical procedures and result descriptions into Materials and Methods or Figure legends in order not to disrupt the flow of the text.

      Comments for improving discussion:

      (4) When citing and discussing the previous literature that claimed a function for GAGA factor (GAF/ Trl) in Polycomb repression (i.e., refs. 12, 14, 49-52) on page 15/16 and again on page 25, the authors may also want to cite earlier studies that failed to observe a function of GAF/Trl in Polycomb repression. Specifically, previous studies (Brown et al, Development 2003) had investigated a possible role of GAF/Trl in Polycomb repression in larvae using stringent tests (analysis of HOX gene expression in Trl null mutant cell clones in imaginal discs, removing Trl in a sensitized pho null mutant background, mutation of GAF/Trl binding sites in HOX-LacZ reporter genes) and had found no evidence for a role of GAF/Trl in Polycomb repression in larval tissues. It is fair to say that in most of the studies cited (i.e., references 12, 14, 49-52), the effect of GAF/Trl had been analyzed using PRE-miniwhite reporter gene activity as read-out, a much less stringent assay.

      Thank you for pointing this out. Brown et al (reference 15) was overlooked when entering citations to these sections. We have added a sentence to the discussion to highlight the lack of necessity for Trl/GAF for imaginal disc silencing of Hox targets.

      (5) The authors report that removal of GAF/Trl had almost no impact on H3K27me3 domain formation. Although this supports a model where PRC2 recruitment to PREs is unaltered after GAF/Trl depletion, it does not eliminate the potential caveat that PRC1 binding may be affected. Do the authors have binding profiles of PRC1 subunits or the H2Aub1 profile in embryos lacking GAF/Trl? It seems that the current paper would be a great opportunity to include such data and get them published. There is no need to generate PRC1 or H2Aub1 profiles if they don't already exist.

      Unfortunately, we do not as of yet have any PRC1 reagents that are compatible with ChIP. Following the lack of effect of GAF on H3K27me3, we did not pursue the measurements of H2Aub in these knockdown conditions, given the likely dependence of H3K27me3 on H2Aub deposition at ZGA.

      (6) Previous studies found that Polycomb group protein complexes bind to PREs in HOX genes both in cells where genes are OFF but also in cells where genes are ON but that the H3K27me3 profile is different in the two states, with H3K27me3 decorating the gene in the OFF but not in the ON state (Papp and Müller, Genes Dev 2006; Bowman et al, eLife 2014). In this study, the H3K27me3 profiles at target genes represent the sum of ChIP signals coming from cells where the gene is OFF and cells where the gene is ON. Is it known how eve and zen are deregulated in Zelda-depleted embryos? Could it be that the increased H3K27me3 enrichment at eve and zen is an indirect effect caused by loss of eve and zen expression in a large fraction of cells and consequently a gain of H3K27me3 ChIP signal at the gene in those cells? It may be worth at least discussing such scenarios in light of the studies mentioned above, and referring to them.

      Thank you for raising this question. We have a manuscript in preparation that specifically addresses this issue, namely the relationship of on/off states to H3K27me3 deposition in the early embryo. In short, it is likely that the increase in H3K27me3 at eve reflects significantly reduced eve expression in zelda mutant embryos. We have added a clarifying sentence to the Discussion and added the indicated references.

    1. Author response:

      We would like to thank the editorial team and the reviewers for their thoughtful assessment of our manuscript. We are highly encouraged that the reviewers found our closed-loop OMR virtual reality assay to be a valuable new behavioural paradigm, and that our findings regarding the early behavioural phenotypes in adgrl3.1 mutant zebrafish provide solid, high-quality data of broad interest to the community. We also appreciate your constructive feedback regarding the structural flow of our manuscript, the need for tighter conceptual framing, and the request for more rigorous methodological explanation and improved presentation. In our upcoming revision, we plan to directly address the suggestions raised to ensure clarity of our work.

      For our revised manuscript, we will ensure to focus on fully contextualising our conceptual rationale and unique utility of our behavioural assay, ensuring a clear link to clinical relevance. This means we will provide further clarity on the rationale and relevance of the paradigm and interpretations of behaviours, while ensuring we are not overstating the absolute separation of anxiety-like states from hyperactivity and contextualising our findings within the broader literature. In addition, as suggested by the reviewers, we will improve the Abstract to explicitly articulate the unique advantages of the closed-loop OMR virtual reality paradigm over standard static assays and provide definitions. In addition, we will also revise our interpretations in the Discussion, such as around swim speed, exploration, and visual sensitivity, to ensure they are fully grounded and supported with clear flow.

      Equally important to our revision is to clarify genetic validation, methods and statistical rigour as recommended by the reviewers. Therefore, we will update the Methods to explicitly detail further husbandry details, such as the embryo medium used, breeding protocols, and the exact timings in hours post-fertilisation. To resolve any ambiguity surrounding our genetic model, we will provide further explanations and include details on genotype and sequencing protocols. Where needed, we will also update the statistical analysis and expand the Methods accordingly to detail all statistical reporting for clarity.

      Finally, as suggested, we will also improve the overall structure, flow, and accessibility. In the Results, we will ensure the core behavioural phenotypes take primary focus throughout the writing. Furthermore, we will also improve the Discussion to ensure it is a unified, cohesive narrative that smoothly integrates our main behavioural findings with genetic model validation, limitations, and future directions. Additionally, we will update all visual presentations, figure formatting, and labelling to ensure accessibility and improved readability.

      We are incredibly grateful to the reviewers for their insightful recommendations. We are confident that by integrating their revisions, we will substantially strengthen the clarity of the work and better demonstrate its impact.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revisions:

      Overall, I think the authors made many improvements to their manuscript. There are, however, still a number of concerns that first need to be addressed, since it is still not currently possible to fully evaluate the analyses, results, and conclusions presented in the paper. I list these points below:

      (1) The authors still use causal language where they must not use causal language. This is true for many places in the manuscript; I am highlighting here just a few places, but the authors nevertheless have to go carefully through the whole manuscript to change these instances.

      We sincerely thank the reviewer for this critical and well-taken point. We fully agree that our design (manipulating DLPFC excitability while measuring procrastination, task value, and aversiveness) does not directly measure or manipulate self-control, nor does it rule out alternative neurocognitive mechanisms. Accordingly, we have conducted a comprehensive, line-by-line revision of the entire manuscript to systematically replace causal claims with cautious, hypothesis-consistent language. In response, we have replaced all the wordings that may imply causal inferences, such as impair, cause, boost, by association-consistent phrasing. Furthermore, as you clearly raised below, we explicitly reframed self-control as a hypothesized theoretical construct rather than an empirically verified mediator throughout the whole re-revised manuscript. Please see specific revisions below:

      Abstract Section (Page 2, Line 53-59)

      “... a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement. In conclusion, these findings are consistent with the hypothesis that enhancing DLPFC function may reduce procrastination by selectively amplifying the valuation of future rewards, not by simply reducing negative feelings about the task.”

      Introduction Section (Page 3, Line 83-86)

      “... Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease (Sirois, 2015; Sirois, 2016).”

      Introduction Section (Page 4, Line 143-145)

      “Consistent with this framework, the left dorsolateral prefrontal cortex (DLPFC)—a region frequently implicated in value-based decision-making and top-down regulation—has been associated with procrastination. ...”

      Introduction Section (Page 4, Line 135-139)

      “... Also, given the integrative nature of prefrontal regulatory functions, we hypothesize a third pathway whereby both decreased task aversiveness and increased task-outcome value may jointly contribute to reduced procrastination, potentially reflecting coordinated downstream effects on valuation and affective processing.”

      Introduction Section (Page 5, Line 183-186)

      “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.”

      Results Section (Page 11, Line 534-536)

      “... Thus, these findings are consistent with the view that neuromodulation of the left DLPFC is associated with reduced task aversiveness and increased task-outcome value.”

      Results Section (Page 11, Line 579-584)

      “... In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      Discussion Section (Page 12, Line 617-620)

      “On balance, our findings provide evidence consistent with the hypothesis that neuromodulation of the left DLPFC is associated with reduced procrastination, primarily through increasing task-outcome value rather than merely reducing task aversiveness. ...”

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019).”

      Discussion Section (Page 14, Line 755-762)

      “... Moreover, this study did not collect data for assessing participants' self-control at either baseline or post-neuromodulation. Accordingly, we explicitly note that self-control was not directly measured or manipulated in this study; the observed associations between DLPFC neuromodulation, task-outcome value, and procrastination are consistent with theoretical models positing a role for top-down regulatory processes, but do not constitute direct evidence that self-control mechanisms were engaged. This limitation precludes definitive conclusions about the unique contribution of self-control-related pathways versus alternative neurocognitive mechanisms.”

      Some examples:

      (a) In response to my comment (1) in the previous round, where the authors adjusted their text, the authors still use causal language in their last sentence "... procrastination behavior has been observed to impair general health..." Unless the cited study truly allowed causal conclusions, the causal language should be removed here as well.

      Thank you for pointing out this inappropriate phrasing. As you kindly suggested, we have reworded it as “Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease” (Introduction Section, Page 3, Line 83-86).

      (b) The authors still make (causal) claims about the involvement of self-control in their observed results. To reiterate from the previous round of revisions: The authors cannot make any strong claims about the role of self-control processes because they do not directly measure self-control nor do they directly manipulate self-control or have a design that would rule out alternative mechanisms other than self-control. Therefore, their claims about self-control have to be toned down. It is laudable that the authors have added a statement towards the end of their discussion about not being able to make strong conclusions about the role of self-control. But the authors need to use similar careful wording not just at the end of the discussion but throughout the manuscript.

      We appreciate you reiterating this concern. In the re-revised manuscript, we have thoroughly removed or rewritten all the statements implying causal inferences, and have substantially toned-down claims for the roles of self-control in procrastination reduction from this neuromodulation. Please see instances for what we have replied to the Comment #1.

      (i) In the abstract, the authors use the formulation "...conceptualized roles of self-control on procrastination..." -- this wording is still too strong, suggesting that you actually studied self-control.

      Thank you for providing this specific instance. This inappropriate sentence has been removed.

      (ii) In the introduction (page 4, lines162-169), the way the authors formulate these sentences suggests that they directly measured self-control. Again, the authors need to make it explicit that they are not directly measuring self-control but its hypothesized down-stream consequences on valuations/behavior.

      Many thanks. This statement has been removed, and we have reworded it as “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.” (Introduction Section, Page 5, Line 183-186).

      (iii) In the discussion, for example, on page 11, lines 555 and following, the authors write: "One major contribution this study has made is to disentangle the neurocognitive mechanism of procrastination by demonstrating that self-control could increase task-outcome value so as to reduce procrastination."

      As you kindly instructed, we have rewritten this statement as “One contribution of this study is to provide empirical evidence partially consistent with the temporal decision model (TDM), showing that increased task-outcome value—rather than decreased task aversiveness—was statistically associated with reduced procrastination following DLPFC neuromodulation.” (Introduction Section, Page 12, Line 625-628), which no longer implies any conclusions for the role of self-control in this study.

      Again, please be aware that you are NOT demonstrating that self-control does anything, since you only measure procrastination rates, outcome values, and task aversiveness. It is possible that mechanisms other than self-control might be relevant for this. Perhaps neuromodulation directly increases outcome values, without involvement of self-control processes. You simply cannot know that and therefore you cannot make those claims in the form that you are making them. You can write that the observed results are consistent with the idea that neuromodulation might have had an effect on self-control and this in turn might have affected outcome values. But you also need to make it explicit that, to substantiate these claims, you would need more direct evidence that indeed self-control was involved. These more careful formulations would not at all reduce the value of your work, but indeed they would rather demonstrate your carefulness in interpreting the results you obtained.

      We sincerely thank the reviewer for this exceptionally clear and constructive guidance. We fully agree that our study design does not measure or manipulate self-control, and therefore we cannot demonstrate that self-control processes are causally involved in the observed effects. As you correctly note, it is entirely possible that neuromodulation directly modulates outcome valuation or engages alternative neurocognitive pathways (e.g., attentional allocation, feedback learning, or affective processing) without invoking self-control mechanisms.

      In direct response, as we replied above, we have completely rewritten the whole revised manuscript to remove any assertions that we “identified” a role of self-control. The revised text now explicitly states as follow: (1) our findings merely are consistent with the theoretical hypothesis that DLPFC neuromodulation might engage prefrontal self-regulatory functions, which in turn influence outcome valuation; (2) we explicitly acknowledge that substantiating this specific pathway would require more direct evidence. Rather than single sentence, we have applied this careful, hypothesis-consistent framing systematically across the Abstract, Introduction, Results, and Discussion. As you suggested, these revisions more accurately reflect the interpretative boundaries of our data and demonstrate our commitment to rigorous, transparent scientific reporting. Please see specific cases for this revision above.

      (2) I am still puzzled by the power analysis. In the text, you write that a sample size of 18 participants (i.e., 9 per group) would be sufficient to achieve 80% power. I still feel this seems far too optimistic and hard to believe, but that is not my point here. While in the text, you write that you need 18 participants, the G*power output seems to suggest a sample size of 34, not 18. Why this contradiction? Or is it not contradictory? If it is not, then please explain it more fully.

      We appreciate you pointing out this critical typo. In the last round of revision, we mean that 18 participants per group are required to achieve at least 80% statistical power, rather than a total sample size, as shown by the GPower software. We are sorry for this critical typo to confuse you. As you correctly pointed out, the GPower indicated that the minimum sample size to reach 80% power is 34 (i.e., 17 per group). Thus, we selected 36 (i.e., 18 per group) participants as minimum sample size in case of potential drop-out. We have thoroughly corrected this typo, and double-checked no such numeric issues:

      Methods Section (Page 5, Line 234-237)

      “... statistical power was predetermined by G*Power at a relatively medium effect size (1-β err prob = 0.80, f = 0.25), indicating the total sample size at 34 (17 per group) to reach acceptable power. To account for potential attrition, we determined to recruit 36 participants, at least.”.

      (3) I have several comments about the mixed-effects analysis.

      First of all, I want to thank the authors for adding more details, things have become much clearer now. However, I still have a few questions and comments related to these analyses:

      (a) The variable Emotions was within-subjects, as far as I understood. Accordingly, Emotions should most likely be modelled with random slopes varying over participants (in addition to being modelled as a fixed effect).

      We thank you raising this reasonable concern on the mixed-effect linear modeling. Yes, the Emotions reflect daily baseline affect, which is modeled as covariates of no interests to adjust for daily emotional fluctuation (if any). In this vein, this baseline emotion score is included for each participant across all the sessions. Therefore, it should be modeled with random slopes as you assumed indeed.

      In response, we have remodeled this mixed-effects analysis by including the daily baseline emotion as random slopes varying over participants. After centering the variables, we estimated the revised models as “Procrastination Rate ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)” and “Task execution willingness ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)”. Notably, as you correctly assumed, fitting this complicated random-effect structure is likely to result in convergence failure, given the limited sample size in the present study. Therefore, we hypothesized the independence among random effects for model simplification. Consistent with this assumption, model comparisons indicated that the simplified models fit better than original ones (Procrastination Rate model, ∆AIC = -1.0, ∆BIC = -11.8, LRT, χ<sup>2</sup>(3) = 5.04, p = .17; Task-execution willingness model, ∆AIC = -5.5, ∆BIC = -16.4, LRT, χ<sup>2</sup>(3) = 0.51, p = .91).

      Taken together, as you kindly suggested, we have rebuilt the mixed-effect models by adding daily baseline emotion as random slopes varying over participants, and have demonstrated the consistent findings with the original one:

      Methods Section (Page 8-9, Line 412-419)

      “... Given the risks of convergence failure with the two correlated random-effects structure (i.e., treatment days and self-reported emotions), we hypothesized that the random effects are independent, leading to model simplification. Consistent with this assumption, model comparisons favored the simplified independent structure over the full correlated structure for both outcomes. For the actual procrastination model, the simplified model showed lower AIC (∆ = -1.0) and BIC (∆ = -11.8), with a non-significant likelihood ratio test (χ<sup>2</sup> (3) = 5.04, p = .17). For the task-execution willingness model, the simplified model was also preferred (∆AIC = -5.5, ∆BIC = -16.4; LRT: χ<sup>2</sup> (3) = 0.51, p = .91).”

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (b) The analyses still cannot fully be evaluated as I cannot access the scripts and data. The authors mention that the scripts and data should be available via a link they provide (https://doi.org/10.57760/sciencedb.35140). However, when I try to access these materials via this link, no page opens; it seems the link is dead?

      Thank you very much for bringing this case to us. We checked this link and found it to be still active.

      To ensure accessibility for your evaluation, we have uploaded scripts and data into this online submission system. Please do let us know if you are still unable to access them. We are glad to send them to you by other available pathways. This link is a private access to you, and the repository would be openly available for other users upon the final publication.

      (c) What are the results and conclusions if you do not include the covariates of no interest? I.e., please re-run your main models without age, gender, SES, Emotions.

      Thank you for raising this question. As you clearly instructed, we have rerun main models without all those covariates. As shown in the table below, the results for the key predictors of interest (Group, Treatment day, and their interaction) remained largely unchanged in terms of effect size, direction, and statistical significance:

      Author response table 1.

      Comparison to statistics derived from model with covariates (i.e., age, gender, SES, Emotions) and without covariates

      (d) The authors mention that they use GLMMs, which would suggest generalized mixed-effects models, but they do not describe what family/distribution they used. Since they mention lmerTest and seem to report F-tests, my guess is that they used Gaussian models. However, both their DVs (procrastination rates and their ratings) are bounded variables and at least procrastination rates hit the lower boundary. That can mean that their analyses suffer from inflated Type 1 and/or Type 2 rates. Therefore, please repeat the analyses with an appropriate generalized mixed-effects model (perhaps a beta regression type of model?).

      We are very grateful to you for raising this crucial statistical point. As you correctly pointed out, we used the Gaussian distribution in estimating this model. We are sorry to confuse you due to the absence of reporting family/distribution we used. In the original manuscript, we meant “general” linear mixed-effect model, rather than “generalized” one. As you clearly and correctly raised, procrastination rates and willingness are technically bounded, and that procrastination rates frequently reached the lower boundary (0%) in the present study, which are in high risks to be inflated for Type 1 and/or Type 2 error.

      Thus, as you kindly suggested, a beta family distribution with logit function is used to reanalyze those main effects of interest. Results are tabulated in Author response table 2.

      Author response table 2.

      These convergent results confirm that the critical main effect (i.e., Group and Treatment day) and their interaction remain statistically significant across distributional specifications, and that our primary conclusions are not artifacts of the Gaussian assumption. Taken them together, as you kindly suggested, we have repeated the analyses with beta regression family distribution, which replicated our main findings, potentially supporting their statistical robustness.

      Following your suggestion, we have added those results derived from such sensitivity analyses into the revised manuscript:

      Methods Section (Page 9, Line 430--436)

      “... To examine whether our findings were sensitive to the distributional assumptions of the dependent variables, we re-analyzed the main models using an alternative distributional specification. Given that both procrastination rates (ranging from 0% to 100%) and task-execution willingness (measured on a 0-100 visual analog scale) are bounded continuous outcomes, and that procrastination rates frequently reached the lower boundary (0%) in the present study, a Beta regression model with a logit link function was employed for a sensitivity analysis.”

      Results Section (Page 10, Line 506-512)

      “... Furthermore, as a sensitivity analysis, we reran the main LMMs using Beta regression distribution with a logit link function, which is appropriate for the both bounded outcomes mentioned above (i.e., procrastination rate and procrastination willingness). The main effects (i.e., Group and Treatment day) and their interaction remained significant for both procrastination rate and willingness (see SI Results and Tab. S5), confirming that our findings are robust to alternative distributional assumptions.”

      (e) When reporting the results of the mixed-effects models, the authors report the regression coefficient, standard error, DFs and p value, but not the actual test statistic. Please add the information about the test statistic and report all degrees of freedom (in case of F tests that would be the degrees of freedom of the test and the residual degrees of freedom).

      We truly thank you for this nuanced reminder. As you suggested, we have added actual test statistics, including t-values and all degrees of freedom (DF). Please see specific instances below:

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (f) Thank you for adding the analysis where you remove the last two sessions. But currently you present them in the manuscript without explaining/motivating why you do this. Please add this motivation, as otherwise it will be puzzling for the reader why you conduct these analyses.

      Thank you for this very practical requirement to clarify the motivation of reanalyzing main models from removing the last two sessions. Please see specific explanation as follow:

      Results Section (Page, Line 499-506)

      “... To systematically test whether such effects are biased by extreme data points or patterns, we reran the main LMMs by iteratively removing data from the last two sessions, which showed extraordinarily high effectiveness from neuromodulation (e.g., all the participants in the neuromodulation group had no actual procrastination behavior in session #6 and #7). Results showed the significant group*neuromodulation sessions interaction effects across all those nested models (removing session #6, #7 or both, all p < .05; see SI Results and Tab. S3-4), potentially indicating a statistical robustness from the data pattern.”

      (4) Mediation analysis

      In your manuscript, you present some mediation analyses. Please be aware that such mediation analyses cannot establish causality and they suffer from extremely high Type 1 error rates (see, e.g., https://datacolada.org/103). My suggestion would be to completely remove all mediation analyses. However, if you want to keep them, then you need to be extremely careful in how you present the results. You need to explicitly mention that you cannot derive any causal conclusions from them and that simulation studies have shown that such mediation analyses suffer from extremely high Type 1 errors.

      We sincerely thank you for this exceptionally important methodological guidance. We fully agree that mediation analyses, especially those based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates, as rigorously demonstrated in recent simulation studies (https://datacolada.org/103).

      As you kindly suggested, please allow us to retain those mediation analyses upon explicitly highlighting that this mediation statistical model cannot generate any causal conclusions. In response, we have systematically replaced all instances of “causal mediation” with “statistical mediation” or “exploratory mediation analysis”, and removed causal verbs (e.g., “depends on”, "drives", "explains") in favor of association-consistent phrasing (e.g., “is statistically mediated”, “aligns with the hypothesis that”) throughout the abstract, introduction, methods, results and discussion sections. Furthermore, in the Discussion section, we explicitly reiterated the limitations of extending this mediation associations to causal conclusions. Please see specific modifications underneath:

      Abstract Section (Page 2, Line 52-56)

      “... While the intervention is significantly associated with both decreased task aversiveness and increased perceived task outcome value, a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement.”

      Methods Section (Page 9, Line 448-453)

      “... As these mediation analyses are based on observational measures rather than experimentally manipulated mediators, they do not establish causal pathways. Simulation studies have shown that such analyses can suffer from inflated Type 1 error rates. Results should therefore be interpreted as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms.”

      Methods Section (Page 9, Line 438-440)

      “... the Quasi-Bayesian mediation analysis was used to model the association between the effects of tDCS, task aversiveness/outcome and decreased procrastination.”

      Results Section (Page 11, Line 568-573)

      “As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. ...”

      Results Section (Page 11, Line 582-584)

      “... Nevertheless, all mediation findings are now explicitly labeled as “exploratory” and framed as quantifying statistical associations consistent with the TDM pathway, not causal mediation.”

      Discussion Section (Page 14, Line 747-755)

      “Notably, we explicitly acknowledge that exploratory Quasi-Bayesian mediation analyses, based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates as demonstrated in recent simulation studies. These findings should be interpreted strictly as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms. Substantiating the precise neurocognitive pathway will require future studies employing stronger causal designs, such as experimental manipulation of task valuation or longitudinal cross-lagged modeling. ...”

      As an example (but the mediation results are mentioned in several places, for example, also in the abstract): On page 10, lines 501-503: What you can causally conclude is that neuromodulation affects your measured variables (outcome values, procrastination rates, task aversiveness), but you cannot conclude that the effect of neuromodulation on procrastination rates causally operates via outcome values. Thus, please adjust the formulation accordingly. The same applies to the mediation section that follows right afterwards (page 10, lines 505-522).

      Thank you for offering those specific instances. As we replied above, they have been revised accordingly:

      Results Section (Page 11, Line 557-559)

      “... Collectively, these findings provide statistical evidence consistent with the hypothesis that the outcome-value pathway may contribute to procrastination reduction.”

      Results Section (Page 11, Line 563-582)

      “Increased task outcome value is specifically associated with reduced procrastination in the context of neuromodulation

      To explore the potential neurocognitive pathways of procrastination, the Quasi-Bayesian mediation analysis was undertaken, with increased task outcome value as a statistically mediated variable. As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. To ensure the statistical robustness and specificity of these findings, the sensitivity analysis was implemented by changing sampling parameters and outcome variables. By doing so, those findings were validated statistically robust, as shown by replicated observations across bootstrapping sampling subsets (see SI Results and Tab. S6-7). Moreover, the results of the control analysis further validated the specificity of these findings by showing a null statistically mediated effect of this model to predict one’s task aversiveness (see SI Results and Tab. S8). In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      (5) In the introduction, the authors introduce several theoretical procrastination frameworks (TMT, mood repair, TDM). Do the results of the current paper help to decide which framework might be the most appropriate, at least for the authors data set? It might be of interest to address this explicitly.

      We do thank the reviewer for this insightful theoretical question. We agree that explicitly positioning our findings within the broader theoretical landscape strengthens the conceptual contribution of our work. Upon careful consideration, we believe that our results provide the strongest empirical support for the TDM over alternative frameworks (TMT, mood repair). TDM uniquely posits procrastination as contingent on the dynamic trade-off between task aversiveness and task-outcome value. In the present study, neuromodulation was identified to be associated with both pathways but only increased outcome value statistically predicted reduced procrastination. This aligns precisely with TDM’s hypothesis that value-based processes may dominate aversiveness-avoidance processes in driving behavioral change. Neither TMT (which emphasizes temporal discounting of utility per se) nor the mood repair perspective (which prioritizes short-term affect regulation) explicitly predicts this dissociable pattern. As you kindly suggested, we have extended the Discussion section to explicitly contextualize our findings into this theoretical landscape (Discussion Section, Page 13, Line 669-687).

      (6) The language is sometimes hard to understand and seems in quite some places grammatically incorrect. Thus, I think the paper would profit very much from thorough English proofreading.

      We sincerely thank you for this practical suggestion. We fully agree that the original manuscript contained grammatical inaccuracies and awkward phrasing that could hinder readability. In response, we have engaged a professional academic editing service to thoroughly proofread and polish the entire manuscript. All sentences have been revised for grammatical correctness, syntactic clarity, and academic tone, while strictly preserving the original scientific meaning and technical terminology. We believe this language improvements have substantially enhanced the readability and overall quality of the paper.

      Reviewer #2 (Public review):

      Summary:

      Chen and colleagues conducted a cross-sectional longitudinal study, administering high-definition transcranial direct stimulation (HD-tDCS) targeting the left DLPFC to examine the effect of HD-tDCS on real-world procrastination behavior. They find that seven sessions of active neuromodulation to the left DLPFC elicited greater modulation of procrastination measures (e.g., task-execution willingness, procrastination rates, task aversiveness, outcome value) relative to sham. They show that HD-tDCS reduces task aversiveness and increases task-execution willingness on real-world tasks as quantified by intensive experience sampling methods, providing causal evidence for the role of DLPFC in modulating contextual features to delaying or completing one's goals.

      Strengths:

      • This is a well-designed protocol with rigorous administration of high-definition transcranial direct current stimulation across multiple sessions. The intensive experience sampling approach which probes and assesses self-relevant task goals is innovative and aims to address an important question regarding the specific role of DLPFC in modulating specific features of chronic procrastination behavior (e.g., task-execution willingness, task aversiveness).

      • The quantification of task aversiveness through AUC metrics is a clever approach to account for the temporal dynamics of task aversiveness, which is notoriously difficult to quantify.

      Weaknesses:

      • While the findings that neurostimulation reduces procrastination behavior is compelling, there remain several alternative interpretations for these effects. For example, it could be that the task-execution willingness isn't increased per se, but rather that the goal completion becomes more valuable as participants learn from feedback or become more aware of their successful attainment of or failure to complete task goals. It is unclear whether the effects could be driven by improved working memory or attention to the reported tasks (and this limitation is addressed by the authors). In short, it is also difficult to examine the temporal dynamics of how these goals are selected across time.

      We sincerely thank you for raising these thoughtful and methodologically important points. We fully agree that the observed reductions in procrastination could reflect multiple neurocognitive pathways beyond the value-based mechanism emphasized in our primary analysis.

      In response, we have thoroughly removed claims on the “unique mechanism” of value-based pathways to procrastination reduction, and fully substituted language implying exclusive mediation by “value amplification” with more cautious phrasing (e.g., “statistically consistent with a value-based pathway”; “one plausible mechanism among several processes”). In the revised manuscript, we reiterated that the pattern of results, showing increased outcome value predicting reduced procrastination while decreased aversiveness did not, aligned with the Temporal Decision Model, yet does not rule out concurrent contributions from attention, learning, or executive processes. To clearly bring this interpretative boundary of our primary findings for audiences, we have explicitly warranted such cautions in the Discussion Section. Please see specific modifications underneath:

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019). ”

      Discussion Section (Page 13, Line 676-687)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Similarly, improvements in working memory for task maintenance, attentional allocation to reported goals, or strategic shifts in goal selection across sessions could plausibly mediate the intervention effects. While our double-blind, sham-controlled design and inclusion of daily emotional covariates help mitigate some non-specific confounds, the present study did not incorporate direct measures of these alternative processes. Consequently, we cannot definitively isolate the value-based pathway posited by the TDM from concurrent contributions of attention, learning, or executive functions.”

      • It is unclear whether the current evidence support long-retention of this neurostimulation intervention. The study includes one 6-month timepoint after the study to examine the long-term retention of the neural stimulation effect. Future studies that evaluate the long-term effects across multiple time points would strengthen the evidence for the robustness of this intervention.

      We genuinely appreciate you for this insightful and methodologically reasonable point. We fully agree that a single 6-month follow-up assessment, while valuable, provides only preliminary evidence for long-term retention, and that multiple follow-up timepoints would substantially strengthen claims about the durability of neuromodulation effects. To carefully address this point, we have rephrased the “long-term retention” as “long-term after-effects” throughout the whole revised manuscript, and overall toned down the claims on the retention effects. Moreover, this limitation has been explicitly elaborated in the Discussion section:

      Abstract Section (Page 2, Line 49-50)

      “... we assessed the effect of anodal HD-tDCS on real-world procrastination behavior at offline after-effect (2-day interval) and long-term after-effect (6-month follow-up).”

      Results Section (Page 12, Line 606-608)

      “... Therefore, beyond short-term effects, the benefits of ms-tDCS neuromodulation on reducing procrastination were still detectable at a 6-month follow-up, providing preliminary evidence consistent with long-term after-effects.”

      Discussion Section (Page 14, Line 721-725)

      “... Thus, the detectable effects at 6 months are consistent with the hypothesis that repeated neuromodulation may induce neuroplastic changes in the DLPFC that support sustained behavioral change. However, we explicitly note that a single follow-up timepoint cannot establish the stability or trajectory of these effects; future studies with multiple longitudinal assessments are required to substantiate claims about long-term retention.”

      Discussion Section (Page 15, Line 771-775)

      “... Finally, while our 6-month follow-up provides preliminary evidence for sustained effects, the use of a single follow-up timepoint limits our ability to characterize the temporal trajectory of retention. Future studies incorporating multiple follow-up assessments (e.g., 1-month, 3-month, 6-month, 12-month) would strengthen evidence for the robustness and durability of this intervention.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please see my detailed comments above (6 points; several of them with subpoints a, b, c, etc).

      Thank you for listing those specific and helpful recommendations above. Please see our detailed response posed above, point-by-point.

    1. Author response:

      The following is the authors’ response to the original reviews

      Reviewer #1 (Public review):

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups. 

      According to this suggestion, we have expanded our discussion to discuss why osm-5 and daf-2 worms fall into different longevity groups despite the fact that disruption of DAF-16 decreases both of the their lifespans. Please see lines 363-376.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses an elegant visual-anagram approach to test whether perceived animacy structures visual working memory and attention while controlling for many low-level image properties. The evidence is solid, with converging results across seven preregistered experiments, but the central claim that animacy itself is represented independently of visual features should be tempered, as residual mid-level configural cues, ensemble or category structure, and broader semantic differences may also contribute to the effects. The work will be of interest to researchers studying high-level visual representation, attention, and working memory.

      We thank the Editors and Reviewers for this careful and informed assessment. We appreciate that every Reviewer found our approach to be elegant, our findings to be solid, and our question to be of broad interest. We respond to each Reviewer’s specific comments in more detail below; but we thought to summarize some of the highlights - especially the specific comments that come up in this Assessment - here.

      (1) The Reviewers make the insightful point that, even if our stimuli effectively control for many low-level features, there may be other high-level features that explain performance in our experiments (R2: “Although the anagram paradigm effectively controls low-level visual features […] these stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories.”). We are happy to embrace this possibility. If our results were explained by high-level visual representation of the natural vs. manmade distinction, rather than the animate vs. inanimate distinction, this would still be an appeal to a (not altogether unrelated) high-level property being represented independently from its lower-level features, which was the primary motivation for our study. We framed our work specifically around animacy given the persistent debates regarding perceived animacy, as well as the fact that our stimuli do quite saliently vary along that dimension; but we are certainly open to other nearby high-level categories being at play. We also think this is an empirical question that could be tested in future work. For example, objects like rocks and lakes are natural but inanimate. If they behave more like dogs than like boots in our paradigms, then Reviewer #2 may be right that naturalness was the relevant property all along; but if they behave more like boots than like dogs, then perhaps it really was animacy doing the work. We now discuss this explicitly in our paper, and we appreciate the opportunity to not only clarify our claims but also spur discussion for future work.

      (2) Multiple Reviewers raise the question of whether semantic factors that go beyond the images themselves may be driving our effects. Reviewer #3 raises a particularly interesting question along these lines: “if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results?” We have now taken this question quite literally and run this experiment exactly as described. Of course, much research already explores cognitive processing of animate/inanimate words, finding (for example) stronger memory for animate objects than inanimate ones (e.g., Nairne et al., 2013; Nairne et al., 2017). However, such tasks do not invoke effects of visual processing, whereas the question at issue here is specifically whether the visual system prioritizes animacy independent of its lower-level features. To this end, we conducted a new, pre-registered experiment (now Experiment 8) where participants search for animate/inanimate words on some trials, and animate/inanimate pictures on others. Given the nature of visual search tasks, we should expect to find no search advantage for words (as their meanings are not processed in vision per se) — and we should also expect to replicate (once again) our search advantage for pictures. This is exactly what we found. In other words, linguistic stimuli alone failed to produce the effect, while anagrams did produce the effect. We believe this rules out the strongest form of the semantic labeling account.

      (3) Finally, Reviewers #2 and #4 raise some concerns regarding residual mid-level features such as configural shape and ensemble statistics, which lie somewhere between animacy itself and more basic properties like contrast or spatial frequency. In our paper, we now clarify each of these concerns in greater detail. In short: We think that our stimuli and experiments indeed control for these residual cues. For example, rotating an image preserves its configural shape; and, as we argue below, the specific ensemble statistics argument fails to get off the ground without appeal to animacy itself.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Evidence for visual representation of animacy.

      Strengths:

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

      Weaknesses:

      I thought this was an excellent submission. I have two suggestions for revision:

      We are glad to hear this Reviewer recognizes the broad challenge we are tackling in this work (separating high-level from low-level visual features) and found our submission to be “excellent”.

      (1) I thought that experiment 7 should have been described in more detail, with the upshot explained better. What exactly do the authors take it to show?

      Sorry for the lack of clarity here. We think the Reviewer actually gets this right earlier in their review; many feature differences exploit orientation, and our silhouettes control (Experiment 7) shows that those differences alone fail to explain our effects. For example, one might worry that our search effects merely reflect oddities in the aspect ratio or center of mass of the images. Converting the anagrams into silhouettes preserves these features. Thus, the fact that we found no search advantage with silhouettes suggests that these features on their own fail to produce the relevant effects; put the other way around, the effects we observed earlier must go beyond those features. We now discuss this in greater depth in our paper.

      (2) There should be a candid discussion of what the loose ends are and how they might be addressed. It would be good to have some examples like the square/diamond case with some indication of what would address such challenges.

      We agree with this, though we are somewhat limited by the space constraints of the Short Report format. A primary loose end we see is the possibility that high-level properties other than animacy explain our results (as raised by other Reviewers). We have added some discussion of this possibility to the paper.

      We would like to thank this Reviewer for their thoughtful feedback.

      Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention. While this is a thoughtful design and is well executed across seven pre-registered experiments, it remains unclear whether the reported effect is truly driven by animacy, as opposed to broader differences in ensemble statistics or semantic structure across the "mixed animacy" versus "uniform animacy" conditions. As such, the interpretation of a "pure" animacy effect may be overstated.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

      We are glad to hear this Reviewer found our work to be “creative” and believes it offers an “important methodological advance”.

      Weaknesses:

      (1) Specificity of the animacy effect vs. category-level ensemble structure

      The central claim is that animacy itself drives the observed effects. However, the key manipulation ("mixed animacy" versus "uniform animacy") also introduces differences in category-level ensemble structure. For example, in Experiments 1-2, cross-category change detection (e.g., dog to chair) may be easier not because of animacy per se, but because of a change in overall ensemble statistics (Brady & Alvarez, 2011, 2015). In addition, since each display contains five objects (two in one category and three in the other category), cross-category changes may also alter category balance in a way that further facilitates detection. In contrast, within-category changes preserve both ensemble structure and category composition, making them more difficult to detect.

      Brady, T. F., & Alvarez, G. A. (2011). Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychological Science.

      Brady, T. F., & Alvarez, G. A. (2015). Contextual effects in visual working memory reveal hierarchically structured memory representations. Journal of Vision.

      We appreciate the opportunity to clarify our claims and the support for them. Our claim is indeed that animacy (or a closely related high-level property; see below) drives our effects, over and above its lower-level correlates — i.e., that the explanation for differences in change detection or search across conditions will invoke a high-level property of the images. As we understand the Reviewer’s concern(s), they either (a) are already addressed by our novel methodology, or (b) would still fall perfectly in line with our claim as stated above.

      Consider the Reviewer’s concern that cross-category change detection “may be easier not because of animacy per se, but because of a change in overall ensemble statistics”. Which ensemble statistics change across categories in our stimulus set? Take as an example the case depicted in our figure, where a rabbit changes into either a dog (within-category) or a boot (cross-category). The dog and the boot are the very same image, just rotated; thus, they have the same luminance, curvature, area, spatial frequency, and so on. So if the change from rabbit to dog changes the array’s ensemble statistics with respect to any of those properties, it does so in the very same way as the change from rabbit to boot — and yet detection is still better for rabbit → boot than for rabbit → dog. Indeed, for nearly any ensemble statistic, the difference between the rabbit-display and the dog-display will be identical to the difference between the rabbit-display and the boot-display. To engage with the specific cases discussed in the two cited papers (Brady & Alvarez, 2011, 2015): The dog and the boot are the same size (because they are the same image), so average size is identical (just as average luminance, curvature, area, and spatial frequency are identical). And the very few properties left over (e.g., aspect-ratio) are addressed by later experiments.

      To be clear: We are not saying that there are no differences in ensemble statistics between the rabbit-display and the dog-display; across those displays, we replace one image with a different image, so there are likely all kinds of corresponding differences in ensemble statistics. The key question is whether that change in ensemble statistics differs across trial types in ways that might explain our effect - i.e., whether there is any difference between the rabbit-display and dog-display that is not also present between the rabbit-display and the boot-display. We don’t see how the answer could be yes, at least with respect to the statistics typically considered. A similar logic applies to the search tasks, with the silhouette control (Experiment 7) providing especially strong evidence that certain ensemble statistics or lower-level features cannot explain our effect.

      Now, it’s possible the Reviewer is referring to properties other than the low-/mid-level properties we mention above. Perhaps, for example, many animate stimuli on a display at one time have a striking collective appearance (all these animals are looking at me!) that lots of inanimate stimuli do not (this might be related to the Reviewer’s concern about “category balance”). But as we see it, this explanation just invokes animacy all over again, and so is the sort of explanation we would embrace.

      We now say more about this concern in the paper to be as clear as possible about our claims.

      (2) Limited stimulus set and potential learning effects

      The relatively small stimulus set (six anagram pairs) and repeated exposure raise the possibility of learning or familiarity effects. Does performance change over time? e.g., are there meaningful differences between early and late trials (e.g., first 10% vs. last 10%)? If such differences are present, they could suggest the development of task-specific strategies or increased efficiency with repeated exposure, rather than stable effects driven by the experimental manipulation itself.

      This is an interesting question, and we recognize this analysis absent from our initial submission. To be fair, stimulus sets of this size are not unusual in change-detection and search tasks, which often involve red, green, and blue squares repeated over the course of several hundred trials. Still, we certainly take the Reviewer’s point here and also embrace their analytical approach to addressing it. We’ve now run the “familiarity effects” analyses the Reviewer suggests (as well as some they did not suggest). The top-level headline is that learning or familiarity effects cannot explain our results, and if anything most of these analyses not only fail to support this alternative account but actively point against it. Below are more details.

      First, we worry that the Reviewer’s concern about “the development of task-specific strategies … rather than stable effects driven by the experimental manipulation itself” isn’t actually addressed by the suggested analysis of comparing the last 10% of trials to the first 10%. One reason for this is simply that it’s possible that both mechanisms are at play - i.e., that there is a baseline difference even without any familiarity that is then enhanced by some learning mechanism. (There are other issues as well: For example, one might imagine that participants get quite good at the task during the middle 80% of trials, but then get fatigued at the end. If this were true, then comparing the first 10% to the last 10% of trials could make it seem like there is no learning or familiarity, even if there were such effects. And on top of all this there is just the issue of statistical power, since far fewer trials go into these analyses than into our primary, pre-registered analyses). Nevertheless, we ran the Reviewer’s proposed analyses (using the first and last 10 trials of each type, which offers the best chance to find the pattern the Reviewer is concerned about). If anything, this analysis points in the opposite direction to the Reviewer’s prediction: 4/6 experiments (Experiments 1, 3, 4, and 5) revealed numerically weaker effects at the end of the task than the start, while only 2/6 experiments (Experiments 2 and 6) revealed numerically stronger effects at the end of the task than the start. Moreover, most of these results were non-significant, with only one marginal result (Experiment 5, p< = 0.08) and one significant result (Experiment 2, p = 0.01), and this is before any correction for multiple comparisons, which would make all of these results non-significant. So even though our account could easily accommodate learning effects, it’s not clear that they even exist here in any consistent or reliable way.

      Second, however, we think a more informative way to answer the Reviewer’s question is to ask not about learning over the course of the experiment but rather whether the key effects arise very early in the task. If they do, then any learning effects arising later couldn’t fully account for our results. Now, again, these tests are underpowered and only exploratory (to do this analysis properly, we would want to run entirely new experiments designed for this purpose), but we in fact did find evidence that our key effects arise early. In 5/6 experiments (Experiments 1, 3, 4, 5, and 6), the key effect was significantly (or in one case marginally) present even at the beginning of the experiment (Experiment 1, p = 0.07; Experiment 3, p = 0.01; Experiment 4, p < 0.001; Experiment 5, p < 0.001; Experiment 6, p < 0.01), and most of these results would survive correction for multiple comparisons. (In only one experiment, Experiment 2, was there a numerical disadvantage, but it was not significant; p = 0.34.) So even though our experiments were not designed or powered for this purpose, they do seem to suggest that the effects arise even without much familiarity at all.

      All told, we think these analyses suggest quite strongly that learning alone fails to explain our key effects. There is no evidence that the effects in general are stronger at the end of the experiment than the beginning (if anything it is the opposite); and there is evidence that most of the effects we investigated can be detected even very early in the experimental sessions. We have added discussion of these new analyses to our manuscript.

      (3) Role of semantics

      Although the anagram paradigm effectively controls low-level visual features, it still relies on high-level semantics (e.g., "dog" vs. "boot"). These stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories. From a semantic standpoint, it remains unclear whether the observed effects can be uniquely attributed to animacy or whether they reflect broader conceptual distinctions.

      We agree with the Reviewer here. While we feel comfortable interpreting our effects in terms of a high-level property like animacy as opposed to a lower-level property like curvature, it remains possible that the observed effects reflect some other, closely related high-level distinction (like natural vs. manmade). Our primary concern was to tease apart high-level properties from low-level features, which the Reviewer’s question does not threaten — if attention and memory are sensitive to the natural/artificial distinction, that’s interesting too, and a near neighbor of our actual claim. Still, we agree that this could be addressed, and we even see it as an empirical question testable in future work. Perhaps the most relevant departures between animate/inanimate and natural/manmade include objects like clouds, plants, and rocks — objects that are natural but not “animate” in the sense often used in this literature. If something like our paradigm revealed that rocks behave more like dogs than like boots, that would suggest that naturalness, rather than animacy, was driving the effects; but if rocks behave more like boots than like dogs, that would point to animacy even more strongly. We remain open-minded about this possibility, but it would of course require multiple new experiments with a brand new stimulus set and so goes beyond the present contribution. In any case, we have added a discussion of this issue to the paper and have adjusted our claims accordingly.

      Reviewer #3 (Public review):

      Summary:

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

      Strengths:

      The main strength is the elegant use of stimuli that control almost perfectly for low-level image features.

      Thank you for this kind feedback. This summary perfectly captures both our empirical contribution and the claims we are making.

      Weaknesses:

      My only real concern about the study is whether the findings truly provide evidence for a high-level visual representation of animacy independent of the low-level stimulus characteristics, or whether, instead, the effects are essentially semantic priming, which is independent of visual processing per se. For example, if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results? Words can also access semantic representations of the animacy of objects, and also don't suffer from low-level visual confounds. It would be helpful to add a discussion of this possibility to the article.

      Wow, we love this question! And so we’ve now conducted exactly the experiment the Reviewer suggests here. In a new pre-registered study (Experiment 8), we presented participants with a present/absent search task (as in Experiments 5–7). One half of trials consisted of the anagram stimuli (such that we could, once again, replicate the mixed-animacy search advantage); but the other half of trials consisted of the words describing the anagrams (e.g., “dog”, “boot”, “sheep”, “car”, etc.). The experiment worked beautifully: We found no effect with the words, but replicated the search advantage with the pictures — and also found a significant difference between the effects elicited by the two stimulus types.

      We agree with the Reviewer that this now rules out the possibility that semantic representations alone explain these visual effects. Thank you! 

      Reviewer #4 (Public review):

      In this article, the authors investigate whether perceived animacy influences visual processing independently of lower-level visual features by using "visual anagrams." Across seven experiments, they test whether animacy, isolated from many lower-level visual properties, structures visual working memory and guides visual attention. The central claim is that the visual system may represent animacy itself, rather than animacy emerging solely from associations among low-level visual properties.

      I find this investigation compelling. The experiments described provide strong control over several lower-level visual features, including curvature, texture, and related image properties. However, the visual anagrams are not pixelwise-identical across orientations. Because the images are rotated, the retinal configuration of pixels and the spatial organization of some low- to mid-level shape features also change. As a result, the configural arrangement of mid-level visual features may still contribute to perceived animacy.

      We are glad to hear the Reviewer finds our investigation “compelling”.

      I encourage the authors to discuss how independent perceived animacy is in this context from the contribution of mid-level visual features, such as configural shape cues that are diagnostic of animacy. This distinction would help sharpen the interpretation of the results and more precisely define the level of visual representation isolated by the visual-anagram approach.

      This is a helpful point, and it also echoes a sentiment expressed by Reviewer #2. While configural shape is diagnostic of animacy writ large, it can’t account for our observed effects here because rotating an image does not vary its configural shape. We now mention this in our work, and we agree that it helps sharpen the interpretation of our studies.

      Additionally, previous studies have argued that low- and mid-level curvilinear features may contribute to animate/inanimate categorization, and may in some cases be sufficient to support such distinctions (e.g., PMID: 33798259; PMID: 28654965). I encourage the authors to clarify how these previous findings on curvilinearity and rectilinearity fit with the overarching claim of the current study, namely that the visual system may represent animacy itself rather than animacy emerging solely from associations among lower-level visual properties.

      Yes, many studies from exactly that corner of the field actually motivated the present work, which is why we cited them in our submission. In a way, we are approaching this issue from the other side of the equation. Whereas the papers the Reviewer points to (along with many others) ask whether mid-level features (such as curvilinearity and rectilinearity) are sufficient to support perceived animacy, we ask whether these and other features are necessary to support perceived animacy. Prior work is relatively split on this issue, leaving the question wide open. We take our work to show that differences in curvature are not necessary for differences in perceived animacy, because our anagrams have identical curvature yet differ in animacy — and the visual system capitalizes on that difference. Put the other way around, representation of animacy can and does go beyond representation of its low- and mid-level correlates. Thank you!

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This study approaches an important topic providing insight into the neuronal circuitry that interconnects memory consolidation and sleep. The data were collected and analysed using a solid methodology, contributing new findings for neurobiologists working on how memories are stored and the roles of sleep. However, the data is incomplete to support the proposed role of the PAM-DPM circuits as the link between sleep state and long-term memory consolidation.

      We sincerely appreciate the editor and reviewers’ thoughtful and constructive comments on our study. Your insightful feedback has not only affirmed the significance of our work on the interplay between memory consolidation and sleep, but also provided valuable inputs for improving the clarity, rigour, and impact of our study.

      We have carefully addressed all the comments raised by the reviewers and revised the manuscript accordingly. We have also streamlined the paper with the goal of making it more accessible to readers. We feel this revised version strengthens our conclusion that the PAM-DPM circuits as the link between sleep and memory consolidation.

      The main improvements in terms of data addition are three complementary sets of circuit-specific experiments:

      (1) To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. This experiment addresses the activity of the microcircuit in a much more relevant time frame than the CRTC data in the previous version of the paper which looked only at the first hour after training. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, both PAM-α1 and DPM neurons displayed similar neural activity over the 3-hour recording period. The first hour was characterized by a synchronous decrease in activity for both neuron types, with hours 2 and 3 achieving a stable baseline. Since the decrease in the first hour is also seen in the trained condition, we think it is likely a reflection of the animals becoming acclimated to the recording tubes.

      Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit in the LTM consolidation time window. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support the hypothesized role of the inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      (2) To further support the functional connectivity of the PAM-DPM microcircuit, we conducted in vivo experiments to complement the dissected brain prep P2X2 data. Optogenetic activation of PAM neurons in intact flies via the red light-gated cation channel CsChrimson (Klapoetke NC et al., 2014) resulted in a significant reduction in GCaMP signals within DPM neurons (revised Figure 2B). These findings strongly confirm that PAM neurons exert direct inhibitory control over DPM neurons in the intact brain.

      (3) Further, we investigated how dopamine signaling to the DPM inhibits its activity, and issue which has not been investigated previously. We conducted a series of experiments:

      Firstly, we verified which dopamine receptors (Dop1R1, Dop1R2, DopEcR, and Dop2R) express on the DPM neurons via double-labeling with gene-embedded GAL4 lines. We found that DPM neurons have expression of both Dop1R1 and Dop1R2 (revised Figure 10A).

      Secondly, to clarify which receptors on DPM neurons respond to dopamine and how they signal, in addition to EPAC experiments in the first submission, we recorded neural activity changes when we knocked down Dop1R1 and Dop1R2 in DPM neurons. DPM neurons exhibited a significantly reduced GCaMP level with DA application, regardless of whether Dop1R1 or Dop1R2 was intact or knocked down knockdown in comparison to the no-DA control condition (revised Supplemental Figure 3C-E). These data suggest that either residual Dop1R1 and Dop1R2 remaining in the RNAi condition is sufficient or that the two receptors may coordinate to mediate the inhibition of neural activity.

      Finally, we investigated the behavioral contributions of Dop1R1 and Dop1R2 in DPM neurons to sleep and memory processes (revised Figure 10C-H). Dop1R1 knockdown resulted in a marked reduction in daytime sleep and a significant impairment of 24 h memory expression. In contrast, Dop1R2 knockdown selectively compromised 24 h memory without affecting sleep.

      When integrated with our EPAC assay findings from the initial submission, which demonstrated, that Dop1R1 is the primary receptor mediating dopamine-induced cAMP elevation, these new data collectively delineate a more complex mechanistic framework: dopamine signaling in DPM neurons coordinates the dual regulation of sleep and memory predominantly via Dop1R1. Meanwhile, Dop1R2 are engaged in the selective modulation of memory.

      All newly generated experimental datasets, comprehensive statistical analyses, and their corresponding figure panels (revised Figures 2B, 10, 11 and Supplemental Figure 3) have been fully incorporated into the revised manuscript.

      In addition to adding the experiments described above, we have reorganized and streamlined the paper. First, the CRTC data have been replaced by the Tric-luc data. The CRTC data were taken in the first hour after training and do not shed light on the bulk of the consolidation window. Since the behavioral and sleep effects we see with manipulation of the PAM/DPM microcircuit all occur with a time delay, examining later times in consolidation is more relevant. Additionally, the first hour post-training is quite complex since there are sensory changes and STM processes overlaid on the processes we want to study. Second, we have moved the data in Figure 8 to supplemental (revised Supplemental Figure 2) since they are basically a control for the experiments in Figure 7 validating known requirements for appetitive LTM.

      We have also substantially expanded the Discussion section to contextualize the PAM-DPM circuit within the broader framework of well-characterized memory-regulatory pathways, such as the intrinsic circuits of the mushroom body, and to explicitly delineate the hierarchical interplay between sleep-dependent synaptic plasticity and LTM consolidation.

      We contend that these complementary experimental assays and targeted revisions markedly strengthen the causal evidence underscoring the role of the PAM-DPM circuit as a pivotal regulatory node bridging sleep states and LTM consolidation. We are confident that these revisions essentially address the concerns raised by the reviewers.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behavior, imaging, and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      From a Drosophila sleep researcher's perspective, the investigation follows a clear and logical strategy to collect a huge dataset of sleep, appetitive memory, and live imaging. The authors clearly identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation-dependent manner. The authors also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the authors applied a new way of sleep statistics to demonstrate hour-by-hour changes between treatment and genotypes. Importantly, the authors demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP during the memory consolidation period after training.

      Weaknesses:

      Two investigatory gaps relate to the misalignment between circuital activity and behaviors, due to the nature of large circuital functional analysis like this. Firstly, the central observation of the study indicates that PAM alpha1 activation causes DPM inhibition which disrupts sleep and memory consolidation. Therefore one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found that the endogenous CRTC::GFP reported neuronal activity for PAMalpha1 and DPM are both increased after memory training (Figure 9). This can be due to the difficult functional demarcation among the 14 PAMalpha1 projections. Secondly, the authors acknowledged the contradicting finding that memory defect is detected in PAMalpha1 inactivation (Figure 7C), yet suggested a tight link between sleep and memory consolidation; it is clear loss of PAM subset activity can disrupt memory consolidation without affecting sleep (cf Figure 7C and 7I).

      Thank you for your insightful analysis and the relevant possibilities you've raised. We agree that given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter may not fully reflect the overall dynamics of neural activity in this microcircuit. To better characterize the dynamics of the PAM-DPM circuit in the consolidation window following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. However, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      Regarding the second question, the core finding underlying the link between sleep and memory elucidated in the present study lies in the whole PAM-α1-DPM microcircuit rather than the specific DANs alone. MB299B and MB043B, the two split-GAL4 drivers employed to target PAM-α1 neurons, were originally characterized previously (Aso et al., 2014). However, these drivers also exhibit non-specific labeling of additional cells, and we can not rule out the possibility that such off-target labeling may have masked the subtype-specific necessity in sleep or memory processes.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain poorly understood. The authors present novel findings implicating two small neuronal groups with inhibitory connections, PAM-a1 to DPM, in sleep regulation and LTM consolidation. However, whether the PAM-a1 to DPM microcircuit promotes LTM consolidation through sleep regulation requires further investigation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-a1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly at night. Interestingly, disruption of PAM-a1 and DPM neurons impairs sleep and appetitive memory consolidation only under starvation conditions, and pharmacological induction of sleep during the night rescues the LTM defects. These findings suggest that PAM-a1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important findings that advance our understanding of the link between sleep and memory consolidation.

      Weaknesses

      Some claims lack sufficient evidence or clarity:

      (1) All sleep experiments are conducted under the "training" (temperature-change) condition. While genotypic controls are helpful, additional no-training controls are required to confirm that the observed differences are due to training rather than unknown genotype-related factors. The fact that experimental genotypes exhibit significantly altered sleep even before "training" (e.g., Figs. 7H, J, K, 8A, B, D) highlights the necessity of these controls.

      Thank you for raising this important question. We have re-examined the sleep profiles recorded over two acclimation days and one day of baseline sleep, which preceded the implementation of the “training” paradigm (temperature manipulation) and thus served as a valid no-training control. As shown in Author response images 1-4, subtle yet discernible genotype-dependent differences were indeed observed under baseline conditions. However, when animals were subjected to starvation, the experimental manipulations (activation or inactivation of the target cells) elicited marked, statistically significant alterations in sleep patterns that cannot be accounted for by the baseline genotype differences. Collectively, these data confirm that the observed sleep phenotypes are attributable to the “training” intervention, rather than to confounding, pre-existing genotype-related factors.

      Author response image 1.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under starvation conditions.

      Author response image 2.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under non-starvation conditions.

      Author response image 3.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under starvation conditions.

      Author response image 4.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under non-starvation conditions.

      (2) Previous studies on disrupted memory due to sleep reduction have primarily examined conditions with severe sleep deprivation. In contrast, this report claims that relatively small decreases in total sleep accompanied by sleep fragmentation are responsible for impaired memory consolidation. It remains unclear whether sleep fragmentation at this level is truly critical for memory consolidation. The authors should cause sleep loss and fragmentation of similar magnitude through other means and determine whether it can impair LTM.

      We appreciate the reviewer’s insightful suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      (3) The authors employed a neural activity reporter to show that starvation increases the basal activity of PAM-a1 but not DPM neurons in untrained flies (Figures 9C-E). They observed small increases in the activity of both neuron groups immediately after training but not one hour later. Given the inhibitory connection from PAM-a1 to DPM, it is unclear why both neuron groups show increased activity after training. Additionally, as the authors acknowledge, it is puzzling how the inactivation of PAM-a1 produces similar effects on sleep and memory as DPM inhibition and PAM-a1 activation. Further experiments are needed to clarify these findings, such as manipulating PAM-a1 activity during the one-hour post-training period and evaluating the effect on DPM activity. Including data from training under fed conditions would provide a more comprehensive understanding of state-dependent neural activity. Even if certain experiments are not feasible, these issues warrant further discussion. It is also important to clarify that the term "synchronized" does not imply single-spike-level synchrony.

      Thank you for raising these critical questions. To deepen our understanding of these issues, we have conducted additional experiments and have incorporated them into the revised manuscript. Below are our specific responses to each of your points:

      (1) Regarding the contradiction between "PAM-α1 inhibition of DPM" and a transient increase in the activity of both neurons immediately after training:

      PAM/PAM-α1 neurons are well-documented to respond to reward signals (Liu et al., 2012, Ichinose et al., 2015), while DPM neurons have been shown to respond to both olfactory stimuli and electric shocks, and to form delayed olfactory memory traces (Yu et al., 2005). Thus, the concurrent increase in the activity of PAM-α1 and DPM neurons immediately following training is likely a response to the olfactory and/or sucrose stimuli in the assay. Given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter likely does not fully reflect the overall dynamics of neural activity in this microcircuit. Additionally, this time window overlaps with the period in which the animals are adapting to the new tubes and is likely contaminated with other sensory information.

      To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 9F-I). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics post-training. We have replaced the CRTC data with this more relevant data set.

      (2) Regarding the state-dependent neural activity:

      We agree that investigating state-dependent neural activity would be an interesting extension of our study. However, this falls beyond the scope of the current study and will be considered in future research. Our primary findings, including sleep disruptions and the associated memory impairments, were specifically observed under starvation conditions, which align with the appetitive memory paradigm employed here. Delving into neural activity changes under non-starvation state would not yield direct evidence to support the core conclusions of the present work, as the study’s focus is on the starvation-dependent interplay between sleep, neural circuitry, and appetitive memory consolidation.

      (3) Regarding the terminology of “synchronization”:

      We believe that the use of the term “synchronization” in our study is appropriate. In the context of neural circuitry, synchronization refers to the process by which distinct neurons or neural populations achieve temporal alignment of their activity, a phenomenon that supports neural communication and information integration. In the present work, this specifically describes how PAM-α1 and DPM neurons exhibit phase-related temporal coordination of their activity to regulate the interplay between sleep and memory consolidation.

      (4) The authors considered that PAM-a1 and DPM might function in parallel, independent pathways for sleep and LTM. They rejected this possibility based on the lack of additive effects when both neuronal groups were simultaneously inactivated. However, they found that MB299B-labelled neurons exert stronger memory effects than MB043B-labelled neurons, while MB043B neurons have stronger sleep effects. If sleep is a primary driver of memory consolidation, a stronger correlation between memory and sleep effects would be expected. This observation merits further discussion.

      We appreciate the reviewer’s constructive suggestions. We have performed additional experiments to explore a well-characterized memory-related PAM-α1 recurrent loop in sleep regulation. The new data, along with further discussion, have been incorporated into the revised manuscript.

      The two split-GAL4 drivers (MB299B and MB043B) used to target PAM-α1 neurons were originally characterized previously (Aso et al., 2014). However, these drivers exhibit non-specific labeling of additional neuronal populations, a technical limitation that may have masked the subtype-specific functional requirements of PAM-α1 in sleep and memory processes.

      In addition, we assessed sleep and LTM following the thermoactivation of DPM neurons (revised Supplemental Figure 1), and no significant changes were observed in either phenotype.

      PAM-α1 has previously been demonstrated to drive appetitive LTM formation and consolidation via a recurrent loop with MBON-α1 (Ichinose et al., 2015). To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions. Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Combined with our observation that inhibition of MBON-α1 during the memory consolidation phase also impaired 24 h LTM, these new data indicate that MBON-α1-mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies.

      Author response table 1.

      (5) Given prior knowledge that PAM neurons are heterogeneous and that the R58E02 driver is broadly expressed, data in Figures 1-5 concerning PAM are outdated. The use of more restricted PAM-a1 drivers from the outset would make the manuscript easier to read and interpret.

      We sincerely appreciate the reviewer’s point of view regarding the selection of PAM drivers. While we acknowledge the well-characterized heterogeneity of PAM neurons and the broad expression profile of the R58E02 driver, and fully agree that employing subtype-restricted drivers enhances the precision of functional interpretation, this set of experiments serves as an essential foundational step and logical basis for subsequent subtype-specific investigations and thus merits retention in the manuscript. As detailed above, the more specific drivers also have some drawbacks in terms of additional expression, making the broad driver critical for setting the stage.

      (6) Some figures lack relevant data, certain experiments are missing necessary controls, and anomalies are present in some data sets.

      We sincerely appreciate the reviewer’s detailed suggestions, and we have revised the manuscript comprehensively in accordance with them.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence that dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strengths:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuits interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of dopamine's role in cognitive functions.

      Weaknesses:

      While the study is well-designed and presents compelling findings, some aspects require further clarification. The interpretation of dopamine receptor signaling remains incomplete, particularly regarding inhibitory pathways. The role of DPM in memory consolidation is not entirely conclusive, as different genetic approaches yield variable results. Additionally, some inconsistencies in neuronal activity patterns and experimental variability, especially regarding sleep patterns or pharmacological rescue, should be addressed to strengthen the mechanistic framework.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopamine circuits interact to regulate memory consolidation. While the findings are compelling, addressing the points above-particularly receptor signaling and the specific role of DPM and its activity patterns within the microcircuit would further solidify the study's conclusions.

      We sincerely appreciate the reviewer’s constructive feedback and useful suggestions, which have been instrumental in enhancing the rigour and completeness of our study.

      To address these points, we have performed a series of additional experiments that we believe strengthen the mechanistic framework of our work. The key new findings are summarized below:

      (1) Regarding the dopamine receptor signaling

      To define the dopamine receptor (DAR) signaling mechanisms underlying DPM neuron activity and its regulatory roles in sleep and memory, we first characterized DAR expression profile of the DPM neurons. Using double-labeling assay, we detected robust expression of Dop1R1 and Dop1R2 in DPM neurons, whereas no detectable colocalization was observed for DopEcR and Dop2R (revised Figure 10A). Accordingly, we refined our FRET-based EPAC data by removing the DopEcR knockdown group, and now present cAMP changes in DPM neurons following Dop1R1 and Dop1R2 knockdown, in direct comparison with the intact receptor control group (revised Figure 10B). These data conform that Gαs-coupled Dop1R1 is the primary receptor mediating DA-dependent cAMP elevation in DPM neurons.

      To further identify the DARs responsible for transducing DA-induced inhibitory effect on DPM neural activity, we quantified GCaMP levels in DPM neurons with targeted knockdown of individual DARs. Knockdown of either Dop1R1 or Dop1R2 failed to abolish DA-induced Ca<sup>2+</sup> decrease; only Dop1R2 knockdown exhibited a trend toward attenuating this Ca<sup>2+</sup> decrease (revised Supplemental Figure 3C-E), suggesting that the two receptors cooperate to modulate DPM neural activity.

      Finally, to dissect the specific contributions of DARs in DPM neurons to sleep and/or memory regulation, we performed sleep monitoring and memory assays in animals with DPM-specific knockdown of distinct DARs (revised Figure 10C-H). Knockdown Dop1R1 in DPM neurons resulted in statistically significant sleep reduction, decreased arousal threshold, and impaired 24 h LTM memory (revised Figure 10C-E). In contrast, knockdown Dop1R2 in DPM neuron selectively impaired 24 h LTM memory with no effect on sleep (revised Figure 10F-H). Collectively, these findings demonstrate that coupling sleep and LTM requires Dop1R1 in DPM neurons through the modulation of both cAMP signaling and neuronal activity, while Dop1R2 specifically mediates LTM regulation, likely through modulating DPM neural activity alone.

      (2) We have additionally characterized the role of MBON-α1 in sleep, which has been previously shown as a PAM-α1-related recurrent feedback loop in the regulation of memory formation and consolidation (Ichinose et al., 2015).

      To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions (revised Figure 11A-D). Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Moreover, inhibition of MBON-α1 during the memory consolidation phase significantly impaired 24 h LTM (revised Figure 11E-F). These results indicate that MBON-α1mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies (See Author response table 1).

      We have modified the schematic diagram in the revised manuscript to illustrate the mechanistic framework (revised Figure 12).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here I listed details for potential clarification or further investigation related to the weaknesses:

      (1) Line 145-147: I suspected the authors used previously verified RNAi lines, but it would be informative to include a citation or their own validation for the effectiveness of these RNAi lines.

      We sincerely appreciate the reviewer’s suggestion. As our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2 and #3), we refined the revised data to focus exclusively on these two DARs. Corresponding revisions have been made to the Materials and Methods, Results and Discussion sections. Additionally, we conducted qPCR analysis to verify the knockdown efficiency of these DARs, providing further support for our findings that Dop1R1 and Dop1R2 are functionally required in DPM neurons for the regulation of sleep and memory (revised Supplemental Figure 3).

      (2) Line 169: Moving from describing Figure 3A/B to Figure 3C, it is not immediately clear from 3C-H, the authors follow the training paradigm of 3B?

      To enhance clarity, we have added the referenced figure citations in the “Memory assay” section: “For all 24 h sucrose-odour memory, a single training session of sucrose paired with an odour for 2 min was employed (Figure 3A).”

      (3) Line 179: before the PAM inactivation data are shown in Figure 7, the authors seem to be getting ahead of themselves by stating "suggesting that activity of PAM neurons is necessary for the consolidation window or that heterogeneity in the subsets of PAM neurons masks any phenotype." when the Figure 3 data collectively indicate that "suppression" of PAM is necessary.

      This statement is based on our observation that inactivation of the majority of PAM neurons labeled by R58E02 results in sleep disruption but leaves memory intact, and we stand by this conclusion.

      (4) Line 186-188: The labelling of DP1 is not entirely aligned between the figures and the text for a reader to follow which time period is described, as DP1 is embedded within the dark phase in the figures.

      These experiments spanned two consecutive days. LP1 and DP1 denote the light and dark periods on the first day, respectively, whereas LP2 designates the light period on the second day. As only one full dark phase was monitored across the experimental interval, we characterized the relevant phenotype using the general terms dark phase or nighttime, rather than specifying DP1. We thank the reviewer for this thoughtful observation; nonetheless, we consider the original description correct and unambiguous, and thus appropriate for inclusion in the manuscript.

      (5) Line 321: The statistics for Figure 9 CRTC::GFP measurement is crucial for interpretation but the referring and labelling for this on Figure 9 is poor: it is not apparent which comparisons are indicated. There is inconsistency between Table 1 Figure 9D and Table 3 Figure 9D entries: no significant between train and untrain indicated in Table 1 but it is described as significant in the text and Table 3?

      Thank you for this observation. As described above, we have removed these data from the paper and replaced them with Tric-Luc data that capture the consolidation window more completely.

      (6) Line 423: The starvation-mediated sleep suppression is not clear in this manuscript, can the author comment on this? The response to this may also alter the summary concept cartoon.

      This is an important point. To directly address the reviewer’s question regarding starvation-mediated sleep suppression, we have generated a representative response figure comparing sleep duration under starvation versus non-starvation conditions (Author response image 5). This figure clearly demonstrates that sleep is suppressed under starvation, providing straightforward evidence to address this concern.

      Author response image 5.

      Examples of starvation-mediated sleep suppression.

      However, the key focus of our study is that changes in neuronal activity disrupt sleep under starvation conditions but not under non-starvation conditions. To emphasize this critical distinction, we have incorporated additional discussion focused specifically on this point.

      “It is well established that starvation induces sleep suppression (MacFadyen, 1973; Thimgan et al., 2010; Melnattur and Shaw, 2019; Keene et al., 2010; He et al., 2020; Yangkyun et al., 2022), and our results are consistent with these previous findings: all genotypes exhibited less sleep under starvation than under fed conditions (i.e. Figures 4A-B, 5A-B and 11). Under normal appetitive memory training, starvation-induced sleep loss does not necessarily impair memory processing (Thimgan et al., 2010; Chouhan et al., 2021). PAM-α1 neuronal activity is higher in starved, trained flies than in fed or untrained flies (data not shown), suggesting that these neurons act as a critical node for integrating internal motivational and arousal states, as well as conveying positive valence for the normal appetitive memory process, independently of starvation-induced sleep loss. While DPM neurons are less sensitive to starvation, they still exhibit training-induced elevated activity (data not shown), indicating coherent responsiveness to upstream signaling. In the present study, we found that under fed conditions, sleep remained intact even when excessive changes in neural activity occurred within the PAM(-α1)-DPM circuit; in contrast, under starvation conditions, significant sleep reduction and fragmentation were observed. These observations indicate that starvation may trigger a transition from a physiologically normal brain state to an unstable, abnormally active state, which consequently elicits negative behavioral outputs.”

      (7) Line 1121: The data points for Figure 9 D-E are surprisingly low considering there are 14 PAMalpha1 labelled, the data presented here indicated potentially only 1-2 neurons were counted per fly brain. Can this contribute to the large variation and the contradiction of PAM's memory-suppressing role?

      We sincerely appreciate the reviewer’s critical comments regarding the sample size of labeled PAM-α1 neurons in Figure 9D–E. We have revisited our raw data, incorporated additional brain samples, and reanalyzed the dataset. For this updated analysis, we included all clearly distinguished neurons, excluded overlapping ones, and calculated a single NLI per brain for statistical analysis. The key conclusions remain consistent with those in the original submission, confirming the robustness of the observed phenotype.

      Memory consolidation is a time-dependent process. To further elucidate the link between neural activity and behavioral outputs, we performed additional experiments with an extended recording period. A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (8) Line 345: the effect size and data spread of THIP restored memory is different from the controls in Figure 10, perhaps warranting a more conservative interpretation of the role of sleep in memory consolidation.

      We appreciate this critical comment. We fully agree that the role of sleep in memory consolidation requires cautious interpretation, a point we have integrated into the revised manuscript.

      Drug treatment in Drosophila, particularly for group-based assays, can introduce substantial variability at both the individual and group levels. To account for this, we employed a statistically valid sample size for our analyses to ensure robust conclusions. While minor quantitative discrepancies exist in the data, this technical consideration does not significantly alter the core conclusions of the study.

      Reviewer #2 (Recommendations for the authors):

      (1) As mentioned in the public review, all data using the broad PAM-DAN driver should be removed. Concerns regarding the experiments involving the broad driver are not included here.

      A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (2) In GCaMP experiments (Figure 9B), the ΔF/F traces for the AHL and AHL+ATP conditions start diverging before the addition of ATP. The quantification shows they are not significantly different in the first 30 seconds, but the fact that in two separate experiments (2A and 9B), they diverge in the same direction makes me wonder whether the AHL condition is different from the +ATP condition even before the ATP treatment. Also, the traces should include standard errors.

      We observed the same diverging trend in the first 30-second baseline as the reviewer. We reviewed the raw data for each sample and found that this divergence is likely attributable a small number of outliers. Given the absence of a statistically significant difference, this divergence does not affect our conclusions.

      We have also added standard errors to the revised figures.

      (3) Figure 9B. The authors need to show data for a control genotype. +>P2X2; VT064246-LexA > GCaMP6f that does not include MB299B-Gal4 is crucial to demonstrate that expression of P2X2 in PAM-α1 is responsible for the inhibitor effect, as LexA-P2X2 may be leaky.

      One of the UAS-P2X2 lines was found to exhibit leaky expression, so we instead used a non-leaky UAS-P2X2 line for all related experiments. To address the reviewer’s comments and further validate our findings, we have added complementary experiments with a control genotype. In addition, we also added a control to confirm the non-leaky expression of LexA-P2X2 under the driver of R58E02-LexA. As shown in revised Figures 8B, application of ATP in the absence of MB299B-GAL4 failed to induce a significant inhibitory effect. These data strongly and convincingly support our conclusion.

      (4) Figure 9B. Some of the individual data show values lower than -100% ΔF/F0. By definition, ΔF/F cannot be less than -100%, as this would require negative fluorescence, which is physically impossible. The calculation of fluorescence changes using ΔF/F should be carefully reconsidered.

      We thank the reviewer pointing out this potential confusion. We used a standard method of calculating the change in fluorescence over time using △F/F = (Fn-F<sub>0</sub>) / F<sub>0</sub>×100% as we previously described (Liu et al., 2019). Changes of greater than +100% of △F/F would not be unusual, since the reported value is a ratio to the initial level of fluorescence, not a subtraction of the baseline value from the signal (which obviously could not go below 100%). We have included a sentence in the results explaining this (page 7): “As previously described, we used the percent change in fluorescence over time as a ratio to the initial level, △F/F = (Fn-F0)/F0×100% for quantification (Liu et al., 2019).” And we have carefully reviewed our raw and processed data and confirmed that our analysis was correct.

      (5) Figure 2B. The number of UAS transgenes should be controlled, as Gal4 could be diluted with 3 UAS constructs in experimental conditions compared to only 1 UAS construct in controls. Are Dop1R2 and DopEcR significantly different from wt? Why do they present an average ΔF/F in 2A and a maximum in 2B?

      We appreciate the reviewer’s careful observations and valuable comments.

      As the reviewer noted, the EPAC imaging experiment utilizes three UAS transgenes, which enable Gal4 enhancement via Dicer, targeted manipulation of dopamine receptor expression levels, and neural activity monitoring in DPM neurons. All other imaging experiments in the study employ only one or two UAS transgenes. Given the robustness of the observed phenotypes, the potential dilution effect is not a major concern. Knockdown of Dop1R2 and DopEcR showed no significant differences relative to the WT control group; the maximum values presented in Fig. 2B are included solely to illustrate statistical significance. While the EPAC (CFP/YPF) signal reflects an obvious cAMP elevation, no differences were detected in the averaged signal across groups.

      Notably, in the revised manuscript, our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2). Accordingly, we have refined our data analysis to focus exclusively on these two DARs.

      (6) Figures 7H, J. Why is almost every MB299B>TrpA1 fly sleeping at ZT0?

      To align the starvation protocol for sleep analysis with that used in the memory assay, MB299B>TrpA1 flies and their genetic controls were transferred to fresh sleep tubes containing starvation food during the ZT0–1 time window. This transfer resulted in no detectable locomotor activity during this period, a pattern indicative of sleep in all flies.

      (7) The number of episodes and P(wake) should be presented for all sleep data.

      We have added these two parameters as new panels to all relevant sleep figures. The corresponding statistical analyses have also been included in the supplemental tables.

      Reviewer #3 (Recommendations for the authors):

      The study's findings provide compelling insights into the neural circuits connecting sleep and memory and the role of dopamine in general. While the anatomic dissection of the microcircuit and its overall involvement in sleep and memory is convincing and of high interest to the field and beyond, some statements of the study need further clarification, particularly the interpretation of receptor signaling and the role of DPM.

      Major Points

      (1) Figure 2: cAMP Imaging and Dopamine Receptor Involvement

      The authors present calcium and cAMP imaging to support the inhibitory connection between PAM and DPM neurons. While using both sensors is a robust approach, I am not entirely convinced that cAMP imaging is the ideal approach for identifying the dopamine receptors involved. To my knowledge, only Dop1R1 is classically linked to Gs-mediated cAMP signaling. Dop1R2 is typically coupled to Gq (PLC and DAG), while DopEcR is non-canonical and can engage both pathways. Additionally, these receptors are classically excitatory, yet the authors did not analyze Dop2R, the primary inhibitory dopamine receptor - which would represent the most relevant candidate for an inhibitory PAM-DPM connection.

      We have addressed this point in our response to the public comments, so will not reiterate here.

      While dopamine receptor functions can vary by neuronal context, I would appreciate clarification on the following points:

      (a) Why was Dop2R not tested? Was it omitted or found to have no effect?

      We sincerely appreciate the reviewer’s critical questions. This point has been addressed in our response to the public comments. Briefly, Dop2R is not colocalized with DPM neurons; instead, only Dop1R1 and Dop1R2 are detected in DPM neurons, which is why we focused exclusively on these two receptors in the revised manuscript.

      (b) Why was cAMP imaging chosen for receptor identification? Was calcium imaging performed, and if so, what were the results?

      This is an excellent point, and we sincerely appreciate the reviewer’s valuable input, which has helped to strengthen the logical framework of our analysis on receptor-mediated neural activity. These dopamine receptors are well-characterized as members of the Gas-coupled protein receptor family, and cAMP signaling serves as a reliable readout of their functional activity. To strengthen the logic flow of our analysis on the target inhibitory circuit, we have made the following key revisions to the manuscript: 1) defined the expression profile of dopamine receptors in DPM neurons; 2) refined our cAMP imaging data analyses based on specific receptor subtypes; and 3) assessed DPM neural activity via calcium imaging under conditions of targeted receptor knockdown. For further details, please refer to our response to the public comments.

      (c) Since the data suggest multiple receptor involvements and complex interactions, I encourage a more detailed discussion of the working hypothesis, particularly regarding the unexpected finding that classically excitatory receptors contribute to an inhibitory connection.

      We appreciate the suggestion to elaborate on our working model. Accordingly, we have revised the schematic diagram and refined the manuscript to clearly illustrate the underlying mechanistic framework. For further details, please refer to our response to the public comments.

      (2) Figure 3: DPM Involvement in Memory Consolidation

      The authors show that PAM activation and DPM inhibition during consolidation impair appetitive LTM. However, the role of DPM is critical. While the c316-GAL4 driver yields strong effects, VT064246 inhibition shows only slight significance, requiring more than twice the sample size of other experiments. Given that c316-GAL4 is not DPM-specific and also labels MB Kenyon cells, I suggest using MB-GAL80 to restrict expression - or commenting on the possibility that other neurons like MB-KCs could directly participate in the phenotype. This is particularly relevant since VT064246 efficiently modulates sleep, indicating that it is generally effective in altering behavior. These issues weaken the claim that DPM plays a crucial role in linking sleep and memory, and should be addressed. Minor comment on this Figure: In the Figure legend, the driver and "n" are not mentioned for 3C, while this is the case for all other panels. Moreover, the DPM schematic only depicts the MB, making it somewhat confusing. DPM innervates the entire MB, still, it would be helpful to shade the DPM projections more distinctly within the MB for clarity.

      We thank the reviewer for the suggestion to improve the precision of our figures.

      Regarding the expression specificity concern, in all experiments using c316-GAL4, we had eyeless-GAL80 and MB-GAL80 co-expressed to restrict GAL4-driven expression to DPMs. While complete suppression of expression of MB-KCs was not achievable, we largely eliminated the potential confounding effects from majority of these cells. VT064246-GAL4 is known to exhibit weak expression (Jenett et al., 2011; Haynes et al., 2015), but high relative specificity. Importantly, the overall conclusion derived from experiments using c316-GAL4 with GAL80s and VT064246-GAL4 are consistent, which strongly supports the role of DPM neurons in mediating the link between sleep and memory.

      As suggested, we have added sample sizes for all panels and refined the depiction of DPM projections in revised Figure 3C.

      Minor Comments

      (1) Introduction:

      The authors introduce dopamine's role in forgetting but focus on aversive rather than appetitive memories. To avoid confusion, this distinction should be mentioned explicitly (likewise in the discussion). Regarding references: Zhang et al. (line 95) do not discuss DPM or APL. Donlea et al. (line 97) do not cover dopamine - I think Pimentel et al. (2016) would be a more appropriate citation.

      This is a good point. We have removed Zhang et al. (2013) and replaced Donlea et al with Pimentel et al. 2016 as suggested.

      (2) Figure 9: DPM Activation During Consolidation:

      The authors show that PAM neurons are activated by starvation and further enhanced by appetitive training. Surprisingly, DPM neurons also increase activity post-training, despite the proposed inhibitory connection between PAM and DPM. The authors state that "PAM-α1-DPM microcircuit exhibits synchronized neural activity changes during the consolidation window" (line 326), yet they do not address this apparent contradiction. If I have not overlooked key information, this should be clarified/addressed e.g. in the discussion.

      This is an excellent point. We have addressed this in our response to point (3) from Reviewer #2 in the public comments, so we will not reiterate it here.

      (3) Figure 10D/E: THIP Rescue of LTM Deficits:

      Some inconsistencies in the THIP rescue experiments need clarification:

      (a) In Figure 10D, MB299B activation with THIP appears not to significantly restore memory relative to zero, nor to differ from untreated conditions in Figures 7A or 10E.

      (b) In Figure 10E, MB299B activation +/- THIP shows a much clearer effect.

      (c) Are Figures 7A, 10E, and 10D independent experiments, or were they conducted together?

      (d) Should the left bar in 10D and the right bar in 10E be identical? If not, I do not fully understand the discrepancy and suggest discussing the variation.

      Upon revisiting the raw datasets and conducting a one-sample t-test to analyze the group differences, the experimental group in Figure 7A showed no significant difference from the theoretical mean (set at zero). This group also did not differ from the two genetic controls, indicating that the restored memory was comparable to control levels. In Figure 10E, the group with MB299B activation plus THIP treatment exhibited a significant difference from the theoretical mean (one-sample t-test) and from the non-THIP control group, confirming a significant restoration of memory function. Owing to our laboratory relocation, the starvation duration at the new facility was adjusted based on a recalibrated starvation curve; the higher overall 24 h memory index in Figure 10E is likely attributable to a relatively longer starvation period. However, this experimental parameter variation does not alter the study’s overall conclusions.

      (4) Sleep Phenotypes and Starvation Effects:

      Sleep scores are shown under starvation/fed conditions but not under baseline conditions (without inhibition/activation). Could the authors indicate whether they observe basal starvation-induced sleep changes? The authors frequently state that PAM-DPM effects on sleep are context-dependent, yet mild but significant changes occur under fed conditions. I suggest rewording to clarify that the effect is enhanced in a context-dependent manner rather than strictly context-dependent.

      Starvation-induced sleep reduction is a well-characterised phenotype. Our study focused on the key question of whether altered neuronal activity modulates sleep under innate starvation conditions. Accordingly, all comparisons were made between the experimental and control groups under both starvation and fed conditions. We appreciate the reviewer’s suggestion to improve clarity and have revised the text as suggested.

      (5) Starvation Duration in Methods:

      The authors use 30-46h of starvation, which is longer than the ~20h typically used in appetitive memory studies. Could the authors explain why such extended starvation times were necessary?

      Determining starvation levels via survival curves is a well-established and relatively objective method, one that has been widely adopted in prior studies. For the memory test, we standardized the total starvation duration for each genotype to the time point at which mortality reached 20%. Owing to inherent differences in to starvation resistance across distinct genotypes, the final starvation durations ranged from 20 hours to 46 hours.

      (6) Variability in PAM-α1 Sleep Effects:

      (a) The extent and timing of sleep effects differ across PAM-α1 drivers (e.g. night vs. light-period effects). Could MBON co-targeting by these drivers contribute to the variability?

      We have supplemented additional experiments to investigate the effects of MBON-α1 neurons on 24h memory and sleep. For detailed findings, please refer to our response to your public comments.

      (b) Even within the same driver, results differ (e.g., Figure 7H vs. 10A). A general comment on these differences would be important, e.g. regarding the relevance of day and night sleep for memory consolidation.

      We sincerely appreciate the reviewer’s incisive observation regarding these details. The discrepancy stems from the timing of neuronal activity inhibition, during which a laboratory relocation led to adjustments in starvation duration for memory experiments, which in turn indirectly altered sleep patterns.

      (c) Technical note: Similar y-axis scales for sleep plots (Figures 10A and B) would make comparison easier.

      We have unified the y-axis scales to the same range.

      (7) Discussion, Line 376:

      The phrase "sleep deprivation is important for memory consolidation" is misleading, as it could imply that deprivation aids memory formation. Please clarify.

      We appreciate the reviewer’s suggestion. We have revised the text to: “These results demonstrate that preserving unperturbed sleep during the critical memory consolidation window is essential for stabilizing appetitive long-term memory.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds upon a major theoretical account of value-based choice, the 'attentional drift diffusion model' (aDDM), and examines whether and how this might be implemented in the human brain using functional magnetic resonance imaging (fMRI). The aDDM states that the process of internal evidence accumulation across time should be weighted by the decision maker's gaze, with more weight being assigned to the currently fixated item. The present study aims to test whether there are (a) regions of the brain where signals related to the currently presented value are affected by the participant's gaze; (b) regions of the brain where previously accumulated information is weighted by gaze.

      To examine this, the authors developed a novel paradigm that allowed them to dissociate currently and previously presented evidence, at a timescale amenable to measuring neural responses with fMRI. They asked participants to choose between bundles or 'lotteries' of food times, which they revealed sequentially and slowly to the participant across time. This allowed modelling of the haemodynamic response to each new observation in the lottery, separately for previously accumulated and currently presented evidence.

      Using this approach, they find that regions of the brain supporting valuation (vmPFC and ventral striatum) have responses reflecting gaze-weighted valuation of the currently presented item, where as regions previously associated with evidence accumulation (preSMA and IPS) have responses reflected gaze-weighted modulation of previously accumulated evidence.

      A major strength of the current paper is the design of the task, nicely allowing the researchers to examine evidence accumulation across time despite using a technique with poor temporal resolution. The dissociation between currently presented and previously accumulated evidence in different brain regions in GLM1 (before gazeweighting), as presented in Figure 5, is already compelling. The result that regions such as preSMA response positively to |AV| (absolute difference in accumulated value) is particularly interesting, as it would seem that the 'decision conflict' account of this region's activity might predict the exact opposite result. Additionally, the behaviour has been well modelled at the end of the paper when examining temporal weighting functions across the multiple samples.

      In response to reviewer comments, the authors have explicitly tested for the effects of gaze-weighting over and above any main effect of value, and convincingly shown that these effects are both present in the main regions of interest - namely |SV| and gazeweighted |SV| in the vmPFC, alongside |AV| and |AV_gaze| in the pre-SMA. This provides clear evidence in support of the notion of gaze-weighting of value signals in these regions.

      We thank the reviewer for their comments.

      Reviewer #2 (Public review):

      Summary:

      In this paper the authors seek to disentangle brain areas that encode the subjective value of individual stimuli/items (input regions) from those that accumulate those values into decision variables (integrators) for value-based choice. The authors used a novel task in which stimulus presentation was slowed down to ensure that such a dissociation was possible using fMRI despite its relatively low temporal resolution. In addition, the authors leveraged the fact that gaze increases item value, providing a means of distinguishing brain regions that encode decision variables from those that encode other quantities such as conflict or time-on-task. The authors adopt a region-of-interest approach based on an extensive previous literature and found that the ventral striatum and vmPFC correlated with the item values and not their accumulation whereas the preSMA, IPS and dlPFC correlated more strongly with their accumulation. Further analysis revealed that the pre-SMA was the only one of the three integrator regions to also exhibit gaze modulation.

      The study uses a highly innovative design and addresses an important and timely topic. The manuscript is well-written and engaging, while the data analysis appears highly rigorous.

      Weaknesses:

      With 23 subjects the study has relatively low statistical power for fMRI although the within-subjects design and relatively high trial count reduces these concerns.

      We thank the reviewer for their comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Something seems to have gone slightly wrong in (I think) the labelling of new figure 7 (the correlation matrix between the different regressors). There are five variables on both the x- and y-axes of the figure, and they are the same five variables - meaning the diagonal of the matrix would be all equal to 1 (being the correlation of a regressor with itself - e.g. |AVgaze| with |AVgaze|). In the figure legend, six variables are mentioned, including lagged |deltaAVgaze| - but this doesn't appear on the x or y-axes. I suspect that the authors may need to check that this matrix has been calculated correctly, and isn't being mislabelled?

      We thank the reviewer for noticing this issue. We have now corrected Figure 7.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank the Reviewers for the favourable feedback. There is no additional comments from Reviewers 1, 3, and 4, and we address the minor concerns from Reviewer 2 as follows.

      (1) I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      We have revised our manuscript to state: “The DC has been proposed to exist as biomolecular condensates.” Additionally, we have included a time-lapse image showing that AXIN1-GFP puncta exhibit dynamic fusion behaviour in cells (Fig. S8), suggesting that the DC may be liquid-like, at least with AXIN1 overexpression.

      (2) I thank the authors for including the additional data comparing tankyrase binding by IWR and IWRPOMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC<sub>50</sub> values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Added.

      (3) Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue? Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

      We now specify in the Introduction that TNKS1/2 are encoded by TNKS/TNKS2, and are collectively referred to as TNKS in this manuscript.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editor and the reviewers for their comments on the revised manuscript. Based on the comments, we decided to go for a minor revision which will address all the comments of reviewer 1.

      Towards the comments of the reviewer 2, we would like to state that we already provided the results from the new experiments and the reasons why we did not perform some of the suggested ones. Interestingly, we have not deviated from standard practices in the field in our approaches. Yet, the reviewer is not convinced and raised concern about the robustness of our observations. We therefore decided to carry out a few more control experiments which in our opinion are redundant as they already were carried out multiple times by us as well as by the other field experts under identical conditions and using the identical cell lines.

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix promoting cancer cell invasion, a process here investigated using gelatine-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endodomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metallo protease MT1-MMP.

      The delivery of MET from the recycling compartment to invadopodia is mediated by RCP which facilitates the colocalization of MET to RAB14 endosomes. On this compartment, HGF induces the recruitment of the motor protein KIF16B promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface.

      This pathway is critical for the HGF-driven invasive properties of TNBC cells as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well organized and executed using state of the art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosisdefective.

      Data analyses are rigorous and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well supported by the results and the data and methodology are of interest for a wide audience of cell biologists.

      Previous Weaknesses:

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings including Triple Negative breast cancer cells. The novelty of the present study mostly consists in the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGFmediated invasion detectable in other cancer cells is not addressed or at least discussed.

      Comments on revised version:

      The authors have partially replied to my previous concerns.

      We sincerely thank the reviewer for careful evaluation of our manuscript and recognizing the strength of our study. We are grateful for the positive assessment about the well-executed methods, rigorous data analysis, usage of appropriate controls, and for acknowledging that we have addressed, at least in part, the concerns raised in the previous round of review. We appreciate the reviewer’s constructive comments, which have helped us further clarify the scope and significance of our findings.

      Reviewer #1 (Recommendations for the authors):

      The authors have partially replied to my previous concerns. The following points still have to be addressed.

      (1) Despite many TNBC tumours present high expression of the EGFR, trials with EGFR inhibitors have been very disappointing (please read PMID: 41651315). EGFR inhibition has been extensively attempted with negative outcome, and this is well known.

      In the clinical practise, Triple Negative breast cancer patients are commonly treated with chemotherapy, not with EGFR or MET inhibitors.

      Line 31-38 are misleading at best and should be removed and the incipit of the study modified. The biology of this study is sound there is no need to push the clinical relevance with instances that are notoriously not applicable.

      We are thankful to the reviewer for bringing up this point. We will modify the manuscript as per the reviewer’s suggestion.

      (2) In reply to point 4, the authors claim that they checked MMP2 and the experiment is shown in FigS5I, which is not......

      We sincerely apologise for uploading an incorrect file. The correct file will be uploaded.

      (3) I noticed that, in the previous version of the manuscript in Fig. S4A the panel showing the mutation frequency was erroneously indicated in the legend as referring to MET.

      The authors replied that they fixed this but actually the legend is still wrong...

      We thank the reviewer for bringing this error into our attention. We will attentively correct the legend in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligandinduced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights.

      Comments on revised version:

      I appreciate the authors' efforts to revise the manuscript and address the reviewers' comments. While the revised version includes additional experiments and several improvements in data presentation, the major methodological and conceptual concerns raised in the initial review remain largely unresolved. In my opinion, these issues critically undermine the central mechanistic conclusions of the study.

      We thank the reviewer for critically re-evaluating our manuscript and acknowledging the novel insights and relevance of our study.

      We thank the reviewer for pointing out the study by Rajadurai et al., Journal of Cell Science, 2012, which provided crucial evidence about the role of MET signaling in invadopodia formation [1]. However, the experimental system largely used by Rajadurai et al. is fundamentally different from the receptor trafficking mechanism investigated in the present study. The group have mostly used overexpression of Tpr-MET, which is a cytosolic MET mutant, that does not undergo the canonical ligand-induced RTK endocytosis and subsequent degradation or recycling. Our study did not only establish another link between MET signaling and invadopodia formation; rather, we identified a trafficking-dependent mechanism whereby HGF stimulation regulates the spatial redistribution and recycling of full-length MET to invadopodia, thus providing more physiologically relevant insights.

      We also respectfully disagree that the methodological concerns raised critically undermine our mechanistic conclusions. Although other approaches as suggested by the reviewer could provide complementary information, we believe that the methods used in our study are appropriate for assessing invasive behaviour of the TNBC cells. These methods have been used by us and other research groups in the filed as reflected from the existing literature [2–6].

      In this context, we would also like to add that the study referred by the reviewer above, has used a more off target prone approach (SiGenome Smartpool) compared to ONTARGETplus (chemically modified SiRNA pool for minimizing off target effect) in addition to the same small molecule inhibitor used in our study. Additionally, we also used a shRNA-based silencing to verify the phenotype. Since the silencing was not as pronounced as the siRNA-mediated knockdown, the reviewer has expressed concern.

      We would like to point out that the antibody we used in IF for MT1-MMP (MMP14) have been published in multiple peer-reviewed journals by various research groups using the identical cell lines (MDA-MB-231, ATCC- HTB-26) [6–8]. MET antibodies used in the study (CST, D1C2 XP & L6E7) has been validated in MDA-MB-231 and MET-depleted cells [9-12]. Since these antibodies have been utilized for IF since a long time across various research groups under identical laboratory/experimental conditions, the exercise of validation suggested by the reviewer is surprising.

      (1) Inappropriate experimental design for studying MET trafficking

      A major concern remains the use of prolonged HGF stimulation times (2-6 hours) to study MET endocytosis and recycling. This is not an appropriate experimental design for investigating receptor tyrosine kinase trafficking dynamics. Ligand-induced internalization of MET occurs within minutes, with maximal endosomal accumulation typically observed within 5-15 minutes, whereas recycling occurs over approximately 15-60 minutes.

      Importantly, the authors have not included short stimulation time points or any kinetic analysis that would allow a proper assessment of MET internalization or recycling. The additional surface biotinylation experiment does not address this issue, as it still does not provide temporal information regarding receptor trafficking.

      Therefore, the current data do not support the conclusions regarding MET endocytosis or recycling, and this major methodological concern has not been adequately addressed in the revised manuscript.

      We want to clarify that, our prime objective is to determine how MET trafficking is regulated at the time points at which we observe the functional effects of HGF on invadopodia formation and matrix degradation. Since our functional assays were performed following 2-3 h of HGF stimulation, we specifically examined MET localization and trafficking at these same time points.

      Though shorter time points could provide information on the kinetics of MET trafficking, but their absence does not invalidate our conclusions regarding the role of MET trafficking in HGF-induced invasive function. In other words, our conclusions are made for the time points for which we have conducted the experiments. It is needless to add that different cargo molecule will show different kinetics.

      In summary, we wanted to study the MET trafficking at the late hours in accordance with our functional assays and accordingly designed our experiments. Also, we have supported our results through biochemical methods which is considered to be one of the gold standards in the field.

      (2) Insufficient validation of antibody specificity in immunofluorescence

      The validation of antibody specificity for MET, phospho-MET, and MT1-MMP in immunofluorescence experiments remains insufficient. While the authors demonstrate knockdown efficiency by immunoblotting and show some reduction in fluorescence signal, they do not provide rigorous evidence that the immunofluorescence signal is specifically abolished upon gene silencing under identical imaging conditions. Such validation is essential, particularly because the manuscript relies heavily on imaging-based localization and colocalization analyses. Without these controls, it cannot be excluded that the observed signal represents non-specific staining.

      Importantly, the authors attempt to justify antibody specificity primarily by citing previous publications that used the same antibodies. However, this is not an adequate substitute for experimental validation within the current study. Previous reports do not guarantee specificity under the present experimental conditions, particularly in immunofluorescence, where staining patterns can be strongly influenced by fixation procedures, antibody concentrations, imaging settings, and cell type. Moreover, those studies may themselves lack sufficiently rigorous validation of antibody specificity. Therefore, antibody specificity should be demonstrated directly in the experimental system used in this manuscript, especially given that the principal conclusions rely extensively on the subcellular localization of MET, phospho-MET, and MT1-MMP.

      We understand the reviewer’s concern regarding antibody specificity and agree that appropriate validation is important for imaging-based analyses. However, we strongly disagree with the statement that the MT1-MMP, MET antibody were not adequately validated. The specificity of the MT1-MMP antibody was independently validated by both siRNA- and sgRNA-mediated gene silencing, where we observed a substantially diminished MT1-MMP signal by immunoblotting (Fig S5I, M’). Moreover, the antibody has been used for IF in the same cell line by multiple research groups [6–8]. So, in our opinion, this validation is completely redundant.

      We have validated the MET staining/ signal using the antibody in gene-silenced cells by immunoblotting and immunofluorescence (Fig S1I, K, L), as also acknowledged by the reviewer in comment-4. In addition, we would also like to clarify that the references cited in support of antibody specificity were not selected simply because they used the same antibodies. They include studies that provide experimental validation of the antibodies by gene silencing.

      To further confirm the antibody specificity, we will add immunofluorescence images of MET or MT1-MMP silenced cells stained with respective antibodies. However, we may not want to add these data to the manuscript as they do not carry any additional values to the manuscript.

      (3) Questionable MET localization in TIRF microscopy

      The presence of punctate MET signal in TIRF microscopy under unstimulated conditions raises additional concerns. Under basal conditions, MET is generally expected to exhibit a predominantly diffuse distribution at the plasma membrane, whereas prominent punctate structures are typically associated with ligand-induced clustering, endocytosis, or trafficking events.

      The observation of numerous MET-positive puncta in unstimulated cells, together with the insufficient validation of antibody specificity, raises the possibility that at least part of the observed signal represents non-specific staining or imaging artefacts rather than bona fide MET localization. This concern is further compounded by the lack of rigorous immunofluorescence antibody validation discussed above and significantly undermines the interpretation of all TIRF-based trafficking analyses presented in the manuscript.

      We would like to highlight the apparent similarities between Fig 2A, B and the unstimulated condition in Fig. 2H. In figure 2A, B, MET is detected using an anti-MET antibody, whereas in Figure 2H, GFP-MET is imaged under live cell condition. We believe the reviewer would agree that imaging GFP-MET in live cells avoids fixation- and antibody-related artifacts. The comparable localization observed using these two independent approaches therefore provides additional support that the MET distribution shown in Fig. 2A, B reflects genuine receptor localization rather than an imaging or staining artefact.

      (4) The evidence supporting a MET-specific role in invadopodia remains unconvincing

      The authors argue that the role of MET in invadopodia formation is validated using three independent approaches: shRNA-mediated knockdown, SMARTpool siRNA-mediated knockdown, and pharmacological inhibition with PHA665752. However, I do not agree that these constitute three independent orthogonal validations of MET function.

      First, the shRNA-mediated knockdown presented in this study achieves only modest depletion of MET protein. The authors themselves acknowledge this limitation and therefore selected cells with visibly reduced MET staining for imaging. Consequently, the shRNA experiments cannot be considered a robust or independent validation of MET function.

      Second, although pooled SMARTpool siRNAs are widely used to improve knockdown efficiency, they cannot exclude off-target effects, as each individual guide RNA contributes its own potential off-target profile. Therefore, pooled siRNAs cannot by themselves establish that an observed phenotype is specifically attributable to depletion of the intended target and do not replace validation using independent individual siRNAs or rescue experiments.

      Third, the pharmacological data should also be interpreted with caution. Throughout the manuscript, PHA665752 is presented as a MET inhibitor supporting the specificity of the observed phenotype. However, there is essentially no such thing as a truly selective receptor tyrosine kinase inhibitor. PHA665752 inhibits multiple kinases in addition to MET, particularly at concentrations commonly used in cell-based assays. Consequently, the inhibitor cannot be considered an independent validation of MET-specific function.

      Importantly, the newly added siRNA experiments do not resolve my original concern regarding the role of MET in invadopodia formation. Although siRNA-mediated MET depletion is substantially more efficient than the shRNA-mediated knockdown presented in the original manuscript, this marked difference in MET depletion is not accompanied by a correspondingly stronger inhibition of invadopodia formation or ECM degradation. If MET were indeed the principal driver of the observed phenotype, one would expect the magnitude of the biological effect to correlate with the efficiency of MET depletion. This inconsistency raises the possibility that the observed phenotype is not solely attributable to MET depletion and calls into question the specificity of the proposed mechanism.

      Taken together, the three perturbation approaches used by the authors cannot be regarded as independent orthogonal validation of MET function. One approach provides only modest target depletion, another relies on pooled RNAi reagents that cannot exclude off-target effects, and the third employs a multi-kinase inhibitor rather than a MET-specific compound. Collectively, these limitations substantially weaken the conclusion that the reduction in invadopodia formation is specifically attributable to loss of MET. A convincing demonstration of MET-specific function would require rescue experiments or another truly orthogonal validation strategy.

      We had adopted three independent approaches to validate the phenotype. All three approaches are well practiced in the field. The small molecule inhibitor has been used in the study by Rajadurai et al, J Cell Science, 2012 and it is very much accepted in studying cellular kinases [1].

      We agree that even though the smart pool has always chance of off-target effects, the OnTargetPlus Smart pool has the minimum chance of off-target effects because of the patented chemical modifications, compared to the individual oligos and SiGenome SMARTpool, which was used by Rajadurai, et. al. in their study [1].

      The shRNA mediated silencing resulted in less reduction in the MET level (~50-60%) but is it scientifically not acceptable, particularly when it showed similar phenotype over n=3 sets of experiments?

      The arguments made by the reviewer in this context seems to be harsh. However, we decided to carry out MET silencing using two independent oligos from the SMART pool.

      (5) Weak evidence for MET-MT1-MMP interaction

      The evidence supporting a physical interaction between MET and MT1-MMP remains unconvincing. The newly added co-immunoprecipitation experiment does not reveal a convincing MET-MT1-MMP interaction, and I am unable to appreciate a specific coimmunoprecipitated MT1-MMP signal in the presented blot. As presented, these data do not convincingly demonstrate a specific or functionally relevant interaction. Given that this interaction constitutes a central component of the proposed mechanistic model, this remains a major weakness of the study.

      We have detected the interaction in both GFP pulldown assay in 4 different cell lines and the corresponding reverse His-Ni-NTA pulldown assay, providing complementary evidence for their physical association (Fig 6E, S5F, G). We have also clearly stated in the manuscript that this interaction is weak in nature and have not claimed it to be a strong interaction. While we acknowledge that the signal is modest, disregarding reproducible positive results would not be an appropriate interpretation of the pulldown assays. We believe the reproducibility of these findings supports a genuine MET and MT1-MMP association, and we have reported it accordingly.

      (6) Overinterpretation of the data

      Taken together, the study proposes a mechanistic model linking MET trafficking to MT1MMP localization and invadopodia function. However, the experimental evidence largely supports correlative observations rather than demonstrating a direct mechanistic relationship.

      Specifically, MET endocytosis and recycling are not properly demonstrated because of the inappropriate temporal resolution of the trafficking experiments; the localization data remain uncertain owing to insufficient validation of the immunofluorescence reagents; and the proposed interaction between MET and MT1-MMP is not convincingly demonstrated. Consequently, the manuscript establishes correlation rather than causality, and the central mechanistic conclusions appear to be substantially overstated relative to the presented data.

      We agree that there is scope for further investigation of MET and MT1-MMP cotrafficking, however, we have provided preliminary evidence demonstrating the cotrafficking of MET and MT1-MMP at the cell surface (Fig 6F). Importantly, MET and MT1-MMP co-trafficking represents only one component of the manuscript and not a central mechanistic conclusion. As the title of the study indicates, the major component of the study is focused on MET trafficking and its implication in invadopodia-associated TNBC invasion, for which we have provided direct experimental evidence. Therefore, we believe that describing the overall study as primarily overstated and correlative underestimates the extent of the experimental evidence supporting our mechanistic conclusions.

      Conclusion:

      While the manuscript addresses an interesting and biologically relevant question, the current experimental evidence does not adequately support the proposed mechanistic model. The combination of inappropriate experimental design for trafficking studies, insufficient validation of key imaging reagents, questionable interpretation of the localization data, lack of convincing evidence for the proposed MET-MT1-MMP interaction, and the absence of a clear relationship between the degree of MET depletion and the biological phenotype substantially limits the reliability of the conclusions.

      In my opinion, these issues cannot be addressed by further revision of the current manuscript, as they require substantial additional experimentation, including appropriately designed trafficking assays with short kinetic time points, rigorous validation of antibody specificity for immunofluorescence, and stronger mechanistic evidence linking MET trafficking to MT1-MMP-dependent invadopodia function.

      We thank the reviewer for outlining the remaining concerns. We will address the points raised by performing additional antibody validation and independent oligo-mediated MET silencing experiment, providing further support for the specificity and robustness of our findings.

      However, we respectfully disagree, that the conclusions require the extensive additional experimentation suggested by the reviewer. As clarified above, our trafficking experiments were designed around the time points at which the functional invasive phenotype is observed, rather than to define the kinetics of MET internalization.

      In conclusion, we would expect that the views of the reviewer 2 towards the manuscript should change and the reliability of our manuscript to the public should improve.

      References:

      (1) Rajadurai CV, Havrylov S, Zaoui K, Vaillancourt R, Stuible M, Naujokas M, Zuo D, Tremblay ML, Park M. Met receptor tyrosine kinase signals through a cortactinGab1 scaffold complex, to mediate invadopodia. J Cell Sci. 2012 Jun 15;125(Pt 12):2940-53. doi: 10.1242/jcs.100834. Epub 2012 Feb 24. PMID: 22366451; PMCID: PMC3434810.

      (2) Sharma P, Parveen S, Vinod Shah L, Mukherjee M, Kalaidzidis Y, Joseph Kozielski A, et al. Title: SNX27-retromer assembly directs MT1-MMP trafficking to invadopodia and promotes breast cancer metastasis.

      (3) Mader CC, Oser M, Magalhaes MAO, Bravo-Cordero JJ, Condeelis J, Koleske AJ, et al. Molecular and Cellular Pathobiology An EGFR-Src-Arg-Cortactin Pathway Mediates Functional Maturation of Invadopodia and Breast Cancer Cell Invasion [Internet]. doi:10.1158/0008-5472.CAN-10-1432

      (4) Parveen S, Khamari A, Raju J, Coppolino MG, Datta S. Syntaxin 7 contributes to breast cancer cell invasion by promoting invadopodia formation. J Cell Sci. 2022 Jun 15;135(12). doi:10.1242/jcs.259576 PubMed PMID: 35762511.

      (5) Joffre C, Barrow R, Ménard L, Calleja V, Hart IR, Kermorgant S. A direct role for Met endocytosis in tumorigenesis. Nat Cell Biol. 2011 Jun 5;13(7):827–37. doi:10.1038/ncb2257 PubMed PMID: 21642981.

      (6) Monteiro P, Rossé C, Castro-Castro A, Irondelle M, Lagoutte E, Paul-Gilloteaux P, et al. Endosomal WASH and exocyst complexes control exocytosis of MT1-MMP at invadopodia. J Cell Biol. 2013 Dec 23;203(6):1063–79. doi:10.1083/jcb.201306162 PubMed PMID: 24344185.

      (7) Wenzel EM, Pedersen NM, Elfmark LA, Wang L, Kjos I, Stang E, et al. Intercellular transfer of cancer cell invasiveness via endosome-mediated protease shedding. Nat Commun. 2024 Feb 10;15(1):1277. doi:10.1038/s41467-024-45558-8

      (8) Pedersen NM, Wenzel EM, Wang L, Antoine S, Chavrier P, Stenmark H, Raiborg C. Protrudin-mediated ER-endosome contact sites promote MT1-MMP exocytosis and cell invasion. J Cell Biol. 2020 Aug 3;219(8):e202003063. doi: 10.1083/jcb.202003063. PMID: 32479595; PMCID: PMC7401796.

      (9) Duan Q, Jia HR, Chen W, Qin C, Zhang K, Jia F, Fu T, Wei Y, Fan M, Wu Q, Tan W. Multivalent Aptamer-Based Lysosome-Targeting Chimeras (LYTACs) Platform for Mono- or Dual-Targeted Proteins Degradation on Cell Surface. Adv Sci (Weinh). 2024 May;11(17):e2308924. doi: 10.1002/advs.202308924. Epub 2024 Feb 29. PMID: 38425146; PMCID: PMC11077639.

      (10) Wei J, Wang J, Guan W, Li J, Pu T, Corey E, Lin TP, Gao AC, Wu BJ. PlexinD1 is a driver and a therapeutic target in advanced prostate cancer. EMBO Mol Med. 2025 Feb;17(2):336-364. doi: 10.1038/s44321-024-00186-z. Epub 2025 Jan 2. PMID: 39748059; PMCID: PMC11822115.

      (11) Yamasaki A, Miyake R, Hara Y, Okuno H, Imaida T, Okita K, Okazaki S, Akiyama Y, Hirotani K, Endo Y, Masuko K, Masuko T, Tomioka Y. Dual-targeting therapy against HER3/MET in human colorectal cancers. Cancer Med. 2023 Apr;12(8):9684-9696. doi: 10.1002/cam4.5673. Epub 2023 Feb 7. PMID: 36751113; PMCID: PMC10166911.

      (12) Soonnarong R, Putra ID, Sriratanasak N, Sritularak B, Chanvorachote P. Artonin F Induces the Ubiquitin-Proteasomal Degradation of c-Met and Decreases AktmTOR Signaling. Pharmaceuticals (Basel). 2022 May 21;15(5):633. doi: 10.3390/ph15050633. PMID: 35631459; PMCID: PMC9145792.


      The following is the authors’ response to the original reviews.

      We sincerely thank the editor and reviewers for thoroughly evaluating the manuscript. Following the comments from the reviewers we caried out four major sets of experiments and added the results and the conclusions derived from them in the revised manuscript. We also modified the abstract and the introduction. As suggested by the reviewers, we have rewritten the discussion. All the mislabelling and typing errors have been corrected, and representative graphs has been replaced as suggested. The list of the newly carried out experiments are -

      (i) In the original submission, we carried out a microscopy-based study to investigate the recycling of MET. We now added surface biotinylation approach to show the RCP or KIF16B-mediated surface delivery of MET (Fig. 5J).

      (ii) To demonstrate the functional effect of KIF16B silencing on TNBC invasion we have performed ECM degradation assay with depleted KIF16B cells (Fig. S4K-L). Further, the MET degradation in KIF16B-silenced cells has also been investigated using immunoblotting (Fig. S4H).

      (iii) To rule out the off-target effect of the siRNA used in this study, we have validated the invadopodia-associated function using 2 individual siRNAs for RAB14 and RCP (Fig. S4K-L).

      (iv) We have now introduced MMP2 as a positive control as a substrate of MT1-MMP to show the effect of silencing of the protease on its cleavage (Fig. S5I).

      We also incorporated following changes, largely additions of new plots, data in the revised manuscript.

      (i) RAB4 and RAB14 colocalization with MET in BT-549 cell line has also been added in Fig. 4 (A, B, C).

      (ii) Data showing MET silencing using siRNA and its effect on invadopodia has been added to Fig 1D.

      (iii) Graph showing percentage of cells forming invadopodia has added to Fig. S1G.

      (iv) Line intensity plots of MET-containing invadopodia has been added in Fig. 2G’.

      (v) We have added quantification of all the blots to the figures.

      (vi) A graph representing MET degradation kinetics with HGF over 3 experiments has been added to Fig S2F.

      (vii) The full field of view of Fig 1F has been added in the Fig S1L. A quantification of the gelatin degradation has been added to Fig 1F.

      (viii) The blot for loading control of Fig S1K has been changed from Actin to Vinculin.

      (ix) The blot showing expression of MT1-MMP in the SCR and knockout cells has been added to Fig S5M’.

      (x) The survival plot for patients with altered or unaltered MET and RCP has been removed. The graph showing frequency alteration of MET has also been removed.

      (xi) Additional immunoblots associated with all the figures are now provided in a newly added supplementary figure (Fig S6).

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix, promoting cancer cell invasion, a process here investigated using gelatin-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endosomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metalloprotease MT1-MMP. The delivery of MET from the recycling compartment to invadopodia is mediated by RCP, which facilitates the colocalization of MET to RAB14 endosomes. In this compartment, HGF induces the recruitment of the motor protein KIF16B, promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface. This pathway is critical for the HGF-driven invasive properties of TNBC cells, as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well-organized and executed using state-of-the-art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied, taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosis-defective.

      Data analyses are rigorous, and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well-supported by the results, and the data and methodology are of interest for a wide audience of cell biologists.

      We sincerely thank the reviewer for the positive feedback and for considering our study to be well executed and rigorous. The valuable suggestions and comments certainly improved the understanding of the role of the RAB14-RCP-KIF16B axis in MET trafficking and breast cancer invasion.

      Weakness

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings, including triple-negative breast cancer cells. The novelty of the present study mostly consists of the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGF-mediated invasion detectable in other cancer cells is not addressed or at least discussed

      We thank the reviewer for raising this point. We would like to clarify that in TNBCs, the overexpression of EGFR and MET in the null background of the hormonal receptors; progesterone receptor, estrogen receptor, and HER2 is considered to be very crucial in terms of prognosis and treatment (PMID: 27655711, 25368674). Hence study of MET signalling and trafficking is more relevant for TNBCs compared to other cancer cells. In the current study, we therefore focused on two TNBC cell lines. We have added this in the first paragraph of introduction. Line no: 31-38.

      Reviewer #1 (Recommendations for the authors):

      Major points

      (1) My major concern refers to the clinical data presented in this study. Different from the mechanistic findings, the quality of these analyses is too low, the description in the figure legends is scant, and is absent in the Method section. The results concerning the prognostic value of RCP are the most problematic. What is shown in the Kaplan Meier? Is it the correlation between the RCP mRNA levels and the patient's survival? More importantly, to study the correlation between genetic alteration and prognostic outcome, multivariable analyses should be performed comparing the genetic alteration with known prognostic factors (sex, age, tumor size, node status, ER/PrR, HER2, Ki67 if available, tumor grade). This is because breast cancer prognosis depends on multiple interrelated factors, and only multivariate models can adjust for confounding and identify which variables independently predict outcome-providing far more accurate and clinically useful prognostic information than univariate analysis. Furthermore, overexpression of RTKs has been extensively reported and studied in TNBCs. Similarly, the relevance of MET in cancer cell invasion has been firmly established. Therefore, the data presented here, whose quality does not match the mechanistic part of the study, can be considered unnecessary. I would, therefore, recommend removing the "clinical" data. The introduction should be modified accordingly.

      We thank the reviewer for this insightful comment. To generate the graphs, we selected studies available in the publicly accessible database cBioportal (cbioportal.org/). The graphs generated by the database, representing the genetic alterations of genes has added in the Fig. S4A. However, as suggested by the reviewer, we have removed the survival plots for MET and RCP, the gene alteration frequency of MET and modified the text accordingly.

      (2) Overexpression of KIF16B has been shown to inhibit the degradative pathway, stimulating the recycling one. In agreement, silencing of KIF16B accelerates EGFR degradation (PMID: 15882625). Does this also apply to the MET receptor? In the present setting, does the overexpression of KIF16B result in prolonged MET expression?

      We thank the reviewer for raising this question. Although we did not analyze the expression level of MET in the KIF16B overexpressed cells, we have analyzed the total MET levels in the control and the KIF16B silenced cells using immunoblotting and did not observe any significant changes in the MET protein levels (Fig S4H). Line no: 390-391. This led us to believe that the depletion or overexpression of KIF16B may not have any direct effect on MET expression.

      (3) The contribution of MMP2 and MMP9 to the degradative properties of HGFstimulated TNBC cells should be investigated and compared to MT1-MMP.

      We believe that this is a relevant note from the reviewer. However, MT1-MMP is one of the most well-established metalloproteases, known till date for its role in invadopodia-associated activities in breast cancer (PMID: 35008569, 27501444). Moreover, it is the best-known candidate protease, which is a transmembrane metalloprotease could be the model in studying membrane recycling to invadopodia (PMID: 20605060, 19692588). However, as pointed out by the reviewer, we completely agree that MMP2 and MMP9 also contribute significantly to invadopodia-associated functions in TNBCs (PMID: 23902685, 25699257). HGF is also known to promote the expression and activity of MMP2 and MMP9 (PMID: 23320110, 26259977). Interestingly, the cellular machineries involved in their enhanced activity due to HGF stimulation may be distinct from what was observed in the current study and may require a distinct, elaborated study, which is beyond the scope of the current one.

      (4) The authors appropriately tested the possible shedding effect of MT1-MMP on MET. They should repeat the experiments, adding a positive control of shedding. Furthermore, the legends referring to these experiments, shown in Supplementary Figure S6, seem to be wrong (or mislabeled).

      We thank the reviewer for the suggestion. MT1-MMP is known to proteolytically cleave and initiate the activation of MMP2 (PMID: 11161720, 15095267). We now carried out the experiment with MMP2 as a control (Fig S5I). We observe an increase in the unprocessed MMP2 level in the MT1-MMP silenced cells, whereas the MET levels are unaltered. Line no: 467-468.

      We sincerely apologize for the oversight in the mislabelling. We have now corrected it in the revised manuscript.

      (5) Does altered expression of RAB14 and/or KIF16B affect MT1-MMP delivery to invadopodia in the TNBC cell lines? Does KIF16B silencing affect invasion?

      The role of RAB14 and KIF16B in MT1-MMP delivery to podosomes, a structure similar to invadopodia in macrophages has been studied by Hey S. et al. (PMID: 37696580). The study suggests that KIF16B silencing reduces the invasion of macrophages. However, the effect is not known for TNBCs. Thus, we have conducted the ECM degradation assay in TNBC cell lines to show the effect of KIF16B gene silencing on breast cancer invasion to the revised manuscript (Fig S4K-L). In both MDA-MB-231 and BT-549 cells we observed reduced ECM degradation activity upon KIF16B depletion, corroborating the observation from Hey S. et al. Line no: 392-400.

      Minor points

      (1) I recommend authenticating cell lines and stable populations by STR profiling.

      All the cell lines used in the study have been purchased from ATCC, and STR is a standard practice followed by ATCC. Further, to avoid any alteration, cells were discontinued after 20 passages. We have added this statement to the methods section in the revised manuscript. Line no: 643-644.

      (2) In the legend to Supplementary Figure S4A, the panel is described as the frequency of alterations of MET, while, if I correctly interpret it, the bar graph refers to RCP. As mentioned above, these data could be removed.

      We thank the reviewer for pointing out the mistake. We have rectified this in the revised version.

      (3) Check for typos. Sometimes invadopodia is written with the capital: "Invadopodia", in other instances it is not. The authors should be consistent throughout the manuscript. English language editing would help.

      We sincerely apologize for the inconsistency in the writing. We have removed the unnecessary capitalization of invadopodia in the revised manuscript. Line no: 142, 159, 280, 436.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14-KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1-MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligand-induced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights. However, some sections require clearer and more concise writing (details below). In addition, the quality, reliability, and robustness of several data sets need to be improved.

      Strengths:

      A key strength of the study is the novel demonstration that HGF-mediated, RAB4- and RAB14-dependent recycling of MET delivers this receptor, together with MT1MMP, to invadopodia -highlighting a previously unrecognized mechanism, regulating the formation and proteolytic function of these invasive structures. Another strong point is the breadth of experimental approaches used and the substantial amount of supporting data. The authors also include an appropriate number of biological replicates and analyze a sufficiently large number of cells in their imaging experiments, as clearly described in the figure legends.

      We greatly appreciate the positive assessment from the reviewer, who also acknowledged the novelty and relevance of our study. Below, we have carefully addressed the comments/concerns raised regarding this study and that have strengthened the reliability and robustness by revisiting the data, providing additional analyses where required, and clarifying methodological details.

      Weakness

      (1) Inappropriate stimulation times for endocytosis and recycling assays. The experiments examining MET endocytosis and recycling following HGF stimulation appear to use inappropriate incubation times. After ligand binding, RTKs typically undergo endocytosis within minutes and reach maximal endosomal accumulation within 5-15 minutes. Although continuous stimulation allows repeated rounds of internalization, the temporal dynamics of MET trafficking should be examined across shorter time points, ideally up to 1 hour (e.g., 15, 30, and 60 minutes). The authors used 2-, 3-, or 6-hour HGF stimulation, which, in my opinion, is far too long to study ligandinduced RTK trafficking.

      We understand the reviewer’s concern regarding the HGF stimulation time point for endocytosis and recycling. We want to highlight that to study the recycling/surface delivery of MET in response to HGF, we performed TIRF microscopy-based imaging, where images were taken within 1h of HGF addition (Fig. 2I). Additionally, we have incorporated surface biotinylation to show the recycling of MET as suggested in comment-7 (Fig. 5J). For this experiment we have used 30 min of HGF stimulation. Line no: 382-391.

      Moreover, we have observed the functional effect of HGF on ECM (gelatin) degradation and invadopodia formation after 3 h of HGF stimulation. We were curious to know where does the MET localises with prolonged ligand stimulation. Hence, to study the localization of MET to invadopodia or the endocytic markers, the cells were stimulated with HGF for 2-3 hours.

      (2) Low efficiency of MET silencing in Figure S1I. The very low MET knockdown efficiency shown in Figure S1I raises concerns. Given the potential off-target effects of a single shRNA and the insufficient silencing level, it is difficult to conclude whether the reduction in invadopodia number in Figure 1F is genuinely MET-dependent. The authors later used siRNA-mediated silencing (Figure S5C), which was more effective. Why was this siRNA not used to generate the data in Figure 1F? Why did the authors rely on the inefficient shRNA C#3?

      We understand the concern raised by the reviewer. We want to emphasize that we have employed three different approaches to investigate the effect of MET silencing/inhibition on invadopodia formation. (i) A MET kinase inhibitor, PHA665752, which shows reduced invadopodia formation (Fig. 1E, E’). (PMID: 21973114, 41009793) (ii) Silencing with shRNA: Since the level of silencing of MET with the shRNA was not sufficient, cells were stained with MET as a readout for MET silencing, and images of the cells with reduced MET expression were captured. ECM degradation activity and invadopodia numbers were found to be reduced in the MET-depleted cells (Fig. 1F). (iii) We have now added the data showing the effect of siRNA-mediated MET depletion on invadopodia formation to the revised figure 1D. Line no: 123-125. To draw a robust conclusion regarding the role of MET on invadopodia-associated TNBC invasion, we have integrated all three complementary approaches.

      (3) Missing information on incubation times and inconsistencies in MET protein levels. The figure legends do not indicate how long the cells were incubated with HGF or the MET inhibitor PHA665752 before immunoblotting. This information is crucial, particularly because both HGF and PHA665752 cause a substantial decrease in the total MET protein level. Notably, such a decrease is absent in MDA-MB-231 cells treated with HGF in the presence of cycloheximide (Figure S2F). The authors should comment on these inconsistencies. Additionally, the MET bands in Figure S1J appear different from those in Figure S1C, and MET phosphorylation seems already high under basal conditions, with no further increase upon stimulation (Figure S1J). The authors should address these issues.

      We apologise for the unintentional omission of experimental detailing about HGF or drug incubation time, which we have incorporated into the figure legend appropriately. Regarding the decreased MET level in the drug-treated condition: literature suggests that the MET inhibitor PHA665752 also promotes MET degradation, corroborating our result shown in Fig. S1J (PMID: 15788682, 18327775). Further in Fig. S1J, the relative phosphorylation of MET when compared to the total MET level in the HGF-treated condition is higher (~2-fold). Quantification of the blot has been added now.

      Further, addition of HGF for 3 h leads to 40±15% reduction in the MET protein levels as seen in the Author response image 1 representing quantification of different immunoblot. The degradation of MET in the Fig S1J is 65% which nearly fall in the range for HGF-mediated MET degradation.

      Author response image 1.

      Quantification of immunoblots showing MET signal intensity in the presence or absence of HGF normalized with the loading control. N=6.

      Next, in the fig. S1A, K the rabbit anti-MET (CST, D1C2 XP) antibody has been used, which binds to a C-terminal motif of MET and identifies both the 170kDa as well as 140kDa protein representing the uncleaved and cleaved form of MET. In Fig. S1J, the mouse antiMET (CST, L6E7) antibody has been used, which binds to an N-terminal motif of MET and recognizes only the 140kDa protein.

      (4) Insufficient representation and randomization of microscopic data. For microscopy, only single representative cells are shown, rather than full fields containing multiple cells. This is particularly problematic for invadopodia analysis, as only a subset of cells forms these structures. The authors should explain how they ensured that image acquisition and quantification were randomized and unbiased. The graphs should also include the percentage of cells forming invadopodia, a standard metric in the field. Furthermore, some images include altered cells - for example, multinucleated cells - which do not accurately represent the general cell population.

      We thank the reviewer for raising this point. The single-cell images are shown for clarity and to visualize the subcellular features; however, the conclusions are made based on the quantitative analysis of multiple cells collected from multiple fields of view (Frames). At least 30 such frames per condition having 4-7 cells/ frame has been acquired and analysed for the quantification throughout this manuscript. We would like to highlight that the image acquisition has been done over random fields on a coverslip. In the revised manuscript, for a better representation of the population of cell-forming invadopodia, a graph showing the percentage of cells forming invadopodia have been added (Fig S1G). Line no: 118-119. The percentage of cells forming invadopodia increased upon HGF stimulation in MDAMB-231.

      (5) Use of a single siRNA/shRNA per target. As noted earlier, using only one siRNA or shRNA carries the risk of off-target effects. For every experiment involving gene silencing (MET, RAB4, RAB14, RCP, MT1-MMP), at least two independent siRNAs/shRNAs should be used to validate the phenotype.

      We would like to clarify that we are using SMARTPool siRNA, which contains 4 individual siRNAs for the target gene. Literature suggests that using a pool of siRNA has reduced off-target effects compared to using single oligos for gene silencing (PMID: 14681580, 33584737, 24875475).

      While SMARTpool siRNA minimizes the off-target effect, it does not eliminate the possibility of it. To confirm that the observed phenotypes are specifically attributable to the genes investigated in this study, we now performed functional experiments using two independent siRNAs targeting RCP and RAB14. The results have been added to figure S4K-L. Silencing of RCP or RAB14 using single oligos resulted in decrease in the degradation index comparable to SMARTpool siRNA, thus phenocopied the SMARTpool siRNA. Line no: 392-400.

      Further, RAB4 is well established to be associated with MET trafficking and it served as a positive control in our study (PMID: 21664574, 30537020). Additionally, a recent study by Hey et al. have used individual oligos for KIF16B to demonstrate the effect of KIF16B silencing on gelatin degradation, which corroborate with our observation from the KIF16B silencing using the SMARTpool siRNA (PMID: 37696580).

      For MET, we used siRNA, shRNA and an inhibitor to show the effect of MET inhibition/perturbation in the invadopodia-associated activity, which validates the observations of siRNA-mediated gene silencing (detailed in point 2).

      We did not perform any experiments using single oligos targeting MT1-MMP, since in our MT1-MMP siRNA-based study, now we have taken an appropriate positive control to validate the efficacy of MT1-MMP silencing (Fig. S1I). In addition, we have shown the effect of MT1-MMP depletion on invadopodia formation using a CRISPR-based gene knock-out study, and another study from our group has shown a similar effect using siRNA (PMID: 31820782), which supports our MT1-MMP KO cell observation.

      (6) Insufficient controls for antibody specificity. The specificity of MET, p-MET, and MT1-MMP staining should be demonstrated in cells with effective gene silencing. This is an essential control for immunofluorescence assays.

      The anti-MET antibody (CST, D1C2 XP) has been used in several studied (PMID: 41166312, 41152910, 39748059). The CST L6E7 anti-MET antibody has been used in studied by Radke et al, Wang et el., Kong et al. (PMID: 36435874, 38262412, 32214092). In our study, immunoblots demonstrating depletion of MET in the siRNA or shRNA-treated cells has been provided in Fig. S1I, K respectively. Further, we have demonstrated MET silencing using immunofluorescence. We also have now added the entire field of view in Fig S1L of showing cells treated with control or shRNA against MET. In the shRNA-treated condition, the cell at the centre shows low MET fluorescence intensity indicating depletion of the RTK, while the surrounding cells have MET staining similar to control.

      Tyr 1234/35 are present in the active site of MET kinase domain and upon binding of the ligand promotes their autophosphorylation (PMID: 17667909). Earlier studies have established the specificity of the phosphor-MET antibody by immunoblotting and immunofluorescence using MET inhibitors (PMID: 21973114, 41009793). In our study we have shown that the inhibition of MET kinase activity using PHA665752 abolished the MET phosphorylation at the Tyr 1234/35, as shown in Fig S1J which revalidates the specificity of the antibody.

      Additionally, in a previous study Joffre et al. have shown that an oncogenic mutant form of MET, M1250T is highly phosphorylated at the Tyr 1234/1235 (PMID: 21642981). Using the phospho-MET antibody, we have shown in Fig 3C, S2I the increased Tyr phosphorylation of M1250T MET mutant as reported by Joffre et al.

      The anti-MT1-MMP antibody is also a very well-established antibody reported in multiple studies (PMID:32479595, 31820782, 35762511). In our study, we have shown the specificity of the antibody using immunoblot analysis. Immunoblots showing significant depletion of MT1-MMP protein level following the SMARTpool siRNA and sgRNA-mediated gene silencing has been provided in Fig. S5I, M’, respectively. Further MT1MMP silencing has been also validated by immunofluorescence in the following studies. PMID: 22291036, 21571860, 20505159.

      (7) Inadequate demonstration of MET recycling. MET recycling should be directly demonstrated using the same approaches applied to study MT1-MMP recycling. The current analysis - based solely on vesicles near the plasma membrane - is insufficient to conclude that MET is recycled back to the cell surface.

      We appreciate the reviewer’s suggestion for an alternative approach to show MET trafficking. We have demonstrated MET trafficking using surface biotinylation, where we have shown that the RCP and KIF16B depletion affect the surface delivery of MET (Fig 5J). Line no: 382-391.

      In addition, to study the surface delivery of cargo, TIRF is a widely used reliable approach and it is highly sensitive technique for detection of surface delivery events (PMID: 24344185, 20971701). We have also tried to investigate the trafficking of MET using antibody uptake approach; however, it could not be established as the binding of the antibody hindered the ligand binding and vice versa.

      (8) Insufficient evidence for MET-MT1-MMP interaction. The interaction between MET and MT1-MMP should be validated by immunoprecipitation of endogenous proteins, particularly since both are endogenously expressed in the studied cell lines.

      We thank the reviewer for pointing out the insufficient evidence for MET-MT1-MMP interaction at the endogenous level. We now carried out the immunoprecipitation of endogenous MET to validate the interaction with MT1-MMP (Fig S5H). A light (low intensity) band corresponding to MT1-MMP was detected in the anti-MT1-MMP immunoblot for the immunoprecipitated sample. We believe that the interaction between MT1-MMP and MET may be weak in nature, resulting in limited co-immunoprecipitation of the endogenous MT1-MMP by MET. The immunoblot is now added to the revised manuscript. Line no: 460-461.

      (9) Inconsistent use of cell lines and lack of justification. The authors use two TNBC cell lines: MDA-MB-231 and BT-549, without providing a rationale for this choice. Some assays are performed in MDA-MB-231 and shown in the main figures, whereas others use BT-549, creating unnecessary inconsistency. A clearer, more coherent strategy is needed (e.g., present all main findings in MDA-MB-231 and confirm key results in BT549 in supplementary figures).

      MDA-MB-231 and BT-549 are two well-characterized TNBC cell lines that readily form invadopodia. These cell lines have been extensively used to study invadopodia-associated breast cancer cell invasion (PMID: 32697977, 35915226, 31533971). These two cell lines also show overexpression of MET, making them suitable model cell lines for our study (PMID: 36139568, 20687930, 27502396).

      Overall, most of the conclusions reported in this manuscript are derived from multiple experimental approaches using two TNBC cell lines for generalization.

      We agree with the reviewer that showing the results from one type of cell line in the main figure would have been better, and wherever possible, we now provided the observations from a single cell line in the main figures and the data from the other cell line in the supplementary figures. However, some of the overexpression studies were performed in BT-549 cells to derive robust statistically meaningful conclusions. Therefore, we could not avoid adding results from both the cell lines in some of the figures. We would like to add that the legends for these figures have been edited to clearly mention the cell lines associated with each of the figure panels to avoid any confusion or inconsistency.

      (10) Inconsistency in invadopodia numbers under identical conditions. The number of invadopodia formed in Figure 1E is markedly lower than in Figure 1C, despite identical conditions. The authors should explain this discrepancy.

      We sincerely thank the reviewer for pointing out the inconsistency in invadopodia numbers across 2 experiments. Fig. 1C has 2 conditions: UT and the HGF-treated condition. The Untreated condition has the serum-free media without any stimulation. Whereas we have added vehicle (DMSO) in Fig. 1E, E’, since the drug is resuspended in DMSO. This difference in the treatment is likely to be responsible for the decreased numbers of invadopodia in Fig. 1E. In different studies it has been shown that DMSO is not biologically inert and can affect invasive properties of cells by perturbing actin dynamics and metalloprotease activity (PMID: 33552397, 22529897, 7188610).

      (11) Questionable colocalization in some images. In some figures - for example, Figure 2G - the dots indicated by arrows do not convincingly show colocalization. The authors should clarify or reanalyze these data.

      As suggested by the reviewer, we have now re-analyzed the data for figure 2G. The apparent visual lack of colocalization is likely due to the relatively lower fluorescence intensity of MET at these structures. We have now added the line intensity plots for the indicated puncta to show the intensity of both channels at the ‘dots’ in the figure 2G’ and they show correlation in their intensity distribution.

      We would also like to elaborate that to quantify the colocalization of two channels, we have used the automated image analysis software Motiontracking (motiontracking.mpi-cbg.de) (PMID: 16143105), which has been detailed in the method section. Briefly, the algorithm works on object-based co-localization. If the fluorescence intensity distributions at a given object corresponding to any two different channels (fluorophores) show 35% or more overlap (Author response image 2), the object is considered as a multi-colour object and the overlapped area value is used to calculate the degree of co-localization. The calculation is carried out over all the objects in a given field of view (frame) and over all the field of views (frames) acquired for a given condition. Also, the apparent colocalization is corrected for random colocalization, which is the random permutation of object colocalization. This makes object-based colocalization more reliable than intensity-based colocalization.

      Author response image 2.

      Image showing the object identification and contour of the multicolour object identified by Motiontracking. The plot shows the intensity distribution of these two objects as analyzed by Motiontracking.

      (12) Abstract, Introduction, and Discussion require substantial rewriting.

      (a) The abstract should be accessible to a broader audience and should avoid using abbreviations and protein names without context.

      (b) The introduction should better describe the cellular processes and proteins investigated in this study.

      (c) The discussion currently reads more like an extended summary of results. It lacks deeper interpretation, comparison with existing literature, and consideration of the broader implications of the findings.

      We thank the reviewer for this suggestion. We have substantially modified the abstract, and the introduction following the reviewer’s suggestion. The introduction has been edited to describe the cellular processes investigated in this study and some of the key associated molecular machineries. In the discussion section, we have avoided redundant descriptions of the results but retained some of them wherever necessary for interpretation and relevant discussion in the light of existing literature.

      Reviewer #2 (Recommendations for the authors):

      (1) Quality of charts. Several charts (e.g., Figure 1B, 1C, 1F) are of poor visual quality. The authors should provide higher-resolution graphs with clearer axis labels, consistent formatting, and properly scaled data.

      We thank the reviewer for pointing out the insufficient visual quality of some of the charts. We believe the resolution of the charts/graphs have changed while converting to PDF, due to image compression. We will provide the uncompressed charts with much improved visual quality, provided they are not restricted by file size limitation.

      (2) Full protein names on first mention. Whenever a protein appears for the first time in the manuscript, its full name should be provided, if possible, before using the abbreviation.

      We have incorporated the full name of the protein while reporting for the first time in the manuscript.

      (3) Correct use of "invadopodium" vs. "invadopodia." Invadopodia is the plural form; the singular is invadopodium. The sentence "Invadopodia, an actin-rich membrane protrusion decorated with proteases, is a tool for ECM and basement membrane degradation during cancer cell invasion" should be corrected accordingly.

      We are thankful to the reviewer for pointing out the grammatical error. We corrected the error in the revised version. Line no: 44.

      (4) Unclear sentence about resistance and invasion.

      The sentence "However, often patients develop resistance to EGFR-targeted therapies due to overexpression of MET; yet, the mechanistic understanding of MET-dependent cancer invasion is unclear" is confusing because the shift from drug resistance to invasion is abrupt. The authors should revise this sentence for clarity and logical flow.

      We are thankful to the reviewer for highlighting this sentence. We have rewritten the sentence as follows “Since one of the receptor tyrosine kinases (RTK), EGFR is often amplified in TNBC patients, they are usually targeted for its treatment. However, often patients develop resistance to EGFR-targeted therapies due to overexpression of another RTK MET”. Line no: 35-38.

      (5) Incorrect figure reference. In the paragraph describing the results related to RAB proteins, there is an incorrect reference to Figure 3 instead of Figure 4. This should be corrected.

      We sincerely apologize to the reviewer for the incorrect figure reference. We have corrected the reference to figures in the revised manuscript. Line no: 251, 261.

      (6) Ambiguous sentence regarding MET activation. The sentence "MET, upon activation by HGF, triggers the activation of the RTK that induces cancer cell invasion" is unclear and should be rewritten for precision and clarity.

      We are thankful to the reviewer for highlighting the unintentional mistake. We have now added a clearer sentence “MET-HGF signalling axis are reported to promotes invasion in gastric cancer cells and melanoma cells”. Line no: 96-97.

      (7) Questionable wording of figure legend. The phrase "Immunoblotting of indicated cell lines with MET and Tubulin" is an informal shortcut. The authors should rephrase it.

      We are thankful to the reviewer for highlighting this sentence. We have modified the figure legend with appropriate text. The modified text is as follows: Lysates of MDA-MB-231, BT-549 and MCF10A DCIS were separated by SDS-PAGE and analyzed by Western blot. Membranes were probed with anti-MET and anti-Vinculin antibody.

      (8) Unnecessary capitalization. Terms such as invadopodia and cortactin should not be capitalized. The authors should correct capitalization throughout the manuscript.

      We are thankful to the reviewer for raising this point. We have modified this accordingly in the revised version. Line no: 142, 159, 280, 436.

    1. Author response:

      We thank the editors for sending our work for review and the reviewers for their thorough and constructive evaluations. In response to their feedback, we will submit a revised version of the manuscript soon. Below, we provide clarification on several of the concerns raised and outline the changes planned for the revised manuscript.

      Reviewer 1, weaknesses

      (1) The two-neuron circuit model is a good choice for the analytics, but it may have hidden a covariance effect of the "rate-dominated" symmetric spike-based kernel that would appear when several inhibitory neurons, each sharing a different spike correlation with the postsynaptic neuron, converge onto it. The rate homeostasis achieved by the rate-dominated model arises from adjusting inhibitory weights according to their initial correlation with the output neuron, so that after learning, the weights are distributed such that these correlations are cancelled out (Vogels et al., 2011). In other words, even the rate-dominated rule is covariance-driven under the hood: with a single inhibitory input, the two-neuron circuit cannot expose this, but with several differently correlated inputs, the covariance dependence should reappear.

      Reviewer 1  highlights that the rate-homeostatic rule by Vogels et al. (2011) also includes a covariance-dependent component. Thus, in a network with multiple excitatory units, differences in pre-postsynaptic correlations can drive a redistribution of weights, resulting in stronger inhibitory weights for higher correlations. Our two-neuron circuit, which contains only one plastic inhibitory input, cannot reveal this competitive effect. We note, however, that although in the Vogels rule the covariance-dependent term is non-zero, it is typically much smaller than the rate-dependent term (see Methods, Section 4.3). Consequently, covariance-dependent organization may emerge on a slower timescale (a similar effect was also shown by Lagzi and Fairhall 2024; https://doi.org/10.1126/sciadv.adi4350). In the revised manuscript, we will clarify this point and analyze small motifs with multiple excitatory units and heterogeneous correlations, focusing on both the timescale of weight redistribution and the resulting steady-state weights. We will also revise our interpretation of Supplementary Figure S9: within the simulated time window, the rate-dominated rule produces uniform inhibitory connectivity, but this does not exclude slower covariance-dependent reorganization.

      (2) It is unclear whether the distribution of inhibitory weights has stabilised after 25 minutes of simulation time (Figure 3C), given that a considerable proportion of (mutual) weights reach the maximum allowed weight while (unidirectional) weights appear to vanish. Without a maximum-weight bound, and given sufficiently long simulations, the weights might diverge to infinity or decay to zero, so the apparent stationarity may be imposed by the bound rather than reflecting a true steady state. This could also be a finite-size effect, given the small number of excitatory connections per neuron.

      The reviewer raises the possibility that the apparent stationarity in Figure 3C is influenced by the imposed weight bounds. In the revised manuscript, we will discuss this point and present extended versions of the simulations in Figure 3 where we remove the weight bounds. We will also test whether the observed behavior depends on network size or excitatory connection density. We note that, under different external input regimes, the weights stabilize without reaching the hard bound (Figure S11), suggesting that saturation is not a necessary outcome.

      (3) The connections from excitatory neurons to the two inhibitory populations are different in the ring model (exc to PV is wider than exc to SST according to Table 3), and it is not clear whether this width difference, rather than the plasticity rules themselves, is responsible for the emergence of the Mexican hat.

      We recognize that we did not fully justify the parametrization used in the ring model. The pre-existing ring architecture determines the spatial correlations available to iSTDP and therefore contributes to the learned connectivity. In the revised manuscript, we will add control simulations in which the excitatory inputs to the PV and SST populations have either identical or markedly different widths, while the plasticity rules are kept fixed. These controls will allow us to assess the relative contributions of these factors to the formation of a Mexican-hat effective-connectivity profile.

      Reviewer 2, weaknesses

      (1) The main limitations concern the extent to which the learned motifs are fully self-organized and how broadly the results generalize. In particular, the ring-network results rely on a pre-specified ring-like excitatory architecture and on two inhibitory populations with distinct plasticity rules, making it important to clarify which aspects of the Mexican-hat effective connectivity emerge from iSTDP itself. The conclusions would also be strengthened by intermediate plasticity rules. Finally, the ring-network simulations provide an interpretable proof of principle, but the authors should clarify whether the PV/SST effects depend on this specific architecture or would also arise in a more generic recurrent or cortex-like connectivity motif.

      This comment raises an important distinction between the components that are specified and those that emerge through plasticity. In our simulations, the excitatory architecture is fixed, whereas the initially weak and unstructured inhibitory-to-excitatory connections are learned through iSTDP. Thus, in the ring network, the spatial organization of excitation is prescribed, but the inhibitory connectivity and resulting Mexican hat-like effective connectivity emerge from the interaction of this architecture with the two iSTDP rules. Importantly, the central result that symmetric and antisymmetric rules promote distinct reciprocal and lateral E/I motifs is not restricted to the ring network, but is also observed in sparse randomly connected spiking networks and in networks with intrinsically generated irregular activity. The ring network is therefore used to demonstrate how these general motif-forming mechanisms can support specific circuit computations.

      Our framework parametrizes a broader family of pairwise iSTDP rules rather than relying exclusively on isolated, preselected rules, allowing the contribution of rule shape and rate-dependent terms to be understood analytically. Intermediate or “mixed” plasticity rules, including those considered by Yang and Doiron (2026) https://doi.org/10.1103/9nv2-y63v, as well as additional network architectures, are valuable directions for extending the framework. We will revise the Discussion to clarify which components are prescribed, which emerge through plasticity, and which conclusions apply beyond the ring-network implementation.

      (2) The authors largely achieve their aim [...] by demonstrating that different temporal forms of iSTDP lead to distinct learned E/I motifs and can shape effective connectivity and cortical-like response patterns. However, the broader biological interpretation remains more suggestive because some results depend on specific assumptions for the network architecture and plasticity rules.

      This comment points to an important distinction between biological generality and mechanistic insight. Our models are deliberately simplified to isolate how the temporal structure of iSTDP interacts with internally generated correlations to select distinct E/I connectivity motifs, and to make this relationship analytically tractable. The resulting predictions are then reproduced in large conductance-based spiking networks with different connectivity structures and sources of irregular activity. Thus, although we do not claim that the specific biological implementations considered here capture the full diversity of cortical circuits, the conclusions are not restricted to a single minimal model or network architecture. Adding further biological detail would introduce additional parameters and architecture-specific assumptions, but would not by itself establish greater generality or provide the same mechanistic understanding. In the revised Discussion, we will clarify the distinction between the general mechanistic principles established by our framework and the more specific biological interpretations that remain to be tested.

      (3) [...] At present, the work identifies rules that are sufficient to generate these motifs in model networks, while the mapping of these rules onto specific interneuron types remains for future experimental testing.

      Our use of the PV and SST labels is intended as a biologically motivated implementation rather than as a universal assignment of plasticity rules to these interneuron classes. The symmetric and antisymmetric kernels were motivated by in vitro measurements from PV and SST interneurons in mouse orbitofrontal cortex, respectively (Lagzi et al., 2021; https://doi.org/10.1101/2021.09.06.459211). Because inhibitory plasticity can vary across brain regions, developmental stages, and experimental conditions (Feldman, 2012; https://doi.org/10.1016/j.neuron.2012.08.001), the general conclusion of our work concerns the mapping from the temporal structure of an iSTDP rule to the E/I motif it promotes, rather than a fixed correspondence between a particular rule and an interneuron identity.

      Our models therefore establish more than the sufficiency of two isolated rules: the analytical framework explains how features of the plasticity kernel and rate-dependent terms determine whether reciprocal, lateral, or blanket inhibitory connectivity emerges. The specific association of these mechanisms with PV and SST interneurons in other circuits remains an experimentally testable prediction. We will clarify this distinction in the revised manuscript and emphasize that, once cell-type-specific plasticity rules are measured in a given circuit, the framework can predict the E/I connectivity motifs that those rules are expected to promote.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Laura Korobkova and Brian Dias describes an interesting study of the role of GABAergic neurons in the zona incerta (ZI) in incentive motivation for reward.

      The authors report that DREADD inhibition of ZI neurons reduced the effort breakpoint in a progressive ratio task, which measures the intensity of incentive motivation to obtain food rewards. In other tests, chemogenetic inhibition did not alter food consumption or memory.

      Conversely, DREADD excitation of ZI neurons increased incentive motivation in the progressive ratio task, expressed as a higher breakpoint for food rewards.

      Korobkova and Dias report that prior stress exposure to a series of stressors (e.g., forced swim & water submersion, restraint, mild footshock) by itself reduced the breakpoint for food reward under vehicle, though it did not impair the ability to learn an instrumental response. However, DREADD excitation of ZI neurons in previously stressed mice increased the breakpoint to normal levels equivalent to the never-stressed group. This important finding indicates the ability of ZI stimulation to rescue the incentive motivational deficit induced by prior stress.

      In fiber photometry studies using vGAT-CRE mice to specifically identify GABA neurons, Korobkova and Dias report that ZI GABA neurons are excited by sensory signals, including neutral cues. However, after reward conditioning, ZI GABA neurons increase their activation to the CS+ cue that predicts reward, but not to the CS- cue that doesn't. ZI neurons also respond in an instrumental reward task during both lever press and reward delivery. The authors conclude that ZI neurons respond to sensory stimuli, but specifically code the motivational significance of reward-related stimuli.

      In optogenetic studies, the authors find that ZI GABA neuron stimulation during a reward CS+ enhances motivated responding to obtain reward, particularly in females, but not stimulation outside the CS+. This suggests the ZI stimulation in females may specifically enhance the incentive salience of the CS+, namely the cue's ability to trigger an increase in 'wanting' for the reward. However, that effect was not found here in males.

      Altogether, this is a fine contribution to the literature, and the authors deserve congratulations on their study and manuscript.

      Strengths:

      This is a powerful and creative set of studies that clarifies the roles of ZI neurons in sensory processing and especially in incentive motivation for rewards. The use of multiple methods and test situations to triangulate on reward motivation functions gives a well-rounded perspective on ZI function. The discovery of incentive motivation roles for ZI neurons is intriguing and improves understanding of ZI, which traditionally has been a relatively understudied brain structure. The finding that ZI stimulation may rescue stress-induced deficits in motivation is especially notable and may have therapeutic implications.

      Weaknesses:

      Minor: This version of the manuscript focuses the introduction and discussion specifically on ZI GABA neurons. The ZI may be primarily GABAergic, but also contains other neurons, and DREADD studies may have used the hSyn promoter, which would impact all types of ZI neurons. Other studies here did more specifically target GABA neurons using vGAT Cre mice and specific targeting. The manuscript might be slightly improved by distinguishing in the discussion a bit more clearly which effects implicate GABA neurons specifically, and which effects might include other neurons too, to more clearly parse out the relative roles of GABA vs broader neuronal populations in ZI.

      The current version of our manuscript notes “Our chemogenetic manipulations targeted all GABAergic ZI neurons, and emerging evidence suggests molecular heterogeneity within this population that may map onto distinct functional roles (Arena et al., 2024; Wilt et al., 2025). Future work to first profile the molecular heterogeneity of GABAergic cells in the ZI and then using intersectional strategies to target defined ZI sub-populations will be essential to providing a more nuanced view of ZI GABAergic influences on motivation.” Our revision will make sure to discuss newer literature demonstrating cellular heterogeneity of ZI and emphasize that this heterogeneity will provide more nuanced contributions of ZI cells beyond the studied GABAergic population on motivation.

      Reviewer #2 (Public review):

      Summary:

      This paper describes a study that uses a combination of observational and experimental techniques to investigate the hypothesis that the zona incerta is a neural loci where sensory information is integrated to interpret the motivational value of reward-associated cues. They show that manipulation of GABAergic neurons in this region bidirectionally modulates responding during a progressive ratio test, that activating these neurons recovers motivational deficits incurred by chronic stress, and that they fire in response to reward-associated visual or auditory cues. They also showed that activity in these neurons is not necessary for incentive salience of reward-associated cues, because inactivating them did not prevent Pavlovian-instrumental transfer. However, activating them did enhance responding during the presentation of reward-associated cues in females but not in males.

      Strengths:

      The study has a very systematic and elegant approach to assess how this region responds first to intrinsic motivation and then to motivation-enhancing effects of reward-associated cues.

      Weaknesses:

      Males and females are used throughout, but sample sizes are generally too small to make a meaningful interpretation of sex differences (which is not the focus of the study, but is worth bearing in mind). In the last experiment, the lack of discrimination between CS+ and CS- conditions across training for males confounds any interpretation of sex-differences in the outcomes.

      The goal of our study was not to determine sex differences in motivation mediated by the ZI but rather demonstrate a general role of GABAergic cells in the ZI on motivation. As such, while all our experiments used male and female mice and our statistical analyses did not uncover any sex differences in most experiments, our sample sizes of each sex are too small to definitively make any statements about sex differences. As noted already in our Discussion, the sex-specific cue-utilization behavioral strategy in our optogenetic experiment warrants further investigation as a contributing factor to motivation that may or may not be influenced by GABAergic cells in the ZI. It bears mentioning that, to our knowledge, none of the recently published literature on the role of the zona incerta in learning, memory and appetitive behavior that is cited in this manuscript (including our own prior work) has uncovered sex differences in the contributions of the zona incerta to these behaviors.

      The ZI is known to be a region where there is notable convergence of neural inputs from a diverse and heterogenous range of sensory and other cortical inputs. To my knowledge, this is the first study that has directly tested whether it may serve to encode motivational/incentive properties of reward-associated cues. The outcomes are not definitive - it appears that they are sufficient but not necessary. However, this study represents an important first step - the ZI also has notable heterogeneity in the genetic identity of neurons, and properly dissecting the function of ZI microcircuits will likely require characterising function based on more than one molecular marker. This is addressed by the authors in the discussion.

      In summary, this study will have a significant impact on our understanding of how motivation is calculated based on complex environmental signals.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the zona incerta in motivation and cue-reward associations. Using chemogenetic and optogenetic manipulations of the ZI, they altered motivation in cued and uncued variants of the progressive ratio task and rescued deficits in motivation induced by chronic stress. They further use fiber photometry to demonstrate that the ZI tracks the formation of cue-reward associations.

      Strengths:

      (1) The authors fill an important gap in the literature linking sensory input to motivation via the zona incerta.

      (2) The authors demonstrate that ZI tracks cue value rather than just tracking sensory input.

      (3) The authors demonstrate that the ZI excitation rescues stress-induced suppression of motivation.

      (4) The authors perform several important control tasks, demonstrating that their findings are not a result of alterations in locomotor activity, food consumption, or memory.

      Weaknesses:

      In Figure 1D and E (inhibitory vs excitatory DREADDS), the control groups in the Gi group appear to have more elevated breakpoints than the control groups in the Gq group, although a statistical comparison between the two is not reported. It is not clear if this is because the two groups were given a different reinforcement schedule, this should be made clearer.

      In our revision, we will be sure to insert language re-emphasizing that the Gi and Gq experiments were performed using different reinforcement schedules, FR3 and FR1, respectively.

      In Figure 1E, it is important to note that although the authors found a significant planned comparison between Gq VEH and Gq CNO, the interaction was not significant, nor were comparisons to mice injected with control virus. Thus, activation of ZI GABA neurons appears to be a relatively weak effect.

      In our revision, we will insert language to acknowledge that reducing motivation after inhibiting GABAergic ZI cell activity is stronger than increasing motivation seen after stimulating the activity of these cells.

      In Figure 5, the authors see what is likely a significant difference in lever presses during acclimation between the Gi and GFP groups, which they state is an expected difference. However, it is difficult to see why this would be expected. While Gi:CNO manipulation yielded lower breakpoints in Figure 1D, it did not yield lower FR1 responding for food in Fig S3 (although this was FR1 for food dispenser visits rather than lever press). One reason I ask is that the authors highlight the differences in CS+/CS- between groups, but the biggest difference between groups appears to be in acclimation, which may be driving the group x block interaction.

      In a revision, we will revise the language to state that the acclimation difference that is more pronounced in the GFP control group vs the Gi group is to be expected because we had already shown that inhibition of GABAergic cells in the ZI would reduce lever pressing. We will also report a planned comparison of CS+ versus CS− responding that omits the acclimation block, which shows discrimination between sound (CS+) and light (CS−) in the Gi group but not in the GFP group, confirming that the cue effect is not driven by the acclimation difference. We will also note that acclimation responding is non-reinforced and effortful, whereas FR1 dispenser visits (Fig. S3) are reinforced and low-effort, which is why the two measures dissociate.

      In Figure 6, the authors demonstrate that optogenetic stimulation during cue light increases the breakpoint in females, but not in males. They suggest that this may be because the males did not sufficiently discriminate the cue light before optogenetic manipulation began. If this were the case, then the authors would need to use "cue discrimination" as a factor to determine if it is a better predictor than sex.

      In a revision, we will add an analysis that includes cue discrimination as a factor, to test whether it is a better predictor of the optogenetic effect on breakpoint than sex.

      The authors' work demonstrates that chemogenetic inhibition of GABAergic ZI cells reduces uncued motivation for reward but enhances cued responses under extinction. The authors state that this is a paradoxical finding that suggests that the ZI operates within a redundant motivation network. However, a critical difference between the two tasks is that one measures motivation for food while the other measures persistent responding under food extinction, which are not the same process. Thus, a simpler explanation is that ZI inhibition reduces motivation and impairs extinction.

      Our revision will include discussion of this important point and tie it into our previous work (Venkataraman et al. 2019 and 2021 – cited in this version) that included extinction-like protocols, albeit in classical (not operant) conditioning protocols.

  3. Jul 2026
    1. Author response:

      General Statements

      Please find below a revision plan for our manuscript entitled ‘Engulfment by brain macrophages in a short-lived vertebrate’ that was reviewed by Review Commons and that we wish to be considered for publication in eLife.

      We were very happy to see that all three Reviewers were interested in our study and we are sincerely grateful to all of them for their constructive suggestions, which we believe can largely be addressed and will improve our manuscript.

      We would like to inquire whether you would consider publishing our manuscript at eLife in its current form, along with the reviews. We will then provide an updated version of the manuscript, based on our revision plan below when we are able.

      Thank you so much for your consideration.

      Description of the planned revisions

      Reviewer #1:

      Major comments

      Comment 1: Although the authors describe and analyze their data from the viewpoint of engulfing macrophages, the paper would benefit from a broader perspective and a comparison to other studies on microglia in different species. Along this line, the title does not really seem to cover the data presented here very well, and the introduction lacks a proper explanation of terminology on microglia/brain macrophages and their known roles, cell types versus cell states and the current state of the art in fish versus other model species in the context of aging.

      We thank the Reviewer for these suggestions. We will expand the introduction to add more information and references on microglia/brain macrophages across fish and other species. We will also edit the title to more closely reflect the data in the manuscript.

      Comment 2: The result that nearly all myeloid cells in the killifish brain are of the engulfing macrophage type is somewhat surprising. This appears to differ from other studies in for instance zebrafish (e.g. ref 80, that describes the heterogeneity of the myeloid cells in detail). There are two questions we like to raise: (Q1) What is the evidence towards this homogeneity? and (Q2) Could there be a technical bias?

      We agree with the Reviewer, and we were also surprised to find a fairly homogenous myeloid population even in wildtype (non-transgenic) killifish brains, especially given that all these 3 different populations of myeloid cells (microglia, non-microglial macrophages, and dendritic-like cells), could be found in the adult zebrafish brain by Rovira et al. using single-cell RNA sequencing of cd45+ cells (ref. 80).

      We will better describe our evidence towards this homogeneity among killifish brain myeloid cells in the manuscript by adding a paragraph in the text at the end of the section referring to Figure 2. We will also provide additional experiments and analyses (see also below, our detailed responses to Q1 and Q2):

      i) In our and Ayana et al.’s single-cell RNA sequencing datasets, we could not find distinct populations of myeloid cells corresponding to microglia, non-microglial macrophages, and dendritic-like cells, even in wildtype (non-transgenic) killifish young adult and old brain, and even when subclustering only the myeloid cells. However, cd45+ cells were not specifically enriched in these experiments.

      ii) When using PCA, we also could not separate killifish brain macrophages into microglia, border-associated macrophages, and monocyte-derived macrophages (whereas we could separate mouse brain macrophages into these three groups as a positive control for this approach).

      iii) As noted by the Reviewer below, our analysis of the Ayana et al. wildtype dataset was limited to the young adult timepoint from this dataset. In the revised version of our manuscript, we will also include analysis of the old timepoint from this dataset.

      iv) As described in our responses to Q1 and Q2 below, we will provide more detailed subclustering that highlights the heterogeneity we do observe among killifish brain myeloid cells. We find this heterogeneity to be subtle and likely corresponding to different cell states rather than functionally distinct cell types. We will also use hybridization chain reaction (HCR) to test whether cd74, a key marker we use to propose a monocyte-derived macrophage-like identity for almost all detected killifish brain macrophages, labels apoeb+ cells in young wildtype brains in situ (as would be predicted based on our single-cell RNA sequencing analysis). We will additionally use HCR to test whether batf3, a key marker used by Rovira et al. to identify dendritic-like cells, labels any cells in the killifish brain (based on our single-cell RNA sequencing analysis, we do not expect to find batf3+ cells in the killifish brain).

      We agree that there could be technical confounds that could lead to the observed homogeneity of the myeloid population and lack of canonical microglia and dendritic-like cells even in wildtype killifish in our analysis. These could include our/Ayana et al.’s protocol for brain dissociation, our FACS-sorting step, or the step of loading cells into the 10x Genomics chip for single-cell RNA sequencing – each of these could be a step where canonical microglia and/or dendritic-like cells could be lost. In addition, we did not enrich for cd45+ cells, so it is possible that rare populations of immune cells in the killifish brain may not be represented in our dataset. As described below in our responses to Q1 and Q2, we will perform additional analyses that we hope will help address some of these key points, and we will more extensively discuss all possible technical confounds leading to potential cell type bias in the Discussion section.

      In addition to discussing potential technical confounds, we will also better highlight biological possibilities that could explain the differences between our and Rovira et al.’s findings. For example, the observed lack of microglia and dendritic-like cells in killifish brains could reflect a species (killifish vs. zebrafish) difference. Alternatively, even though both studies analyzed young adult timepoints, there could also be differences in biological age between the young adult killifish and zebrafish analyzed, which may be additionally influenced by husbandry conditions (e.g., feeding, pathogen exposure).

      Regarding (Q1): What is the evidence towards this homogeneity? The markers used are overlapping with markers for microglia. It would be helpful to clarify how canonical microglia populations are represented in the dataset. What is the heterogeneity of the oScarletHIGH cells? On several plots (Fig1f, Fig2d, Fig4a) this population of cells seems more heterogeneous than described. Are there different cell states or types? What is the percentage of myeloid cells that is oScarletLOW? To what extent do these cells compare transcriptionally to the oScarletHIGH cells?

      The Reviewer makes a series of excellent points. To address them:

      i) We will more prominently highlight markers not expressed by canonical microglia (e.g., mrc1, cd74) that are expressed by nearly all detectable killifish brain myeloid cells, even in wildtype datasets. We could not find canonical microglia (e.g., tmem119+, sall1+) in any of our killifish datasets. We will highlight these results in the text and will move some of the relevant graphs in the main figure.

      ii) We agree that oScarlet<sup>HIGH</sup> cells are somewhat heterogeneous. To better understand the source of this heterogeneity, we will present UMAPs/Seurat FeaturePlots focusing on cell state markers (e.g., cycling [mki67+] and activated [iba1+] cells) as well as specific cell type markers from mammalian/zebrafish literature (canonical microglia, non-microglia macrophages, dendritic-like cells, etc.). We expect this analysis to show distinct cell states, but fairly uniform expression of cell type markers across all oScarlet<sup>HIGH</sup> cells in our datasets.

      iii) The percentage of detected myeloid cells that are oScarlet<sup>LOW</sup> is very low (<1%). However, as noted in the next comment from the Reviewer, our sorting protocol de-enriches for oScarlet<sup>LOW</sup> myeloid cells, so the actual percentage of all myeloid cells that is oScarlet<sup>LOW</sup> may be higher. We will include additional panels to illustrate this point.

      iv) oScarlet<sup>LOW</sup> myeloid cells are largely transcriptionally similar to oScarlet<sup>HIGH</sup> cells (which are almost all myeloid). The main differentially expressed genes in oScarlet<sup>LOW</sup> vs oScarlet<sup>HIGH</sup> myeloid cells are markers of non-myeloid genes (e.g., elavl3, mpz). These could reflect phagocytosis of non-myeloid cells by oScarlet<sup>LOW</sup> myeloid cells but could also reflect ambient RNA contamination, since oScarlet<sup>LOW</sup> myeloid cells (but not oScarlet<sup>HIGH</sup> cells) were sequenced alongside non-myeloid cells. We will include the differential expression analysis and add commentary in the text to discuss these possibilities.

      Fig1f-i depict an enriched oScarletHIGH group alongside oScarletLOW cells. This representation is a bit misleading since it seems to indicate that really all myeloid cells are of the engulfing macrophage type whereas it is the majority, but not all.

      We agree with the Reviewer. We will also include different representations that provide more accurate estimates of the proportion of all myeloid cells that are oScarlet<sup>LOW</sup>/not engulfing macrophages.

      Line 52: The authors describe that the oScarletHIGH cell group is "enriched for signatures characteristic of macrophage functions". This finding is logical, as the isolation procedure of this population of cells was based on the endocytic and phagocytic properties of the cells. This result appears more consistent with a validation of the isolation strategy than with definitive evidence for myeloid cell identity.

      We agree and we will reframe this result as a validation of the isolation strategy.

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We agree with the Reviewer that brain dissociation, FACS, and single-cell RNA sequencing could all represent technical biases that could all influence the representation of different myeloid cell types in our dataset. We will clearly note this potential confound in the Results and the Discussion.

      We agree that it would be valuable to test some of these issues experimentally. We will perform non-dissociative HCR (in situ) to test (1) whether in young wildtype brains, apoeb+ cells (macrophages) express cd74 (a key marker for our argument re: MDM-like cell type identity) and (2) whether we can find cells expressing batf3 (a key marker used by Rovira et al. to identify dendritic-like cells, which we cannot detect in the killifish brain in our single-cell RNA sequencing data). We will also acknowledge that these are only individual markers and would not provide as comprehensive a picture of cell type identity as single-nuclei RNA sequencing.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for these excellent suggestions. We will do the following in our revised manuscript:

      i) We will include comparisons to all immune (CD45+) cells from mice (Barr et al.).

      ii) We will include comparisons to all immune (cd45+) cells in zebrafish (Rovira et al.).

      iii) We will also analyze the old age time point from the Ayana et al. wildtype killifish dataset.

      The percentage of oScarletHIGH cells in the aged condition is 8% (Fig1-suppl1) compared to 4% at young age. On the other hand, a lower number of cells was isolated at old age compared to young age (Figure4a). Can the authors elaborate on this difference? Later on, it is stated that the engulfing capacity declines with aging, but could this be linked to the lower or potentially biased recovery of cells?

      We appreciate the Reviewer’s point. We will elaborate on the percentage differences in oScarlet<sup>HIGH</sup> cells recovered in the different experiments, and more clearly indicate what they could originate from and whether they could influence the engulfment results:

      i) We will indicate more clearly that in each single-cell RNA sequencing experiment, we loaded the entire set of Live oScarlet<sup>HIGH</sup> cells collected from each condition onto the 10x Genomics chip. In the young vs. old comparison, more Live old oScarlet<sup>HIGH</sup> cells were loaded than Live young oScarlet<sup>HIGH</sup> cells (66K vs. 42K, based on the FACS sorter counts, though these numbers are likely overestimates). One possibility is that the lower recovery of successfully sequenced old oScarlet<sup>HIGH</sup> cells might be explained by differences with age in oScarlet<sup>HIGH</sup> cells’ ability to remain intact during loading onto the 10x Genomics chip.

      ii) In the ex vivo experiments where we find that engulfment capacity declines with age (current Fig. 5), the proportion of oScarlet<sup>HIGH</sup> cells among all Live cells slightly increased with age. We will include the data illustrating this, and indicate more clearly in the text that the differences in engulfment capacity observed in these experiments are unlikely to be explained solely by lower recovery of oScarlet<sup>HIGH</sup> cells from old brains.

      iii) We will more clearly acknowledge in the Discussion the possibility of potentially biased cell recovery with age.

      Figure4a: Transcriptional differences are stated between young and old (line 226), can a relevant selection be shown in e.g. a dotplot or heatmap?

      This is another great point from the Reviewer. We will show these genes in a dotplot/heatmap in a Supplemental Figure. We will also clearly indicate in the main text that many of the genes most strongly enriched in old oScarlet<sup>HIGH</sup> cells in our young vs. old comparison experiment were not strongly expressed in an independent old oScarlet<sup>HIGH</sup> cells dataset (which was compared to old wildtype cells, without a young counterpart). 

      The UMAP clustering does seem to indicate batch effects on panels a and g. Can the authors provide subclustering and show that young and old/ FACS sorted high and low cover similar cell types/states? The PCA plot (panel f) and marker analysis is not fully convincing, as PC1 and 2 alone do not suffice to explain all the variance in these cells, and the markers are common ones for many microglia/macrophage cell types (and thus likely to be expressed similarly).

      The Reviewer has another excellent suggestion. We will provide subclustering, which indeed shows that young and old oScarlet<sup>HIGH</sup> cells generally represent similar cell types and states. This analysis also shows a subcluster specific to old oScarlet<sup>HIGH</sup> cells, but this subcluster is far less pronounced in a second old oScarlet<sup>HIGH</sup> dataset (see our previous comment), so we will clearly indicate in the main text that this subcluster may not be robust. We will also remove the PCA plot.

      Figure 5: It would be informative to include the corresponding aged condition for panels c and e.

      We agree and we will include this.

      Minor comments

      The authors use the oScarlet fish in a heterozygous state. Is the homozygous line not viable? Or what is the reason to use the heterozygous state? Too high expression of OScarlett? Is apoptosis of neurons checked for this line? Compared to wild type?

      The homozygous SP-oScarlet line has not been characterized, and we will include this information in the Results. We did not check apoptosis of neurons in the SP-oScarlet line or in wildtype killifish, and we will indicate this in the Methods and Discussion.

      Is the secretion of OScarlett completely proven? Maybe the signal is engulfed by the clearance of cell debris from apoptotic neurons?

      This is a great point. We have not directly tested the secretion of oScarlet in the SP-oScarlet line, and it remains possible that engulfed oScarlet also comes from apoptotic neurons that are being engulfed by myeloid cells. We will acknowledge this possibility in the Discussion.

      We will also more clearly highlight references that adding a signal peptide to proteins (as we did for the SP-oScarlet line) is sufficient to induce their secretion in several contexts in different species (although this does not prove secretion in this particular case).

      We also do have another line (built for a separate project) in which oScarlet is expressed cytoplasmically in elavl3+ cells. In this line, in pilot data, we did not observe oScarlet<sup>HIGH</sup> cells by flow cytometry. While these experiments are not of publishable quality, these observations also suggest that the presence of a signal peptide on oScarlet is necessary for the engulfment phenotype.

      Line 329. What exact difference between teleosts and mice are you pointing at?

      We will edit the text to clarify that we are referring to our inability to find a cell population with a transcriptional profile like that of canonical adult mouse microglia in the adult killifish brain.

      Suggestions for Figures

      General remark IHC/HCR figures: Please add overview figures to guide the reader to ROI shown.

      Fig1.

      1.b Please add overview figures to show the overall distribution of the signal in the brain. Are there any hotspots or low-abundant regions?

      1.d/1.e Please add numbers of cells on figure panels.

      1.d: overview figure is unclear. Telencephalon seems to be missing? Can annotation be added to regions of the structure?

      Additional in vivo evidence of the homogeny/heterogeneity of oScarlet protein+/RNA- cells would be informative, e.g. double labeling (HCR) with some of the markers of panel Fig. 2a). This also would exclude location bias. One could expect myeloid cells that are not in close proximity to secreting neurons, for instance in dense neurogenic niches

      Fig. 2

      2.c Please add number of cells on graph

      2.d Please add brain regions to graphs of published datasets like was done for the own dataset.

      Subtext: Why is there a different number of cells for Barr in c and d? Please add number of cells of own killifish dataset in legend.

      Supplement fig. 2 Please add additional microglia-specific markers such as for instance HEXB, GPR34, SELL1, C1Q.

      Fig.3

      Please add overview pictures.

      Fig.4

      4.g grey color not very visible

      It would be informative to have the markers of 4.h plot on top of the UMAP to show the distribution/differences, maybe in supplementary information.

      We will implement all of these changes. We will add cd74 HCR labeling of oScarlet<sup>HIGH</sup> cells that do not express oScarlet RNA transcripts in situ. The different numbers of cells from the Barr et al. dataset in the current panels 2c and 2d do not reflect differences in brain regions from which macrophages were isolated (in both cases, the same whole-brain dataset was used). Instead, in this dataset, there is a small population of interferon-responsive macrophages that are not microglia, BAMs, or MDMs; these cells were included in the analysis in 2d but not 2c. We will clarify this and update the analysis in current panel 2d to include all immune cells from the Barr et all. dataset, as suggested by the Reviewer above.

      Reviewer #2:

      Red fluorescent proteins are notorious for being prone to aggregation. Are oScarlet proteins being internalized by macrophages aggregates or soluble proteins? This distinction is important as the clearance of extracellular molecules could be mediated by most cells yet aggregates could be removed specifically by macrophages. Can experiments be conducted to distinguish between these two possibilities? We realize this may be challenging. If not feasible, the discussion should be tempered to reflect this possibility.

      The Reviewer makes a great point, and we were also interested in this question. mScarlet and its derivatives (including oScarlet) would generally be expected to remain in the monomeric state to a greater extent than other red fluorescent proteins like mCherry (Bindels et al., Albakri et al.). We will indicate this with references in the Results section.

      But this does not preclude the possibility of oScarlet aggregation in SP-oScarlet brains. We did collect soluble and insoluble fractions from SP-oScarlet brains using a gentle lysis method that would be expected to preserve aggregates in the insoluble fraction (Avar et al.). However, we found that the levels of oScarlet in both fractions were below the limit of detection by western blot and by a plate reader-based fluorescence assay, thereby precluding us from directly testing how much oScarlet was aggregated. We will address the possibility of oScarlet aggregation, and its potential impact on engulfing macrophages, in the Results and Discussion sections.

      Brain dissociation tends to generate a lot of debris, especially from sheared neurons. Therefore, the high level of oScarlet inside macrophages could be an artifact of dissociation rather than a reflection of in vivo clearance. Authors should use internalization inhibitors during dissociation (CytoD, Dynasore, and pitstop) to exclude this possibility. Alternatively, if they have a transgenic killifish that expresses another fluorescent reporter in neurons (and preferably at a similar level to that of oScarlet), authors should dissociate brains together and quantify how many oScarlet+ cells are now also positive for that other fluorescent reporter. This could give an idea of how much engulfment is occurring due to the dissociation processes. It is not ideal, as macrophage eating could be happening during dissociation but before cells are in single cell suspension. However, given that RNAseq is needed to identify macrophages, this reviewer would be satisfied by this alternative approach if the aforementioned pitfall is also presented in the discussion.

      This is another excellent suggestion, and we have already done the following experiments:

      i) We have performed bulk RNA sequencing from FACS-sorted oScarlet<sup>HIGH</sup> cells from middle-aged SP-oScarlet brains dissociated with five inhibitors (cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin) in addition to transcription/translation inhibitors. Deconvolution analysis of these bulk RNA sequencing data showed that oScarlet<sup>HIGH</sup> cells in the presence of cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin are also almost all macrophages. While these data are not at single-cell resolution, they indicate that engulfment by macrophages is unlikely to solely occur during the dissociation process. We will include a revised Supplementary Figure with these experiments.

      ii) We have also already performed a pilot co-dissociation experiment (in a different genetic background) as suggested by the Reviewer using a ubb:GFP reporter as the second transgenic and observed very few GFP+ oScarlet<sup>HIGH</sup> cells (see Author response image 1). We will include co-dissociation data in the revised version of the manuscript.

      Author response image 1.

      We believe that the results of these experiments are consistent with the notion that oScarlet<sup>HIGH</sup> cells have engulfed oScarlet in vivo, prior to dissociation step. Nevertheless, we will also make note in the Discussion of the pitfall mentioned by the Reviewer.

      Related to the above, it appears based on the scRNAseq that dissociation heavily enriched for brain macrophages. Therefore, the claim that clearance is mostly macrophage mediated could be due to an enrichment of this population during dissociation rather than this cell type being responsible for most of the extracellular waste disposal. Authors should quantify the % of total oScarlet that is specifically in macrophages in the brain sections they already have that are stained against oScarlet and CSF1R/ApoEB transcript.

      The Reviewer’s point is well taken and we agree that experiments in intact brain sections are important to orthogonally test whether macrophages are responsible for engulfment. As suggested by the Reviewer, we will quantify the percentage of total oScarlet that is in apoeb+ macrophages in the brain sections we already have (current Fig. 3a). We note that it may not accurately reflect the actual in vivo percentage, as we have found that it can be challenging to draw cell boundaries around macrophages due to their irregular shapes.

      It is concerning that dextran and oScarlet are almost perfectly colocalized in the image presented (Figure 3a). It raises the possibility, among others, that dextran is sticking to potential oScarlet aggregates and then being internalized by macrophages. Therefore, it could be an artifact of the transgenic line. Authors should repeat the experiment in wildtype fish and use HCR against CSF1R/ApoEB to address this issue.

      We agree with the Reviewer. We have already injected dextran into non-transgenic killifish brains and observed dextran engulfment by apoeb+ cells. We will include this experiment in the revised version of the manuscript.

      Reviewer #3:

      Major comments

      The SP-oScarlet model enriches cells based on phagocytic capacity - by design, any phagocytic cell, including microglia, can be labeled. Only 0.5% of oScarlet<sup>LOW</sup> cells were myeloid cells, confirming that this method captures virtually the entire myeloid population. The transcriptional resemblance to BAMs/MDMs is therefore a post hoc characterization of brain phagocytes broadly, rather than evidence for a selectively labeled subset. The authors show examples of apoeb<sup>+</sup> cells near vasculature (Fig. 3b), but do not provide a comprehensive quantification of the full spatial distribution of oScarlet<sup>HIGH</sup> cells. Importantly, neither the SP-oScarlet macrophages nor previously published wild-type killifish brain macrophages could be transcriptionally separated into three subgroups analogous to mammalian microglia, BAMs, and MDMs by PCA. This suggests that fish brain macrophages may not exist as subpopulations that correspond with their mammalian counterparts. The authors should therefore describe these cells as brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics, rather than implying they are a population equivalent to mammalian BAMs/MDMs.

      We agree with the Reviewer and will describe killifish brain macrophages as “brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics.”

      As described in our responses to Reviewer 1’s comments, we will also provide/more prominently highlight analysis of killifish brain myeloid cell datasets showing:

      i) Markers (e.g., mrc1, cd74) that are shared between mammalian BAM/MDMs and killifish brain myeloid cells.

      ii) The heterogeneity we observe among killifish brain myeloid cells, which we believe corresponds to different cell states, but is not likely pronounced enough to correspond to functionally distinct cell types.

      iii) The presence of oScarlet<sup>LOW</sup> killifish brain myeloid cells, which are de-enriched by our sorting strategy.

      Minor comments

      The authors acknowledge that oScarlet is relatively resistant to degradation and could affect the state of cells that engulf it. The independent single-cell experiment comparing old SP-oScarlet and old wild-type macrophages addresses this concern at the transcriptional level. However, it remains possible that chronic accumulation of degradation-resistant protein in lysosomes could itself impair phagocytic capacity. If so, the age-related decline in ex vivo engulfment might be partially attributable to oScarlet burden accumulated over a lifetime, rather than to aging per se. While a direct functional comparison with old wild-type macrophages is experimentally challenging, the authors should at least acknowledge this interpretive caveat in the Discussion.

      The Reviewer makes a great point, and we will acknowledge this caveat in the Discussion.

      The transcriptional data show only modest changes in engulfment- and lysosome-related pathways with age. The authors should offer one or more specific, testable hypotheses (e.g., which phagocytic receptors or lysosomal components might be affected) to guide future investigation.

      We will provide a list of top engulfment-relevant genes that change the most with age, which also show only modest changes. We will better acknowledge that the age-related transcriptional changes in engulfment-relevant genes observed in our single-cell RNA sequencing experiment are modest and are unlikely to directly serve as a resource for testable hypotheses. We will also move our single-cell RNA sequencing data to the Supplementary Figures.

      In the text, gene names such as APOEB are capitalized, whereas the convention for fish gene nomenclature is lowercase italics (e.g., apoeb). The authors should ensure gene formatting follows species-appropriate conventions throughout the manuscript.

      We thank the Reviewer for this suggestion, and we will change killifish gene names to be lower-case italicized throughout the manuscript.

      Description of analyses that authors prefer not to carry out

      Reviewer #1:

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We believe that generating our own single-nuclei RNA sequencing dataset would be beyond the scope of this study. However, a preprint by Williams et al. from the Benayoun Lab (https://doi.org/10.64898/2026.04.09.717549) includes single-nuclei RNA sequencing data from wildtype killifish brains and could provide an excellent test of our findings. We will point readers to this preprint in the Discussion.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for the SAMap suggestion. We did perform SAMap with a mouse reference from the Allen Brain Atlas. Our SAMap analysis identifed distinct microglia-, BAM-, and dendritic-like myeloid populations. However, we were not confident in these SAMap results because the differences between the populations were subtle and we thought they would be unlikely to represent differences between functionally distinct cell types. Of note, the dendritic-like cells we identified by SAMap in the killifish brain did not express canonical markers of dendritic/dendritic-like cells from mammalian and zebrafish literature (e.g., batf3). Additionally, other SAMap studies from the Wang Lab analyzing non-neural cell types in other vertebrates do not separate microglia from non-microglial macrophages (Kalakuntla et al., in preparation).

      Comment 4: Regarding the comparison with the aged brain:

      Figure 4: It would be nice to include the same comparisons as for young fish (cfr Fig.1 panels F-I).

      We agree with the Reviewer. Unfortunately, we did not sequence oScarlet<sup>LOW</sup> cells from old brains, for cost reasons, and this precludes the comparison requested by the Reviewer. We will more clearly indicate that we do not have the oScarlet<sup>LOW</sup> cells from old brain in the Results section.

      Reviewer #2:

      The flow cytometry strategy used does not distinguish between oScarlet protein that has been internalized versus that which is sticking to the surface of macrophages. Authors should stain non-premeabilized and permeabilized cell suspensions with a flow antibody against mCherry/RFP to get a sense of how much oScarlet is inside versus outside of the macrophage. For most antibodies this can be done on the same sample sequentially if the antibodies have a different fluorophore.

      The Reviewer makes an important point, and we will address this caveat in the Methods. We have performed a pilot of the flow cytometry experiment suggested by the Reviewer and, encouragingly, observed a higher proportion of cells labeled by the antibody (the same anti-mCherry antibody we used to label oScarlet in situ) in the permeabilized condition. However, we do not wish to publish these results as even in the permeabilized condition, the antibody labeled <0.2% of cells – meaning it likely did not label most oScarlet<sup>HIGH</sup> cells, making it difficult to draw strong conclusions about internalized vs. surface oScarlet in these cells.

      Reviewer #3:

      The age-related decline in oScarlet fluorescence in oScarlet<sup>HIGH</sup> cells in vivo could reflect either reduced phagocytic capacity of macrophages, or reduced oScarlet secretion by neurons, as the authors have discussed (Fig. 5a). The ex vivo assay addresses this by standardizing substrate concentration, which is a strength, but an in vivo functional assessment would provide a more physiologically relevant complement. The authors have already established the methodology for in vivo substrate injection (Fig. 3a, dextran). A similar experiment comparing substrate uptake in young and old fish would circumvent potential artifacts of the ex vivo approach, such as enzymatic dissociation altering surface receptor availability, and would directly test whether engulfment declines in the native brain environment.

      We agree with the Reviewer and performed a pilot of the suggested experiment, which did not show age-related differences in in vivo engulfment of injected dextran. However, in developing the brain injection procedure, we concluded that this procedure works well for comparisons within a sample, but not for comparisons across samples and conditions (e.g., young vs. old). This is because the injection site is determined by sight (aiming for the most medial point on the telencephalon/optic tectum border, with no standardized way to control injection depth), and not stereotactically with coordinates. With our current procedure, it is challenging to perform reproducible injections across ages due to known age-related differences in fish/brain size and skull thickness.

      We are interested in developing stereotactic injection approaches that would be compatible with comparisons across ages and other conditions, but we believe that this is beyond the scope of this manuscript. We will discuss this possibility as a future direction in the manuscript.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      The evidence described for the claim that this technique improves the alignment of the reconstruction of small complexes compared to standard techniques is incomplete. The authors could better evaluate the effects of model bias on the reconstructed densities, as suggested by reviewer #1.

      We thank the editors for highlighting this remaining concern. To better evaluate the effects of model bias, we have performed the FSC-based analyses suggested by Reviewer 1, including a map–model FSC of the omit map and an FSC between the half-maps (though the latter is unreliable for the composite), and added the results to the revised manuscript (new Figure 5—figure supplements 1–3; detailed under Requested FSC Analysis below). We have additionally revised the text to avoid overstating the reconstruction as “unbiased” and to soften comparisons with other reconstruction methods (detailed below).

      Public Reviews:

      Reviewer #1 (Public review):

      In the revised version, the refinement of atomic occupancies in the 2DTM-generated maps has been insightful: densities only come back at values ranging from 0.55–0.80, whereas residues included in the template remain at 1, suggesting that the 2DTM-reconstruction does suffer from model bias. Their newly added Omega calculations, which are helpful, also suggest that model bias is present in the 2DTM-based reconstructions. These observations therefore contradict the first subsection heading of the Results, which claims “unbiased reconstruction of omitted residues”.

      We agree that the previous subsection heading was too absolute. We have changed the heading from “Unbiased reconstruction of omitted densities in a 43 kDa protein kinase” to “Recovery of densities omitted from the template in a 43 kDa protein kinase”. The opening sentence now states that we evaluated the ability of 2DTM to “recover omitted ligand densities” (L123–125).

      We also revised nearby statements to describe the observations without claiming that the entire reconstruction is free of template bias. The manuscript now states that, because the corresponding features were omitted from the search template, the recovered densities cannot result from direct inclusion of those features in the template (L150–153).

      For the omitted alpha-helical turn, we now state simply that its density was recovered despite its absence from the search template (L191–192). We also revised the interpretation of the occupancy-refinement results from “confirming partial, unbiased recovery” to “supporting partial recovery of density in the omitted regions” (L198–199). The same wording has been applied to the Supplementary file 1 caption on page 27.

      Finally, the summary of the ligand-deletion experiments now states that a ligand and nearby residues can be deleted to reduce template bias while retaining sufficient signal for their density to be recovered (L248–251). In the Methods, we retain the description that the composite omit map was constructed to avoid template bias at the omitted locations (L953–954).

      We have also toned down comparative statements about reconstruction accuracy, including the relevant subsection heading (“Comparison of 2DTM and RELION reconstructions from the same particle stack” at L252–253) and the surrounding discussion (L274–279).

      Requested FSC Analysis

      The measurement of how much model bias is present in this OMIT map by FSC calculations is still pending. This could be done in two ways. My original suggestion was to calculate a mapto-model FSC for the OMIT map and the full reference. This should be compared with a similar map-to-model FSC on the map where only the ligand was omitted. Alternatively, they can use the cisTEM FSC uncorr procedure on the OMIT half-reconstructions and compare the resulting curve with the one presented in Figure 1b.

      We have now completed both analyses and added them to the revised manuscript (new Figure 5—figure supplements 1–3, with accompanying Results and Methods text).

      (1) Map–model FSC. We computed the map–model FSC between the composite OMIT map and a density simulated from the full 1ATP model, and compared it with the equivalent FSC for the Figure 1 reconstruction, in which the ligand and residues 222–227 were omitted from the template (Figure 5—figure supplement 1). The composite OMIT map crossed FSC = 0.5 at 3.5 Å and FSC = 0.143 at 3.0 Å, compared with 3.0 Å and 2.4 Å, respectively, for the Figure 1 reconstruction. Because each local region of the composite map was taken from a reconstruction in which the corresponding residues were absent from the template, this agreement reflects genuine recovery rather than direct inclusion of those local features in the template. The lower FSC values relative to the Figure 1 reconstruction are expected. In the Figure 1 reconstruction, most of the protein remained in the template, whereas the composite is assembled from disjoint local omit regions. This stitched construction introduces holes and mask boundaries that affect Fourier-space agreement across the curve, in addition to the weaker, partial recovery of locally omitted density. Thus, the map–model FSC provides a conservative Fourier-space assessment of recovered omit-region density. The composite map–model FSC also shows a negative dip at the lowest spatial frequencies, which does not indicate failed recovery. Radial binning around the omitted atoms (Supplementary file 3) shows that the atom-centred shells (≤2 Å) recover positive but weakened density upon omission, whereas the peripheral shells (2–3 Å) are negative and nearly identical whether the residue is present or omitted, leading to net-negative density within the molecular envelope. We therefore interpret the low-frequency dip as a consequence of the composite construction rather than as evidence for failed recovery of omitted density.

      (2) Half-map FSC<sub>uncor</sub>. We also computed the half-map FSC of the individual OMIT reconstructions (Figure 5—figure supplement 2): the 36 individual omit reconstructions cross FSC = 0.143 at a median resolution of 2.8 Å, comparable to the Figure 1b reconstruction (∼3.0 Å). The composite map’s own half-map FSC appears higher, but this value is not a reliable resolution estimate because the two composite half-maps are assembled using the same voxel-assignment masks. This shared support introduces artificial correlations, which we demonstrate with a phase-randomization control (Figure 5—figure supplement 3). We therefore use the individual omit-reconstruction FSCs and the composite map–model FSC, rather than the composite half-map FSC, to assess Fourier-space agreement.

      Together, these analyses show that the OMIT reconstructions contain high-resolution signal in regions absent from the corresponding search templates, while also identifying the low-frequency dip and the inflated composite half-map FSC as consequences of the conservative stitched composite construction.

      Reviewer #3 (Public review):

      Nor was it compared to more recent strategies for processing SPA data from small molecules, such as Blush regularization or HR-HAIR. [...] This places this method as a complementary technique, and whether it outperforms those methods for a wide variety of molecules is yet to be determined.

      We agree that a systematic comparison with recent small-particle SPA methods such as Blush regularization and HR-HAIR will be important. We have added this point to the Discussion (L831–846). We also note that such as comparison should consider not only particle stack and molecular mass, but also the fidelity of the 2DTM template forward model. In the ideal limit of an accurate template and forward model, 2DTM should provide a strong prior for particle detection and pose determination. In practice, however, current templates remain imperfect approximations to the experimental signal because of inaccurately modelled solvent-boundary effects and atomic scattering factors, bonding and charge redistribution, and conformational mismatch. Improving template generation is therefore an important direction for extending the range of molecular targets and imaging conditions where 2DTM can be applied.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study addresses the mechanism of action of benzoylurea insecticides and explores the metabolic consequences of inhibiting glycogen breakdown in insects. Both reviewers identify major flaws with the premise of the work. The strength of the provided evidence is inadequate as the data do not, or poorly, support several central claims. The significance of the findings is considered marginal.

      The Assessment stated that “both reviewers identify major flaws with the premise of the work” and that “the strength of the provided evidence is inadequate.” We have addressed both dimensions:

      (1) Premise: The Introduction has been substantially restructured to explicitly acknowledge the compelling CRISPR/Cas9 evidence establishing CHS as the primary site of BPU resistance (Reference 1). The study is now reframed as a systematic evaluation of GP as an independent insecticidal target and an investigation of metabolic compensation mechanisms — questions with scientific value independent of the BPU mechanism debate (see details in lines 47-54 of the revised manuscript).

      (2) Evidence: Four new sets of experiments directly address the specific evidence gaps identified by the reviewers: (i) GP enzyme activity measurements in RNAi-treated larvae; (ii) expression analysis of alternative glycogen catabolic enzymes; (iii) molecular docking and MM/GBSA binding free energy analysis; (iv) comprehensive fitness cost assessment including feeding rate, larval weight, pupal weight, adult wing area, and female fecundity.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate whether glycogen phosphorylase is a potential molecular target of benzoylphenylurea insecticides and examine the physiological consequences of inhibiting glycogen breakdown in the diamondback moth Plutella xylostella. The authors express and characterize recombinant glycogen phosphorylase, test its inhibition by a mammalian glycogen phosphorylase inhibitor and by the insecticide diflubenzuron, and assess the physiological effects of glycogen phosphorylase inhibition through chemical exposure and RNA interference. Based on these experiments, the authors conclude that benzoylphenylurea insecticides do not target glycogen phosphorylase and propose that insects compensate for glycogen phosphorylase inhibition through activation of gluconeogenesis, allowing them to maintain glucose homeostasis and complete development despite strong suppression of the enzyme.

      Strengths:

      The study addresses an interesting and long-standing question in insect toxicology regarding the mechanism of action of benzoylphenylurea insecticides. The authors combine several complementary approaches, including recombinant enzyme characterization, inhibitor assays, RNA interference, gene expression analyses, and metabolite measurements. The biochemical characterization of the recombinant glycogen phosphorylase and the demonstration that the tested glycogen phosphorylase inhibitor can strongly inhibit enzyme activity represent important technical strengths. In addition, the study integrates biochemical and physiological observations to explore how insects might compensate for disruptions in central carbohydrate metabolism.

      We are grateful that Reviewer 1 recognized the study's strengths, including the complementary multi-approach strategy, the biochemical characterization of recombinant PxGP, and the integration of biochemical and physiological observations.

      Weaknesses:

      (1) The proposed compensatory mechanism relies on indirect evidence; direct measurements of gluconeogenic flux are lacking.

      We agree that isotopic tracer experiments would provide the most direct evidence for gluconeogenic flux. Such experiments are beyond the scope of the current revision, and we now explicitly acknowledge this as a key limitation and an important direction for future research (revised Discussion: Study limitations and future directions).

      However, we note that the convergent evidence from multiple independent lines collectively supports gluconeogenic activation: (i) transcriptional upregulation of PEPCK and G-6-Pase; (ii) declining protein levels now independently confirmed by our new enzyme activity data showing a 30.78% decrease in total protein concentration at 24 h post-RNAi (new Figure 10A); (iii) altered amino acid profiles; and (iv) maintained trehalose levels. The revised manuscript presents this evidence more cautiously, framing it as “consistent with gluconeogenic compensation” rather than establishing metabolic flux.

      Additionally, we now provide new data on GP enzyme activity (new Figure 10A, B; see response to Recommendation 4 below) and alternative glycogen catabolic enzyme expression (new Figure 10C, D; see response to Recommendation 2 below) that further strengthen the evidence chain.

      (2) Alternative glycogen degradation pathways are proposed but not experimentally examined.

      We have now directly addressed this concern. RT-qPCR analysis of glycogen branching enzyme (GBE) and α-amylase following PxGP knockdown reveals a striking and informative differential response (new Figure 10C, D):

      GBE was significantly upregulated at 24 h (+29.24%, P < 0.05), 48 h (+16.78%, P < 0.05), and 96 h (+44.46%, P < 0.001), indicating transcriptional activation of an alternative glycogen-remodeling enzyme in response to GP suppression.

      α-Amylase showed no significant change at any time point, demonstrating that the compensatory response is pathway-specific rather than a generalized upregulation of all glycogen-degrading enzymes.

      This differential pattern — GBE up, α-Amylase unchanged — provides the first evidence that P. xylostella selectively activates specific glycogen remodeling pathways when GP function is compromised. Upregulation of GBE, which increases glycogen branching and solubility, may facilitate glycogen mobilization through alternative routes even when GP-mediated phosphorolysis is impaired. These data are incorporated as new Figure 10C, D and discussed in the revised Results and Discussion.

      (3) Physiological consequences (fitness costs) are not explored.

      We have now conducted a comprehensive fitness cost assessment (new Figure 11). The results reveal a transient but significant fitness cost confined to the larval stage:

      Feeding rate: no significant difference between dsGP and dsGFP groups at any time point (24–120 h; Figure 11B), confirming that the observed metabolic changes are not attributable to reduced food intake.

      Larval weight: significantly reduced at 24 h (−29.10%, P < 0.05) and 48 h (−25.38%, P < 0.05; Figure 11C), demonstrating a measurable short-term cost of metabolic compensation.

      Pupal weight: no significant difference (Figure 11D), indicating full recovery before the pupal transition.

      Adult wing area: no significant difference (Figure 11E, Figures S5–S6), suggesting no impairment of flight capacity.

      Female fecundity (3-day egg production): no significant difference (Figure 11F), demonstrating no reduction in reproductive output.

      This pattern — transient larval weight loss with complete recovery of pupal weight, wing morphology, and reproductive performance — is consistent with our proposed model: GP suppression triggers protein catabolism to fuel gluconeogenesis (explaining the short-term weight loss), but the compensatory mechanism is sufficiently effective to restore metabolic homeostasis before pupation. These data strengthen the conclusion that GP is functionally non-essential for completing development and reproduction.

      (4) Broader conclusions regarding BPU class may require testing additional compounds.

      We agree. The revised manuscript now explicitly limits the biochemical conclusion to diflubenzuron: “DFB does not inhibit PxGP” rather than making broader claims about the BPU class as a whole. We discuss this as a limitation and note that testing additional BPU compounds would be needed before generalizing.

      (5) Some biochemical and cell-based observations would benefit from confirmation in whole insects.

      We have now provided whole-insect confirmation through: (i) GP enzyme activity measurements in RNAi-treated larvae (new Figure 10A, B); (ii) in vivo fitness assessment showing measurable physiological consequences of GP suppression (new Figure 11); and (iii) expression analysis of compensatory enzymes in intact larvae (new Figure 10C, D). These data bridge the gap between our cell-free biochemical observations and whole-organism biology.

      Reviewer #2 (Public review):

      (1) The central premise — that structural similarity among acylurea compounds implies shared targets — is not supported.

      We agree that the original manuscript overstated the significance of the shared acylurea core as a predictor of common biological activity. The Introduction has been substantially restructured to:

      Explicitly acknowledge the compelling genetic evidence from CRISPR/Cas9 experiments (Reference 5) establishing CHS as the primary site conferring BPU resistance.

      Reframe the study's objective: rather than proposing to “resolve” the BPU target controversy, the revised manuscript focuses on the systematic evaluation of GP as an independent insecticidal target and the discovery of a gluconeogenic compensation mechanism — questions with scientific value independent of the BPU mechanism debate.

      Remove the claim that the study “resolves the primary hypothesis.” The conclusion now states that our biochemical data demonstrate DFB does not inhibit PxGP, adding enzyme-level evidence to the existing genetic framework.

      (2) Target selectivity is determined by side-chain composition, not the shared acylurea core.

      We fully agree, and our new structural data now provide a molecular explanation for this principle at the atomic level. Molecular docking and MM/GBSA analysis (new Figure 12, new Table 1) reveal that both GPI and DFB anchor to PxGP through their common acylurea carbonyl groups (Arg193), but diverge dramatically in side-chain engagement:

      GPI's methoxyphenyl-methylurea moiety establishes extensive contacts with seven residues across both subunits (Asn44 and Val45 from chain A; Trp67, Gln71, Tyr75, Arg193, and Asp227 from chain B), binding at the allosteric site at the dimer interface — consistent with the experimentally determined binding mode of acylurea inhibitors in mammalian GP (PDB: 2ATI).

      DFB contacts six residues primarily from subunit B, and its difluorobenzoyl moiety remains entirely solvent-exposed without productive protein contacts.

      MM/GBSA analysis confirms GPI binds with substantially higher affinity (ΔG = −34.63 vs. −29.29 kcal/mol; ΔΔG = −5.34 kcal/mol), with van der Waals interactions as the dominant driver (Δ<sub>VDW</sub> = −11.49 kcal/mol), reflecting superior shape complementarity of GPI.

      These structural data directly support Reviewer 2's important point and are now presented as new Figure 12 and Table 1.

      (3) References 6–9 characterization.

      We have replaced the original citations (former References 6–9) in the Introduction with references that directly demonstrate the absence of CHS inhibition by BPUs in cell-free preparations: Cohen & Casida (1980) showed DFB did not inhibit Tribolium gut chitin synthetase; Mayer et al. (1981) and Cohen (1985) systematically confirmed that BPU-type insect growth regulators do not inhibit chitin synthase in cell-free assays; and Zhang & Zhu (2013) reported only slight in vitro CHS inhibition by DFB in Anopheles gambiae with no in vivo effect. We also cite authoritative reviews by Matsumura (2010) and Merzendorfer (2013) that contextualize this evidence gap. We thank Reviewer 2 for identifying this important citation issue, which has led to a substantially more accurate and well-supported presentation of the literature.

      (4) The term “dataology” is non-standard.

      This term has been removed and replaced with “data.” In accordance with eLife's policy on AI tools and technology, we have added a statement in the Materials and Methods section declaring that AI-based language editing tools were used for English grammar and style refinement. All scientific content was generated entirely by the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Direct assessment of gluconeogenic flux (e.g., metabolic tracer experiments).

      As discussed above, isotopic tracer experiments are beyond the current scope. We acknowledge this as a key limitation in the revised Discussion (see details in lines 560-565 of the revised manuscript). However, we now provide additional supporting evidence: the 30.78% decline in total protein at 24 h post-RNAi (from our new enzyme activity data, Figure S3) provides independent biochemical confirmation of protein catabolism, consistent with amino acid mobilization for gluconeogenesis (see details in lines 357-359 of the revised manuscript). We have also added discussion of potential future approaches, including measurements of key enzymes in amino acid catabolism (e.g., aspartate aminotransferase, glutamate aminotransferase) and lipid content dynamics (see details in lines 565-569 of the revised manuscript).

      (2) Expression or activity of alternative glycogen degradation enzymes (α-amylase, glycogen debranching enzymes).

      We have measured the expression of GBE (glycogen branching enzyme) and α-amylase by RT-qPCR in RNAi-treated insects. We also attempted to measure glycogen debranching enzyme (GDE), but multiple primer pairs failed to yield amplification products, likely due to sequence annotation issues; this is noted as a limitation.

      Results (new Figure 10C, D): GBE was significantly upregulated at 24 h (+29.24%), 48 h (+16.78%), and 96 h (+44.46%). α-Amylase was unchanged at all time points (see details in lines 377-381 of the revised manuscript). The selective upregulation of GBE but not α-amylase suggests a targeted compensatory response within the glycogen remodeling pathway.

      The absence of glycogen accumulation following GP knockdown may reflect reduced flux into glycogen synthesis (potentially through feedback inhibition of glycogen synthase) rather than activation of alternative degradative routes. This possibility is discussed in the revised manuscript, with glycogen synthase expression identified as a key target for future investigation (see details in lines 501-505 of the revised manuscript).

      (3) Fitness cost assessment (body size, flight capacity, reproductive performance).

      Complete data are now provided (new Figure 11A–F, Figures S5–S6):

      Author response table 1.

      The transient larval weight reduction (24–48 h) with complete pupal and adult recovery demonstrates that metabolic compensation carries a short-term physiological cost but is ultimately effective in maintaining developmental trajectory and reproductive fitness (see details in lines 388-408 of the revised manuscript).

      (4) Enzyme activity measurements in RNAi-treated insects.

      GP enzyme activity (GP-a) was measured in crude extracts from RNAi-treated larvae using a coupled-enzyme spectrophotometric assay kit (Solarbio BC3345) at 24, 48, 72, and 96 h post-injection (new Figure 10A, B).

      Two normalization approaches were used: - Per-protein activity: significant reduction only at 48 h (−10.35%, P < 0.05; Figure 10A). The modest per-protein reduction reflects a concurrent 30.78% decline in total protein at 24 h, which inflates per-protein specific activity when the protein pool shrinks. - Per-larva activity: significant reduction at 24 h (−27.57%, P < 0.05) and 48 h (−29.28%, P < 0.01; Figure 10B), confirming that RNAi-mediated transcript suppression translates to reduced enzyme function in vivo.

      The 30.78% decline in total protein at 24 h provides independent biochemical confirmation of protein catabolism — consistent with amino acid mobilization for gluconeogenesis (see details in lines 349-370 of the revised manuscript).

      (5) Scope of BPU conclusion — clarify whether additional compounds should be tested.

      The revised manuscript explicitly states that the biochemical conclusion applies to diflubenzuron specifically. We have added a Discussion paragraph noting that extending this conclusion to additional BPU compounds would require systematic testing, and that our structural analysis (Table 1) provides a framework for predicting which acyl urea side-chain architectures are compatible with GP binding (see details in lines 455-466 of the revised manuscript).

      (6) RNAi transcript recovery at 96 h and implications for non-essentiality.

      Our new enzyme activity data directly address this concern. GP activity (per-larva) showed partial recovery at 72 h and 96 h (Figure 10B), mirroring the transcript recovery pattern. However, the critical observation is that even during the period of maximum suppression (24–48 h), when per-larva GP activity was reduced by ~27–30%, larvae maintained glucose homeostasis and completed development. This confirms that GP is non-essential even during the period of strongest suppression. The revised Discussion addresses this point explicitly (see details in lines 479-488 of the revised manuscript).

      (7) GPI concentrations in larval exposure experiments and pharmacokinetic considerations.

      We have added a dedicated Discussion paragraph addressing this concern in detail. The GPI concentrations used (250–500 mg/L in diet) encompass a wide dose range; even at the highest concentration (500 mg/L, approximately 409,000-fold above the in vitro IC<sub>50</sub> of 2.96 nM), no toxicity was observed. We discuss several pharmacokinetic factors that may contribute to this apparent discrepancy, including limited oral bioavailability, metabolic inactivation by detoxification enzymes, and sequestration by hemolymph binding proteins. However, we note that the metabolic phenotype observed (elevated trehalose, reduced protein, upregulated gluconeogenic enzymes) provides indirect evidence that GPI does reach its target. The most parsimonious interpretation is that GPI achieves sufficient target engagement to partially suppress GP activity in vivo, but metabolic compensation renders this suppression non-lethal — an interpretation reinforced by our RNAi data, in which direct genetic suppression of GP (bypassing all pharmacokinetic barriers) similarly fails to cause mortality (see details in lines 536-552 of the revised manuscript).

      (8) Structural evidence that GPI binds PxGP comparably to its mammalian target.

      This has been comprehensively addressed through molecular docking and MM/GBSA analysis (new Figure 12, new Table 1). The PxGP homodimer structure was modeled using SWISS-MODEL with the human liver GP–acyl urea co-crystal structure (PDB: 2ATI) as the template. Docking and binding free energy calculations were performed in Cresset Flare V11.

      Key findings: GPI binds at the allosteric site at the dimer interface with ΔG = −34.63 kcal/mol, engaging seven residues across both subunits — a binding mode consistent with the experimentally determined site in mammalian GP. DFB binds with lower affinity (ΔG = −29.29 kcal/mol) and its difluorobenzoyl moiety is entirely solvent-exposed. Van der Waals interactions are the dominant driver of selectivity (Δ<sub>VDW</sub> = −11.49 kcal/mol). See new Figure 12 and Table 1 for complete data.

      (9) Dietary carbohydrate compensation and feeding behavior.

      Our new data directly address this concern. Feeding rate measurements show no significant difference between dsGP and dsGFP groups at any time point (24–120 h; Figure 11B), confirming that the metabolic changes are not attributable to altered food intake. A Discussion paragraph has been added acknowledging the potential contribution of dietary carbohydrates to glucose homeostasis and noting that starvation-challenge experiments would provide additional insight (see details in lines 570-581 of the revised manuscript).

      Minor comments and suggestions

      (1) Terminology.

      “Gluconeogenolysis” has been replaced with “gluconeogenesis” throughout the manuscript.

      (2) Typographical errors.

      A thorough language revision has been performed. “Over over four decades” and other errors have been corrected. The term “dataology” has been removed.

      (3) Metabolite normalization.

      We now present GP enzyme activity using two normalization approaches (per-protein and per-larva; Figure 10A, B) and discuss the implications of protein level changes on per-protein normalization in the Results section. For metabolite data, we have added a note in the Methods explaining our normalization approach and discussing how declining protein levels may influence interpretation (see details in lines 908-920 of the revised manuscript).

      (4) Clarity of pathway descriptions.

      The description of metabolic pathways has been simplified and a revised schematic figure has been included (Figure 13).

      (5) Figure clarity.

      Figures have been added with clearer labeling and simplified schematics. To improve clarity, we used red arrows to show the blocked metabolic flow when GP is inhibited, and green arrows to depict the activated gluconeogenic pathway (Figure 13).

      Cohen E, Casida JE. Inhibition of Tribolium gut chitin synthetase. Pestic Biochem Physiol. 1980;13(2):129-36. doi: 10.1016/0048-3575(80)90064-4.

      Mayer RT, Chen AC, DeLoach JR. Chitin synthesis inhibiting insect growth regulators do not inhibit chitin synthase. Experientia. 1981;37(4):337-8. doi: 10.1007/BF01959848.

      Cohen E. Chitin synthetase activity and inhibition in different insect microsomal preparations. Experientia. 1985;41(4):470-2. doi: 10.1007/BF01966152.

      Zhang X, Yan Zhu K. Biochemical characterization of chitin synthase activity and inhibition in the African malaria mosquito, Anopheles gambiae. Insect Sci. 2013;20(2):158-66. doi: 10.1111/j.1744-7917.2012.01568.x.

      Matsumura F. Studies on the action mechanism of benzoylurea insecticides to inhibit the process of chitin synthesis in insects: A review on the status of research activities in the past, the present and the future prospects. Pestic Biochem Physiol. 2010;97(2):133-9. doi: 10.1016/j.pestbp.2009.10.001.

      Merzendorfer H. Chitin synthesis inhibitors: old molecules and new developments. Insect Sci. 2013;20(2):121-38. doi: 10.1111/j.1744-7917.2012.01535.x

      Preiss J. Bacterial glycogen synthesis and its regulation. Annual review of microbiology. 1984;38:419-58. doi: 10.1146/annurev.mi.38.100184.002223.

      Janeček Š, Svensson B, MacGregor EA. α-Amylase: an enzyme specificity found in various families of glycoside hydrolases. Cell Mol Life Sci. 2014;71(7):1149-70. doi: 10.1007/s00018-013-1388-z.

      Waterhouse A, Bertoni M, Bienert S, Studer G, Tauriello G, Gumienny R, et al. SWISS-MODEL: homology modelling of protein structures and complexes. Nucleic Acids Res. 2018;46(W1):W296-W303. doi: 10.1093/nar/gky427.

      İnak E, De Rouck S, Van Leeuwen T. Molecular mechanisms of pesticide selectivity: Insights from acaricide toxicology. Pestic Biochem Physiol. 2025;213:106537. doi: 10.1016/j.pestbp.2025.106537.

      David MD. Insecticide ADME for support of early-phase discovery: combining classical and modern techniques. Pest Manage Sci. 2017;73(4):692-9. doi: 10.1002/ps.4345.

      Haunerland NH, Bowers WS. Binding of insecticides to lipophorin and arylphorin, two hemolymph proteins of Heliothis zea. Arch Insect Biochem Physiol. 1986;3(1):87-96. doi: 10.1002/arch.940030110.

      Other revisions

      Correction of primer sequences. Upon re‑checking the primer sequences during revision, we noticed that the originally reported primers for dsRNA synthesis of the GP gene (Table S1) were inadvertently copied incorrectly. The correct sequences have now been substituted in the revised manuscript (dsPxGP-F, dsPxGP-R). This correction does not affect any of the experimental data, results, or conclusions of the study. We apologize for the oversight.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have clarified several points that were raised by the reviewers. First, we now state more explicitly that the presence of intact env-containing Ty3/gypsy retrotransposons does not by itself demonstrate their mechanism of transmission, tissue specificity, infectivity, or current activity. We have therefore revised the wording throughout the manuscript to distinguish intact element structure and multicopy genomic expansion from experimentally demonstrated activity.

      Second, we performed targeted host-taxonomy concordance analyses on selected clades of the POL RT tree. These analyses do not exclude local horizontal transfer, particularly between closely related hosts, but they show that horizontal transfer alone is insufficient to explain the broader host-taxonomic structure observed across the dataset.

      Third, we incorporated representative viral and retroelement-associated fusogen proteins into our F-type ENV phylogenetic analysis and HSV/gB-type ENV structural comparison. These additions place the ENV proteins associated with Ty3/gypsy elements in a broader evolutionary context and strengthen the conclusion that these ENV associations are deeply diverged rather than recent derivatives of a single sampled viral lineage.

      Fourth, we added two each of entirely new Supplementary figures (S5 and S7) and Tables (S2 and S3) and substantially modified now Supplementary figure S8. Other figures have also been modified only to increase readability. The four tables from the original manuscript have not been modified although their numbering has changed.

      We believe that the revised manuscript is substantially improved in clarity, terminology and interpretive precision, while retaining the central conclusion that the association between env-like genes and Ty3/gypsy retrotransposons is ancient in metazoan evolution. Sincerely,

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive systematic analysis of envelope-containing Ty3/gypsy retrotransposons (errantiviruses) across metazoan genomes, including both invertebrates and ancient animal lineages. Using iterative tBLASTn mining of over 1,900 genomes, the authors catalog 1,512 intact retrotransposons with uninterrupted gag, pol, and env open reading frames. They show that these elements are widespread present in most metazoan phyla, including cnidarians, ctenophores, and tunicates-with active proliferation indicated by their multicopy status. Phylogenetic analyses distinguish "ancient" and "insect" errantivirus clades, while structural characterization (including AlphaFold2 modeling) reveals two major env types: paramyxovirus F-like and herpesvirus gB-like proteins. Although bot envelope types were identified in previous analyses two decades ago, the evolutionary provenance of these envelope genes was almost rudimentary and anecdotal (I can say this because I authored one of these studies). The results in the present study support an ancient origin for env acquisition in metazoan Ty3/gypsy elements, with subsequent vertical inheritance and limited recombination between env and pol domains. The paper also proposes an expanded definition of 'errantivirus' for env-carrying Ty3/gypsy elements outside Drosophila.

      Strengths:

      (1) Comprehensive Genomic Survey:

      The breadth of the genome search across non-model metazoan phyla yields an impressive dataset covering evolutionary breadth, with clear documentation of search iterations and validation criteria for intact elements.

      (2) Robust Phylogenetic Inference:

      The use of maximum likelihood trees on both pol and env domains, with thorough congruence analysis, convincingly separates ancient from lineage-specific elements and demonstrates co-evolution of env and pol within clades.

      (3) Structural Insights:

      AlphaFold2-based predictions provide high-confidence structural evidence that both env types have retained fusion-competent architectures, supporting the hypothesis of preserved functional potential.

      (4) Novelty and Scope:

      The study challenges previous assumptions of insect-centric or recent env acquisition and makes a compelling case for a Pre-Cambrian origin, significantly advancing our understanding of animal retroelement diversity and evolution. THIS IS A MAJOR ADVANCE.

      (5) Data Transparency:

      I appreciate that all data, code, and predicted structures are made openly available, facilitating reproducibility and future comparative analyses.

      Major Weaknesses

      (1) Functional Evidence Gaps:

      The work rests largely on sequence and structure prediction. No direct expression or experimental validation of envelope gene function or infectivity outside Drosophila is attempted, which would be valuable to corroborate the inferred roles of these glycoproteins in non-insect lineages. At least for some of these species, there are RNA-seq datasets that could be leveraged.

      We added a sentence in the discussion, subsection “The survival mechanism of errantiviruses in the genome”, citing our recent work now published (PMID: 41922845), explaining that the defence mechanism against errantiviruses appears to be conserved in insects beyond Drosophila, indirectly suggesting that their biology dependent on the presence of env—may be more universal.

      (2) Horizontal Transfer vs. Loss Hypotheses:

      The discussion argues primarily for vertical inheritance, but the somewhat sporadic phylogenetic distributions and long-branch effects suggest that loss and possibly rare horizontal events may contribute more than acknowledged. Explicit quantitative tests for horizontal transfer, or reconciliation analyses, would strengthen this conclusion. It's also worth pointing out that, unlike retrotransposons that can be found in genomes, any potential related viral envelopes must, by definition, have a spottier distribution due to sampling. I don't think this challenges any of the conclusions, but it must be acknowledged as something that could affect the strength of this conclusion

      We have added a targeted host-taxonomy concordance analysis for two well-sampled POL extended RT/connection subclades: an Annelida-associated clade from tree position A6 and a Lepidoptera-associated clade from tree position I1 (new Fig S5). Rather than attempting to infer exact numbers of duplication, loss and horizontal transfer events, which is difficult across highly expanded and unevenly sampled transposon families, we tested whether host-taxonomic labels were more clustered on the observed POL extended RT/connection topology than expected by chance. In the Annelida clade, highly supported small subclades showed strong host-family and host species concordance under host-label permutation tests. The Lepidoptera clade showed a more mixed pattern, but still contained several highly supported subclades enriched for related host groups at the superfamily or broader taxonomic level. These results do not exclude rare horizontal transfer, particularly between closely related hosts, but support the conclusion that the observed POL extended RT/connection trees retain significant host-taxonomic structure and are not consistent with frequent broad horizontal transfer between distantly related animal groups. We have added a paragraph in the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” describing these observations, and also revised the Discussion to more explicitly acknowledge the possibilities of lineage-specific loss and the horizontal transfer.

      (3) Limited Taxon Sampling for Certain Phyla:

      Despite the impressive breadth, some ancient lineages (e.g., Porifera, Echinodermata) are negative, but the manuscript does not fully explore whether this reflects real biological absence, assembly quality, or insufficient sampling. A more systematic treatment of negative findings would clarify claims of ubiquity. However, I also believe this falls beyond the scope of this study.

      In the revised manuscript, we have added a targeted analysis of two representative genomes from each phylum. Although we did not detect full-length GAG-POL-ENV elements in these genomes, we recovered multiple full-length, multicopy GAG-POL Ty3/gypsy elements from all four genomes, many of which were flanked by predicted LTR sequences and associated with putative tRNA primer-binding sites. This suggests that the apparent absence of env-carrying elements in these representative Porifera and Echinodermata genomes is unlikely to be due simply to poor assembly quality or a general inability to recover intact Ty3/gypsy-like retrotransposons. We have added these data as the new Supplementary table S3 and revised the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” to clarify that absence in these phyla may reflect true biological absence, lineage-specific loss, or incomplete taxon sampling.

      (4) Mechanistic Ambiguity:

      The proposed model that env-containing elements exploit ovarian somatic niches is plausible but extrapolated from Drosophila data; for most taxa, actual tissue specificity, lifecycle, or host interaction mechanisms remain speculative and, to me, a bit unreasonable.

      We stressed in the Discussion section “The survival mechanism of errantiviruses in the genome” that the mere presence of env gene does not imply the mechanism of transmission of retrotransposons.

      Minor Weaknesses:

      (1) Terminology and Nomenclature:

      The paper introduces and then generalizes the term "errantivirus" to non-insect elements. While this is logical, it may confuse readers familiar with the established, Drosophila-centric definition if not more explicitly clarified throughout. I also worry about changes being made without any input from the ICTV nomenclature committee, which just went through a thorough reclassification. Nevertheless, change is expected, and calling them all errantiviruses is entirely reasonable.

      We have revised the Results section and discussion where we introduced the term "errantivirus" to clarify that we use "errantivirus" operationally to refer to env-containing Ty3/gypsy retrotransposons identified in this study, rather than as a formal taxonomic proposal. We also now state explicitly that bona fide infectivity and amplification through the Drosophila-like ovarian somatic-cell route have not been experimentally established for most non-Drosophila elements. Our use of the term is therefore intended to distinguish env-containing Ty3/gypsy elements from related non-env containing Ty3/gypsy retrotransposons, while acknowledging that their biology outside Drosophila remains to be determined.

      (2) Figures and Supplementary Data Navigation:

      Some key phylogenies and domain alignments are found only in supplementary figures, occasionally hindering readability for non-expert audiences. Selected main-text inclusion of representative trees would benefit accessibility.

      We agree that clearer navigation between the main text and supplementary figures would improve readability. Although we considered moving selected supplementary phylogenies and alignments into the main figures, the main figures are already data-dense and are intended to provide representative summaries across many host groups and ENV types. We therefore retained the detailed trees and alignments as supplementary figures, where they can be shown at readable scale, but revised the manuscript to improve navigation. Specifically, we added signposting sentences in the Results where supplementary figures are mentioned, expanded the relevant figure legends, and clarified how each supplementary tree or alignment supports the corresponding main-text conclusion.

      (3) ORF Integrity Thresholds:

      The cutoff choices for defining "intact" elements (e.g., numbers/placement of stop codons, length ranges) are reasonable but only lightly justified. More rationale or sensitivity analysis would improve confidence in the inclusion criteria. For example, how did changing these criteria change the number of intact elements?

      We agree with the reviewer that the rationale for the ORF integrity thresholds should be stated more clearly. We have revised the Methods section "Identification of intact genomic copies of Ty3/gypsy errantiviruses" to clarify that the initial length, gap and stop-codon thresholds were deliberately permissive screening criteria, designed to avoid excluding divergent or non-canonical elements at the discovery stage. These initial filters were not used alone to define the final “intact” set. Candidate elements were subsequently subjected to multiple additional curation steps, including confirmation of Ty3/gypsy POL identity, recovery of full-length RT and Integrase domains within continuous ORFs, HHpred-based domain annotation of GAG, POL and ENV, and removal of elements with large domain truncations. Thus, the final set of intact elements is substantially more refined than would be implied by the initial stop codon or length thresholds alone.

      A full sensitivity analysis varying each threshold across the entire iterative discovery and manual-curation pipeline would be difficult to interpret, because changing early permissive filters would alter the candidate pool that then undergoes downstream structural and phylogenetic validation. Instead, we have clarified in the Methods that the early thresholds were intended as inclusive prefilters, whereas final inclusion required intact domain architecture and phylogenetic/domain support.

      (4) Minor Typos/Formatting:

      The paper contains sporadic typographical errors and formatting glitches (e.g., misaligned figure labels, unrendered symbols) that should be addressed.

      We now fixed these issues in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The authors first surveyed metazoan genomes to identify homologs of Drosophila errantiviruses and classified them into two groups, "insect" and "ancient" elements, supporting the hypothesis of an early evolutionary origin for these retrotransposons. They subsequently identified two distinct types of envelope proteins, one resembling the glycoprotein F of paramyxoviruses and the other akin to the glycoprotein B of herpesviruses. Despite differences in their primary amino acid sequences, these proteins display notable structural similarity in their predicted domain architectures. The congruence between the phylogenies of the envelope and pol genes further supports the ancient origin of the envelope genes, challenging earlier hypotheses that proposed recent recombination events with baculoviruses. Additional analysis of the Pol "bridge region" corroborated the divergence among these elements, consistent with a pattern of limited cross-species recombination. Finally, by comparing these elements with non-envelope-containing Gypsy retrotransposons, the authors concluded that errantiviruses originated from multiple elements independently.

      Strengths:

      The conclusions of this study are based on a comprehensive collection of errantiviruses identified across a wide range of metazoan genomes. These findings are further supported by multiple lines of evidence, including phylogenetic congruence and the diverse evolutionary origins of envelope genes. AlphaFold2-assisted protein domain structure analyses also provided key insights into the characterization of these elements. Together, these results present a compelling case that errantiviruses arose independently through multiple evolutionary events, extending well beyond previous hypotheses.

      Weaknesses:

      It would be beneficial to emphasize in the Abstract the potential impact of this work by more clearly articulating the current knowledge gap in the field. While the second paragraph of the Introduction briefly touches on this point, highlighting the broader significance in the Abstract would better capture readers' interest. Additionally, some methodological choices would benefit from clearer justification and explanation. For instance, in Figure 6, the selection of the bridge region/RNase H domain is not explicitly explained, leaving the rationale for its choice unclear. As a minor point, some figure labels and texts are too small and difficult to read, and improving their legibility would enhance overall clarity.

      We have revised the Abstract to more clearly state the knowledge gap addressed by this study: although env-containing Ty3/gypsy elements were known from Drosophila and sporadically reported in other animals, whether their association with env-like fusogen genes reflected recent, lineage-specific acquisitions or a much deeper evolutionary relationship remained unclear. We now highlight this broader significance in the Abstract and frame our results as evidence that env-containing Ty3/gypsy elements represent deeply diverged, genome-resident retroelements rather than a recent insect-specific phenomenon.

      We have also revised the Results, Methods and Figure 6 legend to explain why the

      RNase H-containing bridge region was analysed. Specifically, we now distinguish the Pol extended RT/connection region used for phylogenetic analysis from the RNase H-containing bridge region analysed structurally in Figure 6. We define the bridge region as the canonical RNase H domain together with the C-terminal region between RNase H and Integrase, and explain that this region was selected because RNase H-related and adjacent RNase H-like domains vary among LTR retroelement lineages. The bridge region architecture therefore provides an independent structural feature for comparing the “insect errantivirus” and “ancient errantivirus” groups.

      Finally, we have revised the figures and figure legends to improve readability. In particular, we enlarged labels where possible, clarified figure annotations, corrected cross-references between main and supplementary figures, and added signposting sentences in the Results so that readers can more easily connect the main conclusions to the supporting supplementary trees and alignments.

      Reviewer #3 (Public review):

      Summary and Significance:

      In this work, Cary and Hayashi address the important question of when, in evolution, certain mobile genetic elements (Ty3/gypsy-like non-LTR retrotransposons) associated with certain membrane fusion proteins (viral glycoprotein F or B-like proteins), which could allow these mobile genetic elements to be transferred between individual cells of a given host. It is debated in the literature whether the acquisition of membrane fusion proteins by non-LTR retrotransposons is a rather recent phenomenon that separately occurred in the ancestors of certain host species or whether the association with membrane fusion proteins is a much more ancient one, pre-dating the Cambrian explosion. Obviously, this question also touches upon the origin of the retroviruses, which can spread between individuals of a given host but seem restricted to vertebrates. Based on convincing data, Cary and Hayashi argue that an ancient association of non-LTR retrotransposons with membrane fusion proteins is most probable.

      Strengths:

      The authors take the smart approach to systematically retrieve apparently complete, intact, and recently functional Ty3/gypsy-like non-LTR retrotransposons that, next to their characteristic gag and pol genes, additionally carry sequences that are homologous to viral glycoprotein F (env-F) or viral glycoprotein B (env-B). They then construct and compare phylogenetic trees of the host species and individual encoded proteins and protein domains, where 3D-structure calculations and other features explain and corroborate the clustering within the phylogenetic trees. Congruence of phylogenetic trees and correlation of structural features is then taken as evidence for an infrequent recombination and a long-term co-evolution of the reverse transcriptase (encoded by the pol gene) and its respective putative membrane fusion gene (encoded by env-F or env-B). Importantly, the env-F and env-B containing retrotransposons do not form a monophyletic group among the Ty3/gypsy-like non-LTR retrotransposons, but are scattered throughout, supporting the idea of an originally ancient association followed by a random loss of env-F/env-B in individual branches of the tree (and rather rare re-associations via more recent recombinations).

      Overall, this is valuable, stimulating, and important work of general and fundamental interest, but still also somewhat incompletely explored, imprecisely explained, and insufficiently put into context for a more general audience.

      Weaknesses:

      Some points that might be considered and clarified:

      (1) Imprecise explanations, terms, and definitions:

      It might help to add a 'definitions box' or similar to precisely explain how the authors decided to use certain terms in this manuscript, and then use these terms consistently and with precision.

      (a) In particular, these are terms such as 'vertebrate retrovirus' vs 'retrovirus' vs 'endogenized retrovirus' vs 'endogenous retrovirus' vs 'non-LTR retrotransposon' and 'Ty3/gypsi-like retrotransposon' vs 'Ty3/gypsy retrotransposon' vs 'errantivirus'.

      We agree with the reviewer. We inserted a paragraph at the end of the first Results section, explaining how we define endogenous retroviruses (ERVs), Ty3/gypsy retrotransposons and errantiviruses.

      (b) The comment also applies to the term 'env' used for both 'env-F' and 'env-B', where often it remains unclear which of the two protein types the authors refer to. This is confusing, particularly in the methods, where the search for the respective homologs is described.

      We revised the manuscript and now used F-type env/ENV and HSV/gB-type env/ENV throughout the text. We also modified the method section where we explained the tBlastn search to clarify which ENV proteins were used initially for the search and how we classified them in later analyses.

      (c) Other examples are the use of the entire pol gene vs. pol-RT for the definition of the Ty3/gypsy clade and for the generation of phylogenetic trees (Methods and Figure S1), and the names for various portions of pol that appear without prior definition or explanation (e.g., 'pro' in Figure 1A, 'bridge' in Figure S1C, 'the chromodomain' in the text and Figure 7).

      We revised the manuscript and explained ‘pro’, ‘bridge’ and ‘the chromodomain’ in the Results section or figure legends when they first appear. Please refer to other sections of the response for pol-RT definition.

      (d) It is unclear from the main text which portions of pol were chosen to define pol-RT and why. The methods name the 'palm-and-fingers', 'thumb', and 'connections' domains to define RT. In the main text, the 'connection' domain is called 'tether' and is instead defined as part of the 'bridge' region following RT, which is not part of RT.

      We agree that our previous terminology around Pol domains was imprecise and could confuse readers. We have revised the manuscript to distinguish the region used for phylogenetic analysis from the region analysed structurally in Figure 6. The phylogenetic analysis used an extended RT/connection region, comprising the RT polymerase core together with the downstream connection subdomain. This connection subdomain is treated as part of retroviral RT in structural studies, but corresponds to a partial RNase H-like fold and has been interpreted evolutionarily as a degenerated RNase H-like tether domain. It is therefore broader than the RT polymerase core alone, but it is not the complete canonical RNase H domain.

      We now define the Figure 6 “bridge region” separately as the region spanning the canonical RNase H domain and the C-terminal region between RNase H and Integrase. Figure 6 shows that the invertebrate errantiviruses analysed retain an intact canonical RNase H domain immediately downstream of the extended RT/connection region, but differ in the additional downstream RNase H-like or mini-domain structures before Integrase. We have revised the Results, Methods and figure legends accordingly. We also acknowledge that a phylogeny based strictly on the RT polymerase core alone could differ in some local branch relationships, but the major conclusions are supported independently by the Integrase tree, host-taxonomic structure, ENV-type distribution, Pol bridge-region architecture and ENV structural features.

      (2) Insufficient broader context:

      (a) The introduction does not state what defines Ty3/gypsy non-LTR retrotransposons as compared to their closest relatives (Ty1/copia retrotransposons, BEL/pao retrotransposons, vertebrate retroviruses). This makes it difficult to judge the significance and generality of the findings.

      (b) The various known compositions of Ty3/gypsi-like retrotransposons are not mentioned and explained in the introduction (open reading frames, (poly-)proteins and protein domains, and their variable arrangement, enzymatic activities, and putative functions), and the distribution of Ty3/gypsi-like retrotransposons among eukaryotes remains unclear. The introduction does not mention that Ty3/gypsi-like retrotransposons apparently are absent from vertebrates, and Figure 7 is not very clear about whether or not it includes sequences from plants ('Chromoviridae').

      We agree that the Introduction needed more context on Ty3/gypsy retrotransposons. We have revised it to briefly state that LTR retrotransposons include several major lineages, including Ty1/copia, BEL/Pao, Ty3/gypsy and retrovirus-related elements, and that Ty3/gypsy elements are classified primarily by POL similarity and domain organisation. We also now explain that Ty3/gypsy retrotransposons typically encode GAG and POL proteins, with POL providing the enzymatic activities required for reverse transcription and integration, while noting that ORF arrangement and accessory domains can vary between lineages.

      Please note that we stated that our screen did not identify intact env-containing Ty3/gypsy elements in vertebrate genomes that were homologous to the invertebrate errantiviruses analysed here. This is not to say that non-env-containing Ty3/gypsy elements are also absent in vertebrate genomes. Finally, we revised the Figure 7 legend to make clear that the comparison includes representative non-env-containing Ty3/gypsy elements from animals, fungi and plants, including chromovirus or chromovirus-related elements.

      (c) The known association of Ty3/gypsi-like retrotransposons from different metazoan phyla with putative membrane fusion proteins (env-like) genes is mentioned in the introduction, but literature information, whether such associations also occur in the context of other retrotransposons (e.g., Ty1/ copia or BEL/pao), is not provided. The abstract is somewhat misleading in this respect. Finally, the different known types of env-like genes are not mentioned and explained as part of the introduction ('env-f', 'envB', 'retroviral env', others?)

      We expanded the introduction to introduce literature information of known env-associated retroelements, including Ty1/copia and BEL/pao and explained which ENV types are known to be associated to these elements.

      (d) Some key references and reviews might be added:

      - Pelisson, A. et al. (1994) https://www.embopress.org/doi/abs/10.1002/j.1460-2075.1994.tb06760.x (next to Song et al. (1994), for the identification of env in Ty3/gypsy)

      - Boeke, J.D. et al. (1999) In Virus Taxonomy: ICTV VIIth report. (ed. F.A. Murphy),. Springer-Verlag, New York. (cited by Malik et al. (2000) - for the definition and first use of the term 'errantivirus')

      - Eickbush, T.H. and Jamburuthugoda, V.K. (2008) https://doi.org/10.1016/j.virusres.2007.12.010 (on the classification of retrotransposons and their env-like genes)

      - Hayward, A. (2017) https://doi.org/10.1016/j.coviro.2017.06.006 (on scenarios of env acquisition)

      Thank you. We included these references in the introduction.

      (3) Incomplete analysis:

      (a) Mobile genetic elements are sometimes difficult to assemble correctly from shortread sequencing data. Did the authors confirm some of their newly identified elements by e.g., PCR analysis or re-identification in long-read sequencing data?

      Most newly identified elements are found in contigs/chromosomes that are longer than 100kb. The information of the contig/chromosome size, in which the representative copy of the identified elements are found, can be found in the column “CONTIG_SIZE” in supplementary table S1.

      (b) The authors mention somewhat on the side that there are Ty3/gypsy elements with a different arrangement (gag-env-pol instead of gag-pol-env). Why was this important feature apparently not used and correlated in the analysis? How does it map on the RT phylogenetic tree? Which type of env is found with either arrangement? Is there evidence for a loss of env also in the case of gag-env-pol elements?

      We agree that the non-canonical GAG-ENV-POL arrangement is an important feature that was insufficiently integrated into the analysis. We have revised the Results and figure annotations to make this clearer. Specifically, we now indicate GAG-ENV-POL elements in the POL tree in Fig S4 and in the HSV/gB-type ENV alignment/architecture figure S8. These elements are found in Nematoda, Bryozoa and Platyhelminthes and all carry HSV/gB-type ENV. They do not form a single monophyletic group in the Pol tree, but instead occur in distinct host-associated clades. They also show different HSV/gBtype cysteine-bridge architectures. Thus, the GAG-ENV-POL arrangement is unlikely to represent a single recent rearrangement event shared by all such elements; rather, it appears to be associated with several deeply diverged HSV/gB-type errantivirus lineages.

      We have not inferred specific env-loss events for GAG-ENV-POL elements, because doing so would require a separate analysis of related non-env-containing elements.

      (c) Sankey plots are insufficiently explained. How would inconsistencies between trees (recombinations) show up here? Why is there no Sankey plot for the analysis of env-B in Figure 5?

      We agree that the Sankey plot was insufficiently explained. We have revised the Figure 4 legend to clarify that the Sankey plot was used as a qualitative visual summary of global congruence between the Pol extended RT/connection phylogeny and the F-type ENV ectodomain phylogeny. We now state that ribbon crossing alone should not be interpreted as recombination, because tree drawings can be rotated without changing topology and the relative order of clades in the two displayed trees may differ. Instead, the relevant signal is whether Pol-defined clades map mostly to corresponding F-type ENV-defined clades. Strong discordance, potentially reflecting recombination, env exchange or poor phylogenetic resolution, would be expected to appear as extensive splitting or many-to-many connections between Pol and ENV clades.

      We did not include an equivalent Sankey plot for HSV/gB-type ENV in Figure 5 because we did not construct a global HSV/gB-type ENV phylogeny comparable to the F-type ENV ectodomain tree. Instead, HSV/gB-type ENV proteins were analysed by predicted structural organisation and cysteine-bridge architecture, which are shown in Figure 5 and Supplementary Figure S8.

      (d) Why are there no trees generated for env-F and env-B like proteins, including closely related homologous sequences that do NOT come from Ty3/gypsy retrotransposons (e.g., from the eukaryotic hosts, from other types of retrotransposons (Ty1/copia or BEL/pao), from viruses such as Herpesvirus and Baculovirus)? It would be informative whether the sequences from Ty3/gypsy cluster together in this case.

      We agree that comparison with homologous fusogens outside Ty3/gypsy retrotransposons is informative. We have therefore added an expanded F-type ENV ectodomain phylogeny that includes representative viral and retroelement-associated F-like proteins, including baculovirus F proteins, paramyxovirus and pneumovirus F proteins, and the BEL/Pao-associated Drosophila Roo F-like protein. In this expanded tree, the added viral sequences formed family-level clades within the broader F-type ENV diversity. Errantivirus F-type ENV proteins did not cluster as a shallow Ty3/gypsyspecific group or as a recent derivative of a single sampled viral family; instead, they spanned a level of diversity comparable to that separating major viral F-protein groups.

      For HSV/gB-type ENV proteins, we did not generate an equivalent global phylogeny because the primary sequences and domain organisations of viral class III fusogens and errantivirus HSV/gB-type ENV proteins were too divergent for reliable full ecto domain multiple-sequence alignment. Instead, we added a structural comparison with representative viral and retroelement-associated class III fusogens, including herpesvirus gB, rhabdovirus G, orthomyxovirus GP75/GP64-like proteins, baculovirus GP64 and BEL/Pao-associated gB-like proteins. This analysis showed that viral class III fusogens often retained family-specific cysteine-bridge architectures despite low primary-sequence identity. Errantivirus HSV/gB-type ENV groups showed a comparable pattern, retaining lineage-specific cysteine-bridge architectures despite extensive sequence divergence. We have revised the Results, Methods and supplementary figure legends to clarify these analyses and to distinguish the phylogenetic analysis of F-type ENV from the structural comparison of HSV/gB-type ENV.

      (e) Did the authors identify any other env-like ORFs (apart from env-F and env-B) among Ty3/gypsy retrotransposons? Did they identify other, non-env-like ORFs that might help in the analysis? It is not quite clear from the methods if the searches for env-F and envB - containing Ty3/gypsy elements were done separately and consecutively or somehow combined (the authors generally use 'env', and it is not clear which type of protein this refers to).

      We agree that this was not sufficiently clear. We have revised the Methods to clarify that the search was designed to identify Ty3/gypsy elements carrying ORFs structurally resembling known envelope/fusogen proteins. In the iterative tBLASTn searches, bait sequences representing both F-type ENV and HSV/gB-type ENV were included together in each round, rather than being searched as two entirely separate pipelines. Candidate elements were then annotated and classified by ORF structure, HHpred/domain similarity and structural prediction.

      Among intact Ty3/gypsy candidates recovered by this strategy, we identified two recurrent classes of env-like ORFs: F-type env and HSV/gB-type env. We did not identify an additional recurrent class of env-like ORF among the intact Ty3/gypsy elements analysed here. We also did not identify other recurrent non-env accessory ORFs that were informative for the phylogenetic analyses beyond the GAG, POL and ENV features described in the manuscript.

      (f) Why was the gag protein apparently not used to support the analysis? Are there different, unrelated types of gag among non-LTR retrotransposons? Does gag follow or break the pattern of co-evolution between RT and env-F/env-B?

      We agree that the role of GAG in the analysis should be clarified. GAG ORFs were used during element annotation to identify intact GAG-POL-ENV or GAG-ENV-POL retrotransposon architectures, but we did not use GAG as a major phylogenetic marker because GAG proteins are less conserved and less reliably alignable across deeply diverged Ty3/gypsy elements than the enzymatic POL domains. Our central question was the acquisition and long-term retention of env-like ORFs by POL-defined Ty3/gypsy retrotransposons. We have revised the Methods to clarify that GAG was used for structural annotation and intactness assessment, whereas phylogenetic analyses were based on the Pol extended RT/connection region and Integrase domain.

      (g) Data availability. The link given in the paper does not seem to work (https://github.com/RippeiHayashi/errantiviruses_2025/tree/main). It would be useful for the community to have the sequences of the newly identified Ty3/gypsy retrotransposons listed readily available (not just genome coordinates as in table S1), together with the respective annotations of ORFs and features.

      The GitHub repository that contains suggested data is made public. Please check the link again.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional Analyses That Could Strengthen Claims (but I concede might be well beyond the scope of this study):

      (1) Reconciliation and Gene Tree-Species Tree Analysis: Implementing explicit genetree-species-tree reconciliation (e.g., using Notung or ALE) could formally test the frequency of horizontal transfers vs. vertical inheritance in errantivirus evolution.

      We performed targeted host-taxonomy concordance analyses in representative well-sampled clades to address this question, as described in the revised manuscript.

      (2) Functional Validation in Non-Insect Hosts: RNA-seq or proteomics from representative non-insect hosts could reveal whether env genes are expressed, and-if possible-experimental assays for envelope function would move beyond computational inference. This is the one thing that might be done in a revision.

      We agree that direct functional validation in non-Drosophila species would substantially strengthen the conclusions. Such experiments are beyond the scope of the present revision. However, we now cite our recent work (PMID: 41922845) showing conserved piRNA-mediated defence against errantiviruses across insect orders, providing indirect evidence that these elements remain biologically active outside Drosophila.

      (3) LTR Age Dating: Estimating insertion ages using LTR divergence across clades would contextualize the timing of expansion events and help test hypotheses about ancient vs. recent proliferation.

      We agree that LTR divergence-based insertion dating would be informative for estimating the timing of recent expansion events. However, our main evolutionary conclusions concern the deeper history of env acquisition and long-term retention across metazoan lineages, rather than the precise insertion age of individual genomic copies. Because the dataset includes elements from highly divergent genomes with variable assembly quality and many multicopy families, systematic LTR dating across all clades would require additional curation and is beyond the scope of the present revision.

      (4) Comparative Host Defense Analysis: Surveying host antiviral or transposon defense systems (e.g., piRNA, APOBEC) in lineages rich in errantiviruses could test for signatures of recurrent molecular arms races.

      We agree that comparative analysis of host defence pathways, including piRNA and antiviral systems, is an important future direction. However, a systematic survey of host-defence evolution across all errantivirus-rich lineages is beyond the scope of the present revision.

      Reviewer #2 (Recommendations for the authors):

      Apart from my comments in the Public Review, I have a few additional minor recommendations:

      (1) The authors frequently use the term "active" to describe complete retrotransposons. However, in transposon biology, "active" implies recent or ongoing transpositional activity and should therefore be used with caution. Terms such as "complete" or "fulllength" would be more appropriate in this context.

      We agree that intact ORF structure should not be over-interpreted as evidence of recent mobilisation. We have therefore revised the manuscript to distinguish element intactness from evidence of recent expansion. Specifically, we now use “intact” or “full-length” to describe element structure. We also clarify that the presence of multiple highly similar copies in the same genome, defined as >98% nucleotide identity across >98% of the three ORFs, is evidence consistent with recent or ongoing genomic expansion, but not definitive proof of current transposition. These changes have been made in the Abstract, Results, Methods and Discussion.

      (2) In the third paragraph of the Results, the authors conclude that many identified errantiviruses were mobilized recently. However, since the search strategy specifically targets complete and uninterrupted elements, this may introduce a bias toward younger elements, making the conclusion somewhat circular.

      We agree and have revised the wording to avoid over-interpreting completeness as evidence of recent mobilisation. We now distinguish between intact/full-length element structure, which was part of our search strategy, and independent evidence for recent or ongoing mobilisation, such as the presence of multiple highly similar copies. We have therefore softened statements that previously implied that intact ORFs alone demonstrate recent activity.

      (3) The first paragraph of the Introduction requires appropriate references.

      We have now added references to enhance the readability of the first paragraph of the introduction.

      (4) It would benefit readers if the retrotranspositional process were briefly explained, including a description of what the tRNA primer binding site (PBS) is and its role in retrotransposition.

      In the same part of the introduction, we now describe the role of the tRNA primer binding site in retrotransposition.

      Reviewer #3 (Recommendations for the authors):

      Suggestions and requests for clarification currently are part of the public review, as they also point the reader to critical open questions if the authors decide not to amend their version of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this important study, the authors have performed a zebrafish drug screen to identify suppressors of atherogenic lipoproteins. They utilize a well-established LipoGlo assay to find molecules that modulate these lipoproteins, identifying 49 potential hits. They perform some validation experiments, including studies linking enoxolone to its likely inhibitory effect on a specific transcription factor, HNF4alpha. Overall, the results are convincing and robust, and will open up new areas of exploration for those investigators interested in in vivo lipid biology.

      We appreciate the manuscript assessment and provide new data (see below) that significantly increases the “strength of the evidence”.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The authors performed a whole-organism chemical screen with over 3000 agents. Such screens are challenging, and the authors used strict criteria for determining hits. The conclusions of this study are well supported by the presented data.

      Weaknesses:

      There are areas within the study and writing that can be improved and extended, specifically within the gene expression studies.

      We appreciate the Reviewer recognizing the strength of the data and the challenging nature of performing a whole animal small molecule screen. With regard the “Weaknesses”, as the Reviewer suggested we improved and clarified the text and “extended” the study by adding analysis and discussion of six additional hits (see Figure 3 and Sup. Fig.1) and modified the results section with the addition of over a page of text describing the additional phenotyping results:

      “Secondary characterization of selected validated B-lp lowering hits.

      To further evaluate the biological relevance and potential mechanisms of selected validated hits, we performed secondary analyses assessing total B-lp levels, larval morphology, and lipoprotein size distribution. Our lab previously defined the method to measure total B-lp levels from whole animals by using the homogenate of a single zebrafish larva [35]. We confirmed that treatment of animals with 4 µM pomiferin significantly reduced total B-lp levels measured from homogenates collected from whole animals after treatment (p = 8.6x10<sup>-4</sup>; Figure 3A). However, pomiferin treatment (4 µM) produced animals with reduced body length and lethality at higher doses, suggesting developmental toxicity may confound interpretation of its B-lp-lowering effect.

      Treatment of animals with riboflavin tetrabutyrate (Figure 3B) and calcipotriene (Figure 3C) reduced total B-lp levels (p < 2x10<sup>-16</sup>) but did not affect larval morphology. A key feature of B-lps is their size, often a proxy for the total amount of lipid in the particle [35]. Particle size can impact the particle's lifetime (e.g. in metabolically healthy humans, small particles are cleared rapidly by the liver) [54–56]. Thus, we also assessed whether these compounds alter B-lp size distribution. Animals were treated for 48 h with vehicle, 5 µM lomitapide, or a drug of interest, and whole-animal homogenates were prepared and subjected to native polyacrylamide gel electrophoresis followed by chemiluminescent imaging. B-lps were classified into four classes based on gel migration: zero mobility (ZM), very low-density lipoproteins (VLDL), intermediate-density lipoproteins (IDL), and low-density lipoproteins (LDL) as previously described [35]. Lomitapide treatment effectively reduces VLDL particles and increases LDL particles [35] (Figure 3E), whereas riboflavin and calcipotriene did not affect lipoprotein classes. Thus, riboflavin tetrabutyrate and calcipotriene reduce total B-lp levels without overt developmental toxicity or changes in lipoprotein subclass distribution, suggesting they may act through mechanisms that decrease overall particle abundance rather than altering lipoprotein turnover or catabolism.

      Although doxycycline treatment lowered B-lp levels in whole fixed animals in the primary screen and validation studies, we did not observe a reduction in total B-lps in whole-animal homogenates (Figure 3D). However, we detected a slight increase in VLDL levels (p < 2x10<sup>-16</sup>; Figure 3E), suggesting that doxycycline may alter lipoprotein composition or distribution rather than total particle abundance.

      Alternatively, two structurally related compounds, thiethylperazine and prochlorperazine, at 4 µM significantly reduced (p < 1.4x10<sup>-10</sup> and p < 1.2x10<sup>-6</sup> respectively), B-lp levels measured from whole-animal homogenates (Figure 3F and 3H). Furthermore, both 8 µM thiethylperazine and 8 µM prochlorperazine increased relative LDL (p = 1.2x10<sup>-3</sup> and p = 2.3x10<sup>-3</sup>, respectively) and decreased relative VLDL levels (p = 8.4x10<sup>-4</sup> and p = 9.1x10<sup>-4</sup>, respectively; Figure 3G and 3I) suggesting a shift toward smaller lipoprotein particles and a potential alteration in lipid processing or clearance pathways. Together, these results highlight the diversity of mechanisms among validated hits, ranging from compounds that reduce total B-lp abundance without affecting B-lp class composition to those that shift B-lp class distribution, while also underscoring the importance of secondary assays to distinguish true B-lp modulators from those that likely produce a B-lp effect through generalized toxicity.

      Enoxolone significantly reduces B-lps in the larval zebrafish.

      Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).”

      The Discussion now has the following additional text:

      “Further validation of these hits demonstrated a wide range of potential mechanisms of lipoprotein regulation. We identified hits that affected larval development, some hits that reduced total B-lp levels, and several structurally related compounds that directly reduced B-lp particle size (Figure 3).”

      Reviewer #2 (Public review):

      Strengths:

      The study was methodical and robust, using a published and well-validated zebrafish LipoGlo model. The authors validated the hits from the screen independently and considered the possibility that some drugs may have been detected as false positive results due to effects on the enzymatic activity of NanoLuciferase; only one hit, verteporfin, was shown to be a false positive. Using LipoGlo-Electrophoresis, the authors are able to obtain extra insights into the ApoB-lipoprotein size/subclass distribution. They showed that while enoxolone treatment reduces total B-lps, there are no overt changes in B-lp size distribution compared to vehicle-treated animals, other than a slight increase in the zero mobility (ZM) fraction, which contains very large particles and/or tissue aggregates. In contrast, the positive control, lomitapide, does show a change in B-lp size distribution compared to vehicle-treated animals - an increase in frequency of LDLs (low-density lipoprotein), but a decrease in VLDLs (very low density lipoprotein). This study also assesses the LipoGlo-Electrophoresis profile of HNF4⍺ inhibitors. Work in the zebrafish larvae means that the effect on overall development and an entire vertebrate organism can also be assessed. Finally, the authors applied a thorough statistical measure to define a hit, using the Strictly Standardized Mean Difference (SSMD) method.

      We appreciate that the Reviewer valued the rigour and robustness of our approach.

      Weaknesses:

      While the screen was thorough and well-validated, the authors missed a chance to provide a lot of extra significance to a wide range of readership. While the hits were thoroughly validated and displayed, the authors could have also presented the LipoGlo-Electrophoresis for all validated hits or at least a number of them. This would hugely increase the insights into these compounds. Also, the authors chose to validate and follow up a mechanism for Enoxolone, yet this hit was already known to modulate lipid metabolism through HNF4⍺, therefore, hugely limiting the impact of the paper. So what the authors have shown that is novel is only subtly added to this - consistent in vertebrate models, RNA sequencing of pathways, further validation of the HNF4⍺ pathway, and a profile of resulting B-lp size distribution. It seemed an easy way out to pick such a candidate, and they could have followed up by validating more thoroughly a completely novel drug. Also, the authors' prior paper showing the methodology also depicted complementary EM and LipoGlo-microscopy approaches. The microscopy especially, would have been an easy complementary add-on to the screen to really give extra insights into B-lp metabolism in a whole organism for all candidates. This felt like a missed opportunity.

      Here we agree and added Fig. 3 describing the phenotyping of 6 additional compounds including some LipoGlo-Electrophoresis analyses as suggested by the Reviewer. The text of the Results section was modified as described for Reviewer 1 (see above).

      Reviewer #3 (Public review):

      Strengths:

      The study uses a well-validated in vivo stain (LipoGlo) for measuring lipoproteins in the context of a developing whole organism with a quantitative read-out on a high-throughput platform, allowing for screening of thousands of compounds altering the complex metabolic/physiologic functions necessary for lipoprotein production.

      The use of genetic mutant HNF4alpha to assign the mechanism of action to the prime candidate compound studied (enoxolone) is a powerful approach for this challenging aspect of chemical genetics studies.

      We appreciate that the Reviewer understands how challenging it can be to assign a mechanism to any small molecule and thereby recognizes the power of the zebrafish model combined with our unique lipoprotein phenotyping tools.

      Weaknesses:

      As shown in Figure 5A, the HNF4alpha mutant homozygous -/- already lowers lipoproteins. Is it just that the mutant level is already at a minimum in this homozygous mutant (and thus enoxolone cannot induce even lower lipoprotein levels), or is it true that the enoxolone molecule is primarily acting through this TF (i.e. HNF4alpha homozygous mutant is truly epistatic to enoxolone function) as favored in the text.

      While it is definitely interesting to study enoxolone effects during whole embryo development, the link to HNF4alpha had previously been described in the literature, as pointed out by the authors. The generalizability of the approach to identify truly novel pathways remains to be fully realized, but sharing this available screen data to date will invite further inquiry and be very valuable to the community.

      Here too we agree that a link between enoxolone was proposed in the literature. However, we added quite a lot of additional insight regarding the transcriptional targets shared by HNF4alpha and enoxolone. The goal of identifying the mechanism(s) of action of other novel small molecule hits from the screen is important and that work is ongoing.

      Figure 5 - The same allele of HNF4alpha loss of function/hypomorph (rdu14) is used in both 5A and 5B, but labeled differently in each subpanel. This is explained in the figure legend, but could be updated to use the same nomenclature in both panels to clarify the Figure presentation.

      We thank the Reviewer for catching this and have modified the Figure (now Fig. 6) so the subpanels are labeled identically to avoid any confusion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors describe statistical methods to improve the calling of hits. In the results section, they discuss the use of fold change and of strictly standardized mean difference as criteria to account for the variability that occurs within in vivo screens. The combination of both criteria resulted in a significant reduction in the total number of hits, from 487 (16%) to 49 (1.6%), which is a large decrease in total hits. The authors should comment on whether some of the 487 compounds were randomly tested individually to confirm that they were true negatives and compared to the true hits of 49. For example, did the authors independently test calcipotriene and diphenylboric acid to confirm that these are truly negative?

      Initially, we did not directly retest any of the 487 compounds outside of the top 49 hits with the intention to compare their effect size to the top compound. While we do suspect that there is a chance that some of these compounds may still have a significant effect on lipoprotein levels, we prioritized compounds with the strongest overall biological effect size and largest statistically significant effect. We agree with the Reviewer that it would be interesting to continue to test more of these compounds and examine whether the lipoprotein reduction phenotype correlated with the primary screen effect, this would be a large experimental undertaking, though a randomized subset could be tested. Validation studies of Calcipotriene (Supplemental Figure 2E) showed a small, but significant, lipoprotein reduction at only the 0.25 µM dose across 3 independent experiments. Because the effect is small and not a dose-response, it is not a highly prioritized hit. We did not validate diphenlylboric acid in this study but will in the future.

      (2) The transcriptomics profile studies are not well described. The authors do not provide a detailed list of differential genes for each of the conditions in their dataset. This should be included.

      Supplemental Table 2 includes the fold change and p-values of all genes measured in differential expression analysis. We agree with the Reviewer and added new supplementary tables (Supplemental Table 5 and 7) that contain the differentially expressed genes at each treatment duration.

      (3) The Gene Ontology analysis results are not accompanied by a list of genes that match the GO terms. The authors should include these results. This part was confusing as it was not clear if the cholesterol pathway affected was due to an abundance of genes that were up- or down-regulated.

      We agree with Reviewer 1 and added Supplemental Table 6 that contains each gene ontology term with the associated DE genes that contribute to the significance of that term as well as explicit directionality information.

      (4) Figure 6A shows heat maps of differential gene expression, but there is no key provided in the figure or legend. Are they color-coded for fold change or log2 fold change?

      We agree that Figure 6A can be approved and added a key to the legend as suggested by the Reviewer.

      (5) The overlap of genes that are changed in hnf4a mutants and with enoxolone is provided as percentages, but the actual genes are not listed. Do these genes represent cholesterol biosynthesis pathways? There are other bioinformatic tools that the authors could apply to their dataset for further analyses. For example, enrichr (https://maayanlab.cloud/Enrichr/) is one tool to query for GO Biological processes, cell/tissue type, and even overlap of genes with genetic models and other disease states. An extension of these bioinformatic studies will be useful to determine if other pathways are relevant.

      We agree with the Reviewer and added a table of the genes that overlap between drug treatments and hnf4a mutants (Supplemental Table 7). We also appreciated the Reviewer’s suggestion to deploy Enrichr, which we have done and modified the Results section to now read:

      “Of the 439 differentially expressed genes from 12, 16, and 24 hours post-treatment, 34 differentially expressed genes are shared between all three treatment durations and are associated with gene ontology terms related to carbohydrate metabolism and signaling pathways (Figure 7D). We expanded this analysis using the bioinformatic tool Enrichr [68–70], which largely recapitulated gene ontology results described above. However, Enrichr analysis revealed significant overlap between several late enoxolone-responsive gene sets and transcriptional signatures associated with prochlorperazine, another compound identified in our screen (Figure 2, Figure 3, Supplemental Figure 1AO, Supplemental Figure 2Z). Enrichment of prochlorperazine-associated signatures was observed at 12 (prochlorperazine MCF7 up, 4/58 genes [INSIG1;IRF7;ISG15;ATF3], adjusted p-value = 0.006), 16 (prochlorperazine MCF7 up, 6/58 genes [INSIG1;DDIT4;IRF7;PMAIP1;ISG15;ATF3], adjusted p-value = 0.00008; prochlorperazine PC3 up, 3/29 genes [INSIG1;DDIT4;ATF3], adjusted p-value = 0.006), and 24 hours post-treatment (prochlorperazine PC3 up, 6/29 genes [DUSP5;INSIG1;DDIT3;TRIB3;SQSTM1;ATF3], adjusted p-value = 0.0001; prochlorperazine MCF7 up, 7/58 genes [DDIT3;INSIG1;IRF7;PMAIP1;ISG15;SAT1;ATF3], adjusted p-value = 0.0005). These results suggest that enoxolone and prochlorperazine may perturb overlapping molecular pathways, an observation that warrants further investigation. Ultimately, these data demonstrate distinct early and late responses to enoxolone treatment, and the early response modulates key lipid metabolism pathways.”

      (6) Figures 1E and 1F would benefit if the exact spots/points where enoxolone and other hits mentioned in the text were labelled.

      Great idea, we modified Figure 1F as suggested.

      (7) Figure 3 and Figure 4 graphs should state Fold change on the Y-axis title.

      We are thankful the Reviewer noticed this typo and we have fixed both Figures.

      Reviewer #2 (Recommendations for the authors):

      (1) To boost the impact for more readers, the authors should include the LipoGloElectrophoresis and LipoGlo-Microscopy results from a few more of the validated hits, especially ones that are completely novel (unlike Enoxolone, which already had a known role in lipid metabolism). Results on enoxolone are useful as a validation of the assay, mostly with some minor additional insights.

      We agreed with Reviewer 2 and added more validation testing of for a few hits (see response to Reviewer 1, Strengths and Reviewer 2, Weaknesses). We did not perform these experiments on each drug for technical reasons, mainly because these experiments are low-throughput (especially the Microscopy). Nonetheless, we performed many additional experiments to add phenotyping data for 6 new drugs that included multiple LipoGlo-Electrophoresis panels to an entirely new Figure.

      (2) The authors should include raw data from the screen from all drugs tested in the supplementary and then for which SSMD was calculated for, providing in an excel sheet or similar the values and how these were calculated, i.e. the 487 unique drugs that lower B-lp levels with an SSMD cutoff of < -1.0.

      This information was provided in the supplemental file as separate .csv files with the associated R script, which can be run locally and contains the SSMD functions. Considering the Reviewer comments, we ensured this information in provided in Supplemental Tables 1 and 2.

      (3) Page 3, lines 23-24: What does the 2 to 4 fold chance mean? Perhaps rewrite: Genetic mutations in Lipoprotein(a) increase the chance of heart attack or stroke 2-4 fold greater than without the mutation.

      We agree and the sentence now reads: Patients with genetic mutations in the Lipoprotein(a) encoding gene have a 2 to 4 fold increased risk of sudden heart attack or stroke.

      (4) Figure 1, for C and D, label some of the most significant hits and definitely show where exonolone lies.

      We agree see response to Reviewer 1 Pt6

      (5) Page 6, lines 1-9: I'm a bit confused why this is here if you do not present the data.

      We thought this was relevant information to share for researchers that running drug screens with positive controls and defining hit cutoffs. In light of the Reviewer’s comment, we removed the last sentence from this paragraph.

      (6) Page 6, line 8: This needs better justification of why you are validating enoxolone rather than other hits; otherwise, it could seem like cherry picking. Especially as enoxolone is known to affect lipid metabolism. Otherwise present more details of a couple of validated candidates.

      We agree. As the Reviewer requested, we validated more compounds (described above) and modified the text of the results to elaborate on our justification for selecting enoxolone for further study. The text of the results now reads: Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and prior reports the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).

      (7) Supplementary Figures 1 and 2: The resolution is too low, and the reader cannot even see the charts or the text.

      We agree and now have uploaded higher resolution images

      (8) Page 7, lines 31-35: Needs a higher resolution and magnified image to merit this 'offhand' statement. Also cite reference [59] here.

      We agree with the Reviewer and added magnified insets of the heart. As far as the suggestion of adding Ref 59, we do not see the connection to that paper (A point mutation decouples the lipid transfer activities of microsomal triglyceride transfer protein PLOS Genetics 16:e1008941)

      (9) Figure 3E: In addition to the proportions graph would be useful to also have an absolute amount of lipoprotein in each class graph.

      While there may be changes in total luminescence values from lane to lane in these gels, we have not fully validated the absolute quantitation of a full lane. We typically use plate-based whole-animal assays to determine total absolute lipoprotein levels and calculate the proportion of the whole lane for each lipoprotein class, as described in our prior publication detailing the assay. We do expect that there is some additional variation incorporated into the native PAGE assay due to sample freeze/thaw, dilution, and loading.

      (1) Page 10 lines 27-28: "Continual statin use for more than 1 year reduced circulating Blps and all-cause mortality by ~30% in individuals with high B-lp levels." This sentence doesn't seem right, intimates continual statin use causes death - I don't think that's right.

      We thank the Reviewer for catching this and have corrected the sentence. It now reads: “Continual statin use for more than 1 year in individuals with high B-lp levels reduced circulating Blps and lowered all-cause mortality by ~30% [12,13].”

      (11) Page 12, line 25: "canlikely" is a typo, should be can likely.

      We fixed that sentence and now reads: “Further, the drug screening paradigm we developed using the LipoGlo system is highly scalable and can be deployed to screen large novel drug libraries to identify many additional B-lp-lowering compounds.”

      (12) Figure 1 legend: "An ordered plot of each SSMD score measured from 5 μM lomitapide treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment." This comes across as though it's 5uM Iomitapide/vehicle. But it's the SSMD score of each drug compared to Iomitapide and relative to the respective vehicle (I think) - make it clearer.

      We agree and clarified the legend so it now reads: “…(D) An ordered plot of each SSMD score measured from positive control (5 µM lomitapide) treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment. The solid black line at y = 0 represents the divide in increased and decreased SSMD score, the solid blue line at y = -1.41 represents the curve's inflection point, and the dashed black line at y = -1 represents the SSMD (open circles) cutoff used to define a hit. “

      (13) Figure 1E: What is the x axis?

      Each data point on the x-axis represents each drug at every dose tested, we will clarify the test. The legend now reads: “(E) A plot of SSMD scores measured from each drug at each dose tested, each open circle represents the SSMD score of an individual drug at an individual dose.” In addition, “Compound (each dose tested)” was added to the x-axis of the figure panel.

      (14) Figure 2: Would you not have space to put the drug names in the figures? Where, for example, is enoxolone?

      We agree and have updated the figure accordingly.

      (15) Figure 3A-C: label enoxolone on the x axis.

      We agree and added this text to what is now Figure 4.

      (16) Figure 3D: Looks like delayed development with enoxolone, if left to grow, would the embryos develop normally?

      We did not examine if animals treated from 3-5 dpf develop normally beyond 5 dpf.

      (17) Figure E. Are stars all compared to vehicle control? Perhaps useful to have lines to indicate what are the significantly different relationships.

      Comparison in these experiments are always to the vehicle (negative control) and clarified the legend considering the Reviewer’s comment we modified the legend to now read,”… * <0.05 as compared to vehicle.”

      (18) Figure 3E: As well as the proportion of total lipoprotein, it would also be beneficial to see absolute lipoprotein levels.

      See above response to Reviewer 2 Pt9.

      (19) Figure 4: Does the overall health or size of the animal correlate with the luminescence score?

      While we do know that, in untreated animals, lipoprotein levels vary with age (and, thus, size), we have not examined this more granularly than in 24-hour time points after treatment. Further, we have not examined this in the context of a drug treatment.

      (20) Figure 4D: Again the absolute in each fraction would be meaningful, also the 5078 looks brighter?

      See above response to Reviewer 2 Pt9.

      (21) Figure 4A, 5A: Would the traces (line plots) not be useful here to see the overall dynamics over time?

      We considered presenting Figure 5A this way but decided to keep the plots as is because we wanted to show the individual points which make the line blots very difficult to read. Further, our analysis evaluates individual animals at each time point as it is not possible to follow the same animal over time.

      Reviewer #3 (Recommendations for the authors):

      Figure 4: Consider using standard scientific notation for the very small p values in some of the figure legends.

      We agree and adjusted the p-value notation as suggested.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments following re-submission:

      Overall, I think the authors have done a satisfactory job of addressing most of the points I raised.

      There’s one final issue which I think still needs better discussion.

      I think reviewer 2 articulated better than I have the point I was concerned about: the relationship between JNDs and metamers as depicted in the schematics and indeed in the whole conceptualization.

      I think the issue here is that there seems to be a conflating of two concepts- ’subthreshold’ and ’metamer’-and I’m not convinced it is entirely unproblematic. It’s true that two stimuli that cannot be discriminated from one another due to the physical differences being too small to detect reliably by the visual system are a form of metamer in the strict definition ’physically different, but perceptually the same’.

      However, I don’t think this is the scientifically substantial notion of metamer that enabled insights into trichromacy. That form of metamerism is due to the principle of univariance in feature encoding, and involves conditions in which physically very different stimuli are mapped to one and the same point in sensory encoding space whether or not there is any noise in the system. When I say ’physically very different’ I mean different by a large enough amount that they would be far above threshold, potentially orders of magnitude larger than a JND if the system’s noise properties were identical but the system used a different sensory basis set to measure them. This seems to be a very different kind of ’physically different, but perceptually the same’.

      We are in full agreement with this. Typically, the notion of metamers is about deterministic information loss, which can be modeled as a projection from a high-dimensional physical space to a lower-dimensional perceptual space. This is the topic of the paper, and is analogous to the work on color matching in the 19th century.

      In contrast, sensitivity to small differences within the perceptual space due to internal noise is usually addressed by other methods, such as signal detection theory, a topic which is not the focus of this paper. It is analogous to the work on discriminability of colors such as MacAdam ellipses (MacAdam, 1942). It is instructive to look at progress in the color field. The color matching experiment and the question of metamerism is quite well worked out, whereas the question of how to quantify discriminability within that space has been an ongoing topic of investigation for over a century.

      Here, we aim to develop and test a model of metamers, analogous to the color matching experiments, but we do not attempt to develop a model of discriminability. Nevertheless, while the two types of information loss are conceptually distinct, they are both present in the nervous system of the observer, and both are reflected in the performance vs. scaling plots in our paper. We have added clarifications about this point in the introduction on page 2, starting on line 45, and the discussion, starting on page 16, line 397.

      Finally, regarding physical differences between stimuli: the differences between target images and synthesized metamers are quite large (high mean squared error), as shown in Appendix 5. In no condition did we present subjects with stimulus pairs that were physically similar.

      I do think the notion of metamerism can obviously be very usefully extended beyond photoreceptors and photon absorptions. In the interesting case of texture metamers, what I think is meant is that stimuli would be discriminable if scrutinised in the fovea, but because they have the same statistics they are treated as equivalent.

      The notion of “texture metamers” is perhaps a reference to the work by Freeman and Simoncelli (2011), whose stimuli are similar to ours: when synthesized using a model with sufficiently small scaling, they are indiscriminable, and therefore metamers. The reviewer is of course correct that the stimulus pairs are not metamers when the observers move their eyes due to differences in spatial encoding as a function of eccentricity. That is, they are only metameric under a specific set of viewing conditions, and they are not metameric when those conditions are violated. The same is true for color metamers, as the spectral sensitivity of the cones also differ with eccentricity (Stockman and Sharpe, 2000).

      I think the discussion of this could still be clearly articulated in the manuscript. It would benefit from a more thorough discussion of the difference between metamerism and subthreshold, especially in the context of the Voronoi diagrams at the beginning.

      We agree that a more thorough discussion of the diagrams could help clarify the issues to the reader. We have modified the caption of figure 1 with the goal of clarifying interpretation of the diagrams, and see also our discussion earlier in this note about noise and discriminability.

      It needs to be made clear to the reader why it is that two stimuli that are physically similar (e.g., just spanning one of the edges in the diagram) can be discriminable, while at the same time, two stimuli that are very different (e.g., at opposite ends of a cell) can’t.

      Do the cells include BOTH those sets of stimuli that cannot be discriminated just because of internal noise AND those that can’t be discriminated because they are projected to literally the same point in the sensory encoding space? What are the strengths and limits of models that involve the strict binarization of sensory representations, and how can they be integrated with models dealing with continuous differences? These seem like important background concepts that ought to be included in either the introduction of discussion sections. In this context it might also be helpful to refer to the notion of ’visual equivalence’ as described by:

      This is an important point and we appreciate the reviewer raising it. In brief, as one traverses a region in one of the Voronoi diagrams, the images are changing physically but subject to the constraint that they all project to the same single point in the reduced perceptual space. When one crosses from one region to another, the images now project to a different point in the perceptual space. Whether or not that the two locations in the perceptual space are distant enough to be distinguishable given the internal noise is a question pertaining to the topic of JNDs in the perceptual space, rather than the mapping from physical space to the perceptual space. We do not address that question in detail in this paper, though we do now reference it in the caption of figure 1, as well as in the new sections in the introduction and discussion mentioned earlier in this response.

      We do note that the perceptual space is not discrete: the model outputs are real-valued. The apparent discretization is a limitation of the simplified 2-D schematics.

      Ramanarayanan, G., Ferwerda, J., Walter, B., & Bala, K. (2007). Visual equivalence: towards a new standard for image fidelity.ACM Transactions on Graphics (TOG), 26(3), 76-es.

      Other than that, I congratulate the authors on a very interesting study, and look forward to reading the final version.

      Reviewer #2 (Public review):

      Summary:

      The authors have improved clarity overall and have spoken to most of the issues raised by the reviewers. There are still two outstanding problems however, where issues raised during the review were inappropriately dismissed in the manuscript. These should be explicitly addressed as limitations to the results presented (no eye tracking), and early pilot experiments that informed the experiments as presented (pink noise) rather than brushed off as ’unnecessary’ and ’would be uninformative’.

      Eye tracking:

      It is generally accepted that experiments testing stimuli presented at specific locations in peripheral vision require eye tracking to ensure that the stimulus is presented as expected, in particular, in the correct location. As I stated in the previous round of review, while a stimulus presentation time of 200ms does help eliminate some saccades, it does not eliminate the possibility that subjects were not fixating well during stimulus onset. I am also unclear what the authors mean by ’trained observer’ in this context, though the authors state that an author subject in a different portion of the paper is an ’expert observer’. Does this mean the ’trained observers’ are non-expert recruited subjects?

      Given the conditions tested differ from previous work (Freeman & Simoncelli, 2011) ‘these differences are a main contribution of the paper!’ which DID include eye tracking in a subset of subjects, it is entirely possible to get similar results to this work in the context of non eye-tracking controlled stimulus presentation. The reasons now in the manuscript are not reasons that make eye tracking ’considered unnecessary’.

      I appreciate that the authors now state the lack of eye tracking explicitly, but believe the paper needs to at least state that this is a limitation of the results reported, and eyetracking being ’considered unnecessary’ is unreasonable, nor a norm in this subfield.

      By “trained” observers, we mean people who were recruited from the community of vision science labs at NYU and who have participated in many visual psychophysics experiments. All of the participants are “trained” in this sense, and are thus used to maintaining fixation while performing peripheral tasks. One of these participants, an author, was also an expert in the specific content area of the paper. By “expert”, we mean high familiarity with the stimulus types and models employed in the paper.

      We have now further clarified this in the text in the subsection of the methods on Observers, on page 22.

      We also discuss the issue at greater length in the methods subsection Apparatus, on page 26. We removed the word "unnecessary" and make it clear that while we don’t think our results are undermined, the lack of eye tracking is nonetheless a limitation.

      N=1: The authors now state clearly the limitations of a single subject in the manuscript, and state the expertise level of this subject.

      Large number of trials: The authors now address this and include an enumeration of the large number of trials.

      Simple Models / Physiology comparison: I support the choice to reduce claims regarding tight connections to physiology, and appreciate the explanation of the luminance model.

      Previous Work: I appreciate the author’s changes to the introduction, both in discussing previous work and citation fixes.

      Blurred White, Pink Noise: While the authors now address pink noise, the explanation for such stimuli being expected to be uninformative is confusing to me. The manuscript now first states that pink noise is a natural choice, then claims it would be uninformative, while also stating in the rebuttal (not the manuscript) that they tried it and it indeed reduced the artifacts they note. The logic of the experiments indeed relies on finding the smallest critical scaling value, which is measured by subjects determining if a synthesis is similar or different to a target or second synth. A synthesis free from artifacts would surely affect the subjects responses and the smallest critical scaling measured.

      The statement that the authors experimented with pink noise early on and found this able to address the artifacts should be stated in the manuscript itself, not just in the rebuttal, and the blanket statement that this experiment would be ’uninformative’ is incorrect. Surely this early pilot the authors mention in the rebuttal was informative to designing the experiments that appear in the final paper, and would be an informative experiment to include.

      First, we did render some test stimuli with pink noise seeds, but we did not collect psychophysical data, hence there are no results we could add. Visual inspection of these stimuli was indeed clarifying in the following sense. The pink noise stimuli had fewer high-frequency “artifacts”. If our goal was to synthesize stimuli that are indistinguishable from the original stimulus, as one might do to save compute power when in a device that for foveated rendering, then starting with pink noise would be better than starting with white noise. Our purpose was just the opposite. For our experiments, the artifacts were just what we wanted: the more artifacts, the better. The reason is that a strongest test of a metamer model is whether two stimuli that are as physically different from one another as possible, are nonetheless indistinguishable when their model representations are the same. Stimuli synthesized from pink noise seeds are harder to discriminate from the target stimulus, not easier. Thus using them in an experiment would result in a larger estimate of critical scaling. Since our explicit goal was to estimate the smallest critical scaling window, these stimuli would not bring us closer to our goal. As the reviewer points out, these metamers were “informative” in the sense that they informed our experimental design, but they are “uninformative” (relative to white noise seeds) for estimating the critical scaling.

      We have updated our description in the discussion starting on page 19, line 449, and included a new appendix to demonstrate this point (appendix 2 on page 35).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Typo: p. 19, l. 439: ’Why does asymptotic performance, but not critical scaling, depends on image content?’": remove ’s’ from ’depends’.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      Recommendations: State that the lack of eye tracking to control stimulus presentation is a limitation of the results presented.

      Remove the claim that pink noise or filtered white noise seeds would be uninformative, and mention the fact that the authors in fact experimented with pink noise seeds in an early version of the experiments (which was surely informative to the experimental setup as presented here).

      Addressed as described above.

      References

      Freeman J, Simoncelli EP. Metamers of the ventral stream. Nature Neuroscience. 2011 aug; 14(9):1195–1201. doi: 10.1038/nn.2889.

      MacAdam DL. Visual Sensitivities To Color Differences in Daylight*. Journal of the Optical Society of America. 1942 may; 32(5):247. http://dx.doi.org/10.1364/josa.32.000247, doi: 10.1364/josa.32.000247.

      Stockman A, Sharpe LT. The Spectral Sensitivities of the Middle- and Long-Wavelength-Sensitive Cones Derived From Measurements in Observers of Known Genotype. Vision Research. 2000 jun; 40(13):1711–1737. doi: 10.1016/s0042-6989(00)00021-3.

    1. Author response:

      We greatly appreciate the positive and constructive comments from the reviewers, which recognized the intellectual contributions of our manuscript and also captured its limitations. Below is our response to the three major points from Reviewer #1 and the comments from Reviewer #2.

      Response to Reviewer #1

      (1) Use of transient transfection in non-hematopoietic cells. We appreciate the reviewer’s point regarding physiological relevance. Our goal in these cellular microscopy experiments was to dissect the biophysical principles of NUP98::KDM5A (including its mutants) condensate formation under controlled expression levels (concentration). While AML model systems driven by NUP98::KDM5A are available, they do not offer this possibility because of pre-existing NUP98::KDM5A expression. The suspension culture of hematopoietic cells also brings practical challenges for correlative FISH+IF and high-resolution microscopy analysis. We agree that validating our observed behaviors in AML models would be valuable, e.g. by creating HSPC lines with inducible expression of tagged NUP98::KDM5A and its mutants, but such experiments fall outside the scope of the current study. We will add text acknowledging this limitation and clarifying that our mechanistic conclusions are grounded in biophysical principles that should generalize across cell types.

      (2) Correlation between H3K4me3 and gene activation. We agree that active transcription correlates with H3K4me3, and that this baseline relationship must be considered. Our analysis explicitly uses fold-change between patient cells and healthy controls as the readout. This comparison inherently accounts for the activating effect of H3K4me3 itself. Regarding the reviewer’s comment that “many H3K4me3-positive genomic regions do not show NUP98::KDM5A binding”, we would like to clarify that this point is exactly what our manuscript aims to explain. Our cell line studies demonstrate that, at a patient-relevant expression level, NUP98::KDM5A condensates form preferentially at H3K4me3 locus with high local mark density. Considering that NUP98::KDM5A concentration in the nucleus is lower than the K_D between KDM5A PHD3 and H3K4me3, this means that H3K4me3 loci with lower mark density will not see NUP98::KDM5A binding without the high local concentration of the fusion protein (as a result of condensate formation). This is consistent with our observation that genes with the highest local density show disproportionately stronger upregulation. We will further clarify this point in the revised manuscript.

      (3) Generalization to other NUP98 fusions lacking PHD domains. We appreciate this important conceptual question. We will expand the discussion to note that many NUP98 fusions, despite diverse partner domains, produce similar transcriptional programs. Our current hypothesis is that the highly active status of the HOX cluster genes in HSPC attracts NUP98 oncofusions targeting H3K4me3 (e.g. KDM5A and PHF23), while NUP98::HOXA9 and other transcription factor fusions directly target the HOX cluster via DNA sequence recognition. This convergence and the broad targets of the dysregulated HOX transcription factors explain the transcriptional program similarity, sustained by the previously described enrichment of transcriptional co-activators by NUP98 oncofusion condensates. We will explicitly discuss this hypothesis in our revised manuscript.

      Response to Reviewer #2

      We thank the reviewer for the thoughtful evaluation and agree that the causal link between H3K4me3 recognition, condensate formation, and gene activation remains partly correlative. Experiments such as qPCR or ChIP following expression of WT versus mutant constructs would indeed strengthen the causal chain. However, performing these assays across multiple constructs and loci in a physiologically relevant system is not feasible within the current revision cycle. We have added text acknowledging this limitation and clarifying that our study focuses on establishing the biophysical mechanism of targeting, while functional consequences are inferred from patient datasets rather than new perturbation experiments.

    1. Author response:

      Reviewer 1:

      (1) Major Concerns (highest priority):

      There is a lack of detail about the methodology in the results section/figure legends, which makes it difficult to interpret the data. Sometimes, adequate information is also not included in the methods themselves. For example, Figure 1A: how many doses did each mouse receive? How long after dosing were animals sacrificed? Figures 1E and 2A: is this RNA-seq analysis?

      Thank you for pointing this out. We will provide additional details about the methodology in the results, figure legends, and methods to clarify our findings.

      Using a second EGFR inhibitor for some of the key experiments would increase the rigor of the studies shown.

      Thank you for this suggestion. We agree that a second EGFR inhibitor may increase the scientific rigor of the findings. However, at the time the experiments were completed, osimertinib was the clear standard of care first-line therapy for EGFR-mutated lung cancer and the clinical relevance of using other inhibitors in these experiments is not clear.

      (2) Nice to have experiments:

      Using CRISPR KO of CDK4 in the CDK4-amplified HCC827 and testing response to Osi and presence of replication stress would also increase the rigor of the studies.

      Thank you for this suggestion. We agree that this would increase the rigor of the studies, but these experiments are currently beyond the scope of this manuscript.

      The authors show that in their patient data, some cell cycle regulators which are amplified in NSCLC at similar rates to CDK4/6, such as CCNE1, had no increase in FGA. Overexpressing CCNE1 and testing Osi response in their cell models would be a nice test of their proposed mechanism that it is the genomic instability and FGA that are driving resistance. This wouldn't need to be done in vivo, but could be done using cell culture-based methods.

      Thank you for this suggestion. For this manuscript, we have chosen to focus on the role of CDK4 and CDK6 in osimertinib resistance as they have shown the clearest correlation with decreased responsiveness to EGFR TKI treatment in other studies. We agree that assessing the role of CCNE1 in osimertinib resistance is important and will address this in future studies.

      Similarly, testing the overexpression of some of the proposed target genes, such as STEAP1 and AGR2, on the therapeutic response to Osi in cell culture would also be a nice test of the mechanism proposed.

      Thank you for this suggestion. We agree that these experiments are important, but they are currently beyond the scope of this study.

      Reviewer 2:

      Major Comments:

      (1) Reliance on overexpression models over loss-of-function

      The mechanistic studies mainly rely on overexpression of CDK4 and CDK6 to simulate the amplified state. Although the authors argued that the level of overexpression mimics that observed in resistant tumors, a complementary study in which CDK4/CDK6 were suppressed in a model where CDK4/CDK6 is amplified (such as HCC827 or TH116), and replication stress and osimertinib sensitivity tested would greatly strengthen their observations. Indeed, there is mention of CDK4 constructs to perform knockdown studies in the methods, but those studies are not included in this submission.

      Thank you for this suggestion. We agree that loss of function studies are an important complement to over expression studies. However, we have chosen to focus on the use of pharmacologic inhibitors of CDK4/6, which are more clinically relevant than knockdown studies. We demonstrate that the CDK4/6 inhibitor palbociclib is able to prevent replication stress and restore osimertinib sensitivity in CDK6 amplified TH116 patient-derived xenografts (Figure 5 and Supplementary Figure 10).

      (2) Mechanistic link between genomic instability and specific gene amplifications

      The authors highlight that CDK4/6 activation leads to recurrent copy number gains and transcriptional upregulation of specific pro-tumor genes including AGR2, ASNS, and STEAP1. While the paper establishes that CDK4/6 overexpression causes general genomic instability (increased FGA), it does not mechanistically explain why these specific genes are consistently amplified. The authors should investigate or discuss whether these specific loci are inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if this is a result of random genomic instability followed by strong positive selection under osimertinib pressure.

      Thank you for this suggestion. We will provide additional discussion about whether these specific loci are likely to be inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if it is more likely the result of random genomic instability followed by strong positive selection under osimertinib pressure.

      (3) Discrepancies in tumor mutational burden (TMB) reporting

      There is a slight contradiction regarding the TMB data that needs to be clarified for readers. The manuscript states that in the clinical datasets, "EGFR-mutant LUAD harboring cell cycle gene alterations exhibited significantly elevated FGA and TMB relative to cell cycle-negative tumors" (Line 265-266). However, in the next section, the authors say, "Notably, no corresponding increase in TMB was observed with CDK4 or CDK6 CNA, similar to our findings in preclinical models" (Line 271-273). The authors should clarify or discuss why broad cell cycle alterations correlate with high TMB, while CDK4/6-specific alterations drive structural instability (FGA) without increasing TMB.

      Thank you for this suggestion. We will clarify our findings and provide further discussion about why there may be a difference between the effects of broad cell cycle alterations and CDK4/6-specific alterations on TMB and structural genomic instability.

    1. Author response:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      We thank the reviewer for these positive comments.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

      We thank the reviewer for this review. We do believe that the manuscript has a mechanistic component, as the central experiments involve direct manipulation of neuronal activity, and we show an increase in calcium levels and gene expression changes in dopamine neurons that coincide with the degeneration. However, we agree that deeper mechanistic investigation would strengthen the conclusions of the paper. We have planned several important revisions, including the addition of CNO behavioral controls, manipulation of intracellular calcium using isradipine, additional transcriptomics experiments and further validation of findings. We anticipate that these additions will significantly bolster the conclusions of the paper.

      Reviewer #2 (Public Review):

      Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:

      This is an exciting and important paper.

      The paper compares mouse transcriptomics with human patient data.

      It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      We thank the reviewer for these insightful comments.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      This is an important point. Although we show that CNO does not produce degeneration of DA neuron terminals, we do not exclude a contribution to the behavioral changes. We agree that this behavioral control is necessary, and will address it in revision with a CNO-only running wheel cohort.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      We agree that additional electrophysiology conducted in the VTA dopamine neurons would meaningfully add to our understanding of the selective vulnerability in this model, and will complete these experiments in revision.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

      We will explicitly clarify which mice had access to a running wheel in our revision. Briefly, mice for histology, electrophysiology, and transcriptomics all had access to a running wheel during their treatment. The mice used for photometry underwent about 7 days of running wheel access approximately 3 weeks prior to the beginning of the experiment. The photometry headcaps sterically prevented mice from having access to a running wheel in their home cage.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons.

      We thank the reviewer for the careful and thoughtful review of our manuscript.

      While extensive depolarization and associated intracellular calcium elevations promotes degeneration generally, we emphasize that the process we describe is novel. Indeed, prior studies delivering chronic DREADDs to vulnerable neurons in models of Alzheimer’s disease did not report an increase in neurodegeneration, despite seeing changes in protein aggregation (e.g. Yuan and Grutzendler, J Neurosci 2016, PMID: 26758850; Hussaini et al., PLOS Bio 2020, PMID: 32822389). Further, a critical finding from our study is that in our paradigm, this stressor does not impact all dopamine neurons equally, as the SNc DA neurons are more vulnerable than the VTA, mirroring selective vulnerability characteristic of Parkinson’s disease. This is consistent with a large body of literature that SNc dopamine neurons are less capable of handling large energetic and calcium loads compared to neighboring VTA neurons, and the finding that chronically altered activity is sufficient to drive this preferential loss is novel.

      In addition, we are not aware of prior studies that have chronically activated DREADDs to produce neurodegeneration. Other studies have shown that acute excitotoxic stressors can produce neuronal degeneration, but the chronic increase in activity is central to our approach.

      In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript.

      As discussed in greater detail in the results section below, our data suggests this may not be a prominent feature in our model. However, we cannot rule out a contribution of depolarization block, and will expand on the discussion of this possibility in the revised manuscript.

      The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      We completely agree that evidence of increased dopamine neuron activity from human PD patients is lacking and the existing data are difficult to interpret without human controls. However, as we outline in the manuscript, multiple lines of evidence suggest that the activity level of dopamine neurons almost certainly does change in PD. Therefore, it is very important that we understand how changes in the level of neural activity influence the degeneration of DA neurons. In this paper we examine the impact of increased activity. Increased activity may be compensatory after initial dopamine neuron loss, or may be an initial driver of death (Rademacher & Nakamura, Exp Neurol 2024, PMID: 38092187). Beyond what is already discussed in the manuscript, additional support for increased activity in PD models include:

      - Elevated firing rates in asymptomatic MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488)

      - Increased frequency of spontaneous firing in patient-derived iPSC dopamine neurons and primary mouse dopamine neurons that overexpress synuclein (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060)

      - Increased spontaneous firing in dopamine neurons of rats injected with synuclein preformed fibrils compared to sham (Tozzi et al., Brain 2021, PMID: 34297092)

      We will include and further discuss these important examples in our revision.

      Similarly, in future studies, it will also be important to study the impact of decreasing DA neuron activity. There will be additional levels of complexity to accurately model changes in PD, which may differ between subtypes of the disease, the disease stage, and the subtype of dopamine neuron. Our study models the possibility of chronically increased pacemaking, and interpretation of our results will be informed as we learn more about how the activity of DA neurons changes in humans in PD. We will discuss and elaborate on these important points in the revision.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      We agree that the findings of Hollerman and Grace support compensatory changes in dopamine neuron activity in response to loss of dopamine neurons, rather than informing whether dopamine neuron loss can also be an initial driver of activity. We will clarify this point in our revision. In addition, the results of other studies on this point are mixed: a 50% reduction in dopamine neurons didn’t alter firing rate or bursting (Harden and Grace, J Neurosci 1995, PMID: 7666198; Bilbao et al, Brain Res 2006, PMID: 16574080), while a 40% loss was found to increase firing rate and bursting (Chen et al, Brain Res 2009. PMID: 19545547) and larger reductions alter burst firing (Hollerman & Grace, Brain Res 1990, PMID: 2126975; Stachowiak et al, J Neurosci 1987, PMID: 3110381). Importantly, even if compensatory, such late-stage increases in dopamine neuron activity may contribute to disease progression and drive a vicious cycle of degeneration in surviving neurons. In addition, we also don’t know how the threshold of dopamine neuron loss and altered activity may differ between mice and humans, and PD patients do not present with clinical symptoms until ~30-60% of nigral neurons are lost (Burke & O’Malley, Exp Neurol 2013, PMID: 22285449; Shulman et al, Annu Rev Pathol 2011, PMID: 21034221).

      Other lines of evidence support the potential role of hyperactivity in disease initiation, including increased activity before dopamine neuron loss in MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488), increased spontaneous firing in patient-derived iPSC dopamine neurons (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060), and increased activity observed in genetic models of PD (Bishop et al., J Neurophysiol 2010, PMID: 20926611; Regoni et al., Cell Death Dis 2020,  PMID: 33173027).

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      We agree that a discussion of hyperactivity, calcium, and neurodegeneration would benefit the introduction. While we briefly discuss calcium and neurodegeneration in the discussion, we will expand on this literature in both the introduction and discussion sections. We will carefully review and contextualize our work within existing frameworks of calcium and neurodegeneration (e.g. Surmeier & Schumacker, J Biol Chem 2013, PMID: 23086948; Verma et al., Transl Neurodegener 2022, PMID: 35078537). We believe that the novelty of our study lies in 1) a chronic chemogenetic activation paradigm via drinking water, 2) demonstrating selective vulnerability of dopamine neurons as a result of altering their activity/excitability alone, and 3) comparing mouse and human spatial transcriptomics.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      We do report the input resistance in Supplemental Figure 1C, which was unchanged in CNO-treated animals compared to controls. We did not report the resting membrane potential because many of the DA neurons were spontaneously firing. However, we will report the initial membrane potential on first breaking into the cell for the whole cell recordings in the revision, which did not vary between groups. This is still influenced by action potential activity, but is the timepoint in the recording least impacted by dialyzing of the neuron by the internal solution. We observed increased spontaneous action potential activity ex vivo in slices from CNO-treated mice (Figure 1D), thus at least under these conditions these dopamine neurons are not in depolarization block. We also did not see strong evidence of changes in other intrinsic properties of the neurons with whole cell recordings (e.g. Figure S1C). Overall, our electrophysiology experiments are not consistent with the depolarization block model, at least not due to changes in the intrinsic properties of the neurons. Although our ex vivo findings cannot exclude a contribution of depolarization block in vivo, we do show that CNO-treated mice removed from their cages for open field testing continue to have a strong trend for increased activity for approximately 10 days (S1E).  This finding is also consistent with increased activity of the DA neurons. We will add discussion of these important considerations in the revision.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      We thank the reviewer for this insightful comment, and we agree that this is a caveat of our mCherry quantification. Quantitation of the number of mCherry+ DA neurons specifically informs the impact on transduced DA neurons, and mCherry appears to be less susceptible to downregulation versus TH. As the reviewer points out, it carries the caveat that there is some variability between injections. Nonetheless, we believe that it conveys useful complementary data. As suggested, we will discuss this caveat in our revision. Note that mCherry was not quantified at the two-week timepoint because there is no loss of TH+ cells at that time.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      We agree that the stereology experiments were performed on relatively small numbers of animals. Combined with the small effect size, this may have contributed to the post-hoc tests showing a trend of p=0.1 for both the TH and mCherry dopamine cell counts in the SN at 4 weeks. As part of the planned experiments for our revision, we will perform an additional stereologic analysis to further assess the loss of SNc dopamine neurons. We will also review and ensure the images are representative.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      We thank the reviewer for this comment. We understand that this method of comparing absolute values is unconventional. However, these animals were tested concurrently on the same system, and a clear effect on the absolute baseline was observed. We will include a caveat of this in our discussion. Panel D of this figure shows the raw, uncorrected photometry traces, whereas panel E shows the isosbestic corrected traces for the same recording. In panel E, the traces follow time in ascending order. We will also include frequency and amplitude data for these recordings.   

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focusing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      We will review the expression of activity-related genes in our dataset, although we must keep in mind that these genes may behave differently in the context of chronic activation as opposed to acutely increased activity. We will also include experiments assessing striatal dopamine levels by HPLC in the revision.

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Our mouse model and human PD progress over distinct timescales, as is the case with essentially all mouse models of neurodegenerative diseases. Nonetheless, in our view there is still great value in comparing gene expression changes in mouse models with those in human disease. It seems very likely that the same pathologic processes that drive degeneration early in the disease continue to drive degeneration later in the disease. Note that we have tried to address the discrepancy in time scales in part by comparing to early PD samples when there is more limited SNc DA neuron loss. Please note the numbers of DA neurons within the areas we have selected for sampling (Figure at right). Therefore, we can indeed use spatial transcriptomics to compare dopamine neurons from mice with initial degeneration and patients where degeneration is ongoing during their disease.

      Author response image 1.

      Violin plot of DA neuron proportions sampled within the vulnerable SNV (deconvoluted RCTD method used in unmasked tissue sections of the SNV).

      Control and early PD subjects.

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      Our model utilizes hM3Dq-DREADDs that function by increasing intracellular calcium to increase neuronal excitability, and our results show increased Ca2+ by fiber photometry and changes to Ca2+-related genes, strongly suggesting a causal relation and crucial role of calcium in the mechanism of degeneration. However, we agree that we have not experimentally proven this point, as we acknowledged in the text. Additionally, we have planned revision experiments involving chronic isradipine treatment to further test the role of calcium in the mechanism of degeneration in this model.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      As discussed, we can sample SN DA neurons in early PD (see figure above), and in our view there is great value for such comparisons. We agree that discussion of appropriate caveats is warranted and this will be clearly addressed in the revision.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis.

      As discussed above, our analyses of DA neuron firing in slices and open field testing to date do not support a prominent contribution of depolarization block with chronic CNO treatment. However, we cannot rule out this hypothesis, therefore we will include additional electrophysiology experiments and add discussion of this important consideration.  

      Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      As discussed above, while increases in dopamine neuron activity may be compensatory after loss of neurons, the precise percentage required to induce such compensatory changes is not defined in mice and varies between paradigms, and the threshold level is not known in humans. We also reiterate that a compensatory increase in activity could still promote the degeneration of critical surviving DA neurons, whose loss underlies the substantial decline in motor function that typically occurs over the course of PD. Moreover, there are also multiple lines of evidence to suggest that changes in activity can initiate and drive dopamine neuron degeneration (Rademacher & Nakamura, Exp Neurol 2024). For example, overexpression of synuclein can increase firing in cultured dopamine neurons (Dagra et al., NPJ Parkinsons Dis 2021, PMID: 34408150) while mice expressing mutant Parkin have higher mean firing rates (Regoni et al., Cell Death Dis 2020,  PMID: 33173027). Similarly, an increased firing rate has been reported in the MitoPark mouse model of PD at a time preceding DA neuron degeneration (Good et al., FASEB J 2011, PMID: 21233488). We also acknowledge that alterations to dopamine neuron activity are likely complex in PD, and that dopamine neuron health and function can be impacted not just by simple increases in activity, but also by changes in activity patterns and regularity. We will amend our discussion to include the important caveat of changes in activity occurring as compensation, as well as further evidence of changes in activity preceding dopamine neuron death.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results. The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

      While our model demonstrates classic excitotoxic cell death pathways, we would like to emphasize both the chronic nature of our manipulation and the progressive changes observed, with increasing degeneration seen at 1, 2, and 4 weeks of hyperactivity in an axon-first manner. This is a unique aspect of our study, in contrast to much of the previous literature which has focused on shorter timescales. Thus, while we will revise the discussion to more comprehensively acknowledge previous studies of calcium-dependent neuron cell death, we believe we have made several new contributions that are not predicted by existing literature. We have shown that this chronic manipulation is specifically toxic to nigral dopamine neurons, and the data that VTA dopamine neurons continue to be resilient even at 4 weeks is interesting and disease-relevant. We therefore do not want to use findings from other neuron types to draw assumptions about DA neurons, which are a unique and very diverse population. We acknowledge that as with all preclinical models of PD, we cannot draw definitive conclusions about PD with this data. However, we reiterate that we strongly believe that drawing connections to human disease is important, as dopamine neuron activity is very likely altered in PD and a clearer understanding of how dopamine neuron survival is impacted by activity will provide insight into the mechanisms of PD.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      We agree with the reviewer that the conclusions would be greatly solidified with more direct mechanistic validation. However, we are unable to conduct more experimentation as the grant is finished and the Sernagor lab is in the process of being shutdown, after the unexpected passing of the PI.

      In order to make clearer that this mechanism was found in retinal tissue, not CNS, we have moved any mention of the implications of our work to a broader CNS mechanism to the discussion section. We have also added text into the discussion highlighting the need for more mechanistic investigation to uncover the full extent of the developmental processes described herein, see Line 413.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      We have been clear to only present our findings as correlational as we were unable to fully explore the causational nature within the mechanisms presented. In the discussion, we have used published evidence and experimental papers to bolster our understanding of the causal aspects of this research. We have also included sections of text to address what experimentation is be required to examine the causal interactions more directly, see Line 413.

      Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

      We thank the reviewer for their insightful suggestions for further experimentation to bolster the research. Initially, we utilised pharmacological interventions as they provided acute and quick answering of the research question. At the outset of the research, we were not certain that purinergic release through PANX-1 channels was the mediator for the developmental mechanisms described. We tested a wide variety of specific agonists and blockers before seeing any profound effects on wave generation. These agonists and antagonists have been used before and are proven to deliver reliable results. In addition, since the ACCs had never been reported before we were unsure if a knockout animal would display the same anatomical phenotype. Furthermore, it is known that knockout mouse lines, especially connexin and hemichannel pores, do not lose function but rather have other isoforms or compensation mechanisms which can substitute the original function. For the retina, for example, it was shown that Cx36 can functionally replace Cx45 after Cx45 KO (Frank et al, 2010).

      We agree that while direct mechanistic validation would significantly reinforce the arguments, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      In order to address the omission of mechanistic validation in the paper we have added text into the discussion highlighting the need deeper investigation in the causality of the developmental processes described herein, see Line 413.

      M. Frank et al., Neuronal connexin-36 can functionally replace connexin-45 in mouse retina but not in the developing heart, J. Cell Sci. 123, 3605 (2010).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      General and major comments

      (A) Introduction

      (1) The introduction is currently quite extensive. I recommend streamlining the background information to more directly frame the study's core objectives.

      We have reduced the background information contained in the introduction to better align with the direct outputs of the study. However, as this research paper examines the interactions of multiple complex developmental processes a relatively in-depth introduction is needed to inform the reader of the salient points.

      (2) To improve clarity, it would be highly beneficial to conclude the introduction with a sequential summary of key observations in the order they are presented in the study. This would provide a clearer roadmap for the reader.

      We have reworked the end of the introduction to better align with the key observations in the order they are presented through the figures.

      (B) Results

      (3) The characterization of ACCs (Apoptotic Cell Clusters) would be more effective if separated from the description of their spatiotemporal occurrence with SVPs. Consider moving the microscopic description to a dedicated section or merging it with the subsequent paragraph on molecular identification.

      We thank the reviewer for their suggestion to reorganise the results section. However, we believe that highlighting the integrated nature of the ACC positioning and development of the vascular plexus is an important stepping off point for the reader. It allows us to highlight the original serendipitous discovery of the ACCs and their highly cohesive role in bridging multiple developmental processes.

      (4) Please explicitly state the specific retinal developmental stages (e.g., P3-P6) within the section regarding RNA sequencing, as this timing is critical for interpreting the transcriptomic data.

      We have amended the text to specify the postnatal days used for RNA sequencing, see Line 123

      (5) Regarding the scRNA-Seq data: were ACC-negative samples (isolated via FACS) also processed? A direct comparison between ACC+ and ACC- sampled microglia would significantly strengthen the claim that microglia are specifically attracted to ACCs. If these data are available, they would make an elegant and compelling addition to the manuscript.

      We thank the reviewer for this important suggestion and agree that direct comparison of ACC+ and ACC− microglia would further strengthen the study. Unfortunately, ACC− populations were not processed for scRNA-seq in the current study because of the prioritisation of the rare ACC+ population. Nevertheless, several independent observations support the conclusion that microglia are preferentially associated with ACCs, including: (i) the enrichment of microglia within ACC-containing regions observed histologically, (ii) the spatial proximity analyses shown in Figure 3, and (iii) the distinct transcriptional profile of ACC-associated microglia identified by scRNA-seq.

      We have also added a section to the discussion to highlight the need for a direct comparison of the ACC+ and ACC- transcriptomics profiles in future work, see Line 344.

      (6) For the Ca2+ -imaging experiments, please briefly describe the staining protocol and specify which cell types (e.g., RGCs) were labeled within the Results text to assist the reader's immediate understanding.

      We have added a short description of the labelling technique in the results, see Line 228

      (7) The manuscript notes that MEA waves are evident in the graphs of Figure 7, but the raw wave data or representative traces are not shown. Including these (similar to those of the imaging waves) would provide necessary visual verification of the physiological phenomena described.

      We have added a supplementary figure 3 which details stage 2 retinal waves recorded using MEAs.

      (C) Interpretations and Logic

      (8) The finding ' ...Wholemount staining revealed a broad centro-peripheral gradient of apoptosis; however, this apoptotic annulus was positioned more peripherally than the ACCs, SVP, and Hmox1-positive microglia (Figure 4B)' seems to contradict the HMOX1/Yo-Pro-1 stained microglia. It is not clear whether the authors aim to prove that this particular set of microglia phagocytise RGCs, or another set that lines up better with dying cells and does not show up on the HMOX-1 label. I believe the authors intend to show the time difference of the two events - cells dying and HMOX-1 microglia appear at the site later. I believe the logic is good; it may need a sentence pointing this out at the end of this paragraph.

      We have added a statement in Line 182 which clarifies our intent to show that the Hmox1 microglia phagocytose the dying RGCs after they initiate apoptotic mechanisms.

      (9) If the authors intend to demonstrate a temporal lag between cell death and the appearance of HMOX1+ microglia as evidence of causality, a concluding sentence to this effect would greatly clarify the logic of this paragraph.

      We have added a concluding sentence to the paragraph in Line 190 which indicates a causal link between RGC cell death and appearance of hmox1 positive microglia.

      (D) Figures and Presentation

      (10) The blood vessel staining in the final panel of Figure 1B is currently quite faint. Increasing the brightness/contrast for this panel would allow the reader to better appreciate the underlying architecture.

      We have updated the panel in Figure 1B to match the brightness of the others of that series.

      (11) Given that the peripherality of events in Figure 7 suggests a specific sequence, the authors should consider adding a summary timeline (P3-P6). A plot using curves (mean or median values), color-coded to match the corresponding events, would provide a much-needed visual synthesis of the data.

      We agree with reviewer that the D1/2 metrics would benefit from more clarification to show the timeline of development more clearly. We have added another panel to Figure 7, which shows mean/standard deviation plots for each developmental measure using D1/2 as timelines. This allows the reader to better compare the progression of centrifugal spread more clearly.

      (12) Please ensure that graph labels and axis titles are uniform in size across all figures. e.g., the labels in Figure 6 and several other graphs are currently too small to be legible in the PDF; these should be enlarged for better accessibility.

      We have fixed the labels and axis titles to maintain readability across the paper

      Minor comments

      (1) In line 125, there is a missing closing parenthesis after the reference to Figure 1C.

      We have fixed this error

      (2) The specific algorithm used for the unsupervised cluster analysis has not been identified in the text. Please specify whether k-means, Louvain, or another method was employed to ensure reproducibility.

      We have reworked the section detailing the cluster analysis to make clear we used Louvain-based clustering, see line 475.

      (3) While it is appropriate to leave comprehensive technical details for the Methods section, a brief conceptual explanation of the D1, D2, and D3 metrics should be included in the Results text to aid general comprehension.

      We have added a brief description of the D1/2 and D1/3 metrics when they are first mentioned in the results section, see line 243.

      (4) In the Figure 7 schematic, the representation of the starburst amacrine cell should be revised to more accurately reflect its well-characterized morphology (e.g., thin primary and gradually thickening higher-order dendrites).

      We have changed the SAC representation to better match the characteristics of that cell type.

      (5) Throughout the manuscript (e.g., in lines 293-294), it is claimed that RGC apoptosis promotes the expression of PANX-1 hemichannels. While the data effectively demonstrate the release of purinergic molecules (e.g., ATP) via PANX-1 from dying cells, the evidence for an actual upregulation or increase in PANX-1 protein/mRNA levels is not explicitly shown. Please clarify whether the findings suggest increased activity of existing channels or a true increase in expression. If the latter is not empirically supported, the phrasing should be adjusted to reflect functional activation rather than de novo expression.

      We agree with the author that our research shows a functional increase in PANX-1 and we have adjusted the language to match. In the introduction and discussion, we provide published evidence and experimental papers which describe the upregulation of the PANX-1 molecule in dying RGCs.

      Reviewer #2 (Recommendations for the authors):

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Furthermore, the authors demonstrate the initiation of Ca2+ waves occurs at the leading edge of vascular growth in the developing retina. This manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth. The main weakness of the manuscript is the overreliance on only two pharmacological agents to test their central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP releases (i.e., BAX-KO, Panx-1 KO). In addition, the following comments should also be addressed:

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Specific comments:

      (1) Line 128: Why would ACCs be involved in SVP guidance if they are trailing the leading edge of the vasculature? Would it be the other way around, where the leading edge of the vasculature would be trailing the ACCs?

      At this point in the paper we are suggesting that the highly stereotyped position of the ACCs under the leading edge of the SVP indicates that they have a mechanistic involvement in SVP growth. Not that they are the direct cause of the expansion. As the paper progresses, we make clear that contrary to our original hypotheses which state the ACCs may cause or control the integrated development of the retina, they are a hallmark of the apoptotic RGCs in the periphery which are the chemogenic beacons for vascular growth being ‘decommissioned’ by the microglia which fine-tune vascular growth and create the ACCs.

      (2) Line 137: Please state in the text and figure legend, at what age ACCs were isolated from the retina.

      We have added the relevant information to Line 123

      (3) Line 154: Why didn't RGCs form their own cluster? Why do the RBPMS+ cells appear across the entire dataset (Figure 2G)? Have previous investigations also shown engulfed cell transcriptomes appearing in the microglia clusters using scRNAseq?

      Previous transcriptomic studies have demonstrated that phagocytic microglia can encapsulate transcripts originating from neurons and other neural cell types, which are detectable by RNA‑seq despite not belonging to a common microglial genetic signature. For example, Solga et al. showed that CNS microglia contain neuronal and oligodendrocyte‑specific mRNAs that localise within microglia but are not translated. This research group interpreted that this RNA is acquired through phagocytosis or macropinocytosis of surrounding neural cells (Solga et al., 2015). Similarly, in zebrafish, synapse‑engulfing microglia identified in situ display neuronal and synaptic gene expression in single‑cell RNA‑seq profiles, consistent with engulfed neuronal material contributing to the detected transcriptome (Sliva et al., 2021). In line with these observations, and given our FACS strategy enriching autofluorescent ACCs rather than intact RGCs, we interpret the widespread Rbpms expression across ACC‑associated clusters as an expected consequence of microglial engulfment of apoptotic RGCs, rather than evidence for a distinct population of viable RGCs that failed to form a separate cluster.

      Solga, A.C., Pong, W.W., Walker, J., Wylie, T., Magrini, V., Apicelli, A.J., Griffith, M., Griffith, O.L., Kohsaka, S., Wu, G.F. and Brody, D.L., 2015. RNA‐sequencing reveals oligodendrocyte and neuronal transcripts in microglia relevant to central nervous system disease. Glia, 63(4), pp.531-548.

      Silva, N.J., Dorman, L.C., Vainchtein, I.D., Horneck, N.C. and Molofsky, A.V., 2021. In situ and transcriptomic identification of microglia in synapse-rich regions of the developing zebrafish brain. Nature communications, 12(1), p.5916.

      We have incorporated this information into the discussion in Line 329

      (4) Figure 4C-H: The data in the figure would be strengthened if the authors added quantification for their co-localization images.

      We thank the reviewer for this important suggestion and agree that quantification of the co-localisation would further strengthen the study. Unfortunately, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      (5) Figure 5A-I: The error bars are quite large. The figure legend says they represent SEM, but are you sure they don't represent standard deviation (SD)?

      The large SEM bars reflect substantial biological variability across retinas and across individual microglia, which is expected for morphometric measures such as circularity, perimeter, branch number and total skeleton length, as well as for counts of rare cell populations (Hmox1+ microglia, YO‑PRO‑1+ cells and double‑positive cells). Importantly, despite this variability, the effects of probenecid and PSB‑0739 on microglial morphology and on the frequencies of apoptotic and double‑positive cells remain statistically robust in our non‑parametric ANOVA and post‑hoc tests, as indicated by the reported P‑values.”

      For panels where the distributions were clearly non‑Gaussian, we used non‑parametric statistics (Kruskal‑Wallis ANOVA), reporting medians and 95% confidence intervals, and we retained SEM in Figure 5 for consistency with the original plotting routine while clarifying this choice in the legend and Methods.

      (6) Figure 5: What are these measurements made at? Please add the age to the results section and the figure legend.

      We have added the postnatal day of the animals used.

      (7) Line 236-239: There is no mention of the use of Probenecid in the text and in Figure 6B. This treatment of Ca2+ waves should be mentioned before line 245.

      We have added a brief description of probenecid application in Line 227.

      (8) Figure 6: The order of this figure and the results section may be clearer if panels C-F were switched with panels G-J.

      We thank the reviewer for their suggests to improve the flow of the results section and figure 6. We have swapped the panels as suggested and amended the text to fit the new flow.

      (9) For the discussion section: Why does probenecid only affect the wave initiation at the P3 timepoint, but not later timepoints, P4-P6.

      We have added some text in the discussion to explain our findings of differential effects of PANX-1 blockade across the P3-6 timeline. See Line 404.

      Editorial revisions:

      (1) Line 92-96: Citation needed.

      We have added appropriate citations for this section.

      (2) Figure 1F: Please consider changing the color scheme from green/red to green/magenta for colorblind readers.

      We have changed the image LUT

      (3) Figure 3A: This figure may benefit from separating out each channel separately (Iba1 and ACC) and then having a "merge" panel.

      We have separated out the panels in Figure 3A

      (4) Figure 4E, H: There is no label for the immunomarkers in Figure 4H. There is no inset in Figure 4E as mentioned in the figure legend. However, it appears that Figure 4H is a magnified image of Figure 4G and not Figure 4E.

      We have fixed the error in the figure legend and included an inset indicator in Figure 4G. We have added immunomarkers to Figure 4H

      (5) Line 198: "SAC" should be "SACs".

      We have fixed this error

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e<sup>-</sup>/Å<sup>2</sup>), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1°, 2°, 3°, 5°, and 10°) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3° tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      The conclusions of this paper are mostly well supported by data, but some aspects of data analysis need to be clarified and/or extended, including:

      (1) Line 109: The authors state that the tilt range was kept at ± 60° relative to the lamella plane. Assuming a typical lamella pre-tilt of ~10°, the absolute stage tilt would approach its mechanical limit. Two clarifications would be appreciated: (a) What was the average pre-tilt across all lamellae? (b) How many dark tilt images, if any, were excluded during tomogram reconstruction?

      We thank the reviewer for asking for further clarification. For all our datasets, the pre-tilt of the stage was +8° with the lamella untilted under the e-beam, resulting in a tilt range of -52° to + 68°, thereby not reaching the mechanical limit, which is 70° for our microscope stage.

      Regarding “dark tilt images”, for most datasets, we did not need to remove many tilt images. However, we now noticed notably more absence of images from higher tilt values for the 1° dataset (see SFig 1). When analysing further, we noticed that for this dataset, we did not actively remove many images prior to tomogram reconstruction, but rather that they were not acquired in the first place by SerialEM. During acquisition, SerialEM performs various safeguarding checks that can abort the acquisition of a tilt series (or of a single branch). As this seems predominantly a problem for the 1-degree tilt-increment dataset, we have decided to add this to the manuscript as follows, including the figure as new SFig 1.

      In the main text:

      “For most datasets, image acquisition was largely complete, with the exception of the 1-degree dataset, which showed a markedly higher proportion of missing images at high tilt angles (SFig. 1). Closer inspection revealed that many of these images were not acquired, as SerialEM applies built-in safeguards (e.g. autofocus inconsistency or insufficient image counts) that can abort a tilt-series branch before completion.”

      (2) Line 148: "When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B)." It would be helpful to specify which comparisons are statistically meaningful (e.g. Mann-Whitney U test?). While the difference between 1° and 2° appears pronounced, the differences between 2°, 3°, and 5° seem minimal. From my point of view, reporting the mean SNR values +/- standard deviations for each condition would already indicate some significance. Furthermore, since SNR is expected to depend on lamella thickness, it should be clarified whether the average lamella thickness is comparable across the five datasets.

      We have now calculated the mean and standard deviation of the tomogram SNR, as follows:

      Author response table 1.

      Furthermore, we performed a statistical significance test. Kruskal-Wallis test confirmed significant differences in SNR across tilt increments (H=270.97, p<0.001). Pairwise Mann-Whitney U tests with Bonferroni correction revealed significant differences between all pairs except 2° and 3° (p=0.093), suggesting these two conditions indeed yield comparable SNR.

      Lastly, we have now added data regarding local lamella thickness for all tilt-series used in the study, as displayed in SFig. 4.

      We incorporated this in the manuscript as follows.

      In the main text:

      “When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B and Supplementary Note), whilst showing a similar lamella thickness distribution (see SFig. 4A).”

      and:

      “Firstly, we selected ca. 20 tomograms per condition, based on tomogram content and local lamella thickness [31] (for more details, see Methods and SFig. 4B).”

      As a supplementary note:

      “As the tomogram SNR distribution of particularly the 2° and 3° dataset showed similar SNR distributions, we performed formal significance testing for the data in this panel (see Fig. 2B). Kruskal-Wallis test confirmed significant differences across conditions (H=270.97, p<0.001); pairwise Mann-Whitney U tests with Bonferroni correction revealed all pairs were significantly different except 2° vs. 3° (p=0.093), indicating comparable SNR for these two tilt increments.”

      And, adding the test in the Methods:

      “To quantify differences in signal-to-noise ratio (SNR) across tilt increment conditions, a non-parametric Kruskal-Wallis test was performed as an omnibus test of the null hypothesis that all groups are drawn from the same distribution. Because SNR distributions were not assumed to be normal, and sample sizes differed across conditions, non-parametric tests were used throughout. Following the omnibus test, all 10 pairwise comparisons between conditions were assessed using two-sided Mann-Whitney U tests. To control for multiple comparisons, raw p-values were adjusted using the Bonferroni correction (multiplied by the number of comparisons, n=10, capped at 1.0). Statistical significance was defined as a Bonferroni-corrected p-value below 0.05. All analyses were performed in Python using the scipy.stats module.”

      (3) Line 167: "Indeed, the variation in maximum resolution correlates with lamella thickness across all datasets (see Fig. 2F)." The reported R<sup>2</sup> values of 0.30 (1°), 0.38 (2°), 0.66 (3°), 0.61 (5°), and 0.60 (10°) reveal a notably weak linear relationship for the finer tilt increments. It is also difficult to assess whether the lamella thickness distributions are comparable across conditions from the current figures - visually, the 1° dataset appears to be based on thinner lamellae, while the 10° dataset appears to include thicker samples. A histogram of lamella thickness distributions for each condition, provided as supplementary material, would greatly aid interpretation. Given this thickness dependency, reporting mean +/- standard deviation of lamella thickness per condition is highly appreciated.

      We have added the full lamella thickness distribution per dataset now in SFig. 4.

      The apparent weaker relationship between resolution fit and local lamella thickness for the 1 dataset seems to be largely apparent to the few very thin data points in this data (for more clarity, see the same data plotted separately in Author response image 1). We speculate that this is due to even less signal in these very thin and very low-dose images.

      Author response image 1.

      (4) Figure 4: It should be specified which tomogram subsets were used for the Rosenthal-Henderson analysis, whether lamella thickness was taken into account in the subset selection, and whether ribosomes too close to the lamella edges were excluded. Finally, linear fits should be displayed across the full x-axis range for all tilt increments to facilitate direct visual comparison.

      We have described the process of tomogram subset selection in detail in the Methods section Template matching and 3D classification. To further add clarity, we have incorporated the local lamella distribution plots for the full data, as well as specifically for the tomograms subjected to TM and STA in SFig. 4.

      Regarding the linear fits, we respectfully disagree with this suggestion. Displaying the linear fits only over the range used for their calculation avoids implying that the linear relationship extends beyond the measured data, and in our view produces a clearer figure.

      (5) General: Were ribosomes located at the lamella edges excluded from the analysis? As demonstrated in the authors' own prior work (Tuijtel et al., Science Advances, 2024), Ga-FIB milling induces structural damage at the lamella surfaces. To exclude the influence on the STA results, particles near the lamella edges should be removed prior to analysis, and the criteria for this exclusion should be stated explicitly.

      We have not excluded any ribosomes from close to the surface. As we still treated all data the same for each condition, we anticipate that the results of the comparison reported here still hold true.

      The aim of the authors was to provide the cryo-ET community with an evidence-based recommendation for the choice of tilt increment, and they largely succeeded in this goal. The identification of 3° as a practical optimum - balancing sufficient dose per tilt image for effective per-particle refinement with fine enough angular sampling for accurate tilt-series alignment - is well supported by the data and consistent across the multiple quality metrics employed. The conclusion that coarser increments (5° and 10°) compromise tomogram quality, template matching accuracy, and STA resolution is robust and clearly demonstrated. However, the conclusion rests entirely on a single biological system using ribosomes as the sole molecular target, which are exceptionally favourable due to their abundance, size, and electron contrast. Whether the identified optimum holds for smaller, lower-abundance, or lower-contrast targets remains an open question.

      In future, it would be particularly interesting to test whether emerging tilt interpolation strategies, such as cryoTIGER, which is particularly intriguing, can effectively compensate for coarser experimental angular sampling in post-processing. Here, the optimal experimental increment may shift, and the interaction between these two approaches represents a promising direction for future work. More broadly, as cryo-ET datasets grow larger and public repositories expand, the practical tradeoffs between acquisition time, data storage, and structural quality identified here will become increasingly relevant to the field.

      We agree with the reviewer and thank them for this positive assessment. An interesting note to the use of cryoTIGER in particular is that it uses already aligned tilt-series as an input, and it therefore is unlikely to overcome severe alignment issues associated with large tilt-increments.

      Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Overall, the manuscript is clearly written, and the conclusions are well supported by the data presented. I have no major concerns. There are some minor points that the authors should address, including:

      (1) The phrase "electron optical density distribution" (line 31, Introduction) should be revised to "electrostatic potential" or "Coulomb potential distribution," which more accurately reflects what is measured in cryo-EM/ET.

      We thank the reviewer for this correction and have adjusted it in the text:

      “It captures the 3-dimensional (3D) electrostatic potential of the specimen under scrutiny and enables the structural analysis of macromolecular complexes within their native context.”

      (2) The authors state that the maximum tolerable electron dose is approximately 100-150 e<sup>-</sup>/Å<sup>2</sup> (line 34, Introduction). This is an oversimplification, as bacterial specimens, for example, have been shown to tolerate doses of 200 e<sup>-</sup>/Å<sup>2</sup> or higher (see Breigel et al., PNAS, 2009; https://www.pnas.org/doi/10.1073/pnas.0905181106#T1). The statement should be revised to reflect this variability.

      We adjusted this statement to now read:

      “One of these is the maximum electron dose that can be applied to biological specimens before irreversible damage occurs, which is about 100-150 e-/Å2 for most eukaryotic cells.”

      (3) Lines 56-57: The authors do not cite their own prior work benchmarking tilt-series acquisition strategies on in vitro samples. This earlier study provides important context and should be referenced and briefly discussed.

      We assume the reviewer meant this study: Turonova et al., Nat. Comms. (2020). We have now added and discussed this reference as follows:

      “Since accumulated radiation dose progressively degrades high-resolution information, this motivated the development of the dose-symmetric tilt scheme, which prioritizes acquisition of low-tilt images early to better preserve high-resolution information [11, 16].”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 159: "Surprisingly though, the resolution to which the CTF was fitted was similar for all conditions, despite an 8-fold increase in dose (see Fig. 2B, E)." The reference to Figure 2B at this point is unclear.

      We thank the reviewer for pointing this out, we have removed the reference to panel B.

      (2) Line 162: "As the data shown in Fig. 2D-F pertains to images of the untilted specimen, ..." For clarity, this should also be stated explicitly in the figure caption.

      We have added this to the figure legend; “both estimated with Gctf on projection images of the sample at effective zero-tilt position.”

      (3) Figure 3: These are compelling results, and the 3D classification outcomes provide an excellent visual representation of the quantitative data shown in Figure 3. Including (some or all) of the initial five classes in the figure would further strengthen this already convincing presentation. Additionally, applying a uniform extraction threshold (e.g., z-score of 3.75 or 5) across all tilt increments would facilitate a more direct comparison. But this is really a minor remark, the authors and the editors may judge if the current presentation is already sufficient.

      We have adjusted Figure 3 according to the reviewer’s recommendation:

      We referenced this in the main text as:

      “To ensure similar data processing strategies for all conditions, extraction thresholds were also lowered for the 1-, 2- and 3-degree conditions, and 3D classification was performed to filter out the junk particles (see Fig. 3 A and SFig. 8). “

      In order to directly compare the particle extraction, we have already carried out such a uniform extraction threshold, with a z-score threshold of 5 (apart for the 10-degree data, where this was not possible). This extraction was then used for the 3D classification that led to the TM analysis and further STA investigations.

      Reviewer #2 (Recommendations for the authors):

      (1) Supplementary Figure 1: To improve accessibility for a broader readership, the authors should annotate or highlight the key organelles and protein complexes visible in the tomographic slices.

      We thank the reviewer for this suggestion, but adding arrowheads made this figure too crowded in our opinion. Furthermore, for most of the panels, the mentioned features of interest is centred in the image panel, which should make identification straightforward.

      (2) 'In situ' should be in italics throughout the text.

      We have changed this.

    1. Author response:

      We sincerely thank the reviewers for their time, helpful critiques and overall positive evaluation of our work. We plan to investigate the points raised by the reviewers and look forward to submitting a revised manuscript that addresses concerns about therapeutic relevance, functional distinctions between the CI-intact and CI-suppressed adapted states, and PC regulation in this system.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

      We will improve these aspects of the presentation and reporting of statistical analyses in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the revised manuscript, we will clarify that the possibilities illustrated in figure 1 are not intended as mutually exclusive. We indeed find all of these signals intermingled. In terms of a more focused question amenable to hypothesis testing, this can be expressed as: for each dimension (laterality vs manuality) are the activity patterns in each area closer to those predicted by invariance or dependence, as compared to the other areas? The various statistical analyses in the paper all essentially boil down to testing this question. Broadly, the answer is yes: we see activity closer to the invariant prediction LOM, and activity closer to the dependent prediction in fl-M1 and fl-M2. Testing whether bimanual activity can be explained as a simple linear sum of the left and right unimanual might provide additional insight into this question, and this analysis will be presented in the revised manuscript.

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      The issue of orofacial movement confounds is an important one that we made a point of addressing in the discussion. There are three main points that we believe cast doubt on this as the most likely explanation for the effector-invariant representation in LOM.

      First, while we do not disagree that LOM has an important role in tongue and jaw control, there is plenty of evidence from mapping and behavioural studies (which we cite in the introduction) that it also plays a role in forelimb motor control as well.

      Second, while we cannot observe all orofacial movements, we have previously shown that the jaw is less active when the hands and LOM are most active (Barrett et al., 2024). Conversely, LOM firing is much lower during chewing, when the tongue and jaw are very active.

      Finally, the correlation between LOM firing and forelimb movements is not merely a coarse-grained one on the timescale of active manipulation vs passive holding phases. LOM firing closely tracks the position of the forelimb(s) on fast timescales and with near-zero lag, giving better decoding than from fl-M1 or fl-M2, as we have shown here and previously (Barrett et al., 2022). If we assume that LOM only encodes orofacial movements, then this result implies that orofacial movements correlate with forelimb position better than fl-M1 or fl-M2 firing correlates with forelimb position.

      The revised manuscript will include an expanded discussion to clarify these and related points.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      We will discuss this in the revised manuscript.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

      This limitation applies to the analyses relating to the transport-to-mouth movement (Figures 3-7). The population correlation structure and decoding analyses (Figures 8-10) consider activity throughout the full duration of food handling, which involves a much greater variety of movements (Barrett et al., 2020). Indeed, this was a major motivation for including these analyses. The revised manuscript will clarify this point.

      Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual).

      See our response to reviewer #2 above regarding the variety of movements. We agree that this a limitation for the unit-level and PCA analyses, but one that is alleviated by the population correlation and decoding analyses, which relate to complex ongoing movements.

      This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable.

      Important behavioural confounds include orofacial movements, non-specific movement initiation signals, and arousal. Orofacial movements we have discussed above in our response to Reviewer 2. Movement initiation signals would likely be transient and well-timed to movement onset, hence this may explain some of the effector-independent activity in fl-M1 and fl-M2 (consistent with e.g. (Kaufman et al., 2016)). However, we do not believe this to be the case in LOM as its activity is delayed and sustained relative to movement initiation. Regarding arousal, the mouse is actively engaged in consuming the food even when the hands are stationary, so there is no a priori reason to believe that arousal varies rapidly during the behaviour. Consistent with this, measurements of noradrenergic activity in the locus coeruleus during consumption suggest that arousal varies on slow timescales, on the order of seconds to tens of seconds (Sciolino et al., 2022). Such slow variation in arousal would not explain the rapid but condition-invariant changes in firing in any of the cortical areas studied here. The updated manuscript will include more detailed discussion of these points.

      Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been.

      Behavior tracking was performed with kilohertz temporal resolution and submillimeter spatial resolution.

      Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

      Our recording coordinates are based on the territory of corticospinal neurons retrogradely labeled from C6 spinal cord, medial to any layer 4 labelling, as reported in our previous study (Yamawaki et al., 2021). Thus we are confident in calling this area forelimb M1.

      References:

      Barrett, J. M., Martin, M. E., Gao, M., Druzinsky, R. E., Miri, A., & Shepherd, G. M. G. (2024). Hand-jaw coordination as mice handle food is organized around intrinsic structure-function relationships. The Journal of Neuroscience, 44(42), e0856242024.

      Barrett, J. M., Martin, M. E., & Shepherd, G. M. G. (2022). Manipulation-specific cortical activity as mice handle food. Current Biology, 32(22), 4842-4853.e6.

      Barrett, J. M., Tapies, M. G. R., & Shepherd, G. M. G. (2020). Manual dexterity of mice during food-handling involves the thumb and a set of fast basic movements. PLOS ONE, 15(1), e0226774.

      Kaufman, M. T., Seely, J. S., Sussillo, D., Ryu, S. I., Shenoy, K. V., & Churchland, M. M. (2016). The Largest Response Component in the Motor Cortex Reflects Movement Timing but Not Movement Type. eNeuro, 3(4).

      Sciolino, N. R., Hsiang, M., Mazzone, C. M., Wilson, L. R., Plummer, N. W., Amin, J., Smith, K. G., McGee, C. A., Fry, S. A., Yang, C. X., Powell, J. M., Bruchas, M. R., Kravitz, A. V., Cushman, J. D., Krashes, M. J., Cui, G., & Jensen, P. (2022). Natural locus coeruleus dynamics during feeding. Science Advances, 8(33), eabn9134.

      Yamawaki, N., Raineri Tapies, M. G., Stults, A., Smith, G. A., & Shepherd, G. M. (2021). Circuit organization of the excitatory sensorimotor loop through hand/forelimb S1 and M1. eLife, 10, e66836.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Weaknesses:

      (1) The mechanism underlying spontaneous di-decamer formation in the absence of glutamine is insufficiently explored and lacks quantitative biophysical validation.

      (2) Claims of decamer-only behavior in mutants rely solely on negative-stain EM and are not supported by orthogonal solution-based methods.

      We thank the reviewer for the summary and noting of the strengths. We agree that the evolutionary divergence of metabolic feedback in GS homologs is a fruitful avenue for future studies. With regard to the weaknesses, the di-decamer in the absence of glutamine only forms under high (higher than physiological) concentrations of enzyme. Our primary evidence for the mutant behavior was the lack of crosslinking (Figure 1E), with supplementary support from the negative stain. In the revised version we will soften the language to say “reduced” rather than “did not support” filament formation.

      Reviewer #2 (Public review):

      The authors set out to resolve the high-resolution structure of a glutamine synthetase (GS) decamer using cryo-EM, investigate glutamine binding at the decamer interface, and validate structural observations through biochemical assays of ATP hydrolysis linked to enzyme activity. Their work sits at the intersection of structural and functional biology, aiming to bridge atomic-level details with biological mechanisms - a goal with clear relevance to researchers studying enzyme catalysis and metabolic regulation.

      Strengths and weaknesses of methods and results:

      A key strength of the study lies in its use of cryo-EM, a technique well-suited for resolving large, dynamic macromolecular complexes like the GS decamer. The reported resolutions (down to 2.15 Å) initially suggest the potential for detailed structural insights, such as side-chain interactions and ligand density. However, several methodological limitations significantly undermine the reliability of the results:

      (1) Cryo-EM data processing: The absence of critical details about B-factor sharpening - a standard step to enhance map interpretability - is a major concern. For high-resolution maps (<3 Å), sharpening is typically applied to resolve side-chain features, yet the submitted maps (e.g., those in Figures 1D, 2D, and supplementary figures) appear unprocessed, with density quality inconsistent with the claimed resolutions. This makes it difficult to evaluate whether observed features (e.g., glutamine binding) are genuine or artifacts of unsharpened data.

      (2) Modeling and density consistency: The structural models, particularly for glutamine binding at the decamer interface, do not align with the reported resolution. The maps shown in Figure 2D and Supplementary Figure S7 lack sufficient density to confidently place glutamine or even surrounding residues, conflicting with claims of 2.15 Å resolution. Additionally, fitting a non-symmetric ligand (glutamine) into a symmetry-refined map requires justification, as symmetry constraints may distort ligand placement.

      (3) Biochemical assay controls: While the enzyme activity assays aim to link structure to function, they lack essential controls (e.g., blank reactions without GS or substrates, substrate omission tests) to confirm that ATP hydrolysis is GS-dependent. The use of TCEP, a reducing agent, is also not paired with experiments to rule out unintended effects on the PK/LDH system, further limiting confidence in activity measurements.

      Achievement of aims and support for conclusions:

      The study falls short of convincingly achieving its goals. The claimed high-resolution structural details (e.g., side-chain densities, ligand binding) are not supported by the provided maps, which lack sharpening and show inconsistencies in density quality. Similarly, the biochemical data do not robustly validate the structural claims due to missing controls. As a result, the evidence is insufficient to confirm glutamine binding at the decamer interface or the functional relevance of the observed structural features.

      Likely impact and utility:

      If these methodological gaps are addressed, the work could make a meaningful contribution to the field. A well-resolved GS decamer structure would advance understanding of enzyme assembly and ligand recognition, while validated biochemical assays would strengthen the link between structure and function. Improved data processing and clearer reporting of validation steps would also make the structural data more reliable for the community, providing a resource for future studies on GS or related enzymes.

      We disagree with the reviewer’s overall assessment.

      With regard to sharpening and resolution: we examined sharpened maps and in a revised version will present additional supplementary figures showing these maps side by side. We note that the resolutions reported are global and that the most interesting features are, of course, in the periphery and subject to conformational and compositional heterogeneity. We will include supplementary figures of core side chain densities that are more like what are expected by the reviewer in the revision. With regard to modeling: the apo filament and turnover filament datasets were handled nearly identically. The additional density is therefore likely not artefactual to the symmetry operator - however, the lower resolution in this region noted by the reviewer is worthy of further exploration. The maps are public and we think this is the most plausible interpretation of the density, which we based primarily on the biochemical data and will include more speculation in the version.

      With regard to the biochemical controls: we point the reviewer to Figure S1, which shows that omission of ammonia or glutamate in the wild-type (tagless) system removes any coupling of the reactions. We will perform the additional controls to publication quality in the revised version along with the TCEP control. We note that the reducing agent is present across all experiments, ruling out an effect on any specific result. The inclusion of TCEP is also very standard in other published uses of the Coupled ATPase assay (e.g. PMID: 31778111 and PMID: 32483380 by our first author)

      Additional context:

      Cryo-EM has transformed structural biology by enabling high-resolution analysis of large complexes, but its success hinges on rigorous data processing and validation steps that are critical to ensuring reproducibility. The challenges highlighted here are not unique to this study; they reflect broader issues in the field where incomplete reporting of methods can obscure the reliability of results. By addressing these points, the authors would not only strengthen their current work but also set a positive example for transparent and rigorous structural biology research.

      All the data is public and the reviewer or anyone is free to reinterpret the maps and models - and we encourage that rather than just an interpretation of our static figures. In addition, we will upload the raw micrograph data for the apo filament and turnover filament datasets to EMPIAR prior to submitting the revision.

      Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      (1) In Figure 2D and Supplementary Figure 7, the extra density observed between the two decamers does not appear to have the defining features of a glutamine. A less defined density may be expected given the nature of the complex, but even though mutagenesis assays were performed to support this assignment, none of these results constitutes direct and conclusive evidence for glutamine binding at this site. I would thus suggest showing the density maps at multiple contour thresholds to allow readers to also better evaluate the various small molecules under turnover conditions that cannot be well fitted based on this density map, helping to provide a more balanced interpretation of the results.

      (2) On the same point regarding the density for the enzyme under turnover conditions, more details should be provided about the symmetry expansion and classification performed, and also show the approximate ratio of reconstructions that include this density. Did you try symmetry expansion followed by focused classification, especially on the interface region?

      (3) The interface between the two decamers of the model needs to be double-checked and reassigned, especially for the residues surrounding the fitted glutamine. For example, the side chain of the Lys residue shown in the attached figure is most likely modeled incorrectly.

      We thank the reviewer for the feedback. As noted above, we will include supplemental figures that show maps at multiple thresholds and sharpening schemes. We noted in the manuscript and above that our interpretation here is based on integrating biochemical evidence alongside the density and will make that even more clear in the revised manuscript. The filaments +/- the putative glutamine density were processed nearly identically, but we will attempt various schemes of focused classification/symmetry expansion in the revision as well. However, we point out that there is extensive averaging there that makes modeling a bit trickier than expected given the global resolution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Limitation to Di-decamer Formation:

      Could the authors clarify why hGS, when visualized by cryo-EM, predominantly forms di-decamers rather than extended filaments? Since glutamine bridges the two decameric rings, one would expect this to promote further polymerization. It remains unclear why longer filaments are not observed under the conditions used. Moreover, data in Figure 1B-1D are not mixed with Gln, and the mechanism of the apo-form filament formation is not clearly discussed. If K52 and C53 are in charge of filamentation, it is supposed to form a long filament, not a stack of 2 decamers. The kinetics of wild type and mutations, including K52A and C53A, are different. This confused me as K52 and C53 don't participate in the reactions of glutamine synthesis. Does data reduce catalytic efficiency with K52A or C53A mutation suggest that di-decameric GS exhibits a greater catalytic turnover rate than the pure decameric GS?

      We thank the reviewer for pointing this out and it was indeed filament length that was a point of curiosity during the study. Prior to preprint, we repeated the freezing conditions under identical turnover conditions and with high protein concentration as reported and indeed found much longer filaments - pointing to the capacity of the system to form larger complexes. However, we only captured a couple of ‘screening’ images and did not collect a full second dataset. Therefore, we hypothesize that filament length may be stochastic based on the subtle differences of individual grid vitrification based on the assumption that decamers within a filament can freely and quickly exchange. However, as shown in the time-resolved cryoEM experiment, the fraction of particles that are characterized as participating in a filament form (length-agnostic metric), does not appear to be subject to individual grid vitrification conditions but rather by experimental conditions (concentration of reaction-derived glutamine).

      The filament formation in apo state was not further explored because the concentrations required to achieve filament formation in this case were supraphysiological.

      We have edited the discussion to emphasize this point more clearly:

      “While the enzyme concentrations to achieve robust filamentation in the absence of glutamine are much higher than observed in cells, the protein concentrations used in our time-resolved cryoEM experiments where filamentation is correlated with accumulation of glutamine are within the range of intracellular GS concentrations in S. cerevisiae (Engel et al. 2025) and human cell lines (Wiśniewski et al. 2014).”

      Regarding the ability of apo-GS to form filaments - we identified that these residues were important based on their structural location at the decamer: decamer interface (Figure 1D; away from the active site as pointed out) and because their individual mutation to alanine attenuated the ability to form higher order filaments (Figure 1E). Therefore, these mutants were crucial controls in the steady-state kinetic experiments reported (Figure 2E). Here, we used exogenous glutamine to seed/stabilize filaments because we identified glutamine as serving this function and, crucially, because we did not observe glutamine occupancy in the active site under turnover conditions (which would suggest an orthosteric feedback inhibition mechanism; Supplementary Figure 10). Under these conditions, the wild-type, filament-competent protein displayed a ~3-fold K<sub>M, ammonia</sub> increase with glutamine addition compared to no glutamine, which when considered in the context of the greater E-305 flap conformational heterogeneity speaks to a model of allosteric feedback inhibition. Importantly, K52A and C53A show no difference in K<sub>M, ammonia</sub> between the glutamine and no glutamine condition, suggesting that attenuation of filament formation at this interface via mutation, eliminates the kinetic deficit. Therefore, K52A and C53A are not product-inhibited in the same manner as wild-type GS.

      We have clarified the discussion to emphasize this result:

      “Importantly, point mutations of the interfacial residues do not show a K<sub>M, ammonia</sub> defect in the presence of glutamine, indicative of the importance of the filament form for product feedback.”

      (2) Origin of Di-decamer Formation in the Absence of Glutamine:

      While the manuscript demonstrates glutamine-stabilized filamentation, the spontaneous formation of di-decamers under apo conditions is not mechanistically explained. The observation of 10-mer, 20-mer, and 40-mer species in Figure 1B should be validated against molecular weight standards or through SEC-MALS. The inference of higher-order oligomers based solely on migration is insufficient. Additional characterization (e.g., SEC-MALS, AUC, or mass photometry) would clarify whether these assemblies are biologically relevant or incidental.

      We thank the reviewer for pointing out the low precision of preparatory size exclusion chromatography assignments of GS molecular weight filament depicted in Figure 1. We have included calibration standards and assignment in Supplementary Figure 1 and updated Figure 1 to include the ambiguity of these assignments in panel B. It is important to note that GS has historically been underestimated in size via these methods and was originally assigned as an octamer for which there was previous consensus (PMID: 10708854). The low concentration requirements of mass photometry preclude its use for this purpose and we are not in a position to do SEC-MALS or AUC for this. Hopefully, the negative stain, crosslinking, and cryo-EM results are sufficient to indicate that we have correlated signals with the correct species!

      (3) Validation of Interface Mutants as Decamer-only Species:

      K52A and C53A mutants are used to disrupt di-decamer formation and are shown by negative-stain EM to exist as decamers. While supportive, this is qualitative. The inclusion of quantitative biophysical data (e.g., SEC-MALS or mass photometry) would more convincingly demonstrate that these mutants do not transiently assemble into higher-order oligomers. Furthermore, molecular measurements describing the spatial relationship of interface residues - such as the distance between K52 and E55 or between C53 residues of opposing decamers - would aid interpretation. The use of the term "adjacent" (line 530) is vague and should be made more precise.

      We thank the reviewer for the thoughtful comments. Beyond negative-stain EM we also performed a biochemical validation of the filament interface through bi-functional crosslinking based on the premise that the new filament interface, as defined by the apo-filament structure, presented new/unique pairs of nucleophilic amino acid R-groups in close proximity. We used Bis-sulfosuccinimidyl glutarate (BSG) or bis-maleimoethane (BMOE) to covalently link adjacent primary amines and sulfhydryls respectively (Figure 1E). This experiment defines two key principles of GS filament formation in the absence of glutamine:

      (1) It is concentration dependent. In the wild-type case there is a protein-dependent increase on crosslinking efficiency for both crosslinkers.

      (2) It is dependent on C53 and K52. Mutation of C53 or K52 significantly attenuated crosslinking efficiency.

      To make sure that these results are more prominent, we have now included a table of these intersubunit distances between epsilon amine groups of lysines and gamma sulfhydroxyl of cysteines groups based on the apo-filament structure and labeled this as either participating in the filament interface or not. Furthermore, in line with multiple reviewers comments, we have updated Figure 1D to include a sharpened representation of the map that shows strong side chain density for the amino acid side chains to further support these reported side chain distance measurements.

      We thank the reviewer for pointing out the low precision of the SEC chromatogram interpretation of Figure 1B and the figure has been amended to show filaments of variable length, instead of defined length. We also included a calibration curve to Supplemental Figure 1B and estimated molecular weights. While these estimates of size are lower than ground truth it is important to note that GS has historically displayed smaller than predicted molecular weights via size exclusion chromatography and analytical ultracentrifugation where initial characterization papers defined the oligomeric state as an octamer rather than decamer (PMID: 10708854). These studies and the present indicate potential adherence to resin and/or other factors about the shape of GS that lead to longer retention. Lastly, these SEC procedures were performed as a preparative step rather than for analytical purposes, so resolution was not the ultimate goal.

      (4) Terminology: "Scarless" hGS:

      The term "scarless human glutamine synthetase" is unconventional and potentially confusing. If it refers to the wild-type sequence lacking N- or C-terminal tags or mutations, I recommend using the term "native hGS" for clarity.

      We usually reserve “native” for proteins isolated from the original species and not recombinantly expressed (as here). So we will leave the term scarless in the document.

      (5) Helical Parameters of Filament Assembly:

      The manuscript states a ~26{degree sign} rotation (clockwise or counterclockwise?) between decamers in the filament, yet does not describe how this value was derived. Given that hGS filaments form helices, this parameter could be assessed via helical reconstruction. Is it possible the actual helical twist is ~30{degree sign}, implying 12 stacked decamers per full turn? Please elaborate on how the rotational angle was determined.

      Rotation was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was measured. Helical reconstruction was not pursued in this work owing to the typically short filaments observed in micrographs and the relative ease by which a 20-mer species could be selected via traditional 2D and 3D classification/reconstruction methods.

      We have added to the methods the following to better illustrate this measurement:

      “Decamer: Decamer rotation across the filament interface was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was depicted in Figure 1D.”

      (6) Time-Resolved Cryo-EM and Filament Growth:

      The use of time-resolved cryo-EM is innovative; however, the accessible timescales are relatively short. I suggest complementing this approach with techniques such as dynamic light scattering (DLS) or mass photometry, which allow extended real-time monitoring of filament assembly over longer durations (e.g., hours). These methods can also provide higher temporal resolution and particle size distributions.

      We thank the reviewer for this suggestion and agree that understanding the kinetics of filament formation is critical. While DLS and mass photometry are excellent for monitoring assembly over hours, our data indicates that GS filament formation occurs on a much faster timescale.

      As shown in Figure 2C, when we added ATP and Glutamine directly to GS and vitrified the sample after only 5 minutes, the majority of particles had already formed filaments, indicating that the interaction had reached saturation. This contrasts with our time-resolved experiment, where the kinetics of filament formation were likely rate-limited by the enzymatic generation of glutamine rather than the assembly process itself.

      Consequently, we anticipate that filament assembly occurs on the order of seconds or less—a timescale we interpret as a necessary prerequisite for a rapid and effective cellular feedback mechanism. Therefore, we believe the current cryo-EM data accurately captures the biologically relevant window of assembly.

      (7) Cryo-EM Symmetry Imposition and Loop Flexibility:

      The use of D5 symmetry in cryo-EM reconstructions may obscure conformational heterogeneity in flexible elements, such as the E305 loop. Since the authors used MD simulations to characterize loop dynamics, it would strengthen the study to also perform symmetry expansion followed by non-uniform refinement and alignment-free 3D classification of individual subunits. This could provide experimental validation of the proposed conformational variability.

      We agree with the reviewer that symmetry enforcement can mask conformational heterogeneity, particularly for flexible elements like the E305 loop (the E-flap). To address this, we followed the reviewer’s suggestion and performed symmetry expansion on both the turnover decamer and filament consensus maps. This was followed by focused, alignment-free 3D classification on the asymmetric unit containing the E-flap.

      Our analysis revealed a clear distinction: while 4 out of 12 turnover decamer classes showed partial density for the E-flap (class 3, 7, 8, and 12)—consistent with the flexibility observed in our MD simulations—none of the turnover filament classes demonstrated similar density. To ensure a direct comparison, we utilized a C5-expanded turnover decamer map to maintain an identical asymmetric unit to the D5-expanded turnover filament map. We note here that the turnover decamer consensus volume is different from the deposited map for which no symmetry was applied.

      Despite the different initial symmetries (D5 for filaments vs. C5 for decamers), we utilized a C5-expanded decamer map to maintain an identical asymmetric unit. For transparency, we have uploaded this C5-refined consensus map and all resulting 3D classification maps to Zenodo. We agree that the text is now strengthened given this result and we have added the following to the main text:

      “The differential loop density between turnover-decamer and turnover-filament species was further supported by 3D classification of symmetry-expanded particles, which recovered partial E305-loop density in 4/12 turnover-decamer classes (C5; Supplemental Figure 18) compared to 0/12 classes for the turnover-filament (D5; Supplemental Figure 19).”

      (8) Crosslinking Gel Analysis (Figure 1E):

      The SDS-PAGE gels shown in Figure 1E have molecular weight ladders cropped. For proper interpretation, please include full ladders with size markers and labels. In addition, clarify whether the crosslinked samples were denatured in reducing buffer. Crosslinking efficiency and specificity using BMOE or BSG require verification under reducing conditions (e.g., DTT, β-mercaptoethanol, or TCEP) to confirm covalent linkage between decamers.

      We have added in Figure 1E molecular weight markers estimates to aid in gel interpretation and have included the uncropped gels in Supplementary Figure 1D that contain the full MW ladder. The methods were clarified to indicate that reducing reagent was used in both the crosslinking reaction and all SDS-PAGE samples.

      “Protein samples were diluted to concentrations noted in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.1 mM TCEP) and, reacted with crosslinker to a final concentration of 0.5 mM for 10 mins at room temperature followed by quench in 5X SDS-PAGE sample buffer (225 mM Tris pH 6.8, 50% glycerol, 0.05 % SDS, 0.2 mg/mL bromophenol blue, 1M DTT) supplemented with 100 mM of either NH<sub>4</sub>Cl (to quench BSG reactions only) or DTT (to quench BMOE). Protein concentrations were normalized after quench prior to SDS-PAGE analysis.”

      (9) Missing Reference for NADH-Coupled Assay:

      Line 678-679 refers to an NADH-coupled assay described "previously" without citing a source. Please provide a proper reference to ensure reproducibility.

      The original paper describing the implementation of a coupled-assay to measure ADP production from glutamine synthetase was written by Bennett Shapiro and Eric Stadtman in 1970 and has been included. We will note that the conditions of this assay have been much improved since this time with better buffers, commercially available reagents of combined lactate dehydrogenase and pyruvate kinase, and modern plate readers. We added the following reference:

      “Shapiro, B.M. and Stadtman, E.R., 1970. [130] Glutamine synthetase (Escherichia coli). In Methods in enzymology (Vol. 17, pp. 910-922). Academic Press.”

      (10) Unclear Description of the NADH Assay:

      The stability of NADH is influenced by pH and light exposure. Please specify the pH range used in the assay and whether precautions (e.g., light shielding) were taken. NADH autoxidation at high pH or degradation at low pH could impact assay reliability and should be addressed in the Methods section.

      For clarity and transparency the following text was added to the Methods section.

      “Stocks of ATP, NADH, and phosphoenolpyruvate were made in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.5 mM TCEP) and the pH was adjusted until it reached 7.5 on ice prior to aliquoting, flash freezing, and storage at -80°C in the dark. NADH was only exposed to light upon thawing and assay set-up and no appreciable change in absorbance of control experiments were noted.”

      (11) Ligand Density in Figure 2 and Supplementary Figure 7:

      The density attributed to glutamine, ADP, and phosphate appears broader than expected. Please include cross-correlation (CC) values, estimated occupancies, and Q-factors for ligand fitting. Varying the contour level to assess density consistency would clarify whether the observed volume represents multiple conformations, partial occupancy, or overfitting. A similar concern applies to the cysteine sidechain density.

      We have updated Supplemental Figures to include:

      (1) Globally refined map in comparison to locally refined map where both are sharpened per previous feedback.

      (2) Ligand placement now also include Q-scores and CC values

      We did not include multiple contour levels because these are included in the resolution representative Supplemental Figure and because alternative contours do not influence Q-scores. From this analysis it is apparent that phosphate and ADP are both worse fits to the density.

      Moreover, we have now included Supplementary Table 2 that includes all ligand validation statistics for the reader to evaluate the range of B-factor, CC values, and Q-scores for all ligands in all models.

      (12) Missing Ligand B-factors in Supplementary Table 1:

      The ligand refinement statistics in Supplementary Table 1 are incomplete. Please include B-factors and occupancy values for all ligands.

      We have updated the PDB depositions to include B-factors in .cif files that are now available. We have also included Supplementary Table 2 in the manuscript detailing the ligand statistics for all models including cross-correlation, Qscore, and Bfactor.

      (13) Style and Formatting Issues: format consistently throughout.

      (a) Kinetic Parameters: Please follow the IUPAC and IUBMB-recommended formatting:

      kcat should be italic with subscript.

      KM should be italic K with upright M.

      Use lowercase s<sup>-1</sup>, not uppercase S<sup>-1</sup>.

      Refer to:

      IUBMB enzyme nomenclature guidelines https://iubmb.org/wp-content/uploads/2021/01/Current_IUBMB_recommendations_on_enzyme_nome nclature.pdf

      IUPAC Green Book https://publications.iupac.org/books/gbook/green_book_2ed.pdf

      We have corrected the abbreviations according to the reviewers recommendations.

      (b) Inconsistent Terminology and Typography:

      cryo-EM vs. cryoEM are used inconsistently - standardize throughout.

      FSC 0.143 appears with and without subscript formatting-please unify.

      Line 123: "X-ray" should be capitalized.

      Line 266: CryoEM should cryoEM, lowercase "c"

      Line 571: "100 μg ml<sup>-1</sup>"-use superscript minus; ensure consistency with "mg ml<sup>-1</sup>".

      Lines 605, 606, 626, 628: MgCl<sub>2</sub>-ensure the <sub>2</sub> is subscripted throughout.

      Temperature units (lines 607, 630, 640): Write as "4 {degree sign}C" instead of "4C".

      Microliters (lines 660, 681, 697): Replace "uL" with "μL".

      Line 797: Use superscripts: K<sup>+</sup>, Cl<sup>-</sup>.

      Line 862: CO<sub>2</sub> should appear with subscript.

      We have made all terminology and typography consistent throughout based on these suggestions.

      Reviewer #2 (Recommendations for the authors):

      To strengthen the manuscript and address the methodological and interpretational gaps identified, we recommend the following revisions and additions:

      (1) Data processing and cryo-EM map quality

      (a) B-factor sharpening: Reprocess all cryo-EM maps using standardized B-factor sharpening workflows (e.g., the autoSharpen tool in cryoSPARC or similar methods) to enhance side-chain and ligand density visibility.

      We have updated main and supplementary figures to include sharpened maps. All maps were sharpened using the Autosharpen feature of Phenix, specifically, by half-maps. We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      (b) Document the specific parameters used (e.g., B-factor values, solvent content estimates) in the Methods section to improve transparency.

      See above regarding the additions made to the methods section.

      (c) Map replacement and reanalysis: Replace all figures and supplementary panels displaying raw (unsharpened) maps (e.g., Figures 1D, 2D, 3A/B, 4A, 5B, and Supplementary Figs. 2D, S7B, S10) with the newly sharpened versions. Reanalyze density features (e.g., glutamine binding sites, ATP triphosphate groups) using these revised maps and update results to reflect any changes in interpretation.

      As requested, we have updated figures with sharpened maps and found our original analyses to hold. In particular, we have included multiple metrics of ligand model scoring in Supplementary Figure 7B including Q-score and CC. Additionally, we have included all ligand model statistics in Supplementary Table 2.

      (2) Structural modeling and validation

      (a) Ligand fitting justification: Provide high-resolution ({less than or equal to}3 Å) density slices or side-chain density close-ups (e.g., for phenylalanine rings or glutamine-binding regions) to validate claims of atomic-level detail. For non-symmetric ligands (e.g., glutamine) fitted into symmetry-refined maps, explicitly describe how symmetry constraints were adjusted or applied during fitting (e.g., local symmetry refinement, manual adjustment of ligand orientation) and include validation metrics (e.g., cross-correlation scores, density fit plots) to support the placement.

      We have supplied 5 new supplementary figures to demonstrate the resolution of our sharpened cryo-EM maps (most notably Supplementary Figures 3, 4, 8, 12, 22 and panels in others) . Of particular note is the sharpened map features of R298A decamer under turnover conditions which demonstrates multiple instances of a ring density for aromatic residues.

      See discussion above regarding the placement of glutamine in the interface density and updated handling of symmetry during refinement.

      (b) Ligand density supplements: Include supplementary figures showing representative ligand-density fits (e.g., ATP, glutamine) with clear side-chain or functional group annotations, as is standard in structural biology publications.

      In our revision we have included the following updated figures and figure panels demonstrating ligand density into sharpened maps:

      Turnover Filament Glutamine Ligand: Figure 2D-E (updated representation) and Supplementary Figure 7 (new and updated representations).

      Turnover Filament ATP and Mg(II): Supplementary Figure 10 (updated representation)

      Turnover Decamer ADP and Mg(II): Supplementary Figure 5 (new figure panel)

      Turnover R298A ADP and Mg(II): Supplementary Figure 14 (new figure panel)

      (3) Biochemical assay rigor

      (a) Control experiments: Perform and report the following controls to strengthen enzyme activity claims:

      - A blank control (reaction mixture without GS, ammonia, or glutamate) to quantify background ATP hydrolysis.

      - Substrate omission controls (reactions lacking ammonia or glutamate) to confirm that ATP hydrolysis depends on both substrates and GS catalysis.

      - A TCEP effect control (compare ATP hydrolysis rates with and without TCEP) to rule out reducing agent interference with the PK/LDH coupled assay.

      We have provided blank, substrate omission, and TCEP controls in Supplementary Figure 1. These results demonstrate negligible ATP hydrolysis without complete substrate inclusion and do not indicate any impact from TCEP inclusion.

      (b) Direct activity validation: Consider supplementing the coupled assay with a more direct measure of GS activity (e.g., quantifying inorganic phosphate release via malachite green assay) to cross-validate results.

      On the merits of the PK/LDH coupled assay being used for >55 years to measure steady-state activity of glutamine synthetases and that it is a robust assay as supported by the additional control experiments presented above in Supplementary Figure 1 we have elected not to pursue tedious cross-validation with a non-continuous assay and believe our interpretation of the enzyme kinetic results hold.

      (4) Writing and presentation clarity

      (a) Methods detail: Expand the Methods section to explicitly describe:

      - Cryo-EM data processing steps, including B-factor sharpening parameters, map reconstruction workflows, and any post-processing (e.g., filtering, masking).

      - Criteria used to validate ligand fitting (e.g., density threshold values, manual vs. automated docking).

      See above the revisions made in response to critique from review #1 which we will briefly summarize here:

      We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      Focused masks are represented in Figure 2E, Supplementary Figure 18, and Supplementary Figure 19. The details around focused mask utilization are included in the revised figure captions and the following was included in the Methods.

      “Focused masks were generated in ChimeraX (v.1.7 and above). Focused refinement and 3D classification (3 Å filter resolution, PCA initialization) were performed in cryoSPARC. Strategy of class picking and refinement are noted in Supplementary Figures 18 and 19.”

      Map reconstruction workflows are present in the relevant Supplementary Figures. No post-processing steps beyond map sharpening in Phenix were carried out. In general, human GS represents a straightforward protein to reconstruction via cryo-EM.

      Ligand identification criteria was described throughout the results section. Supplementary Figure 7 was revised to show sharpened density for either globally refined or locally refined maps fit with all three products of the glutamine synthetase reaction (ADP, Pi, and glutamine) individually showing the best CC and Qscore for glutamine. Beyond Supplementary Figure 7 we also combined both biochemical experiments and cryoEM to make this ligand assignment supported by:

      (1) Time-resolved cryo-EM experiments that show increasing filament particles over reaction time (Figure 2B and Supplementary Figures 9 and 10)

      (2) Glutamine+ATP cryo-EM screening (Figure 2C) showing long filaments

      (3) Supplementary Table 2 showing reasonable ligand statistics for glutamine

      To clarify this in the Methods sections we include the following statement:

      “Ligands were placed with ISOLDE (v1.7) and those with >0.5 Qscore and supporting biochemical and/or literature precedent were built.”

      (b) Results framing: In the Results, clearly distinguish between observations supported by sharpened maps and preliminary/unvalidated features. Avoid over interpreting density in unprocessed maps (e.g., referring to "glutamine binding" in Figure 2D without noting current density limitations).

      We have updated our discussion of Figure 2D (and now also Figure 2E) to include discussion of only sharpened maps and noted current density limitations to the interpretation.

      (5) Data and material availability

      (a) Ensure all supporting data are publicly accessible:

      - Upload raw cryo-EM movies, particle stacks, and processed maps to the Electron Microscopy Data Bank (EMDB) with appropriate accession codes.

      - Deposit final atomic models in the Protein Data Bank (PDB) and reference these accession codes in the manuscript.

      We deposited maps and models with accession codes in advance of review. The PDB and EMDB codes are available in Supplementary Table 1.

      Furthermore, for the focused maps and focused classifications that were generated during the review, and for the benefit of not cluttering the PDB/EMDB, we have included these more specific analyses in Zenodo: 10.5281/zenodo.20298855.

      - Provide detailed protocols for biochemical assays (e.g., TCEP handling, enzyme purification) in the Methods or as supplementary information to enable reproducibility.

      See updates above to reviewer #1

      (b) By implementing these revisions, the manuscript will better align with eLife's standards for methodological rigor, transparency, and reproducibility, allowing readers to confidently evaluate the study's contributions to structural and functional biology.

      We agree!

      Reviewer #3 (Recommendations for the authors):

      (1) In line 252, it would be helpful to show negative-stain EM images for each SEC peak, further probing whether any peaks correspond to partially aggregated, as this could affect the measured Kcat and Km.

      We aren’t in a position to do this experiment. We routinely check for aggregation by noting Absorbance at 340nm for non-specific scattering indicative of aggregation and observed no evidence of aggregation in our fractions.

      (2) In Supplementary Figure 6, many of the classes in the "Selected Filament Classes" inset appear to be averages of closely spaced particles, which may bias the calculation and should be excluded. In the "Selected Decamer Classes", I would suggest removing the top-view particle classes, as these particles not only have significantly different ice penetration rates, but are also more difficult to distinguish in 2D classification.

      We agree and have provided an additional, more strenuous cutoff, analysis of the tr-cryo-EM data wherein only classes that show clearly aligned decamers are included and all top views are omitted (Additional Supplementary Figure 6). We are happy to say that even with the more strenuous cutoffs that our main conclusions that filaments increase with forward reaction time holds.

      (3) In line 445, "Figure 5A" should be corrected to "Figure 5B".

      We thank the reviewer for pointing this out and have made the correction.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors performed seqFISH in 26 gastruloids and performed a variety of computational analyses on these novel spatial data sets. Whilst the data is valuable and the computational concepts useful (exposure index, L-metric, ...), the article falls short on novelty and is written using a very clunky language, often with contradictory conclusions.

      We thank the reviewer for their comments about the value of our data and computational concepts. We agree with the reviewer’s critical comments and have endeavored to address them all. We believe the resulting manuscript is greatly clarified and improved.

      Major issues:

      (1) The authors did well in explaining and detailing the provenance of data and the individual experiments performed. However, their 26 gastruloid data still constitute a very limited sampling from their total organoids: one experiment pooled 4 plates at an 80-94% success rate; 6 different aggregation experiments were done, making a total of 1843 gastruloids, sampled 26 (~1-2%). A simple IF stain of 2-3 markers in a bigger sample could have given a more accurate picture of specific domains of interest and their proximity. Regardless, more information should be given about the existing samples: variation across experimental batches, differences between 300-cell vs 100-cell gastruloids that were used.

      This omission was an oversight on our part and we thank the reviewer for catching it. We did the following to address this point:

      (1) Added date labels to Figure S1.2d (now S1.1a) so that the proportion correct for each separate experiment is clear:

      (2) We added the raw images of the gastruloids used in the study taken before fixation.

      (3) We segmented these images and quantified metrics of the masks to address differences across samples in morphology. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation. Elongation was measured as 1-(the ratio of the width and the height of the segmented gastruloid area); the code can be found here:

      https://github.com/arjunrajlaboratory/ImageAnalysisProject/blob/1b2f2119f77083c27f58f8c36b14 c48d97ea706c/workers/properties/blobs/blob_metrics_worker/entrypoint.py

      Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300).

      The literature also supports our assumption that combining experiments with different starting numbers of cells would not dramatically affect the results (Bennabi et al. 2024). In this paper they show that only 35 genes were differentially expressed between gastruloids formed from 100 cells and those formed from 300 cells (compared to 319 for those formed from 1200 cells and 336 for those formed from 50 cells, both compared to 300 cells). The same paper also demonstrates that the positioning of gene expression (as measured by IF staining) for several representative genes (Bra and Foxc1) does not significantly differ between gastruloids formed from 100 and 300 cells when normalized for overall size and AP axis length (as we’ve done in this paper as well).

      We have updated the text and revised Figure S1.1 to reflect these changes.

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. The experiment was performed 3 times on different days, so to ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].

      To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (see also Revised Figure S1.1a-c).

      (2) Language in the manuscript should be revised. Overall the manuscript is very long, descriptive and written "impressions and beliefs" are often not adequately justified and indeed can be contradictory, e.g. in Section 1: the title states "cell types' locations ...are consistent", a few sentences down we find "there was substantial variation" and "within range of what would be considered a 'morphologically normal' gastruloid". "quite consistent", "compelling patterning", "we don't believe"... these types of expressions are best avoided and replaced with data or used and bolstered with quantitative numbers such as percentages when a given cutoff is used. Another example: "location of each cell type relative to gastruloid morphology was quite consistent the posterior region ... mainly consisted in NMPs." Given T expression in the posterior, this result phrased as such appears quite inflated, in fact, looking at cell types in Figures S1, 2a/b/c, this reviewer would state they are all but consistent and indeed it takes sophisticated analyses to find a pattern (of sorts) beyond the coarse domains expected!

      We thank the reviewer for their careful reading of our paper and appreciate that the work would be strengthened by increasing the degree to which quantitative measures are used to justify the statement we make. We have made the following changes to the manuscript to address this criticism:

      (1) We more clearly delineate where we are making qualitative descriptions and have removed summary language (like ‘consistent’, ‘normal’, ‘variable’ etc.) from these sections. For example, the section the reviewer refers to originally read:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we characterized the organization of each by mapping where each cell type was found relative to other types and overall morphology. The approximate location of each cell type relative to gastruloid morphology was quite consistent: the posterior region, although highly variable in size (Figure S1.2a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]....”

      And now reads:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we first qualitatively examined where each cell type was found relative to other types and overall morphology. The posterior region, although variable in size (Figure S1.3a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]...”

      (2) We follow this qualitative description with a quantitative analysis of cell type proportion where we clearly state which variable aspects are statistically significant:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d).”

      (3) We added a summary paragraph at the conclusion of the results from the first two figures which clearly states which aspects of gastruloid organization we find to be variable and which are consistent, with statistical testing:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (3) Figure 6 is one of the most valuable parts of the work, as the authors use the battery of analyses developed to investigate the variable and not-so-robust endothelial clusters in gastruloids. However, this investigation is still very preliminary, and it should be further linked with known biology. It is still unclear what the unique organization of this cell type is (circularity isn't convincing) and whether any signalling cues of adjacent cells could explain it. Is there any evidence that more mature endodermal cell types are generated (like the suggested "liver") to give rise to endothelial cells? It would certainly be interesting to perform IF for this cell type together with mesodermal and endodermal markers to validate seqFISH predictions on a bigger sample.

      We appreciate the reviewer pointing out that the comparisons between different endothelial cell types was interesting, and agree that the clustering methods were insufficiently justified and that a more explicit consideration of the signaling context of the gastruloid could strengthen our findings.

      We have re-evaluated how we calculate differentially expressed genes. We restricted our analysis to only consider genes that are expressed at > 2 counts/cell in at least 50% of the subsets considered. The results are in shown in the revised Figure 6.

      We find that, as the reviewer suggested, some signaling genes are significantly differentially expressed. Specifically, Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in notch signaling in the posterior of the gastruloid. Tek, on the other hand, is more expressed in the somite-associated endothelial cells, and Tek has been annotated to be involved in retinoic acid signaling. These findings align with the reviewer’s observation that signaling from adjacent cells could explain or relate to differentially expressed genes.

      We also did a more thorough review of the literature, and found several papers that reported unique subsets of endothelial precursors, albeit in related systems. In [Rossi 2022] and [Rossi 2021] the authors find a population of endoderm-associated endothelial cells in gastruloids grown with a different protocol that involves Matrigel embedding, treatment with factors that promote blood development, and growth for 168 hours. In [Veenlveit 2020] the authors find a unique somite-associated population of endothelial cells in Trunk-Like Structures, which are similar to gastruloids but model later in development and have more physical organization with discrete somites. To address the reviewer’s request that we further link with known biology we have added the following to the text:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which show more tissue-like organization than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      Finally, we tested several methods of clustering and calculating circularity, and determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and any conclusions drawn from the text.

      (4) Figures 1c and 6b need statistical significance assessments.

      We thank the reviewer for pointing out that without significance testing these plots are difficult to interpret. We have removed plot 6b (see response above about removing the circularity assessments). For plot 1c we appreciate that it is difficult to interpret which cell types vary more than others in their occurrence without significance testing. To address this we did two things: we first calculated the coefficient of variation for the proportion of each cell type across samples:

      Author response image 1.

      To calculate significance, we first considered that since these values are proportions, they must sum to 1 and changes in one cell type will affect at least one other cell type within the same sample. To properly account for this when applying statistical tests, we calculated the CLR-transformed proportion and tested all pairs of cell types for significant variation. The results are shown in the Author response image 2:

      Author response image 2.

      We added a plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We also address said variation in the text:

      “We sought to quantify variability in cell type composition between gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centred log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.3a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.4).

      The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.4d).”

      (5) The article should include an analysis of Hox colinearity expression in these gastruloids as a validation of the system.

      We thank the reviewer for pointing out the importance of these genes in validating the gastruloid system and agree that assessing their expression specifically would help readers assess data quality.

      We analyzed the center of mass of expression along the AP axis for the Hox genes included in our panel (Hoxb6, Hoxc10, Hoxd1, Hoxb9, Hoxc8, Hoxc6, Hoxaas3, and Hoxb1). We highlighted these genes in Figure S1.3: their expression along the AP axis is consistent with previously reported expression in the tomoseq dataset from [van den Brink 2020]. A summary of the correlation coefficients for each individual gastruloid for all genes (blue) and the Hox genes (orange) is shown in the Figure S1.2. The Hox genes have similar correlation coefficients overall, although their variation is higher. This is likely due to differences in gastruloid pseudo-age; in future experiments we plan to include more Hox genes and use their expression to further classify gastruloids (see updated Figure S1.2).

      We have updated the text with these new results:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an ambitious and technically challenging spatial-transcriptomic atlas of 26 gastruloids using seqFISH. The authors introduce quantitative metrics (mixing score, exposure index, L-metric / scL-metric, spatial L-metric, triplets) to characterize spatial organization at multiple scales. The dataset is valuable, and several analyses are original, particularly the rank-based L-metric family for mutual exclusivity.

      Strengths:

      The authors generate one of the most detailed spatial transcriptomic datasets of gastruloids to date. They propose creative computational metrics (L-metric/scL-metric) to quantify mutual exclusivity of gene expression without predefined thresholds, and they explore organizational principles from single-cell topology to cluster-level structure. Many observations align well with known gastruloid biology, such as posterior robustness and anterior variability. The writing is generally clear, and the figures are rich.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      Several central claims rely on metrics whose computation and justification are insufficiently explained, making it difficult to assess how robust or interpretable the results are. Many choices in the analysis appear arbitrary or are insufficiently motivated (normalization schemes, choice of parameters such as the number of neighbors, the distance cutoffs, hierarchical clustering setup, and so on). The interpretations of spatial consistency, gene-program inference, and endothelial heterogeneity are plausible but might be stronger than the evidence currently supports.

      The manuscript would benefit from stronger benchmarking, quantification of uncertainty, and explicit controls for known artifacts in spatial transcriptomics (e.g., spillover, 2D slicing, cell type assignment entropy). The biological insights are promising, but since several depend on methodological assumptions that have not yet been demonstrated to be stable, they would benefit from clearer methodological explanation.

      We thank the reviewer for spending time to give constructive and actionable comments, and we believe the manuscript is greatly strengthened and more consistent and clear as a result of the changes suggested.

      The work is rich and could become a reference dataset. Then, clarifying and validating the quantitative methods will considerably strengthen the impact and reliability of the conclusions.

      Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020), the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas, which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-metric', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can, in principle, be asymmetric; although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus their analysis on two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g., that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract the statistical structure of developmental processes. The L-metric offers a parameter-free tool to analyze transcriptomic datasets that could overcome the pitfalls of other approaches.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their dataset can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      We agree that the developmental processes studied here are not inherently novel, and we hope that by showing sufficient overlap with different, less highly resolved methods we have created a convincing document that highlights the potential for this technique to be used to analyze other systems. We appreciate the reviewer’s comments and that the manuscript is improved after making the suggested changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) S1.2 plates shown individually, but unclear from which experiment.

      The reviewer was right to point out this oversight — we have updated the figure (now S1.1a) with labels for the individual experiments:

      (2) Figure 2 could include a clearer indication of the types of triples/ doublets to make it even more informative.

      We thank the reviewer for pointing out that the types of triples were not clear — we’ve added a color key and more explanatory text to this figure in order which explicitly explains the type of triplets considered.

      (3) Figures should be presented in order. Figure 3c is before 3b, etc.

      We appreciate the reviewer’s attention to detail and have swapped these two panels so that their order in the figure reflects the order they are referenced in the text.

      (4) Figure 3 is interesting, and the L-metric appears useful to pinpoint crucial genes that, when expressed, indicate a type transition has occurred. It would be great to test this with another set of cell types besides NMPs/presomitic/spinal cord.

      We thank the reviewer for their interest in this biological transition, and agree that testing on another transition would be really interesting. There isn’t another set of cell types expected at this stage of gastruloid development that are predicted to have the same type of bifurcating differentiation. However, in an effort to address the spirit of this comment (that looking at other sets of cell types with the L-score would be interesting), we have used our analytical framework on a non-spatial, single-cell dataset from van den Brink et al 2020.

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (5) Figures 4 and 5 felt exploratory, and I would recommend combining them into a single figure highlighting the usefulness of L-metric and its spatial version.

      We thank the reviewer for the suggestion to merge the content of Figures 4 and 5. While we agree that they are thematically related, given the size of the heatmaps generated in the analyses, we were unable to combine them in a way that preserved the readability of the figure and stayed within the space constraints of the page size; thus, we have chosen to keep these as separate figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Quantification methods require clearer formalization and justification

      A key limitation is that the manuscript relies on several spatial metrics whose definitions are not sufficiently formalized.

      (a) To evaluate biological interpretations, the reader needs a precise description of:

      - How each metric is calculated (mixing score, exposure index, scL-metric, spatial L-metric).

      - Why specific choices were made (normalizations, parameter values, distance thresholds).

      - What the expected ranges and interpretations are.

      We agree that these elements are crucial to interpreting quantitative metrics and thank the reviewer for their close read of the work. The reviewer had many comments on the exposure/mixing values calculated for Figures 1 and 2, and for the L-metric (now called L-score) values calculated for Figures 3, 4, and 5. To address the reviewers concerns we have done the following:

      (1) Created a new, unified framework for calculating cell type exposure and mixing.

      (2) Rewritten the methods section for this section with a particular emphasis on including elements the reviewer suggested, including specifically outlining normalizations and what they account for, distances chosen and the biological rationale behind them, and the range of values expected for each measure:

      “Quantification of Cell Type Spatial Relationships: Exposure index

      To characterize the spatial organization of cell types, we computed two related metrics: a pairwise exposure index matrix capturing type-specific spatial relationships, and a scalar mixing index summarizing overall spatial integration. For each cell, we identified neighbours as all cells whose centroids fell within a specified radius r of the focal cell's centroid. We chose a value of r of 16 μm, which gave an average of 5-6 neighbors per cell. We chose this value as it captures local interactions, which was the primary goal of these analyses. For each ordered pair of cell types (s, t), we calculated the exposure index as the proportion of type s cells' neighbors that are type t:

      where Ns→t denotes the count of neighbour pairs in which the focal cell is type ‘s’ and the neighbour is type ‘t’, and Ns denotes the total number of neighbours across all type ‘s’ cells. Each row of the resulting exposure matrix sums to unity and represents a probability distribution over neighbour types for a given source type. To account for differences in cell type abundance, we normalized exposure indices relative to the expectation under random spatial arrangement:

      Where pt is the proportion of cells that are type ‘t’. Normalized values of zero indicate exposure consistent with random mixing, positive values indicate spatial attraction (co-localization), and negative values indicate spatial avoidance, with a minimum of −1 representing complete exclusion.”

      “Quantification of Cell Type Spatial Relationships: Mixing index

      To summarize overall spatial integration across all cell types, we computed a mixing index defined as the fraction of neighbour pairs involving different cell types:

      where N_cross is the number of neighbour pairs involving cells of different types and N_total is the total number of neighbour pairs. For normalization, we compared the observed mixing to the expectation under random spatial arrangement:

      where

      is the expected cross-type interaction rate given cell type proportions. Normalized values of zero indicate random spatial mixing, positive values indicate greater integration than expected (hyper-mixing), and negative values indicate spatial segregation.

      Exclusion of untyped cells. When computing the mixing index, cells lacking confident type assignments were optionally excluded from both the numerator and denominator, ensuring the metric reflects only spatial relationships among typed cells. These cells were retained in the exposure matrix to quantify how typed cells interact with unclassified cells.

      Statistical analysis of variance. To identify cell type pairs whose spatial relationships varied significantly across samples, we computed the variance in exposure indices across samples for each type pair. To account for the expected relationship between mean exposure and variance, we regressed log-variance against log-absolute-mean across all type pairs and computed residuals. Type pairs with residual variance exceeding the 97.5th percentile (two-tailed α = 0.05) were considered significantly variable, indicating spatial relationships that differ across samples beyond what is expected from sampling variation and composition differences.”

      (3) We have also rewritten the methods for how the L-score is calculated, adding emphasis to where we normalize, what ranges of values are expected, and what the interpretation of these values are:

      “Calculating the single-cell L-score

      The single-cell L-score (scL-score) was computed for each ordered gene pair (gene A, gene B) within a single gastruloid. The cell-by-gene expression matrix was filtered to retain only cells with at least 2 detected transcripts for both genes. Cells were sorted in descending order by gene A's expression values; gene A thus serves as the reference distribution, and the score is asymmetric with respect to gene order.

      Three reference distributions were constructed for gene B: (1) perfect coexpression, in which gene B's values were sorted in the same descending order as gene A; (2) perfect mutual exclusivity, in which gene B's values were sorted in ascending order; and (3) independence, in which every cell was assigned the mean expression value of gene B.

      Cumulative sums of expression values were computed for gene A, for the observed expression of gene B, and for each reference distribution. The cumulative sum of each gene B distribution (observed and reference) was then plotted against the cumulative sum of gene A. This cumulative-sum-versus-cumulative-sum representation captures how gene B's expression accumulates relative to gene A's: if gene B's expression is concentrated in the same high-expressing cells as gene A, gene B's cumulative curve rises steeply at first; if concentrated in opposite cells, the curve rises steeply at the end. The area under each curve was calculated using trapezoidal integration and normalized by the product of gene A's and gene B's total expression, yielding four normalized areas: A_observed (observed relationship), A_positive (perfect coexpression), A_negative (perfect mutual exclusivity), and A_uniform (independence). This normalization ensures that scores are comparable across gene pairs with different overall expression levels (Figure S3.2a,b). The scL-score was then defined as follows:

      If A_observed > A_uniform: scL-score = (A_observed − A_uniform) / (A_positive − A_uniform), yielding values in (0, 1].

      If A_observed = A_uniform: scL-score = 0.

      If A_observed < A_uniform: scL-score = −(A_observed − A_uniform) / (A_negative − A_uniform), yielding values in [−1, 0).

      A score of +1 indicates perfect coexpression, −1 indicates perfect mutual exclusivity, and 0 indicates independence.

      The scL-score was computed for all gene pairs in each gastruloid from the 05/07/2025 dataset (n = 18 gastruloids). The other two datasets (n = 8 gastruloids) were excluded due to lower transcript detection quality. To generate an average scL-score matrix, the analysis was restricted to a common set of 202 genes well-detected across all 18 gastruloids, and per-gastruloid matrices were averaged. Unless otherwise noted, a symmetrized scL-score was used: scL-score_sym(A, B) = [scL-score(A, B) + scL-score(B, A)] / 2.

      Calculating the spatial L-score

      The spatial L-score extends the scL-score to spatial regions. For each gene, a kernel density estimate (KDE) was fitted over all detected transcript spots and evaluated on a regular square grid spanning the gastruloid. Bin side length was set to twice the median nearest-neighbour distance between detected spots, calculated separately for each gastruloid. Spatial bins were ranked by KDE-derived density and processed identically to the scL-score calculation. Low-density bins were not filtered, as KDE smoothing produced non-zero density values throughout the imaging area. The spatial L-score was symmetrized as for the scL-score, except when displaying asymmetric heatmaps.

      Hierarchical clustering of L-score matrices

      Gene-gene distances were defined as Euclidean distances between L-score vectors. Agglomerative hierarchical clustering was performed using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean'). This approach operates on L-score vector differences rather than on pairwise L-score values directly, and therefore does not require the L-score itself to satisfy the properties of a mathematical distance metric; the Euclidean distance between L-score vectors is non-negative and symmetric by construction, satisfying the requirements of Ward's method. Heatmaps display pairwise L-score values, not vector distances.

      We applied this clustering procedure to the following gene sets:

      (1) A subset of NMP, presomitic mesoderm, and spinal cord marker genes in one gastruloid (n = 36 genes; Figure 3g).

      (2) All well-detected genes excluding cell cycle genes, averaged across all gastruloids (n = 166 genes; Figure 4 and Figure S4.3a).

      (3) All well-detected genes common to all gastruloids, averaged across gastruloids (n = 202 genes; Figures S4.1a, S4.3b, S5.1a).

      (4) All well-detected genes excluding cell cycle genes in one example gastruloid (n = 171 genes; Figures 5b, S5.2a).

      (5) All well-detected genes common between our seqFISH panel and those that were detected in > 3 cells in scRNA-seq data from [XXX] (n=207 genes; Figure S4.5a).

      (6) All genes detected in > 3 cells in scRNA-seq data from [XXX] (n=19075 genes, Figure S4.5c-f).

      In Figure S4.2a,b a transformation of the L-metric values was used to cluster genes. The scL-scores were averaged across gastruloids as described above, and then each pairwise scL-score was transformed to a distance-like value with(1 - scL)/2. The matrix was then symmetrized as described previously. Agglomerative hierarchical clustering was performed directly on this transformed gene-gene distance matrix using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean').”

      (4) We have added an illustrative figure about how the L-score is calculated which defines expected behaviour for several cases, gives a visual explanation of the process, and shows several extreme behaviours and their biological interpretation (see Revised Figure S3.2).

      (b) For example:

      - Mixing score: The normalization is unclear. Why only 1-nearest neighbor instead of k-NN? Why not consider existing spatial-autocorrelation metrics such as Moran's I, which would also apply to gene-level mixing?

      - Exposure index: The normalization makes the metric unbounded (e.g., exposure > 1 when local frequency > global frequency). It is unclear whether this behavior is intended. Since exposure to self is meaningful, the same metric could replace the mixing score and simplify the framework. The choice of k = 5 is not justified; parameter-free approaches like Delaunay triangulation could avoid arbitrary cutoffs. If k-nn is preferred, then the robustness of the score to change the k value should be studied.

      These issues make it difficult to interpret the magnitude of reported effects or compare them across studies.

      The reviewer makes an excellent point — we have completely overhauled this analysis in the following way to address the issues raised:

      (1) Created one unified metric (see points 1 and 2 above) which considers for every cell, the identity of its neighbors in a 16 um radius (on average 5 or 6 neighbors for each cell in each gastruloid). This value was chosen so that in most cases, the cells in the immediate vicinity of a cell were considered, but not those further out (i.e. the measure is sensitive to close interactions rather than far ones). We made a matrix of all interaction pairs for a given gastruloid, with diagonal elements representing within-type interactions and off-diagonal elements representing cross-type interactions. We have replaced the previous description with the following:

      “Several studies of gene expression in gastruloids have used pooled measurements to infer the AP axis-location of genes and cell types [3,7,13,24,27] and our data are largely consistent with these lower-resolution findings (Figure S1.3a). Yet it is obvious from individual gene staining [1,4,8,28] and our detailed 2D maps of cell identity and location that gastruloid organization is much more complex than the average order of cells along the AP axis. We thus needed an analytical method for quantifying spatial organization beyond distributions along the AP axis. To further characterize spatial organization, we sought to quantify the degree to which cells were mixed in each gastruloid, and how that mixing might vary between gastruloids. For each cell in each gastruloid, we counted the interactions between that cell and all its neighbours within a 16 μm radius (on average 5-6 neighbors per cell), and summarized all these interactions for all cells in the gastruloid in a matrix, normalizing each element by the frequency of the cell type considered to be the ‘neighbour’ in the interaction.”

      (2) To quantify overall mixing, we calculate the sum across types of the frequency of self interactions (normalized to the total interactions) and then take the inverse (1-M). We then normalize this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of that cell type).

      Because the density of cells is fairly consistent across the gastruloids, w is very close to p (the proportion of that type).

      is the expectation of cross-type rate with random mixing.

      Mixing ranges from -1 (totally segregated) to +1 (more mixed than random, i.e. there is attraction between unlike types). 0 is completely random, and negative values indicate that cell types within that gastruloid tend to cluster. We added the following to the text to explain this:

      “To quantify overall mixing, we calculated the sum (across types) of the frequency of self interactions, normalized to the total interactions) and then took the inverse. We normalized this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of the neighbouring cell type). This gave us, for each gastruloid, a value that we call the mixing index that ranged from -1 (totally segregated) to +1 (totally mixed with less frequent self-interactions than expected from chance). A mixing index of 0 indicates a random distribution, i.e., neighbour frequency is exactly what would be predicted by that cell type’s frequency alone.”

      We also edited the following description of the overall distribution of mixing indices:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). All gastruloids had a negative mixing index, indicating that they all, on average, had more like-cell type interactions than would be expected given random mixing of types. However, we note that there is a ~14% difference in the mixing index across gastruloids, meaning some variation in overall mixing is present.”

      (3) The exposure index for a given pair can be found from the off-diagonal elements of the interaction matrix, and the normalization means it represents relative overexposure/clustering (positive values) or underexposure/avoidance (negative values). The minimum value is -1 and the maximum is (1-pt)/pt. We changed the description of how the exposure index is calculated to reflect this unified method of quantification:

      We noted that the off-diagonal elements of the matrix we used to calculate the mixing index were informative about cell type-cell type interactions. Specifically they quantify the degree to which each cell type (source) is exposed to another cell type (neighbours). To assess the overall frequency of cell type-cell type interactions, we first pooled the data from all gastruloids together into one interaction matrix (Figure 2b). The measure can range from -1 (no interactions at all), with higher values indicating a greater frequency of being found in close proximity. It is asymmetric in that the exposure of cell type A to B may not be the same as the exposure of cell type B to A.

      We have updated all of the quantification in Figures 1 and 2 with these new measures.

      (2) L-metric: unclear justification and interoperability

      (a) First of all, even if it is not a major issue, the L-metric is not a "metric" at least in the mathematical sense of a metric since a metric is always positive. The L-metric is central to several major conclusions (gene exclusivity, modules, spatial organization), but its conceptual basis and computational steps need more justification.

      We thank the reviewer for pointing this out and have changed “L-metric” to “L-score” throughout. We have also endeavoured to clarify the conceptual basis and have fleshed out the various computational steps as outlined in more detail in our responses below.

      (b) Several steps (ranking, cumulative curves, area under the curve) are difficult to interpret biologically It is unclear why each transformation is required and how it responds to common scenarios (highly expressed genes, correlated vs mutually exclusive patterns).

      Since the metric is rank-based, two genes that are both highly expressed in all cells may show low L-metric despite being truly correlated.

      We appreciate the reviewer’s comments about both the interpretation of the scL-score calculation and how it behaves in common expression scenarios, particularly for genes that are broadly expressed across many cells. To address these points, we generated a set of simulated examples spanning five scenarios: ubiquitously expressed genes with similarly high average expression, ubiquitously expressed genes with differing average expression, ubiquitously expressed genes with similarly low average expression, genes coexpressed across a subset of cells rather than all cells, and genes generally expressed in opposite subsets of cells. For the first three simulations, we independently sampled two genes across 50 cells using Poisson distributions with mean expression set to 100 or 50 (to simulate a gene with high or low average expression, respectively), without any expression bias towards any subsets of cells. For the latter two simulations of dependent expression relationships, we first sampled gene 1 across 50 cells using a Poisson distribution with mean expression set to 4, then generated gene 2 from gene 1 by sampling from cell-specific Poisson distributions fitted to either generally match or oppose the transcript count obtained for gene 1 in that cell. These simulations show that genes can independently appear broadly coexpressed at the population level (regardless of average expression) simply by being ubiquitously expressed, yet still receive low scL-score values. The simulations of dependent coexpression or mutually exclusive expression relationships receive scL-score values near +1 and -1, respectively. These results align with the reviewer’s prediction, but they reflect why we designed the L-metric to follow a rank-based methodology since they preserve the specificity of the metric’s upper bound (+1) for detecting non-random coexpression relationships rather than chance coexpression relationships resulting from independently ubiquitous expression. The results of these simulations are depicted (See Revised Figure 3.3).

      We have made the following edits to the text to specifically address the case the reviewer raised about highly expressed genes:

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3.”

      (c) Interpretation of L-metric values is ambiguous

      What does 0 represent? Randoms? Co-expression? Is 1 the strongest exclusivity? The manuscript currently mixes "co-expression" and "mutual exclusivity" scales.

      We agree with the reviewer that clearly defining what values of the L-score mean is critical to understanding the text. We have added a more explicit discussion of this in the text (excerpted from the response to 2c):

      “A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes.”

      And made a figure representing visually how the L-score is calculated which shows the behaviour and biological interpretation of several extreme values and an example of how the L-score is calculated. See new Figure S3.2:

      We also edited figure captions where we referred to plots as ‘co-expression’ plots, since in some cases the plots showed genes that were mutually exclusive or not expressed together in most cells. Figure S3.1:

      “c. Spatial distribution of the expression of Nkx1-2 and Rfx4 in an example gastruloid.”

      In all other cases we checked, we used the term “co-expression” to mean the opposite of mutually exclusive, as outlined in the definition above.

      (d) Additional issues also limit interpretability - Benchmarking is missing.

      - No tests on synthetic datasets, negative controls, or curated examples.

      - Prior exclusivity methods (EEI, COTAN) routinely benchmark against ground truth; this is now standard.

      - The code for the L-metric seems to be missing in the repository.

      We appreciate this suggestion offered by the reviewer as benchmarking against prior exclusivity-oriented methods provides an important comparison for clarifying both where the scL-score agrees with existing approaches and where it offers distinct advantages. To address this, we explicitly compared the scL-score to the Exclusively Expressed Index (EEI), which is bounded below by 0 and increases with mutual exclusivity, and to the signed coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework, in which positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. We performed this comparison using seven simulated scenarios as well as four representative gene pairs from one gastruloid sample (2025-05-07_roi2). In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (high_high, high_low, low_low), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained low, and COEX was also positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.

      We then applied the same comparison to four gene pairs from one gastruloid sample (2025-05-07_roi2). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types. Pax6-Eogt, Rfx4-Eogt, and Nkx1-2-Rfx4 all showed opposing expression by scL-score, with values of -0.572 and -0.526, -0.970 and -0.955, and -0.537 and -0.615, respectively. EEI detected exclusivity most strongly for Rfx4-Eogt (0.171), but gave values of 0 or approximately 0 for the other two pairs, while COEX was negative for all three pairs (Pax6-Eogt: -0.235, Rfx4-Eogt: -0.350, and Nkx1-2-Rfx4: -0.106), consistent with opposing expression, with strongest signal for Rfx4-Eogt. By contrast, Cdx4-Cdx2 showed moderate levels of coexpression by scL-score (0.402 and 0.424) and EEI (0), but COEX indicated that they were not coexpressed (-0.331).

      These comparisons also clarify the practical advantage of the L-metric over the EEI and COTAN frameworks. EEI is based on binary zero/non-zero quantification and is therefore designed specifically to measure exclusivity rather than coexpression. COEX provides a signed value and, in our simulations, tracked both exclusivity and coexpression; however, on representative gene pairs from one gastruloid sample, scL and COEX diverged in magnitude for highly exclusive expression relationships (Rfx4-Eogt) and sign for a coexpression relationship (Cdx4-Cdx2), motivating our introduction of a signed measure based on the mutual exclusivity of expression with the scL-score (see New Figures S3.3 and S3.4).

      We have updated the text to address the reviewer’s concerns: we benchmark using simulations of commonly occurring scenarios (such as varying expression levels, degree of mutual exclusivity, and amount of noise present in the relationship between the two genes in question) as the reviewer suggested. We also provided a direct comparison to an earlier exclusivity measure (EEI):

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3 and Figure 3.4.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We compared the scL-score to two existing measures of exclusivity. The Exclusively Expressed Index (EEI) (Nakajima et al. 2021) is bounded below by 0 and increases with mutual exclusivity. EEI is computed from binary zero/non-zero quantification and is designed specifically to measure exclusivity but not coexpression. The coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework (Galfrè et al. 2021) can also be used to quantify relationships between genes: positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (but with varying relative expression levels), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained close to 0, and COEX was positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We calculated the scL-score, EEI, and COEX for four gene pairs from one gastruloid sample (2025-05-07_roi2). The scL-score consistently delineated gene pairs possessing opposing expression profiles, while EEI was not always able to measure those exclusivity patterns (Pax6-Eogt: scL-score=-0.572 and -0.526, EEI=0; Rfx4-Eogt: scL-score=-0.970 and -0.955, EEI=0.171; Nkx1-2-Rfx4: scL-score=-0.537 and -0.615, EEI~0). Only the scL-score was able to detect the coexpression pattern present between the positively associated expression profiles of Cdx4 and Cdx2 (scL-score=0.413, EEI=0, COEX -0.331) (shown visually in Figure S3.4a). Thus, while EEI was informative for measuring gene expression relationships characterized by mutual exclusivity, the scL-score more clearly separated positive, random, and mutually exclusive relationships on a single signed bounded scale. The COEX value trended in the opposite direction than expected, but this may be due to the fact that it cannot be calculated on single-gene pairs and necessarily uses information from the entire count table, which here only consisted of 6 genes. These comparisons combined with the simulations in Figure S3.3, clarify a conceptual advantage of the L-metric over the EEI and COTAN frameworks. In contrast, by leveraging the ranked structure of transcript counts across cells, the L-metric framework does not binarize expression and does not require fitting a parametric distribution. It can be calculated on single gene pairs, and is more sensitive to mutual exclusivity.”

      We also amended the Data and Code Availability section to include a specific reference to the L-metric package that was previously missing:

      “All code used to process the raw data and generate figures, as well as the processed data and figures can be found at the following link :

      https://www.dropbox.com/scl/fo/bchkqlbcjb8ub9m606did/AIudcWZaC566toXzb2L-jXc?rlkey=u0wgtkq8j erxoqb5ump599oip&dl=0

      Additional custom scripts used to process the raw seqFISH data can be found on GitHub:

      https://github.com/arjunrajlaboratory/NimbusImage/

      The code for calculating the L-score can be found on GitHub:

      https://github.com/arjunrajlaboratory/l-metric

      Images of all gastruloids generated for this study, as well as single-channel seqFISH images with segmentation and annotations are available here:

      https://app.nimbusimage.com/#/project/69d3fa8f1be4701f5fab6359 Raw seqFISH images are available upon request.”

      (e) Clustering using L-metric vectors

      In this study, the authors use hierarchical clustering to group genes according to the L-metric. This choice is reasonable: hierarchical clustering provides a natural representation of similarity relationships across multiple scales, and the L-metric captures a form of signed dissimilarity between genes. However, this approach raises an important issue. Standard hierarchical clustering methods typically assume a non-negative metric, whereas, as noted earlier, the L-metric can take negative values, meaning it does not strictly satisfy the requirements of a metric in the mathematical sense.

      To the best of our understanding from both the text and the source code, the authors address this issue by defining the distance between two genes A and B as the Euclidean distance between two vectors: the L-metric values from A to all other genes, and from B to all other genes. Although this procedure is mentioned in the manuscript, it is neither justified nor accompanied by any discussion of how such a distance should be interpreted. It is not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes. These two gene distances are not without overlap, but they are not the same.

      This choice magnifies the interpretability issue: readers must understand two layers of transformations. If the [-1,1] range poses problems for hierarchical clustering, simple transformations (e.g., 1 − L) or alternative clustering methods could avoid these issues.

      Given that the L-metric underlies major biological inferences (novel gene modules, spatial subclusters, endothelial states), clearer justification and benchmarking are essential. We think this can lead to more consistency in the spatial metrics.

      We really appreciate that the reviewer took the time to understand our proposed method thoroughly, and apologize for any confusion resulting from a lack of clarity in how it is calculated, and the language used to describe it. The reviewer is absolutely correct that it is not a metric in the mathematical sense; we have replaced the word ‘metric’ with the word ‘score’ throughout the text.

      The reviewer also raised concern about how the hierarchical clustering was performed, and they were absolutely correct about what the vectors represent—they are, as the reviewer states, “not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes”. The heatmaps in figures 3, 4, and 5 are intended to cluster genes that have similar expression patterns, i.e. similar L-score values with all other genes. The reviewer pointed out that this transformation wasn’t clear, so we have added an explicit explanation of what the vectors represent, as well as an explanation of why we were interested in how these vectors, which represent a ‘fingerprint’ of how the gene interacts with all other genes, clustered (because this is in the section discussing NMP differentiation we focus on a specific subset of genes here, but later apply to the entire panel):

      “For each gene annotated as belonging to any of the three cell types (NMP, PSM, or spinal cord), we calculated a vector of scL-score values with all other genes. Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      We also appreciate the reviewer’s suggestion that alternative transformations of the scL-score may improve clustering interpretability. We tried using the reviewer’s suggestion of doing a simple transform: we averaged the scL-score matrices across the 18 gastruloids using the shared gene panels, transformed each scL-score from the original [-1,1] scale to a [0,1] scale using (1-scL)/2, symmetrized the resulting matrix so that each gene pair was represented by a single value, and then performed hierarchical clustering directly on this gene-by-gene distance matrix using average linkage. We used this analysis to test whether a more direct distance-based approach would change the gene groupings recovered by our original clustering method.

      This alternative approach largely recovered the same cell type-associated groupings, but the separation between branches in the dendrogram was smaller, making fine-scale ordering harder to interpret. We measured this by evaluating the average cell type dispersion, measured in terms of additive branch length, which was 0.682 and 0.671 for the 166-gene and 202-gene panels, respectively. Since the branch separation on our original dendrograms was greater and thus representative of more robust groupings, we chose to continue using hierarchical clustering based on Euclidean distances between scL-score vectors.

      We comment on this alternative transformation we tried in the next section, when considering the clustering of the entire gene panel. We feel this is appropriate as the motivation for looking at the difference vector was derived from expected behaviour of a smaller set of genes, and as the reviewer pointed out it is not clear that that expectation should or would hold for the entire panel.

      “To generate the heatmap shown in Figure 4a and S4.1a, we used the same clustering method as described earlier with the Euclidean distance between scL-score vectors. However, we also tried clustering directly on the scL-scores themselves, by transforming each scL-score from the original [-1,1] scale to a [0,1] distance-like scale using a (1-scL)/2 mapping (Figure S4.3a,b). The results were largely consistent, however the cophenetic distance scale was relatively compressed when the transformed values were used (Figure S4.3a,b). This shallow structure implies that many branches are separated by only modest distances, so fine-scale ordering within the dendrogram should be interpreted more cautiously than the larger-scale cell type block structure. We chose to continue using Euclidean distance-based hierarchical clustering of scL-score profiles, where cell type grouping is observed alongside larger cophenetic separations between clusters” (See New Figure S4.3).

      (3) Claims of "remarkably consistent" spatial organization are stronger than the data currently support

      (a) The manuscript emphasizes reproducible organization across gastruloids, but several factors complicate this interpretation.

      We agree that the distinction between what is reproducible/invariant between gastruloids and what varies was not clear in the original manuscript. To address this, we have updated the text to more explicitly distinguish between the two categories. The other suggestions made by the reviewer to sharpen the quantitative measures to strengthen these claims was very helpful and we appreciate the thought put into them, and have used that framework (emphasizing the statistically significant variations and consistencies) in summarizing our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrate several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior:posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (b) Possible selection bias. Only elongated, QC-passing gastruloids were retained; 18/26 datasets remain. seqFISH runs with uneven housekeeping signals were excluded.

      We agree with the reviewer that our data are elongated, QC-passing gastruloids, although these represent two sources of variation (biological and technical respectively). Our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category, and we have attempted to signal this to readers by consistently including language like “morphologically normal” and “elongated”. We have further updated the language in the manuscript to emphasize this point:

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. To ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].”

      In regards to the exclusion of datasets, the only time 18 out of 26 were used was when calculating the averaged L scores for all genes in Figures 4 and 5. In this case we used all 18 gastruloids from the seqFISH run performed on 4/4/2025; this dataset had the highest spot counts due to protocol improvement between runs, and integrating the datasets with very different spot counts was problematic because a minimum expression level is needed to calculate L scores. We used all 26 samples for the spatial metrics calculated in Figures 1 and 2 (Figures 3 and 6 focus on specific gastruloids). We have added additional labels in Figures 1, 2, 3, and 6 to make clear when all 26 datasets are used and when only a subset is used.

      (c) Gastruloids are known to be variable; restricting to morphologically "normal" samples could inflate apparent regularity.

      We thank the reviewer for this observation and agree that the degree of variability among gastruloids is an important consideration. As stated in response to b), our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category. The rate of occurrence of ‘normal’ gastruloids in our hands is ~80% (Figure S1.1a). We agree that it would be interesting to consider how variations from this baseline affect cell type composition and arrangement, and while we make no claims about it in this paper, we have updated the introduction to highlight this point:

      “To address these gaps, and to create a systematic, high-resolution dataset of gene expression in gastruloids considered to be morphologically normal, we developed a spatially resolved, single-cell molecular map of the location, identity, and gene expression of cells within 26 individual gastruloids with normal morphologies. We found that despite some morphological variability within the qualitative category of elongated and polarized, “normal” gastruloids had largely reproducible cell type composition.”

      (d) Partial lack of statistical validation. The manuscript shows descriptive consistency but no formal tests across runs or batches (e.g., mixed-effects models, ICCs, leave-one-run-out validation).

      We agree with the reviewer that a quantitative comparison between batches is important. We have added the following to the text to address this point:

      “To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (See Revised Figure S1.1a-c).

      (e) 2D sampling limitations. Spatial metrics rely on a single imaging plane chosen as the "midplane," but z-position varies between gastruloids. AP projections, mixing, and triplet analyses could all be sensitive to z-plane choice. Prior work shows that 2D slices can underestimate distances and contacts by large margins (https://pmc.ncbi.nlm.nih.gov/articles/PMC5522766). These limitations should be acknowledged explicitly.

      We agree that sampling in 2D can limit the interpretation of our findings and we thank the reviewer for bringing up this important point. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      “Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.”

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2017].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (4) Uncertainty in cell-type assignment is not incorporated into spatial metrics

      Many spatial measurements depend directly on cell-type calls (exposure, triplets, mixing). However:

      (a) Anterior cell types have higher entropy in their marker-based scores (Figure S1.1a).

      We thank the reviewer for pointing out that several of the cell types in the anterior have high entropy — specifically cardiac mesoderm and paraxial mesoderm. However, we think there is nuance to this point; two of the other prominent anterior cell types (somite and endothelial) have low entropy scores overall, and that spinal cord/neural precursor cells, which are mostly posterior, have somewhat higher entropy; higher entropy scores are not exclusive to the anterior, nor is low entropy exclusive to the posterior. Inspired by the reviewer’s comments, we have re-analyzed our data to include this nuance (see response to point d) below.

      (b) Uncertain labels inflate apparent "mixing" or "disorder," because misclassifications randomly create mixed neighbors and triplets.

      We agree that uncertainty in labels could affect the interpretation of mixing. We appreciate these comments and the reviewer’s suggestions, and we have followed them in our response to point d) below.

      (c) Posterior cell types have low entropy, so comparisons between anterior vs posterior mixing may partly reflect label uncertainty, not biology.

      We thank the reviewer for bringing up this important caveat to our findings. We incorporated discussion of this in our text edits (see point d) below.

      (d) The authors should incorporate confidence measures (e.g., probability-weighted neighbors, entropy filtering, bootstrapping) to confirm that patterns hold independently of classification noise.

      We appreciate these suggestions and have chosen to use entropy filtering to assess whether the spatial organization we observe is highly sensitive to what values are considered ‘low’ entropy. We have updated the text (see below), and added Figure S2.2 to address the reviewer’s comments:

      “The contrast between organized posterior clustering and disorganized anterior mixing matches expectations based on literature that shows that self-organization mechanisms in gastruloids in the anterior vs. posterior are differentially sensitive to culture conditions, with somitic patterning requiring external matrix support [1,4,8,24], distinguishing it from the seemingly more autonomous organization observed in posterior cell types.”

      “One potential caveat to this finding is that differences in uncertainty in cell typing could be the primary driver of mixing and cell type interaction differences, both between the anterior and the posterior within an individual gastruloid, or overall between gastruloid. To control for this, we applied an entropy filter to our dataset. We filtered out cells that had entropy > 1.5 (see plot below for cutoff), which was chosen based on the distribution of entropy values for ‘none’ type cells, which effectively describe the upper limit of random transcript assignment (99.7% of ‘none’ type cells are removed with this filter, and about 50% of cardiac mesoderm cells and paraxial mesoderm cells, see Figure S2.2a). We first examined overall mixing; there was strong correlation between the per-gastruloid mixing indices before and after entropy filtering (Pearson r = 0.809, Figure S2.2b). Globally, mixing indices decreased with filtering, meaning that overall the cell types were more clustered. When we compared the absolute value of the change in mixing index pre and post-filtering to the proportion of each cell type, the only significant correlation was with cardiac mesoderm (Figure S2.2c). Exposure indices were overall quite similar after filtering, although the strength of somite-somite and somite-paraxial mesoderm interactions increased (Figure S2.2d).”

      “We also examined how entropy filtering might affect the exposure index, given that the mixing index is calculated from the exposure index of across cell types. In general, the magnitude of the exposure index values increased when more uncertain cells were excluded, but the directionality and relative ordering was not affected. Although the magnitude of change in the posterior cells types was less than the anterior cell types, the cross-cell type exposure values, particularly between paraxial mesoderm/endothelium and somites, doubled. From these results we conclude that mixing in the posterior is driven mainly by NMP/presomitic mesoderm interactions, and is overall lower than mixing in the anterior, which is driven by rarer cell types like cardiac mesoderm, paraxial mesoderm, and endothelium, being interspersed within somite cells” (See Figure S2.2).

      (e) This leads to reviewing the claim on endothelial heterogeneity, which strongly depend on spatial adjacency and gene exclusivity metrics.

      We have extensively considered claims of endothelial cell heterogeneity, and these are discussed in detail in response to the reviewer’s next point. We have also copied them here for the reviewer’s convenience:

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in the Author response image 3:

      Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although many of the somite-associated genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0). See Author response image 4.

      Author response image 4.

      (5) Endothelial "spatially dependent" gene expression may reflect spillover rather than intrinsic state

      The comparison between anterior-associated and posterior-associated endothelial nuclei suggests two transcriptional states. However, spatial adjacency confounds the interpretation:

      (a) seqFISH assigns transcripts to nuclei in dense tissue; partial-volume effects can mix RNA from neighboring endodermal or somitic cells.

      We thank the reviewer for their attention to detail and agree that a careful consideration of these points is important. We also note that given that some of our differentially expressed genes in endothelial cells are endoderm or somite genes, there indeed may be some transcript misassignment.

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript misassignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for none-typed cells at any dilation considered.

      (b) Endoderm and endothelium are closely intermixed (Figure S6.1d), and their gene expression co-localizes in KDE maps (Figure 5b-c).

      We agree with the reviewer and thank them for their close reading of the manuscript. We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      (c) Without explicitly quantifying spillover, differential expression between these two endothelial subsets cannot be confidently attributed to cell-intrinsic differences.

      (d) You may control for spatial proximity with any of the following:

      - Include adjacency index as a covariate in DE models.

      - Use scL-metric to test the mutual exclusivity of endothelial vs endoderm genes within the same nucleus.

      - Apply local permutation nulls: shuffle transcripts within local windows and recompute DE.

      - Restrict analysis to gastruloids that contain both endothelial subsets.

      We thank the reviewer for bringing up this important point, which we were eager to address. We agree with the reviewer’s point that by only considering gastruloids that contain both subsets of endothelial cells is the correct way to do the analysis. We were already only considering this case (specifically the values calculated in Figure 6 and for 1 gastruloid pictured in Figure 6a). We added an n=1 label to revise Figure 6d to emphasize this point.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although some genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0) (See Author response image 4).

      We have also included in the supplement larger images of some of the top differentially expressed genes, which more intuitively show the differential expression results (Endoderm enriched and Somite enriched).

      Finally, we have referenced several previously-reported instances in the literature where distinct subsets of endothelial precursors with unique gene expression programs were identified. Although in these cases 1) the embryo models were different (in [Rossi 2021, Rossi 2022] gastruloids made with a different protocol and treated with factors designed to promote blood development, and in [Veenlveit 2020] trunk-like structures) and 2) the methods were different (IF and 10x single-cell sequencing) this at least establishes a precedent for the observation of multiple types of endothelial precursors. In the case of [Veenlveit 2020] the authors specifically note that one subset is associated with somites, and we have updated the text to reflect these new results:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b,c). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Pecam1 also clustered uniquely in our L-metric analysis (Figure 4c), suggesting this differential expression is conserved across gastruloids. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behaviour is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6)

      (6) Interpretation of gene-program modules may be overstated

      Claims that the L-metric reveals "novel gene programs" should be softened:

      (a) The seqFISH panel is an approx. 200-gene marker-enriched panel, already biased toward known cell-type markers.

      This is true and we appreciate that this came through in the text since it’s important for the reader to understand the approach we took in this study.

      (b) Strong blocks in Figure 4a and S4.1a may reflect panel design rather than newly discovered programs.

      We agree with the reviewer that the panel design was not sufficiently highlighted, so we have made the following changes to the text to emphasize which patterns would be expected due to the genes we are probing for, and which findings were surprising given the known functional role of the gene.

      At the end of the section titled ‘The L-metric captures the spatial distribution of gene expression despite being calculated without spatial information’:

      “These analyses demonstrate that information contained within the hierarchical relationships between genes, determined by scL-score can reveal novel information about cell states within cell types, although we acknowledge that since cell type is determined by a limited panel of marker genes, results should be further functionally verified. scL-score analysis can also identify distinct spatial locations of cells in this cell state, all without explicit encoding of spatial information, but rather quantifying and clustering the degree to which genes are mutually exclusively expressed with one another.”

      At the end of the section titled ‘Clustering scL-metric vectors clearly resolves cell types and reveals novel genetic interactions’:

      “Finally, although the strong blocks we find in the heatmaps in Figures 4a and S4.1a largely reflect cell types, as is consistent with our panel design, we discovered some novel functions of genes in the panel through their location in the scL-score tree: although Tgfβ was initially included in our panel to generally detect inflammatory and growth signaling, clustering by expression patterns revealed its unique association with endothelial precursors.”

      To further address the concern that the generality of clustering is due to gene selection, we performed random gene drop-out and assessed how well cell types clustered as a function of the number of genes removed:

      Author response image 5.

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      Finally, we performed scL-score analysis on an unbiased, scRNA-seq dataset without any panel selection, and were able to show similar groupings of cell type markers:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (c) Cluster robustness is not assessed (bootstrap, stability).

      We appreciate the reviewer’s suggestion that cluster robustness should be assessed. To address this point, we performed a clustering-stability analysis on the common 202-gene panel that included cell cycle genes by asking to what extent hierarchical clustering could recapitulate cell type-based groupings of genes as the number of genes used for clustering was varied across progressively larger, randomly sampled panel subsets. For each resulting tree, we averaged cell types’ dispersion of genes across the tree using the framework described in our response to suggestion 6m and compared the observed value to a permutation-based distribution generated on the same tree.

      This analysis showed that the observed cell-type dispersion remained consistently lower than the corresponding permutation distribution across all subset sizes examined. In other words, genes assigned to the same annotated cell type remained closer together in the dendrogram than expected by chance even when clustering was performed on reduced random subsets of the panel. We also observed that dispersion values increased as larger gene subsets were included, which is expected as the clustering problem becomes more complex with increasing panel size; however, the separation between the permuted distribution of average cell type dispersion and the observed dispersion value increased as the panel subset size increased. Taken together, these results indicate that the cell type-resolved organization captured by the scL-score derived hierarchy is not dependent on one particular subset of genes, but is instead a stable property of the broader gene panel. New Figure S4.2 addressing cluster stability:

      We have added the following explanation in the text:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      (d) Agreement with cNMF (claimed in text) is not quantified (ARI, Jaccard, hypergeometric overlap).

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See Revised Figure S4.2 (now S4.4)).

      We have updated the text to reflect these quantitative comparisons:

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (e) Testing the scL-metric on larger, unbiased scRNA-seq datasets would help demonstrate generality.

      We agree with the reviewer that this would demonstrate generality, so we applied scL-score analysis to the scRNA-seq dataset from van den Brink 2020 — several gastruloids at the same stage of development as those used in this paper were pooled and sequenced. We added a figure, new Figure S4.5 with the results of this analysis, and a new section in the paper describing them:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (7) Minor Comments

      (a) We suggest including representative raw seqFISH images. The manuscript does not show raw images, which makes it difficult to evaluate the quality of the underlying data that all spatial analyses depend on. A figure showing raw fluorescence channels, detected spots, and nuclei segmentation masks for at least one anterior region, one posterior region, and one dense interface (e.g., endoderm-endothelial) would allow readers to assess spot intensity and background levels, signal-to-noise ratio, segmentation accuracy, and potential over/under-segmentation, channel cross-talk, and transcript crowding or dropouts in dense tissues. A small panel of raw images would improve the transferability to the spatial metrics.

      This is an excellent suggestion and we have included examples of the raw images for a representative gene for all hybridizations, the spots as determined by the spot-finding algorithm distributed with the seqFISH instrument, deconvolved spots, and nuclear segmentation. These images can be found in Revised Figure S1.1d.

      (b) The Introduction mentions 3D organization, which can confuse readers into thinking the seqFISH dataset is volumetric. The data shown and analyzed come from a single 2D plane per gastruloid, not from full 3D z-stacks. Since all spatial metrics rely on true adjacency, the manuscript should explicitly state early on that the dataset is 2D and briefly justify why a single plane is sufficient for the analyses.

      We appreciate the reviewer bringing up this point and we have removed references to 3D so as not to confuse readers.

      We also have included an extensive discussion of the 2D nature of the data. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2018].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (c) Results, first paragraph: "good agreement" along the AP axis should be quantified or defined.

      In the first paragraph we compare the peak in gene expression along the (length-normalized) AP axis of each gene with a similar but orthogonally measured dataset from another group (van den Brink 2020). The text specifically reads:

      “When we compared how gene expression varies along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section.”

      We have revised Figure S1.2 with a summary plot showing the distribution of correlation coefficients for all genes and for the Hox genes in our panel (which are known to be expressed sequentially along the AP axis).

      We have updated the text as follows:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      (d) Clarify what "greater cell type distinction" means when using marker panels, and what metric demonstrates improvement?

      By “greater cell type distinction” we meant that when using traditional clustering methods we were not able to individually resolve some cell types: there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together. Cluster labeling was performed by considering which genes showed up as being differentially expressed in each cluster using the same associations as were used to perform the cell type scoring with the marker gene panel. While we acknowledge that these results are subjective to clustering parameters, this is a general problem with clustering and not specific to this study. We have added additional clarification in the text and removed the phrase “greater cell type resolution” since it wasn’t clear what we were comparing to:

      “To profile the spatial organization of individual gastruloids, we assigned a cell type to each nucleus using a cell type scoring method with known marker genes. We compared these results to those obtained with unsupervised clustering. We found that although clustering did produce clusters, they were not strongly separated and we were not able to individually resolve some cell types we expected to find: specifically, there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together (Figure S3.6a). Therefore, we proceeded with the scoring-based method; see Methods for additional details. A representative gallery of typed gastruloids is shown in Figure 1b; the full dataset is in Figure S1.3. Because each cell received a cell type score for each type, we could use the entropy of the cell type score probability distribution to assess confidence of our cell type assignment: a cell that received a similar score for multiple cell types would have a high entropy distribution, while one which scored highly for one type and low for the rest would have low entropy. On average, the cell type entropies for most cells in a given cell type were low, with the exception of paraxial mesoderm and cardiac mesoderm cells, which had intermediate values (see additional discussion below). The generally low entropies indicate that most of our cell type assignments were high-confidence (Figure S1.4a).”

      (e) The variation in "none-typed" cells across gastruloids should be expanded and shown quantitatively.

      We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      (f) Figure S1.1b: explain the criteria used to decide which tissues "varied significantly." Showing the proportion of "None" per gastruloid would help.

      We thank the reviewer for pointing this out. We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      To justify the use of the phrase ‘varied significantly’ we have calculated the degree to which cell type pairs significantly co-vary (adjusted p value < 0.05) in their proportions and added the plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We address said variation in the text:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids profiled. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.32a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.41dc).”

      “The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.41dc).”

      (g) In Figure 1c, showing the variability of "None" cells is important.

      We made this adjustment and added the plot to Figure S1.3c

      (h) For Figure 1e, consider normalizing to T-expression proportion, not only AP length. This could clarify multimodality in the endoderm and spinal cord.

      We thank the reviewer for this suggestion to normalize by molecular as well as physical features. We agree this could be a useful projection of the data, since expression of T is used in many contexts to define the posterior of gastruloids.

      We incorporated T expression into the length normalization in the following way: we generated a cumulative distribution of all T spots along the AP axis, and when this value exceeded a threshold (specifically 90% of all spots) we defined this as the midpoint of the gastruloid, and linearly normalized space before and after it from 0-0.5 and 0.5-1 respectively. We then re-projected nuclei for each gastruloid onto this new coordinate system, and visualized in the same way as the main text figure (See Figure 1f).

      To assess whether this increased or decreased variability in position, we calculated how much the location of the peak of each cell type in each gastruloid differed from the peak position of that cell type in all samples pooled together. We compared this difference to a bootstrapped null drawn from the pooled distribution. Interestingly, we found that nearly all cell types had statistically significant variation (meaning the average distance from the mean fell outside the bootstrapped distribution in the positive direction), however the effect size for most cell types was very small.

      Notably, as the reviewer suggested, normalizing to T expression decreased the effect size of this variability in all cases but one, and particularly decreased spinal cord variability (despite not being a marker for spinal cord):

      We interpret these findings in the following way — that most of the variation observed in cell type arrangement along the AP axis is due to morphological and molecular variability (likely due to stochasticity in initial cell number and differences in developmental timing between gastruloids), and that once these factors are taken into account, for most cell types the effect size of variability is very small. However for some cell types, notably the location of the differentiation front, the position varies among gastruloids. We hypothesize that this may be due to the rhythmic nature of somite development. We have updated the text to reflect these changes.

      “Given that the proportions of cell types within each gastruloid were fairly consistent, to what degree did their spatial organization vary? We first projected each cell’s expression onto the AP axis and looked at the distribution of AP axis locations across gastruloids (Figure 1f, left-hand side). The most posterior cell types (NMP and presomitic mesoderm) showed wide distributions from 0-30% of the axis. Centred around 30%, spinal cord precursors and endoderm had distinctive peaks (clusters), the exact location of which varied between gastruloids. Using a threshold of T expression to define the midpoint of each gastruloid and uniformly length-normalizing each half caused these cell types to collapse into a single peak (Figure 1f, right-hand side). The differentiation front was similarly located in one peak (between 30-40% of the AP axis), and the variation in the location of this peak decreased with T-expression normalization. To quantify this variability, we bootstrapped a null distribution of cell type locations by pooling each cell type together across samples and using the resulting distribution to define a reference mean. When we compared how much peak variation there was among samples randomly drawn from this null to our observed data, we found that although nearly every cell type (excepting paraxial mesoderm) had more variability than expected by chance, the effect size of this variation was small (Figure S1.5a). Normalizing location to T-expression decreased variability in most cell types, most notably for differentiation front and spinal cord (Figure S1.5b).

      We also asked how the order of cell types along the AP axis varied between gastruloids. When we ranked the peaks shown in Figure 1f per gastruloid, we found that Kendall’s W, an overall measure of rank coherence across independent samples that spans from 0 (no agreement) to 1 (complete agreement), was 0.834 (Figure S1.5c). We found that the cell types most likely to swap rank order were spinal cord and endoderm, and paraxial mesoderm and endothelium (Figure S1.5d).”

      (i) Clarify "mixing coefficient values range from 0.29-0.58" (incomplete sentence). Section "Cell types' arrangement is consistent across gastruloids, but varies by type" second paragraph.

      We appreciate that there was some confusion here and we thank the reviewer for pointing it out. We have updated the text with the new values and a more specific interpretation to aid the reader and address the reviewer’s comment:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). While all values are negative, the range was large. This observation led us to conclude that while in all the gastruloids profiled cell types tended to cluster together, there was variation between gastruloids in the degree of coherent clustering between types.”

      (j) The paragraph claiming organization consistency may be too strong, given the lack of statistical validation. "Overall, these data speak to the consistency of gastruloid organization...".

      We agree that quantification of variability, which we only assessed qualitatively, would help readers better evaluate claims of organizational consistency, and we appreciate the reviewer pointing this out.

      We have taken several steps to add quantification, including:

      (1) Calculating the coefficient of variation for the proportion of each cell type

      (2) Quantifying co-variation and assessing statistical significance

      (3) Quantifying the variation in AP-axis location, including comparison to a bootstrapped null, significance testing, and a new form of normalization which decreases some of the variability.

      (4) Quantifying order along the AP axis and calculating Kendall’s W to quantify how concordant this ordering is across gastruloids

      (5) Overhauling the methods used to quantify local patterning and mixing into one unified metric

      (6) Calculating significance for variation in exposure index across samples

      To address the question of whether claims of organization consistency are appropriate, we have added the following summary paragraph at the end of the results for the first two figures, and have taken the reviewer’s suggestion of only discussing statistically significant or quantified results in listing both consistent and variable features. Now rather than making an argument about whether gastruloids are consistent or not, we merely provide the readers with our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (k) Figure 3c: gene set sizes differ substantially; proportion-based normalization may be more appropriate.

      We thank the reviewer for carefully noting the gene set differences. While the NMP-only and spinal cord-only gene sets each have 10 genes, PSM-only and NMP+PSM have 6 genes, and NMP+spinal cord has 5 genes. Given the relatively small N, we feel that normalization would likely introduce a layer of abstraction that would be more confusing for the reader, especially given the qualitative nature of the claims made about the shape of the plots in question. However, we agree this point is important, so we have now indicated the size of the gene sets on the plot (revised Figure 3b, previously c).

      (l) The purpose of the first two graphs in Figure 3c is unclear.

      We thank the reviewer for pointing out that this is unclear. The purpose of this visualization is to demonstrate correlation between genes shared between cell types and genes exclusive to only one of the cell types. The text reads:

      “Figure 3b shows the total expression of each gene group versus the NMP genes for all cells in all gastruloids that we typed as NMP, presomitic mesoderm, or spinal cord. As expected, there was a clear correlation between the mixed categories and NMP genes, supporting the notion of a continuous differentiation process.”

      We have added a correlation line to revised Figure 3b (former Figure 3c) to emphasize this point.

      (m) Clustering of all genes: quantify whether clusters align with cell types.

      We appreciate this suggestion offered by the reviewer as this analysis will allow readers to quantitatively assess the overlap between cell type-based groupings of genes and our scL-score-determined clusters for this dataset. To address this, we have used a bootstrapping approach to evaluate how well our scL-score-based hierarchies capture cell type-based groupings. Specifically, following hierarchical clustering of genes using the scL-score, for each cell type, we computed the average of the minimum cophenetic distance between each pair of genes associated with that cell type. After obtaining each cell type’s average, we took the mean of these averages which we refer to as the cell type dispersion for that tree. We note that the absolute value depends on the topology of the tree. To establish a reference for this measure, we bootstrapped a null distribution by preserving the same tree topology and randomly assigning genes to leaves and calculating the resulting dispersion. We performed this 10,000 times to create a reference null distribution. We compared the true dispersion value to this null, and computed a bootstrapped p-value. In both cases (with and without cell cycle genes), the observed average cell type dispersion is less than that for all permutations (p-value=0.0001, bootstrapping), indicating that genes associated with the same cell type were significantly more likely to cluster together on the observed scL-score hierarchy than would be expected under random clustering.

      We have updated the text to reflect these quantitative comparisons:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type-specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b).

      Author response image 6.

      (n) Clarify how the distance between genes is computed for clustering; the current method is hard to interpret.

      We appreciate the reviewer’s suggestion to clarify how distances between genes were computed for hierarchical clustering. To clarify this point, we have added the following to the text:

      “Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      (o) Quantify similarity between cNMF and L-metric clusters.

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2, now in updated Figure S4.4c). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See updated Figure S4.4 (formerly Figure S4.2)).

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (p) In Figure 3e, clarify what the orange lines represent.

      We have updated the text:

      “Per-cell expression scatterplots of the two pairs of genes shown in b). The y-axis of each is the per-cell expression of Eogt. The x-axis is the per-cell expression of Pax6 (left) or Rfx4 (right). R is Pearson’s r, scL is scL-score. Count data is shown in black; smoothed 2D densities are shown in orange.”

      (q) Ensure figures and panels follow the text order. For example, Figure S1.3a is referenced earlier than 1.1 and 1.2.

      We appreciate the reviewer’s attention to detail. We have changed the order of these figures so that their reference in the text follows their numeric order.

      (r) Typo in "by covariation with any other cell type (Figure S1.1e)" did you mean S1.1.c?

      We have updated the text with this change.

      (s) For circularity (Figure 6), justify the convex-hull-based measure; thin protrusions can distort interpretation. Consider the volume difference between the convex hull and the original shape.

      We thank the reviewer for this helpful suggestion. We tested the difference method suggested by the reviewer, as well as several other methods of clustering and calculating circularity. We determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and text.

      Summary

      This work delivers a rich spatial dataset and introduces creative computational tools. The main limitations lie not in the data but in the clarity, justification, and validation of the quantitative methods. We suggest strengthening these aspects by adding formal definitions, parameter justification, benchmarking, robustness tests, and controlled interpretations. This will improve the manuscript's impact and reproducibility of the methods described, besides making it easier to understand for the readers.

      We thank the reviewer for their kind assessment and also for their many insightful comments for improvement. We feel the revised manuscript is greatly improved because of them.

      Reviewer #3 (Recommendations for the authors):

      In my view, this manuscript is well-designed and clearly written, and supports all of the claims made. I have no suggestions for major revisions for this manuscript; rather, I would suggest the following as outstanding questions for future investigation:

      We thank the reviewer for a careful reading of our manuscript and the several interesting suggestions and useful references in the literature. Including the discussion of these ideas in the text (see below) has improved the flow and scope of the manuscript.

      (1) On the NMP fate, bifurcation has been studied extensively in gastruloids and related structures (see, for example, Underhill et al 2023; Bolondi et al 2024). Can this dataset from Triandafillou and colleagues reveal new regulatory hierarchies in this process? This seems possible in principle, but I was not able to reach this interpretation (for example, it seems the authors interpret the clustering in Figure 3H as reflecting spatial patterns rather than a regulatory hierarchy).

      We agree that this is an exciting implication of the work, but we feel that with the current panel (which was originally chosen primarily to type cells and not to infer regulatory structure) we would not be able to comment on this. However we feel that with a larger set of genes such inference may be possible, so we took your suggestion in point 3) and applied the L-metric to a previously existing single-cell dataset. We looked for possible regulatory structure, and found at least one interesting case where a transcription factor showed strong co-expression with two previously unconnected genes. While more specific analyses and experimental validation would be required to establish a direct relationship, we feel that this suggestion by the reviewer represents an important potential future application of this methodology, so we’ve included a new figure and explanatory text to address this possibility:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.

      (2) On the formation of endothelial clusters ('blood islands'), this process has been observed in gastruloids (Rossi et al 2022). The observation of endoderm-associated endothelium in gastruloids is interesting, but it was also not clear how the authors interpret this finding. Are there two separate endothelial differentiation paths captured here? Or is there one path, and only some migrate towards the endoderm? The authors seem to raise possibilities, and it was slightly unclear on my reading how they interpreted their findings.

      We thank the reviewer for pointing us to this paper. We think it’s especially interesting that we also see close association between endoderm and endothelial precursors, especially given that the protocol used in the referenced paper was designed to generate blood precursors (treatment with VEGF, bFGF and ascorbic acid). In light of the comments of other we re-did this analysis — Author response image 7 shows the updated list of differentially expressed genes.

      Author response image 7.

      Some of the spatially differentially expressed genes are linked to signalling, and likely reflect overall signalling differences between the anterior (where the somite-associated endothelial cells are) and the posterior (where the endoderm-associated endothelial cells are). For example, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA is higher in the anterior). Pecam1 is involved in adhesion and cell morphology, and perhaps is higher in cells interacting with endoderm due to tighter packing/association; the same could also be true of Cdh5.

      While none of these answer the reviewer’s questions about the origin of the cells (and whether it is common), we found another instance of differential endothelial populations in embryo models: in Veenvliet 2020, they find that some endothelial cells have a mesodermal (somitic) origin. Thus we may be seeing a similar phenomenon in our samples. We have updated the text to reflect these additional lines of evidence and to clarify how we think the two populations may differ (while acknowledging that we lack the tools to confidently assign cell of origin or functional differences with this technique):

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.”

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      (3) Finally, a broader question about the analytical framework. The authors emphasize that the L-metric is parameter-free; however, much of their analysis still appears to rely on baked-in priors about known marker genes for cell type assignment. Is there a way to extend their analysis to infer cell types directly from the structure of the L-metric? The hierarchical clustering in e.g., Figure 3G suggests something like this: the hierarchy mostly (but not exactly) follows the marker gene annotation. Does this suggest that the cell type labeling should be revisited?

      Once of the initial motivations behind the creation of the L-metric (now L-score) was to have a more reliable and specific way of quantifying the interaction between genes that are known in the literature to be cell type markers — with more sophisticated (and noisy) methods of analysis like single-cell RNA sequencing, we found that there was substantial variation in the specificity and ubiquity of so-called ‘marker genes’, and that their usefulness often depended on context. We appreciate that the reviewer raises this point as well, and we think that scL-score analysis, such as that exemplified in Figure S4.4, can help identify new marker genes, or at the very least distinguish the biological context in which a marker gene is useful.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version:

      I thank the authors for their extensive efforts to revise the manuscript. I have no further concerns.

      Reviewer #2 (Public review):

      The authors have made several corrections to the original manuscript. For example, they revised the bootstrapping analysis to avoid arbitrarily inflating the degrees of freedom. However, most substantive concerns remain inadequately addressed.

      (1) The primary issue is still the lack of baseline models against which to benchmark the predictive performance of the proposed DenseNet model. This concern was raised independently by two reviewers. Without such benchmarks, it is difficult to interpret the reported results in the context of prior work on MRI-based cognition prediction.

      Notably, the authors state: "While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study."

      However, I could NOT find any discussion or results related to the CPM model in the manuscript. It is therefore unclear whether the DenseNet model was actually statistically compared with CPM, and, if so, how the comparison was conducted.

      Note that the statement, "While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data," is not entirely accurate. Linear models typically have far fewer parameters than deep-learning models and are therefore often less prone to overfitting. In fact, it is well established that deep-learning models are particularly susceptible to overfitting and usually require substantially larger sample sizes to achieve stable and reliable performance. Although deep-learning models may outperform shallower models once sufficient data are available and training is well controlled, this does not justify the authors' claim as stated. I therefore disagree with the argument put forward by the authors.

      The authors further justify the absence of benchmarking by stating: "In this context, deep learning was employed as a flexible framework capable of modelling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach." However, most shallow models can likewise be applied across different brain states and cognitive targets. This rationale does not establish deep learning as a uniquely appropriate or necessary choice. If deep learning is indeed a better approach in this context, the authors should demonstrate this empirically through appropriate benchmarking against established baseline models.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the primary goal of the present study was not to benchmark predictive algorithms, but rather to compare the predictive utility of different brain states using a consistent modeling framework.

      We did perform an exploratory CPM analysis using a conventional implementation that included correlation-based feature selection (p < 0.01), summarization of positive and negative networks, and robust regression for prediction using 3-fold validation. However, CPM performance can depend substantially on analytical choices, including feature-selection thresholds, treatment of positive and negative networks, cross-validation strategies, and model specification. Although we obtained CPM results (see below), we did not systematically evaluate how these choices influenced performance, nor did we optimize CPM to the same extent as would be required for a rigorous methodological comparison.

      For this reason, we chose not to include a CPM-versus-DenseNet comparison in the manuscript. Any direct comparison could easily be overinterpreted as evidence for the superiority of one approach over another, despite the absence of a comprehensive benchmarking framework. We therefore deliberately avoided such claims and instead focused on the scientific question motivating the study: whether predictive performance differs across brain states when the same predictive framework is applied consistently. We agree that comparisons with CPM and other predictive approaches would be valuable, but we believe such analyses are better suited to a dedicated methodological benchmarking study.

      (2) Additional analysis shows that "BCG is not significantly associated with cognition itself". This is the most perplexing result. This is like saying Brain Age Gap is not related to chronological Age. It is counterintuitive since the Brain Age Gap is calculated by chronological age minus actual age, and most research has shown a strong relationship between the Brain Age Gap and age.

      If the brain cognition gap is not related to cognition, is it possible that the results found are mainly due to the predictive model not fitting well with another dataset? Regardless, the lack of association between BCG and cognition deserves a discussion.

      We thank the reviewer for this comment. The absence of a significant BCG-cognition association might be unexpected. We agree it warrants careful interpretation. Theoretically, when a predictive model is trained and evaluated across samples with differing age distributions, the regression-to-the-mean dynamics that typically create BAG–age dependence may not transfer in the same way to BCG–cognition relationships, particularly in age-homogeneous cohorts such as COBRA. However, we acknowledge that the present findings alone are insufficient to resolve this question fully.

      We have added additional findings as supplementary material (Figure S4).

      (3) I still do not fully understand the rationale of the mediation analysis. The analysis and findings are still not related to aims 1 and 2, since DA and entropy are not part of the prediction models. But I appreciate the explanation that this part is related to the authors' previous work, and that the authors attempted to link to them somehow.

      We appreciate the reviewer's comment. The mediation analysis was not intended as a replication of our previous work, but rather as a mechanistic analysis to examine whether the association between BCG and cognition operates indirectly through brain age. Given the established links between BCG, brain age, and cognitive function, mediation analysis provides a principled framework for testing this hypothesis. Essentially, this mediation analysis supports previous theoretical model and our own empirical data where we showed lower DA contributes to more noise in FC metric, which in turn results in less accurate prediction (i.e., larger gap).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I hope the authors report CPM analysis as they claimed in the response letter, with actual statistical tests to compare the performance of DenseNet vs. CPM. It will be better to compare DenseNet with other models as well.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the goal of the present study was not to benchmark DenseNet against alternative predictive frameworks, but rather to compare the predictive utility of different brain states using a consistent modeling approach. Although we explored CPM as a potential reference model, we found that its performance was sensitive to analytical choices, including feature-selection thresholds, summarization of positive and negative networks, and model specification. A rigorous CPM comparison would therefore require systematic optimization and validation of these choices, which would constitute a separate benchmarking analysis beyond the scope of the present study. We therefore deliberately avoided claims regarding methodological superiority and revised the manuscript accordingly. While comparisons with CPM and other machine-learning approaches would be valuable, we believe such analyses would constitute a separate benchmarking study beyond the scope of the present work.

      (2) "n-back did not significantly predict EM in DyNAMiC, and rest did not significantly predict WM. For this reason, we highlighted only the conditions that showed meaningful predictive power in the original analyses."

      These null results should also be reported; otherwise, this suggests the authors cherry-picked only the "meaningful" results, making the claims overly optimistic.

      We thank the reviewer for this comment. We agree that reporting only significant findings could create the impression of selective reporting. However, this was not our intention. All within-dataset prediction results, including both significant and non-significant findings, are reported in Tables 1 and 2. For the cross-dataset analyses, we chose to evaluate only the best-performing models identified in the DyNAMiC dataset. This decision was made a priori to test whether the most reliable predictive models generalized to an independent dataset, rather than to maximize the number of significant findings. For example, resting-state FC emerged as the strongest predictor of episodic memory in DyNAMiC and was therefore selected for external validation in COBRA. Similarly, the movie-watching model showed the strongest performance for working memory and was consequently carried forward to the cross-dataset validation. We have clarified this rationale in the manuscript.

      (3) I appreciate the correlation plots between BCG and physical activity and cardiovascular risk. The results are much weaker in COBRA (r = .17 and -.10 vs. .40 and -.27 in DyNAMIC). Perhaps this warrants discussion. Note that there are potential outliers in DyNAMIC. Perhaps the authors might like to include Spearman's rank.

      We thank the reviewer for this helpful observation. We agree that although the associations between BCG and both physical activity and cardiovascular risk were statistically significant in COBRA, the effect sizes were smaller than in DyNAMiC. To assess whether the observed association was influenced by potential outliers, we repeated the analysis using Spearman’s rank correlation. The association remained the same after controlling for age using partial Spearman rank correlation - between GAP and physical activity (DyNAMiC: r =0.40, p =0.001; COBRA: r =0.17, p=0.03) and for GAP vs. CVD risk score (DyNAMiC: r =–0.27, p =0.03; COBRA: r = –0.10, p =0.40). Nevertheless, we have now tempered the interpretation in the Discussion (P.12) to clarify that the direction and significance of the associations were consistent across datasets, but that the magnitude of the effects was weaker in COBRA. This difference may reflect cohort differences, including age range, sample composition, and differences in how physical activity was assessed.

      (4) Yes, adding figures comparing BCG and BAG in the main text would be helpful, given BAG's popularity.

      We have added additional findings as a supplementary figure.

      (5) The authors should provide this reason as a justification in the method: "We initially attempted to predict both episodic memory (EM) and working memory (WM). However, EM prediction was only reliable within and across samples for the resting state, whereas WM prediction generalized most strongly from the movie-watching condition. Because COBRA does not include a movie-watching paradigm, we could not evaluate WM prediction across datasets. For this reason, we focused on EM when examining the brain-cognition."

      We thank the reviewer’s suggestion. This clarification is included in the method section of the revised manuscript [P 21].

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study is built on the emerging knowledge of trained immunity, where innate immune cells exhibit enhanced inflammatory responses upon being challenged by a prior insult. Trained immunity is now a very fast-evolving field and has been explored in diverse disease conditions and immune cell types. Earhart and the team approached the topic from a novel angle and were the first to explore a potential link to the complement system.

      The study focused on the central complement protein C3 and investigated how its signalling may modulate immune training in alveolar macrophages. The authors first performed in vivo experiments in C57BL mouse models to observe the presence of enhanced inflammation and C3a in BAL fluid following immune training. These changes were then compared with those from C3-deficient mice, which confirmed the involvement of C3a. This trained immunity was further validated in ex vivo experiments using primary alveolar macrophage, which was blunted in C3-deficiency, and, intriguingly, rescued by adding exogenous C3 protein, but not C3a. The genetic-based findings were supported by pharmacological experiments using the C3aR antagonist SB290157. Mechanistically, transcriptomic analyses suggested the involvement of metabolism-linked, particularly glycolytic, genes, which was in agreement with an upregulation of glycolytic flux in WT but not C3-deficient macrophages.

      Collectively, these data suggest that C3, possibly through engaging with C3aR, contributes to trained immunity in alveolar macrophages.

      Strengths:

      The conclusions reached were well supported by in vivo and ex vivo experiments, encompassing both genetic-knockout animal models and pharmacological tools.

      The transcriptomic and cell metabolism studies provided valuable mechanistic insights.

      We thank the reviewers for acknowledging the importance of the work.

      Weaknesses:

      For the in vivo experiments, the histopathological and other inflammatory markers (Figure 1) were not directly linked to alveolar macrophages by experimental evidence. Other innate immune cells (eg. dendritic cells, neutrophils) and endothelial cells could also be involved in immune training and contribute to the pathological outcomes. These cells were not examined or mentioned in the study.

      We agree with the suggestions from Reviewer 1 that other cell types such as dendritic cells, neutrophils, and endothelial cells can also be involved in immune training. As the focus of this study was on alveolar macrophages, we specifically focused on these cell types. However,

      (1) We have re-analyzed a recently published dataset of human volunteers who received aerosolized BCG exposure compared to saline. We observe that by Day 7, aerosolized BCG exposure alters the expression of C3 and C3aR1 in alveolar macrophages in the human bronchoalveolar lavage (BAL) fluid, compared to saline. We have included this new analysis in a revised Figure 1 to clarify why our focus is on investigating the C3-C3aR1 axis in alveolar macrophages.

      (2) We have conducted new experiments where we train the mice in vivo, collect the alveolar macrophages, and then provide the second stimulus ex vivo. We observe a similar phenotype in the in vivo trained, ex vivo stimulated alveolar macrophages. We have included this new data in a new Figure S2.

      (3) We have updated our Discussion to state “However, we also acknowledge that other immune cells such as dendritic cells and neutrophils, and non-immune cells such as epithelial cells, endothelial cells and fibroblasts can also be involved in immune training (Bigot et al., 2025; Friščić et al., 2021; Moorlag et al., 2020).”

      (2) For the ex vivo experiments assessing immune training in alveolar macrophages, only the release of selected inflammatory factors were measured. Macrophage activities constitute multiple aspects (e.g. phagocytosis, ROS production, microbe killing), which should also be considered to better depict the effect of trained immunity.

      We agree with the reviewer and have conducted additional experiments to assess immune responses influenced by training in alveolar macrophages. Specifically, we show that in addition to impairing the release of proinflammatory cytokines such as TNFα and IL-6, C3-deficient alveolar macrophages exhibit significantly lower phagocytosis and ROS production compared to WT alveolar macrophages post-training with heat-killed Pseudomonas aeruginosa. Results from these additional experiments have been included in new Figure S2.

      (3) The proposed mechanism of C3 getting cleaved intracellularly and then binding to lysosomal C3aR needs to be further supported by experimental evidence. 

      The mechanism of C3 being cleaved intracellularly involves serine protease-dependent cleavage of C3 to C3a and has been experimentally demonstrated previously (Liszewski et al. Immunity 2013; Elvington et al. J Clin Invest 2017). A prior report demonstrated that intracellular C3a interacted with a lysosomal C3aR to promote CD4<sup>+</sup> T cell survival (Liszewski et al. Immunity 2013). Based on the reviewer’s suggestions, we performed confocal microscopy on alveolar macrophages. Although we clearly observed intracellular colocalization of C3a (using a monoclonal antibody to the neo-epitope) with C3aR, we observed only some colocalization with LAMP1, a lysosomal marker (see Author response image 1). Hence, we will refrain from making comments on how C3 binds to lysosomal C3aR intracellularly in alveolar macrophages, as this may be cell type-specific or stimulation-specific. We have now revised the sentence in the manuscript to remove any references to lysosomal C3aR and now state – “Upon internalization, C3 is cleaved to C3a (Elvington et al., 2017), binds to C3aR, and affects cytokine production in CD4<sup>+</sup> T cells (Liszewski et al., 2013)”. We have not incorporated the Author response image 1 in the main manuscript as we would like to explore this further to precisely define the subcellular localization of C3a-C3aR in alveolar macrophages, but have provided it for the reviewer to explain the basis of the rewording in the revision.

      Author response image 1.

      C3a-C3aR colocalization in mouse ex vivo cultured alveolar macrophages (mexAM). mexAMs were harvested and cultured as per the protocol from Gorki et al. (2022). Cells were incubated in a Millicell EZ Slide 8-well glass chamber slide overnight to allow for adherence, then fixed, permeabilized, and incubated with anti-C3a conjugated to AF555 (blue, Hycult HM1072), anti-C3aR conjugated to AF647 (red, Hycult HM1123), and anti-LAMP1 (green, Cell Signaling 99437) overnight at 4°C. Slides were washed 3X in PBS (5 min each) and mounted overnight at 4°C in ProLong Diamond Antifade Mountant with DAPI (white). Images were acquired on a Zeiss LSM 880 confocal microscope at 63X. At least 6 cells per condition imaged. Experiments were conducted in duplicate (technical replicates) and repeated (for biological replicates). Scale bar, 2 μm.

      (4) There was an absence of any validation in human-based models.

      We acknowledge that the observations need to be validated in human-based models. The focus of our manuscript is on training in alveolar macrophages. Unfortunately, we do not have access to an adequate representation of human alveolar macrophages for our ex vivo testing to account for individual-level variation in immune responses. We anticipate this work will form the basis of these future studies. In the interim, we re-analyzed a recently published publicly available dataset of human BAL specimens from human volunteers who underwent aerosolized BCG administration (Marshall et al. Nat Comm 2025). We observe an increase in C3 and C3aR1 expression at Day 2, which persists through Day 7 post-training with aerosolized BCG compared to aerosolized saline specifically in human alveolar macrophages. We have included this data in Revised Figure 1. We also validated C3 uptake in alveolar macrophages using precision-cut lung slices from human donors. We have included this additional data in new Supplementary Figure 3.

      Reviewer #2 (Public review):

      Earhart et al. investigated the role of the complement system in trained innate immunity (TII) in alveolar macrophages (AM). They used a WT and C3 knockout murine model primed with locally administered heat-killed P. aeruginosa (HKPA). Additionally, they employed ex vivo AM training models using C3 knockout mice, where reconstitution of C3 and blockade of C3R were performed. The study concluded that the C3-C3R axis is essential for inducing TII in macrophages in the ex vivo model. The manuscript is well-written and easy to follow. However, I have the following major concerns.

      (1) The secondary challenge to assess the reprogramming of innate cells in the BAL was conducted 14 days after the initial exposure to HKPA. However, no evidence is provided to confirm that homeostasis was re-established following the primary exposure. Demonstrating the resolution of acute inflammation is essential to ensure that the observed responses to the secondary challenge are not confounded by persistent inflammation from the initial exposure.

      We thank the reviewer for giving us an opportunity to clarify this point. We have now included additional data from the bronchoalveolar lavage fluid of these mice to show that the levels of protein leaked into the BAL, levels of proinflammatory cytokines (e.g., TNFα, CXCL1) and the neutrophils (all relevant to the acute phase of inflammation) were similar between the untreated and treated wildtype mice. This new data has been included in Figure S1.

      (2) In Figure 1D, cytokine production by BAL cells from WT and C3KO mice after HKPA exposure and LPS challenge is shown. However, it is unclear whether the reduced response in trained C3KO mice is due to a defect in trained immunity or an intrinsic inability of C3KO cells to respond to LPS. To clarify this, the response of trained C3KO cells should also be compared to untrained C3KO controls after the LPS challenge. This comparison is necessary to determine if the reduction is specifically related to innate immune memory or a broader impairment in LPS responsiveness. Such control should be included in all ex vivo training and LPS stimulation experiments as well.

      We thank the reviewers for their suggestions. We have conducted additional experiments and we observe no significant differences in the BAL cytokine levels between the wildtype and C3-deficient mice post-training in the absence of infection. This new data has been included in Supplementary Figure S1.

      Additionally, we came across several manuscripts, including a recent one in eLife as a part of this Series (Gu et al. Elife 2021; Zahalka et al. Mucosal Immunol 2022; Prevel et al Elife 2025) that have done in vivo training followed by an ex vivo challenge. Hence, we have conducted new experiments to compare the response of in vivo HKPA-trained wildtype (WT) and C3-deficient (C3KO) alveolar macrophages compared to untrained AMs after an ex vivo LPS challenge. This new data has been included in Figure S2.

      (3) The data presented provide evidence of alterations in the functional and metabolic activities of innate cells in the lung, indicating the induction of innate immune memory in a C3-C3R axis-dependent pathway. However, it remains to be established whether such changes can lead to altered disease outcomes. Therefore, the impact of these changes should be demonstrated, for instance, through an infection model to support the claim made in the study that C3 modulates trained immunity in AMs through C3aR signalling.

      We acknowledge this is a Limitation of our manuscript. As this is a Short Report, we focused on how C3, via the C3aR, affects the reprogramming of alveolar macrophages. Recent work demonstrated that systemically administered β-glucan induces peripheral trained immunity and aggravates lung injury (Prével et al., 2025), similar to disease in models of periodontitis and arthritis (Haacke et al., 2025). However, training with β-glucan also reduces bleomycin-induced lung fibrosis (Kang et al., 2024). Hence, our ongoing work involves optimizing relevant intrapulmonary exposures to assess how trained immune responses are modulated by the C3a-C3aR axis. We have included the Reviewer’s critique in our revised Discussion as a limitation, while referencing the abovementioned manuscripts.

      (4) Figure 3, panels B and C - stats should be shown for comparing WT-HKPA-trained and C3KO HKPA-trained.

      These suggestions have been incorporated into Revised Fig 3B and 3C (now Figure 4).

      (5) In Figure 4, where the proper untrained C3KO is included, the data presented in Figure 4C show an increase in basal and maximum glycolysis in trained C3KO compared to their untrained control counterparts. Statistical analysis should be provided for this comparison. Based on these data, it appears that metabolic reprogramming occurs even in the absence of C3. Furthermore, C3KO cells intrinsically exhibit reduced glycolytic capacity compared to WT. These observations challenge the conclusions made in the manuscript. Therefore, without the proper control (untrained C3KO) included in all experimental approaches, it is impossible to draw an evidence-based conclusion that the C3-C3R axis plays a role in the induction of innate immune memory.

      We have included the statistical comparisons for all the groups in Figure 4C (now Figure 5), as suggested by the reviewer. The data suggests that C3-deficient (C3KO) alveolar macrophages have a blunted metabolic response to training, as compared to C3-sufficient (WT) alveolar macrophages. However, the C3KO cells do not have reduced glycolytic capacity compared to WT in the absence of training. The blunted response in trained C3KO AMs is rescued by exogenous C3, but is then reversed by C3aR antagonism. We have also provided new data/analyses with proper controls (untrained C3KO) in the other Figures (for example, in Figures S1, S2 and 3B&C (now Figure 4)). Taken together, the data would suggest that the effects of C3 in AM reprogramming are C3aR-dependent.

      (6) The Results and Discussion sections should be separated, and the results should be thoroughly analyzed in the context of published literature. Separating these sections will allow for a clearer presentation of findings and ensure that the discussion provides a comprehensive interpretation of the data.

      We thank the reviewer for this suggestion. The manuscript has been submitted as a Brief Report, and hence, we adhered to the instructions to authors for this format. However, we have added an additional section towards the end of the manuscript based on the Reviewer’s suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is intriguing that whilst there was a significant elevation of C3a in the cell culture medium of alveolar macrophages, the addition of exogenous C3a failed to rescue the phenotypes of C3-deficient alveolar macrophages. Please help discuss this.

      Alveolar macrophages secrete both C3 and proteases, which can cleave C3 to C3a in the supernatant. However, we used the addition of exogenous C3a to compare it to the addition of full-length C3. C3 is internalized by multiple cell types, including alveolar macrophages (as demonstrated in new Figure 3 of our Revised Manuscript) as C3(H<sub>2</sub>O) in comparison to C3a. Hence, we propose that the internalization of C3(H<sub>2</sub>O) provides an intracellular source of C3a (previously reported in Elvington et al. J Clin Invest 2017), which engages with the C3aR to result in alveolar macrophage reprogramming. In comparison, incubating cells with C3a does not exert similar effects. We have included these comments in a separate section towards the end of the manuscript.

      Please provide details for the statement "cell-permeable C3aR antagonist (SB290157)" (Figure 3E). Could a paracrine-based mechanism also be at play?

      SB290157 does not act selectively on the cell surface, but rather, can also enter cells. Our data, along with previously published reports (Quell et al. J Immunol 2017; Zha et al. Cancer Immunol Res 2019), suggest that C3aR may be intracellular in AMs. However, SB290157 can also block any receptor that may be present on the surface. For this reason, we used exogenous C3a as a way to interrogate surface C3aR signaling, and did not observe significant changes in AM reprogramming with exogenous C3a. However, as this is an indirect approach, we cannot completely rule out a paracrine-based mechanism and have included this limitation in the Discussion section of the revised manuscript.

      For Figure 3, please also provide the statistical analysis results for WT versus C3KO HKCA-trained cells. The statistical tests described in the legend for Figure 3D seem to apply to Figure 3E. Please check the labels.

      These suggestions have been incorporated into Revised Figure 3 (now Figure 4).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are pleased that the reviewers recognize the value of our multimodal single-cell approach and the diversity of model systems used to define gene regulatory networks underlying heterogeneity in Ewing sarcoma.

      Reviewer #1 (Public review):

      The investigators elegantly utilized a single-cell co-assay of RNA and ATAC seq to unveil the heterogeneous gene regulatory networks in Ewing sarcoma. The authors should be commended on their ability to identify multiple unique modules of gene regulation of Ewing sarcoma utilizing complex computational methods between numerous Ewing sarcoma cell lines. Additionally, they complemented their single-cell findings with xenografts as well as primary Ewing sarcoma patient tumors - validating the intratumoral heterogeneous gene regulatory networks of Ewing sarcoma. More importantly, they have revealed that exogenous TGF-β may modify these distinct epigenetic and transcriptional signatures within Ewing sarcoma tumors. Overall, the manuscript highlights an important discovery of the heterogenous gene regulatory programming of Ewing sarcoma and further highlights the role that TGFB plays within the tumor microenvironment of Ewing sarcoma. There are some areas of ambiguity that require clarification to increase the impact of the manuscript.

      We appreciate Reviewer 1's positive assessment of our work and their recognition of the importance of identifying heterogeneous gene regulatory programs in Ewing sarcoma, including the role of TGF-β in the tumour microenvironment. We have addressed the areas of ambiguity noted by the reviewer, including clarifying cluster assignments and the selection of k=3 for module identification, adding statistical comparisons to relevant figures, correcting figure cross-references, and improving figure labeling for clarity. We have also added higher-resolution images of the spatial profiling data and highlighted relevant correlations between CHLA9 and CHLA10 clusters. We believe these revisions improve the clarity and rigour of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      This work by Waltner et. al. provides a comprehensive single-cell multiomics analysis of plasticity in gene regulatory networks present in Ewing sarcoma using single-cell RNA-sequencing (scRNA-seq) and single-cell assay for transposase accessible chromatin with sequencing (scATAC-seq). They find that Ewing sarcoma cell line models have distinct patterns of chromatin accessibility compared to non-Ewing sarcoma models, and that there is significant variability across Ewing sarcoma cell lines, and sometimes within a single cell line. These differences across models are linked to 3 distinct gene regulatory modules, 2 of which are present across the range of model systems studied here. The first modules present across models are activated when the fusion is expressed and include genes enriched for the known EWSR1::FLI1 response element, GGAA microsatellites, along with other neural crest transcription factors. The other module primarily consists of genes repressed by EWSR1::FLI1, which are activated in EWSR1::FLI1-low states. Interestingly, EWSR1::FLI1-low cells have already been tied to more migratory and metastatic phenotypes, and the data here suggest these cells are more responsive to external signals from TGF-β, and this may be mediated through FOSL2-mediated gene regulation. While there are some minor additional validation studies that can be performed to strengthen a few individual analyses, this is a technically rigorous study, with a variety of different analytical techniques used to address similar questions, and this approach elevates confidence in the answers provided. This is further strengthened by the diverse set of model systems used, including patient-derived cell lines, cell line xenograft models, patient-derived xenografts, mining available single-cell data from patient samples, and validation of the gene modules identified in a larger set of patient microarray samples. In whole, this study provides a valuable resource for understanding heterogeneity, plasticity, and gene expression networks in Ewing sarcoma. This may be useful for future studies of metastatic disease and may also provide a framework for similar questions in other fusion-driven sarcomas.

      Strengths:

      There are a few core strengths in this study. First is the number and diversity of Ewing sarcoma models studied, spanning commonly used cell lines, patient-derived xenografts, and patient samples. The second is the large array of rigorous and orthogonal approaches used to uncover the identity and function of various gene modules. This includes an array of informatics techniques, as well as specific modulation of cell line models in culture. A third is confirmation that different gene expression programs are present in the same tumor using spatial transcriptomic analysis. Lastly, the authors have made all of their data and code accessible, enabling continued use of this dataset as a resource for others.

      Weaknesses:

      As highlighted by the authors, this study is somewhat limited by the small number of single-cell data from patient samples that are publicly available. Much of the analysis comes from cell lines. Additionally, they focus only on one type of signal that may modulate cell plasticity, and there are likely to be many others. Lastly, there are a few weak spots in the data. Some of this likely arises from the underlying complexity of the data, the generally sparse nature of scATAC data, and the biological heterogeneity present in the cell lines studied. The most pronounced weakness was in the analysis of transcription factors that dictate gene expression in the distinct modules, as well as the response to TGF-β. While some specific transcription factors showed module-specific expression consistent with the computational prediction in Figure 2, others did not likely due to additional factors not tested here. Likewise, the same transcription factors did not always show consistent enrichment in the gene modules that responded to TGF-β treatment when analyzed across cell lines. On the whole, these are relatively minor weaknesses and do not diminish the value of this study.

      We thank Reviewer 2 for their thorough and balanced assessment. We agree that the study's strengths lie in the breadth of model systems and orthogonal analytical approaches, and we appreciate the reviewer's acknowledgement that the identified weaknesses are relatively minor.

      In response to the reviewer's suggestions, we have made several substantive improvements. First, we have included a new Western blot panel in Figure 1 showing EWS::FLI1 protein levels across cell lines, which provides important context for interpreting the long-read fusion transcript detection data. Second, we have revised text throughout the Results section to improve precision — in particular, clarifying that our chromatin analyses assess accessibility at published EWS::FLI1 binding sites rather than binding per se, and ensuring that our stated hypotheses match the metrics presented in the corresponding figures. Third, we have improved figure color schemes and labeling to aid interpretation and corrected errors in panel labeling.

      Regarding the reviewer's observation about transcription factor enrichment patterns across cell lines (particularly RUNX3 in the TGF-β response analysis and FOSL2’s modest expression at the protein level in CHLA9), we acknowledge that not all TFs showed perfectly consistent module-specific behaviour across every cell line. As the reviewer notes, this likely reflects the underlying biological complexity and additional regulatory factors not tested here. We have made attempts to temper our language accordingly.

      We believe that the revised manuscript, with its additional experimental data, improved figures, and clarified text, addresses the concerns raised by both reviewers and strengthens the overall impact of our findings.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Specific comments:

      (1) Figure 1A: Suggest adding cell labels on the UMAP plot - difficult to tell with different colors - which cluster is which cell line.

      We have added labels to Figure 1A (UMAP plot) as suggested.

      (2) Figure 1B: Although this may have been slightly addressed later in the manuscript, since CHLA9 and 10 are from the same patient, are the authors surprised to see the differences in EWS::FLI1 motif accessibility, or are these findings further reinforcement of inherent heterogeneity? Should one expect the ChromVAR deviation z-score at least to overlap between CHLA9 and 10, since they're from the same patient, if not, can the authors explain why?

      We were indeed surprised to see differences in EWS::FLI motif accessibility inferred from our data particularly from the isogenic lines CHLA9 and 10. While it is impossible to know for sure, we suspect that some treatment related changes increased accessibility in CHLA10. However, when considering global accessibility signatures, a subset of cells from CHLA10 (C23) was most correlated with signatures from CHLA9. We have addressed this by stating “Unsurprisingly, of all cell lines, CHLA10 C23 cells exhibited the greatest correlation to profiles from CHLA9,” in the 11th paragraph of the results section.

      (3) Figure 1B: Can the authors comment on the differences in trends of fusion transcript to chromatin enrichment (e.g., A673 and TC71 inverse trends vs. other cell lines that are direct positive or negative trends).

      We are hesitant to draw too many conclusions from transcript counting vs chromatin enrichment given some outliers, but in general, we found that high-fusion transcribing cell lines A4573, SKNMC, RDES were associated with module 3, while the lower transcribing cells CHLA9/10, TC32 and PDX305 were associated with module 2. We did not choose to emphasize this correlation too strongly in the paper as there may be other factors such as additional mutations (such as BRAF<sup>V600E</sup> in A673) that may have either been present in the original tumour or after serial passaging that may play a role.

      (4) Figure S1D: Please make the pt. smaller in order to better visualize the differences and highlight the stark differences of the ChromVAR binding site score.

      We have adjusted the point size of the embeddings in Fig S1D to enhance visualization.

      (5) Figure S1D: Is it surprising to see a large proportion of PDX305 and minor proportions of CHLA10 and CHLA9 with low ChromVAR binding site score?

      Our analyses indeed show that PDX305, CHLA9, and CHLA10 utilise fusion-repressed gene programs, so lower enrichment of EWS: FLI1 microsatellites is consistent with our findings.

      (6) Figure 2A: Unclear how k = 3 was ultimately selected - was it purely a visualization of how well separated each of the cell lines is? Can the authors further clarify how k = 3 was determined to be the most optimal? Is there a UMAP that demonstrated a difference in clustering between groups 1, 2, and 3? Is there a threshold cut-off to assign groups into 1, 2, or 3?

      Thank you for pointing out this omission. In Fig 1 we show that unsupervised clustering of DA peaks across cell lines grouped EwS cell lines into 2 clusters. When using peak-2-gene linkages (Fig 2), we chose a k =3 to explore the potential genes/CREs explaining the grouping found in Fig 1 because A673 was a clear outlier. We have added the following sentence to the results section paragraph 6. “We selected k = 3 to extend the two-cluster structure observed in Fig. 1F–G, reasoning that A673 represented a clear outlier whose distinct regulatory program would be obscured at lower k.”

      (7) Figure 2G: Understanding the substantial heterogeneity of each cell and TF binding motifs - given that CHLA9 is considered group 2, is it unexpected that FOSL2 was not highly expressed compared to the other EwS cell lines within group 2: (TC32, CHAL10, PDX305).

      Thank you for astutely pointing out that CHLA9 did defy the trend for FOSL2 protein expression compared with other group 2 lines. We specifically did not comment on this finding in the manuscript as we could not account for its status as an outlier. We address this in the public response above.

      (8) Figure S3B: Did the authors perform a similar computational analysis (performed for Figure 4), looking specifically at CHLA9 (Clusters 24 and 25) to determine if there are any overlaps with CHLA10 cluster 23's pathway activity/MSigDB/GO: Biological process terms? If there are potential overlaps, can the authors potentially infer tumor clonal evolution from CHLA9 to CHLA10?

      While we didn’t do pathway analysis, Figure S3 shows strong pearson correlation of C23 from CHLA10 with both C24 and C25 from CHLA9. We have highlighted this with a red box in the figure to make this more obvious to the reader.

      (9) Figure 4: Given the heterogeneity of CHLA-10. Did the authors observe any differences in morphology within CHLA-10 between the predominant modules (modules 2 and 3), given such stark transcriptional heterogeneity?

      We did not observe any morphologic differences within CHLA10 but acknowledge this would be an interesting avenue of further investigation.

      (10) Figure 5B: Can the authors also plot out module 3 gene expression to see if the cluster enrichment is unique from module 2 gene expression with TGFB1 and vehicle?

      We thank the reviewer for this suggestion and have replaced Fig S4B with violin plots so the enrichments are clearer. Module 3 expression in cluster 3 cells from CHLA10 is among the lowest in the cell line.

      (11) Figure 5H: Can the authors provide a higher zoom/resolution of the H&E stain of the ROIs in order to see if there are indeed more stroma/fibrosis in ROI9 and ROI10, and if there are differences in tumor cell morphology within different ROIs that harbor different modules?

      We have now included higher resolution and magnification of IHC panels in Figure 5H. These images are included as Supplemental Fig 4C. It is notable that although no dramatic differences in tumour cell morphology are visualized, ROI9 and ROI10 comprise small islands of viable tumour surrounded by necrosis.

      (12) Figure 6: Within Volchenboum/Lawlor's dataset, of the 46 clinically annotated primary tumors, 10 of the COG samples contained substantial stromal elements, while all the tumors in the European cohort were >70% viable tumors. Can the authors separate out the stromal-rich (n = 10) samples and analyze the 36 tumor-enriched samples to see if the survival curve is the same as what is shown in Figure 6G/H and S4E?

      We thank the reviewer for this suggestion. We observed no differences when stratifying the patients as outlined. Indeed, in the original Volchenboum et al, manuscript it was demonstrated that no prognostic gene signature was identifiable when the stromal-rich tumours were removed from the cohort.

      (13) Figure 6: Additionally, can the authors run a similar computational analysis to determine the predominant modules within bulk sequencing of the Volchenboum/Lawlor dataset between each tumor?

      The computational approach deconvolution was developed for use with bulk RNA-seq and is based on count data. The linear mixture model at the heart of deconvolution methods requires that the measured signal is proportional to abundance across the full range, and Affymetrix microarray data violate that assumption in a gene-specific, nonlinear way that can't be fully corrected post hoc.

      (14) Page 15, line 2: Wrong GSE data cited - currently cited as GSE61357 - should be GSE63157 instead.

      This has been corrected, thank you.

      (15) Please add statistical comparisons for Figure 3E-H.

      We thank the reviewer for highlighting this omission. We used the software package ggpubr to perform Wilcoxon rank-sum tests comparing mean module scores between conditions. We have added statistical labels to the plots. All comparisons were statistically to the level indicated in the figure.

      (16) Page 11 Line 21: Figure S3 D-E is not about EMT, migration, and TGFB signaling - I believe the authors are referring to Figures 4D-E

      Thank you, we have corrected these errors

      Reviewer #2 (Recommendations for the authors):

      (1) Data

      (a) In Figure 1 and the associated text, there is an analysis of cells expressing EWSR1::FLI1 performed using a locus-specific amplification and long-range sequencing. On page 7, lines 1-5, there is some discussion about how some of these track with overall transcript levels, while others don't. Additionally, a very low fraction of cells is shown to be EWSR1::FLI1 positive. This analysis might also be strengthened by a Western blot to show protein levels and how they vary across the cell lines, which may help explain additional differences in the data. The Abcam antibody ab133485 works well for western blotting of EWSR1::FLI1. While this isn't additional single-cell data, the percentage of cells with transcript detected is not really equivalent to the total expression level. This seems particularly valuable to do, as prior publications (Pishas, et. al., Mol. Cancer. Ther., 2018) show that of the cell lines tested here, TC32 has relatively high EWSR1::FLI1 protein levels, while A673 has relatively low expression. This contrasts with the percent of positive cells here.

      We thank the reviewer for this suggestion and have included a new panel, Figure 1E containing the western results for cells harvested in log-phase growth.

      (b) Figures 5D-F are not obviously referenced in the section about Figure 5 (page 12 line 4 through pg. 13 line 20). One question to pose to the authors is how to interpret the data here for TC71 in light of the fact that they were relatively insensitive to TGF-β. This shows up obviously in the Western blot in 5G.

      Thank you, for drawing attention to this omission. We have corrected the figure reference to include Figure 5D-F. In regards to our interpretation of TC71’s lack of response to TGF-B, we direct the authors attention to our statement: “The muted response of the TC71 cell line (module-3 dominant) to the influence of TGF-b (Fig. 5A-C, Fig. S4B) suggests that pre-existing transcriptional states may condition the sensitivity to TGF-β signalling.”

      (c) For the discussion of Figure 5, the authors say on page 12, line 21, that "RUNX3 was enriched in non-responsive clusters." I'm not entirely convinced that the data support this statement as written. This is true in A673s, but there appears to be no difference between the maximally responsive and maximally non-responsive clusters in CHLA10. TC71 was not particularly responsive, but showed the opposite effect.

      We agree this was not clearly written. We have de-emphasized RUNX3 findings here. Our new conclusion in the results section paragraph 15 is: “Correlation analyses revealed a reciprocal enrichment pattern of many key TFs from Figure 2, where correlation of FOSL2 (but not RUNX3) gene expression and accessibility was generally highest in TGF-β responsive clusters.”

      Can data for A673 cells be included for Figure 5G? Like CHLA10, this cell line had the pattern of FOSL2 enrichment that is concordant with that described in the text.

      We thank the reviewer for this suggestion and have included A673 in the western blot.

      (2) Figures

      (a) In Figure 2A, it is very difficult to distinguish the colors for CHLA9 and TC32 in the left panel. Can the color scheme here be changed to make these easier to distinguish?

      We thank the reviewer for pointing this out and we have added labels to the UMAP. We hope this makes the visualization more obvious.

      (b) Similarly, the KLF4 and SP1 lines in Figure 2D are a little close and might benefit from having more distinct colors.

      We thank the reviewer for pointing this out. We have changed the KLF4 to a green hue.

      (c) The lower panel of Figure 5I has 2 samples labeled "11" and no sample labeled "12".

      Thank you for catching this error. It has been corrected.

      (3) Text

      (a) One section early in the response was a little bit confusing and could benefit from some revision to improve clarity. On page 6, lines 9-11, this reads a little bit like they looked at cell-line specific EWSR1::FLI1 binding, but that wasn't the assay that was performed. Perhaps there is a better way to describe this than simply "enrichment of EWS::FLI1 sites."

      We thank the reviewer for their efforts to improve clarity of our work.

      We changed the following sentence in the 2nd paragraph of the result section:

      "We visualized the enrichment of EWS::FLI1 sites across all cell lines and discovered distinct EwS cell-line specific usage (Fig. 1C & Fig. S1D)."

      And revised to:

      "We then assessed chromatin accessibility at these published EWS::FLI1 binding sites across all cell lines and discovered that accessibility at these loci varied in a cell-line-specific manner (Fig. 1C & Fig. S1D)."

      (b) Then at the start of the next paragraph (page 6 line 12), the authors talk about "differences in EWS::FLI1 motif and binding site enrichment" and my first thoughts were whether this was differences in which sites were bound or differences in the strength of enrichment. More precise language would be helpful.

      Here in the 3rd paragraph of the result section we changed:

      "Given the differences in EWS::FLI1 motif and binding site enrichment"

      To:

      "Given the heterogeneity in the magnitude of EWS::FLI1 motif enrichment"

      (c) Related to the comment about EWSR1::FLI1 positive cells vs. protein levels above, the hypothesis on page 6, line 13 says that there was a hypothesis that different Ewing sarcoma lines have different levels of EWSR1::FLI1 transcript. But the metric shown in Figure 2D is the percentage of cells with detectable transcript, not transcript levels. Be specific about what the data are showing here.

      We agree with the need to improve clarity. In the 3rd paragraph, we changed: "we hypothesized that EwS cell lines have different levels of EWS::FLI1 transcript."

      To:

      "We hypothesized that EwS cell lines differ in the proportion of cells expressing high levels of the EWS::FLI1 fusion transcript."

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

      We appreciate the reviewer’s summary of our study and its strengths.

      Weaknesses:

      A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

      (1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

      We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

      (2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

      We have added a definition of power in the main text:

      “Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

      We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

      “(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence r<sub>X, Y</sub>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of r<sub>X, Y</sub> between r<sub>X, Y</sub> = 0 and 0.54 in steps of size 0.01. At r<sub>X, Y</sub> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.”

      We have also added dotted lines connecting the legend in panel A to the curves in panel D.

      (3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

      We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

      (4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

      As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

      The relevant panel and an excerpt from the legend text are copied below.

      “(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

      Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

      We appreciate the reviewer’s summary of key aspects of our manuscript.

      Weaknesses:

      (1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

      We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

      Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/n<sup>!</sup>. As a consequence, any reported perfect match p-value below 1/n<sup>!</sup> would break the distribution-free validity of the test.

      A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

      (2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

      We agree this is an important consideration. We have added the following text to the Discussion:

      “The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For n > 10, a standard permutation test already can report a p-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A few more minor notes/comments:

      As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

      We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

      On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

      As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

      “It seems likely that the tests could be further modified to report an even lower p-value when an X- and Y -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

      I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

      In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

      We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

      Reviewer #2 (Recommendations for the authors):

      (1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

      We have added a section on size to our results section, copied below:

      “The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when α is not one of its possible p-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match test in the setting where n = 3, where α = 0.05, and where r<sub>X,Y</sub>= 0 (independent X and Y). Note that in this case, obtaining a Y -perfect match (and thus p = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the p = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, p < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (Y -perfect match under n = 3 and r<sub>X,Y</sub> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, p = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

      (2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

      We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

      (3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

      We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

      We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

      (4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

      We have added the following to the discussion:

      “No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, ρ could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”

    1. Author response:

      We are glad the reviewers found the tandem fluorescent timer approach valuable and the leading-edge Arm stabilization finding significant.

      We agree with the Assessment that our evidence for three specific claims Wingless-independence, JNK-mediated regulation of Arm stability, and a direct mechanical/force-transmission role for stabilized Arm is currently incomplete, and we will revise the text throughout to reflect this more precisely rather than overstating the current data. In addition, we commit to two new experiments, both using existing reagents and fly stocks, that speak directly to the two most experimentally tractable points raised by the reviewers:

      (1) Re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin (reagent already validated in Figure 4), to test whether JNK acts directly on junctional architecture or only indirectly, via broader epithelial disruption.

      (2) Imaging ArmTimer in a wingless loss-of-function background, to directly test Wingless-dependence of leading-edge Arm stabilization as the reciprocal of our existing overexpression data.

      We address each public review point below and outline the accompanying text revisions.

      On Wingless independence (Reviewer #1; Reviewer #2, Weaknesses #2 and #5; Recommendation #1):

      We agree that our current evidence unquantified Wg overexpression restricted to the amnioserosa (C381-Gal4) and uniform overexpression, alongside the absence of detectable Wg-stripe-associated Arm-Timer signal supports a more limited conclusion than "Wingless-independent" as currently stated. We will revise our language throughout the Abstract, Results, and Discussion to state that canonical Wg overexpression does not detectably enhance leading-edge Arm stabilization or perturb dorsal closure under our conditions, rather than asserting pathway independence. As noted above, we commit to imaging ArmTimer in a wg mutant background to test this directly, complementing our overexpression data with the reciprocal loss-of-function manipulation.

      On the related point that the Timer's failure to detect a Wg-stripe-associated stabilization signal could reflect a sensitivity limitation rather than a true absence of stabilization (Reviewer #2, Weakness #5): we agree and will state this explicitly rather than treating absence of signal as evidence of absence. This does not undermine the positive leading-edge finding, which is not defined relative to the stripe comparison: all embryos, channels, and time points were imaged and rendered using identical laser power and brightness/sensitivity settings, and the leading-edge RFP signal clearly exceeds background under those same acquisition conditions. We also note that detection limits of this kind are a recognized challenge for endogenously tagged reporters of canonical Wnt/β-catenin signaling generally, including in mammalian systems, and cite two studies already in our bibliography that report the same class of limitation: de Man et al. (2021, eLife 10:e66440) and Ambrosi et al. (2022, eLife 11:e64498). We have added this clarification, with these citations, to the Results (paragraph describing Figure 3).

      On the mechanistic link between Arm stability and destruction-complex activity (Reviewer #1):

      We agree that because both ΔArm and Axin overexpression converge on the destruction complex, our data cannot yet fully separate "escape from degradation" from "impaired α-catenin/junctional coupling" as the operative mechanism. We will revise the Discussion to state this ambiguity explicitly and will treat the ArmTimer-AA result (partial α-catenin-binding disruption via phosphosite mutation, independent of destruction-complex regulation) as the strongest current evidence isolating the junctional-coupling mechanism. We also thank Reviewer 1 for pointing us to Pokutta, Choi, Ahlsen, Hansen & Weis (2014, J Biol Chem 289:13589-13601), which structurally and thermodynamically characterized the mammalian cadherin·β-catenin·α-catenin complex and showed that α-catenin binding to β-catenin is a distinct, allosterically regulated interface cadherin binding increases β-catenin's affinity for α-catenin roughly 10-fold, and α-catenin homodimerization independently competes with β-catenin binding. We have added this citation to the Discussion as structural support for treating cadherin engagement, α-catenin coupling, and destruction-complex regulation as mechanistically separable interfaces, and note that the crystallized β-catenin·α-catenin interface provides a structural template for future experiments for example, structure-guided point mutations at the homologous interface residues in Arm, or in vitro binding assays comparing wild-type and threonine-mutant (T111A/T121A) Arm affinity for α-catenin.

      We do not, however, believe a destruction-complex-independent stabilizing allele of Arm is a tractable experiment to close this gap directly: any allele that stabilizes Arm without engaging the destruction complex is, by definition, a Wnt pathway gain-of-function allele, since destruction-complex-mediated degradation is the very regulatory step that canonical Wnt signaling controls. Nor would restricting the allele to a transcriptionally inactive form of Arm cleanly resolve the confound: Wnt/TCF target loci include dedicated repressive TCF-binding sites (Blauwkamp, Chang & Cadigan, 2008, EMBO J 27:1436-1446), so a transcriptionally "dead" stabilized Arm could still alter transcription by disrupting TCF-mediated repression. We therefore treat this as a genuine, currently unresolvable confound of the overexpression approach, and rely instead on the CRY2 optogenetic and ArmTimer-AA results as the strongest available evidence isolating a junctional-coupling contribution. We have added this reasoning, with both citations, to the Discussion.

      On JNK acting on Arm directly vs. indirectly (Reviewer #1; Reviewer #2, Weakness #3 and Recommendation #2):

      This is the most actionable point raised by both reviewers. As noted above, we commit to re-imaging and re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin to determine whether junctional/polarity architecture is broadly disrupted under these conditions (indirect mechanism) or whether E-cadherin localization is comparatively preserved while Arm stabilization is specifically altered (direct mechanism). In the meantime, we note that a direct mechanism is biochemically plausible: in mammalian cells, JNK phosphorylates β-catenin directly and regulates adherens junction integrity, and JNK activity separately controls the binding of α-catenin to the junctional complex (Lee, Koria, Qu & Andreadis, 2009, FASEB J 23:3874-3883; Lee, Padmashali, Koria & Andreadis, 2011, FASEB J 25:613-623). We cite these as precedent that a direct route from JNK to junctional β-catenin/α-catenin regulation exists in another system, while being explicit that this does not establish the same mechanism in Drosophila dorsal closure that will be tested directly by the E-cadherin re-staining experiment. We have added these citations and this caveat to the Discussion.

      On the Dsh-DEP-to-JNK mechanistic link (Reviewer #1, Public Review #2 and Recommendation #5):

      We agree that our data show the Dsh-DEP requirement and the JNK requirement for dorsal closure as parallel, independent findings rather than a demonstrated linear pathway in our system. To provide context for why we consider a DEP-to-JNK connection a reasonable working hypothesis, we searched the literature in both Drosophila and vertebrates and will cite six additional studies establishing this link: Axelrod et al. (1998) and Axelrod (2001), establishing that DEP-dependent membrane recruitment and unipolar localization of Dishevelled are specifically required for planar polarity signaling, distinct from Wingless signaling; Paricio et al. (1999) and Fanto et al. (2000), showing Dishevelled acts through Misshapen and Rac1/RhoA to the same JNK module used in dorsal closure; and Moriguchi et al. (1999) and Yamanaka et al. (2002), showing biochemically in vertebrates that the DEP domain of Dvl-1 selectively activates JNK independent of β-catenin/TCF-LEF activity, and that this JNK requirement is conserved in Xenopus convergent extension, the vertebrate process most functionally analogous to dorsal closure. We will state explicitly that this precedent, while now cross-species, comes from planar-cell-polarity and convergent-extension assays rather than dorsal closure itself, so it supports the plausibility of a Dsh/Dvl-DEP-to-JNK connection without establishing that the identical pathway operates in our system.

      On the Dsh DIX/DEP domain-separability argument (Reviewer #1, Recommendation #4):

      We thank the reviewer for pointing us to Gammons, Renko, Johnson, Rutherford & Bienz (2016, Mol Cell 64:92-104), which showed that the Wnt signalosome itself is assembled by head-to-tail DEP domain swapping between Dishevelled molecules, and that this DEP-dependent oligomerization is directly required for canonical Wnt pathway activity not restricted to the non-canonical/planar-polarity branch as we had implied. We agree this evidence undercuts our previous interpretation of the DshΔDEP dorsal closure phenotype as evidence for a strong non-canonical/polarity-specific role for the DEP-dependent branch of Dsh. We have revised the Discussion accordingly: we now state that the DEP domain is required for the morphogenetic program culminating in dorsal closure, cite Gammons et al. directly, and note that this requirement does not by itself establish a non-canonical/polarity-specific role, since we cannot rule out a contribution from DEP-dependent canonical Wnt signalosome assembly.

      On force transmission (Reviewer #2, Weakness #1):

      We agree that we have not directly measured force or tension at the leading edge, and that our current data (colocalization with actin/E-cadherin, and functional requirement shown via CRY2 optogenetics and mutant analysis) are consistent with, but do not directly demonstrate, a role in force transmission. We do not have the in-house expertise to perform direct force/tension measurements (e.g., laser ablation, junctional tension assays), so we will not be adding such an experiment in this revision. Instead, we have revised the language throughout the manuscript including two Discussion section headings that previously stated a mechanical role for stabilized Arm as established fact to consistently present the mechanical/force-transmission role as a hypothesis raised by our data, not a demonstrated conclusion, and we retain a clear statement that direct force measurement (ideally in collaboration with groups with the relevant biophysical expertise) is future work rather than a claim we are making in this manuscript.

      On the phosphomimetic threonine mutant (Reviewer #1, Recommendation #2):

      ArmTimer-AA (T111A, T121A) was generated with the expectation that the tyrosine phosphosite mutants (ArmTimer-EE, ArmTimer-FF) would be the primary drivers of any dorsal closure phenotype, given their proposed role in E-cadherin binding; the pronounced zippering defect we observed in ArmTimer-AA was therefore an unanticipated finding rather than a predicted result. We have not generated the reciprocal phosphomimetic ArmTimer-EE(Thr) (T111E, T121E) allele. Generating and characterizing this allele is a substantial undertaking we estimate over a year including allele generation, validation, and phenotypic characterization and we will state this explicitly in the Discussion as planned future work rather than part of the current revision.

      On confirmation of myristoylated-Dsh membrane targeting (Reviewer #1, Recommendation #3):

      We cannot confirm that myristoylation localizes all Dsh protein to the membrane. However, this strategy has extensive prior genetic validation using the identical Src-derived myristoylation sequence: it was originally used to tether Armadillo and shown sufficient for constitutive Wnt pathway activation (Zecca, Basler & Struhl, 1996; Tolwinski & Wieschaus, 2001, 2004), and the same approach was subsequently applied to GSK3 and Dishevelled, in each case producing the expected pathway-activation phenotypes (Mannava & Tolwinski, 2015; Kaur et al., 2017). We have added these citations to the Results where the Myr-Dsh constructs are introduced.

      On overexpression-based perturbations versus endogenous regulation (Reviewer #2, Weakness #4):

      We would like to clarify that most of the Arm alleles used in this study including all of the point-mutant Timer alleles (ArmF1a, ArmTimer-FF, ArmTimer-EE, ArmTimer-AA) central to our mechanistic conclusions were generated as knock-ins at the endogenous ‘arm’ locus via MiMIC/RMCE, not overexpressed. The two exceptions are ΔArm and ArmS56A, expressed from UAS constructs because both are gain-of-function alleles anticipated to be lethal if expressed from the endogenous locus, based on prior experience with similarly stabilizing mutations. Axin overexpression was used because no Axin mutant or knock-in allele was generated for this study; we agree an endogenous Axin allele would be the ideal complement and will state this explicitly as a limitation, while noting that our CRY2 optogenetic perturbations of Arm and α-catenin which act acutely on the endogenous proteins provide an orthogonal line of evidence supporting the same conclusions. We have added this clarification to the Discussion.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      The authors have added comprehensive analyses in this revision, and all of my concerns have been very well addressed. Here, I just want to re-emphasize the original points 1 and 3.

      (1) The analysis and clarification are very helpful - thanks! I found that Fig. R1 and R2 are very insightful, as DoRothEA-only returns much worse performance. Please consider adding these two figures to the supp figure and possibly highlighting your setting for edge pruning (down-weights); therefore, the model is more likely to be affected by false negatives than false positives in the TF-target prior.

      We thank the reviewer for the positive feedback and for recognizing the value of the additional analyses. We have added the previous Fig. R1 and Fig. R2 to the Supplementary Information as Fig. S13 and Fig. S14, respectively, and have referred to them in the revised manuscript.

      We have also expanded the description of the TF–target prior used in TSvelo in the “Acquiring Prior Knowledge of Gene Regulatory Relations” subsection of the Methods. As noted by the reviewer, TSvelo is expected to be less sensitive to false-positive TF–target interactions because unsupported edges can be down-weighted during training. In contrast, missing true regulatory interactions are not represented in the prior network and therefore cannot contribute to the learned regulatory dynamics, making the model potentially more sensitive to false negatives.

      (3) Please consider adding some discussion on the challenges in capturing cell cycle transitions.

      We thank the reviewer for this suggestion. We have added a brief discussion in the Discussion section on the challenges of modeling cell-cycle transitions. In particular, cell-cycle progression is often characterized by cyclic dynamics and overlapping transcriptional programs, which can complicate the inference of directional state transitions and regulatory relationships.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version.

      The Authors addressed all my comments suitably. I'd like to thank them for the time they spent addressing them: the revised paper is much more convincing.

      I have 2 very minor follow-up concerns:

      (1) I appreciated the simulation study, however, no null simulation is present.

      We know RNA velocity tools are inclined to provide false positives: trajectories even when the data doesn't have any.

      I'd be helpful to add null simulations where the data has no trajectories and see if methods erroneously identify any.

      We thank the reviewer for this helpful suggestion. We have added null simulations to evaluate TSvelo and baseline approaches on data without underlying dynamic structure. Specifically, we generated a null dataset including 200 genes and 600 cells by independently sampling spliced (S) and unspliced (U) counts, thereby removing any coherent transcriptional relationship between them.

      When applying scVelo and UniTVelo to this data, no genes passed the velocity gene selection step under the default likelihood-based filtering, and no velocity field could be obtained. We further tested TSvelo, Dynamo, and cellDancer on the same null data and observed that all three methods still produce trajectory-like patterns despite the absence of true dynamics (See Supplementary Information as Fig. S17).

      Including TSvelo, many RNA velocity and trajectory inference approaches assume that they are applied to datasets reflecting underlying dynamic biological processes. We agree that incorporating additional checks during preprocessing could help prevent applying velocity analysis to non-dynamic datasets. We have added this discussion to the revised manuscript.

      (2) Several of the novel analyses are only reported in the Supplementary material and only references in the main text (e.g., "A validation of TSvelo on simulated data is provided in Fig. S1 and Fig. S2 in the Supplementary Information."). This is pity!

      If allowed, I'd add some comments about the new analyses (simulations, computational benchmarks, etc...) also in the main text.

      We thank the reviewer for this suggestion. We agree that several analyses presented in the Supplementary Information provide important support for our conclusions. To improve their visibility, we have expanded the corresponding descriptions in the main text and briefly summarized the key findings of the relevant Supplementary Figures instead of only citing them. These revisions have been made for Fig. S1, Fig. S2, Fig. S10, Fig. S12, Fig. S13, Fig. S14 and Fig. S17. In particular, we have incorporated a summary of the simulation results at the end of the subsection “Estimate RNA Velocity with TSvelo” in the Results section, and added a discussion of the computational benchmarking analyses in the Discussion section. We hope these changes improve the accessibility of these results while maintaining a concise presentation of the main findings.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      I suggest the paper to undergo (very) minor revisions as detailed in the Public Review.

      Simone Tiberi, The University of Bologna

      We sincerely thank all reviewers for their thoughtful suggestions, which have helped improve the clarity and overall presentation of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of endothelial toll-like receptor 4 (TLR4) leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wild-type and TLR4knockout endothelial cells, the authors further demonstrate that TLR4-NF-κB signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Finally, the authors use RNA sequencing of wild-type and TLR4-knockout endothelial cells to define the TLR4dependent cell-autonomous transcriptional response to E. coli.

      Strengths:

      (1) The authors address an important, well-motivated hypothesis related to the cellular and molecular mechanisms of leptomeningeal inflammation.

      (2) The authors use model systems (mouse conditional knockouts and cultured endothelial cells) that are appropriate to address their hypotheses. The data are of high quality.

      Weaknesses:

      (1) The authors perform single-nucleus RNA-seq on dissected leptomeninges from control and E. coli-infected mice across three genotypes (WT, Tlr4MKO, and Tlr4ECKO). A major discovery from this experiment, as summarized by the authors, is: "Tlr4ECKO mice exhibited a global attenuation of infection-induced transcriptional responses across all major leptomeningeal cell types, as judged by the positions of cell clusters in the UMAP." This conclusion could be considerably strengthened by improving the qualitative and quantitative analysis.

      Thank you for this comment. We agree that the UMAP-based interpretation would benefit from additional qualitative and quantitative support. We have expanded the snRNA-seq analysis with additional images and supplemental figures (Figure 1 – figure supplement 3, Figure 1 – figure supplement 5, and Figure 1 – figure supplement 6). The first and third of these new supplemental figures show dot plots for each major leptomeningeal cell type, for each genotype, for the two experimental conditions (infected vs. uninfected), and for individual genes in three immune-related gene sets (NF-kB and TNF-α, JAK-STAT, and IFN-ɣ), providing a more explicit comparison of infection-induced transcriptional responses across genotypes. The second of these new supplemental figure shows principal component analysis (PCA) of the individual snRNA-seq datasets (one mouse per dataset) for each genotype and experimental condition, demonstrating that the observed transcriptional shifts are consistent across biological replicates. Finally, Figure 1 – figure supplement 7, which was included in the original submission, shows changes in the most up- and down-regulated genes (based on adjusted p-value or fold change) in endothelial and myeloid cells across individual mice and genotypes/conditions, further supporting the genotype-dependent effects at the level of individual animals.

      (2) The authors interpret E. coli infection-induced increases in leptomeningeal sulfo-NHSbiotin as evidence of compromised BBB integrity (i.e., extravasation from the vasculature) (Results, page 7), but another possible route in this context is sulfo-NHS-biotin entry from the dura across a compromised arachnoid barrier. The complete rescue in Tlr4ECKOs is strongly suggestive that the vascular route dominates, but it would strengthen the work if the authors could assess arachnoid barrier fidelity (e.g. via immunohistochemistry). At a minimum, authors should mention that the sulfo-NHS-biotin signal in this context may represent both vascular and arachnoid barrier extravasation.

      Thank you for this comment. We agree that our data cannot rule out leakage across the arachnoid barrier during infection. While the rescue observed in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice strongly supports a dominant vascular contribution, we acknowledge that the sulfo-NHS-biotin signal may reflect permeability at both the vascular and arachnoid barriers. We do not think there is a clear way to directly test this possibility functionally, since the arachnoid barrier appears intact by confocal microscopy. Subtle differences in barrier cell morphology might be detectable by electron microscopy, but this would not definitively address whether infection permits molecular passage across the arachnoid barrier. We have followed the reviewer’s suggestion and revised the Results section to reflect this interpretation. Specifically, we added the following: “The simplest interpretation of these data is that the site of sulfo-NHS biotin leakage is primarily vascular. However, we cannot exclude some contribution from increased arachnoid barrier permeability.”

      (3) The authors state that "deletion of TLR4 prevented both NF-κB nuclear translocation and Cldn5 internalization in response to E. coli (Figure 4A-D)" (Results, page 9). In Figures 4C and D, however, there is no indicator of a statistical test directly comparing the two genotypes. A comparison of within-genotype P-values should not be used to support a genotype difference (PMID: 34726155).

      Thank you for pointing out this omission. We have updated the figures so that the between-genotype p-values are shown for those panels (including this panel) that had not previously shown them.

      (4) In the first paragraph of the Results, the authors summarize the meningeal layers as (1) pia, (2) subarachnoid space, (3) arachnoid, and (4) dura, and then state "The second and third layers constitute the leptomeninges." This definition of leptomeninges seems to omit the pia, which is widely considered part of the leptomeninges (PMID: 37776854).

      Thank you for pointing out this error, which has now been corrected.

      (5) The Cdh5-CreER/+;Tlr4 fl/- mouse lacks TLR4 in all endothelial cells (i.e., in peripheral organs as well as CNS/leptomeninges), and, as the authors note, the periphery is exposed to E. coli. It would be helpful if the authors could comment in the Discussion on the possibility that peripheral effects (e.g., peripheral endothelial cytokine production, changes to blood composition as a result of changes to peripheral endothelial permeability) may contribute to the observed leptomeningeal phenotypes.

      Thank you for raising this point. We agree that peripheral responses could contribute to the observed leptomeningeal phenotypes in this model. We have added two sentences to the second paragraph of the Discussion to address this: “We note that these experiments do not distinguish between local vs. distal anatomic sources of LPS or downstream effector molecules, such as cytokines, that activate the leptomeningeal inflammatory response (Huang et al., 2021). Histologic observations of RFP-expressing E. coli in the brain, liver, and lungs, together with positive blood cultures, indicate substantial systemic dissemination in this model. Thus, the inflammatory responses of leptomeningeal cells likely reflect exposure to bacterial products and inflammatory mediators derived from both local meningeal and peripheral sources.”

      Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell-type-specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single-cell transcriptional profiling and imaging studies using whole-mount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial-specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown have not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type-specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design, and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysosomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plate upon which barrier-forming cell cultures can be placted, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cells types, likely epithelial cells.

      Weaknesses:

      (1) There are no measures of bacterial burden in peripheral organs, blood, in the LPM or brain in the TLR4 endothelial cKO mice. Lack of TLR4 in endothelial cells could prevent bacterial 'access' into the LPM and brain, essentially preventing meningitis and leading to a lack of inflammatory responses in the LPM-located cells simply because there is no bacteria present. Bacteremia may also be reduced, as might inflammatory responses in peripheral organs with TLR4-deficient peripheral endothelium. Bacterial counts and inflammatory measures in peripheral organs and blood are important to better understand the mechanism(s) underlying the reduced inflammatory profile in LPM cells and no LPM endothelial breakdown in the Tlr4 endothelial cKO mice. In other words, does deleting TLR4 in EC protect against the development of meningitis by somehow blocking bacteria access to the LPM (this would be supported by low or no CFU counts in infected Tlr4 endothelial cKO) or is it what the authors appear to propose in Figure 1J that TLF4 in EC is the only cell responding to the bacteria to trigger the immune cascade in the LPM? More data is needed to resolve this, as this is a major claim of the paper.

      Thank you for this comment. We agree that it is important to distinguish whether the reduced inflammatory response in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice reflects altered bacterial burden versus altered host sensing. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4CKO mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs of infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice in E. coli burden; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and the time of sacrifice 24 hours later (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid leptomeningeal cells does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (2) The authors look at the underlying cortical response (cerebral vasculature for ICAM and immune cells) but do not use markers that could identify microglia (Iba1), the primary resident immune cell (CD206 is not useful, at this stage, in perivascular macrophages that are extremely sparse in the postnatal brain). This would be important to better study the impact on CNS resident immune cell morphological activation.

      Thank you for this comment. In response, we have analyzed Iba1 staining in the cortex in infected vs. uninfected mice. This is shown in Figure 2 – figure supplement 3. These data demonstrate a several-fold increase in Iba1 immunostaining in infected compared to uninfected cortex, consistent with increased microglial activation in response to infection. There is no statistically significant difference between infected WT and infected Cdh5-CreER; Tlr4CKO mice in Iba1 staining in cortex.

      (3) The authors suggest that Cldn5 junctional localization is selectively disrupted upon bacterial exposure, mediated by TLR4 - they suggest this based on studying PECAM, GLUT1, ZO-1 and B-catenin (all normally junction or cell surface located in cultured Bend3.1) in relationship to Cldn5 localization (normally high) - it is possibly these are also impact by bacteria exposure (maybe through different mechanisms?) - a better measure would be to use the similar cyto/PM measure they do for Cldn5 in Fig. 4D and to evaluate this or to use intensity measurements.

      Thank you for this comment. As the reviewer noted, the analysis of Cldn5 localization with vs. without E. coli exposure and in WT vs. Tlr4KO bEnd.3 cells (shown in Figure 4B and D) – uses Cell Trace to partition the image into cytoplasmic vs. plasma membrane territories. For the analyses in Figure 5, we wanted to compare the localization (and potentially re-localization) behaviors of a variety of subcellular markers with the localization and re-localization of Cldn5 following E. coli exposure. By directly measuring the % overlap of the two immunostains, we get that data. We note that the goal of this analysis is to assess relative co-localization with Cldn5 rather than absolute subcellular partitioning of each marker. While this analysis could have been extended to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories, it is clear by visual inspection of Figure 5A-C that beta-catenin, ZO-1, and PECAM1 remain plasma membrane-associated with E. coli exposure, and GLUT1 goes from the part of the plasma membrane not involved in cell-cell contact without E coli exposure to cytoplasmic with E. coli exposure (as judged by the appearance of a nuclear “shadow” after E. coli exposure). Thus, we do not believe that additional cytoplasmic vs. plasma membrane quantification for these markers would alter the interpretation. The main reason that we did not extend this analysis to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories is because that would introduce the Cell Trace localization as an additional variable.

      (4) The discussion could benefit from delving more into the prior literature on E coli mediated breakdown of junctions in cultured human microvascular brain endothelial cell model and critical host-pathogen interactions of the bacteria with ECs (PMID: 14593586), and how this might involve TLR4.

      Thank you for this comment. Two paragraphs addressing the prior literature have now been added to the discussion.

      (5) It would be important to discuss how their results relate to earlier studies on TLR4-/- and TLR2-/- global knockout mice and protection vs vulnerability to development of meningitis (see PMCID: PMC3524395) - this paper showed that TLR4 global KO mice have increased susceptibility to die from meningitis and have much higher CFU counts in the CNS. In this manuscript and their prior work (Wang et al., 2023), this group shown that both global TLR4-/- mutants and their EC-specific KO have reduced barrier permeability, but we don't have any information about CFU or susceptibility to death from meningitis in their models.

      Thank you for these comments. The model we use – subcutaneous injection of E. coli (a clinical isolate from an infant with meningitis) at postnatal day (P)5 – results in the death of the infected mouse within 2 days (shown in Figure 1 – figure supplement 3 in Wang et al. 2023). Our analyses of infected mice were conducted 24 hours after infection. As noted in the reply to comment #1, in the revised manuscript we present a clinical assessment of WT vs. Cdh5-CreER; Tlr4CKO mice 24 hours after infection based on (1) a quantitative microscopic analysis of E. coli burden in the brain (visualized based on RFP fluorescence in the E. coli used here), (2) quantifying CFUs in blood and (3) mouse weights at P5 and P6, a sensitive indicator of overall health since this is a time when mice are normally gaining weight rapidly (~25% weight gain per day). These data (shown in Figure 2 figure supplement 4) indicate that bacterial burden and disease severity are similar between genotypes in our model. In Wang et al., 2023, we did not conduct a quantitative clinical assessment of WT vs. Tlr4-/- mice following infection, but by visual inspection, infected WT and Tlr4-/- mice appeared to have similar downhill clinical trajectories. We have expanded the Discussion to relate these findings to prior studies of global TLR4 and TLR2 knockout mice, noting that differences in experimental models and the distinction between global versus VECadCreER-specific deletion may account for the differing outcomes reported.

      Comment on the paper listed by the reviewer (PMCID: PMC3524395).

      The cited study demonstrates that global TLR4 deficiency leads to increased bacterial burden and mortality, indicating an essential role for TLR4 in host defense and bacterial clearance. In our study of Cdh5-CreER; Tlr4CKO mice, bacterial burden and disease severity at 24 hours post-infection are similar between WT and Cdh5-CreER; Tlr4CKO mice, indicating that Cdh5-CreER; Tlr4CKO does not alter the clinical course at this time point. This difference is noted in the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defence in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells (using Cdh5-CreER) or macrophages (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      Weaknesses:

      The weaknesses of the study were in terms of interpretation and perhaps study design.

      (1) Most importantly, the authors need to provide additional validation of their conditional knockout models. The authors need to confirm that the Cdh5-CreER does not impact leptomeningeal fibroblasts and to confirm gene deletion in macrophages.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclearlocalized GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite correct. The Results section of the revised manuscript includes an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      (2) The authors could also strengthen the paper by providing data on the impact of these conditional knockout models on the course of meningitis and bacterial burden.

      Thank you for this comment. We agree that these additional analyses strengthen the manuscript. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences in bacterial burden between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) Finally, it is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

      At present, it is an open question whether TLR4 plays as a large a role in meningitis caused by other gram-negative bacteria and whether TLR2 plays a similarly large role in meningitis caused by gram-positive bacteria. In the Discussion, the last two sentences under “Limitations of the study” summarize this point: “Finally, the present study focused on E. coli K1, the dominant Gram-negative neonatal pathogen. Future work could assess TLR signaling in response to other bacterial pathogens, such as Group B Streptococcus.”

      (4) There are additional minor issues; for instance, the arachnoid fibroblast 2 population appears to closely resemble dural border cells.

      Thank you for this comment. That is correct, and we have changed the nomenclature to “dural border cells”.

      (5) The cell line model (bEnd.3) is a relatively low-fidelity model of BBB endothelial cells, and this should be acknowledged.

      Thank you for this comment. That is correct. Despite being brain-derived, bEnd.3 cells have lost many BBB-specific attributes. Their responses might best be considered as generic endothelial responses rather than brain-specific endothelial responses. This is now stated in the Results section: “Although they are brain-derived, bEnd.3 cells lack many BBB-specific attributes and, therefore, they likely exhibit generalized endothelial responses rather than brain-specific responses to bacterial exposure.”

      With these caveats, it is difficult to be certain that the endothelium alone is the driver of meningeal immune responses in meningitis, and what the impact of these is.

      We agree with this critique. As noted above, the expression of Cdh5-CreER in essentially all endothelial cells and in a subset of dural border cells and leptomeningeal fibroblasts means that the comparison of TLR4 CKO with Cdh5-CreER vs. Lyz2-Cre is assessing phenotypes driven by TLR4 signaling in endothelial plus a subset of other non-myeloid cells vs. TLR4 signaling in myeloid cells. We have revised the text to reflect this more precise understanding of Cdh5-CreER specificity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Transcriptomic analysis: The analysis and display of the single-nucleus RNA-seq data should be improved. The authors could perform a more granular, unbiased clustering of each cell class in the combined dataset and then compare the proportion of each experimental group (genotype x control/infected) in each cluster. At present, it appears the differentially-expressed genes (DEGs) shown in Figure 1 were identified using the Seurat FindMarkers function with default parameters (Methods). This considers each cell as an independent experimental unit and is therefore not appropriate for a comparison of control versus infected groups (see e.g., PMID 34584091, 35880426. The authors should implement a statistical analysis strategy that considers true biological replicates (mice, as shown in Supplementary File 1).

      We do not fully agree with this critique. We agree that biological replication at the level of individual mice is important for interpreting these data, but within each mouse, the characteristics of individual cells is also of interest, including the degree of heterogeneity, the sample size for a given cell cluster, and the statistical significance of any observed changes in transcript abundance. As requested, we have prepared a new supplemental figure (Figure 1 – figure supplement 5) showing a principal component analysis of the scRNA-seq data for each mouse (one mouse was used for each snRNA-seq dataset) and for each of the principal leptomeningeal cell types. This analysis shows, for example, that the three infected Cdh5-Cre; Tlr4flox/- mice have transcriptomes for each of the six cell clusters that are very similar to the transcriptomes of the two uninfected WT and the two uninfected Cdh5-CreER; Tlr4flox/- mice. Thus, the genotype- and condition-dependent effects are consistent across biological replicates. At the most granular level, Figure 1 – figure supplement 7, which was part of the original submission, shows for the most up- and down-regulated genes (based on adjusted p-value or based on fold-change) in endothelial cells and in myeloid cells how individual transcript abundances change for each mouse and for each genotype/condition.

      (2) The authors use immunohistochemistry to assess claudin-5 "disorganization and redistribution" (Results, pages 7-8 and Figures 3A-B). They state that "Tlr4ECKO mice showed minimal changes in the distribution of Cldn5, implying that cell autonomous endothelial TLR4 signaling regulates tight-junction organization." It is not clear, however, that the quantified parameter (Cldn5+ area relative to total area) would be an accurate readout of claudin-5 organization/distribution (i.e., subcellular localization) as it would also be sensitive to claudin-5 expression, vascular density, and vessel diameter. The authors use a similar assessment of ZO-1 to suggest that changes to claudin-5 are not due to a "generalized disassembly of TJs" and could also use this to argue that the above potential confounds (vascular density, vessel diameter) do not change, but the data in Figure 3 - Figure Supplement 1B, lower panel, show that infection does cause an increase in ZO-1 area relative to total area (P = 0.0004). Thus, the statement in the results "Zonula Occludens-1 (ZO-1) [...] remained unchanged during infection (Figure 3 - figure supplement 1)" is not accurate. The authors should revise this section to ensure their conclusions are aligned with the presented data.

      Thank you for this comment. The reviewer is correct that the Cldn5 area measurement is unable to deconvolve the various factors that might contribute to it (vessel density and diameter, and Cldn5 distribution). This part has been rewritten. “Consistent with prior findings (Wang et al., 2023), both WT and Tlr4<sup>MKO</sup> mice showed an increase in the area occupied by Cldn5 in the leptomeninges following infection, likely referable to both increased vessel diameter and a redistribution of Cldn5 within ECs (Figure 3A-B; Figure 3 – figure supplement 1C).”

      The reviewer is also correct about our initial description of the ZO-1 data. What we meant to write and what the revised manuscript now shows is: “The area occupied by Zonula Occludens-1 (ZO-1), a tight junction scaffold protein, showed a modest but statistically significant increase in WT leptomeningeal vessels but no significant change in Tlr4<sup>VEKO</sup> leptomeningeal vessels during infection (Figure 3 – figure supplement 1A and B).”

      (3) In the Methods, under Mouse Models and E. coli Infection, the authors state, "The Cdh5-CreER line (Monvoisin et al., 2023) was the same line used in Wang et al. (2023)." However, there is no mention of Cdh5-CreER in Wang et al. (2023). Could authors please clarify? Also, because this appears to be an inducible Cre, the authors must include details on the dose and timing of tamoxifen or 4-OHT used in this study.

      Thank you for catching that error. We meant to reference Wang et al (2025), not Wang et al (2023). [Wang et al (2025) is: Wang Y, Rattner A, Li Z, Smallwood PM, Nathans J. (2025) Vascular endothelial-specific loss of TGF-beta signaling as a model for choroidal neovascularization and central nervous system vascular inflammation. Elife 14:RP107018.] This has now been corrected.

      We have now included the details related to 4HT injection in the Methods section “Mouse Models and E. coli Infection”. These are intraperitoneal injection at P2 with 40- 50 µL of 2 mg/ml 4HT.

      (4) The legend for Figure 1A is "Schematic of the leptomeninges", but the figure shows the entire brain-skull interface, including underlying cortex, leptomeninges, dura, and skull.

      Thank you. Corrected.

      (5) Page 7, typo: "In the brain, CD206+ cell were too sparse ..." Should be "cells".

      Thank you. Corrected.

      (6) Page 14, typo: "... could represents a double-edged ..." Should be "represent".

      Thank you. Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform CFU counts from LPM, dura, brain, peripheral organs (liver) in infected v mock mice from control v TLR4 EC-cKO.

      Thank you for this comment, with which we agree. We have addressed this by quantifying bacterial burden and assessing disease severity in WT and Cdh5CreER; Tlr4floxed mice. Specifically, we performed CFU measurements in blood, monitored mouse weights at P5 and P6, and histologically surveyed the E. coli-RFP signal (i.e., E. coli burden) in brain, liver, and lung. These analyses show that bacterial burden and disease progression are comparable between WT and Cdh5-CreER; Tlr4floxed mice at 24 hours post-infection. These data are presented in Figure 2 – figure supplement 4 and described in the Results.

      (2) Lyz2Cre/+ is used to delete TLR4 from macrophages, but recombination efficiency (in LPM BAMs) is described as only partial, suggesting that TLR4-response in LPM BAMs (and potentially macrophages in the dura) is at least partially intact. It undercuts conclusions that can be made using this line.

      Thank you for this comment. We have conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The Results section text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Also, as noted in the reply to comment 4 below, a direct analysis of Tlr4 recombination efficiency is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (3) Inflammatory responses [qPCR] from peripheral organs and also physiological measures in the pups [weight post-infection, time to moribund or death curves] in control v TLR4 EC-cKO and TLR4 mac-cKO.

      Thank you for this comment. We have not conducted a qPCR analysis of inflammatory gene expression in peripheral organs because (1) the dramatic upregulation of these transcripts in the leptomeninges, (2) the presence of E. coli in blood and peripheral organs, and (3) the clinical assessment (cessation of weight gain) all predict that such an analysis would reveal a large up-regulation of inflammatory gene expression throughout the body. More specifically, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (4) The conditional macrophage line is problematic due to the partial recombination. I question the utility of including this unless they can come up with a way resolve the response of recombined TLR4 macrophages vs ones that are not (could they use the single cell data to pick this a part? Are TLR4-null cells and TLR4 'wt' cells transcriptionally similar in the infected condition, suggesting TLR4 is not doing much in the macs, potentially due to alternate TLRs?). There are good BAM Cre lines that have been described [Lyve1-cre would be good for LPM BAMS, the other is Pf4-cre, see https://pmc.ncbi.nlm.nih.gov/articles/PMC7375817/ - just as an FYI for the future].

      Thank you for this comment. As noted in the reply to point 2 (above), we have conducted a more in-depth analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      We agree that, based on Figure 6 in the cited paper [McKinsey et al (2020) A new genetic strategy for targeting microglia in development and disease eLife 9:e54590], the Pf4-Cre line may be superior to the Lyz2<sup>Cre</sup> line that we used for recombination in leptomeningeal myeloid cells. Unfortunately, we missed this paper in our literature searches, probably because it focuses on a microglial CreER line, P2ry12-CreER, and the Pf4-Cre line is not mentioned in the title or abstract. Our decision to use the Lyz2<sup>Cre</sup> line was based on an extensive comparison among myeloid Cre lines showing that Lyz2<sup>Cre</sup> was the most efficient [Abram CL, Roberge GL, Hu Y, Lowell CA. 2014. Comparative analysis of the efficiency and specificity of myeloid-Cre deleting strains using ROSA-EYFP reporter mice. J Immunol Methods 408:89-100.] However, the Abram et al study did not look at the leptomeninges. Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (5) Figure 1 - Figure Supplement 2 - the authors nicely break down the pathway response [NFKB and TNF] in EC and macs, it would be great to have similar information for the fibroblasts (in the main figure or the supplement). Does their inflammatory response show a similar pattern?

      Thank you for this suggestion. We have now done that analysis and present it in Figure 1 – figure supplement 3. For completeness, we also performed the same type of analyses for JAK-STAT signaling and IFN-gamma response and these are shown in Figure 1 – figure supplement 6. The principal conclusion is that across all major leptomeningeal cell types, the Cdh5-CreER; Tlr4floxed samples (i.e., Tlr4 KO’d in non-myeloid cells) show much reduced transcriptome changes with infection.

      (6) What is ASC and Cd206 quantification measuring, and how does this relate to 'activation' - is this the intensity of signal or a morphological change? What is the precedence for using ASC (citations)? In their prior work, they showed no change in CD206 number, so a significant increase upon infection here, it's confusing exactly what is being studied. Also, loss of Lyve1 is a well-accepted measure of activation that they have previously used, adding that it could be helpful. This is not a major issue since they have robust data that the macrophages are not transcriptionally activated. Clarification of what exactly is being measured would be sufficient (in the text).

      CD206 immunostaining, which reveals myeloid cell morphology, shows that, with E. coli infection, myeloid cells convert from a more compact morphology to a more expanded morphology. This is now explained more fully in the Results section.

      Regarding ASC, changes in the state of ASC aggregation and ASC subcellular localization have been used by others to monitor immune cell responses to inflammatory signals (Sester et al., 2016; Franklin et al., 2018). While this change in subcellular localization may explain part of the increase in immunostained area in myeloid cells in the infected mice (Figure 2D), the increase in the area of ASC immunostaining largely reflects a shift of myeloid cells from a compact to a more extended morphology. This is now explained more fully in the Results section. We have also added two references (Sester et al., 2016; Franklin et al., 2018) that described how ASC distribution changes with inflammation.

      Regarding LYVE1, we observe a decrease in LYVE1 transcript abundance in myeloid cells with infection, as predicted. Given the large amount of other data that document myeloid activation with infection, we have elected not to include this.

      (7) The authors suggest the internalization of Cldn5 is not due to NFKB downstream signaling that includes transcriptional mechanisms because it happens as early as 1 hour, prior to NFKB localization to the nucleus. However, a lot of their experiments, including on endosomal-lysosomal protein co-localization are done at 4 hours, when their RNAseq data show robust NFKB-mediated gene upregulation and (though not tested) potentially protein production of factors that can act back on the cells, including to impact endo-lysosomal processing. Without studies at earlier timepoints post-bacteria exposure, separating these two mechanisms is difficult.

      Thank you for this comment. We have explored this question by looking at Cldn5 internalization in bEnd.3 cells at 1 hour after E. coli exposure, and the data clearly show that internalization occurs within 1 hour. Additionally, we have conducted this experiment in the presence of 1 uM ACHP, an IKK inhibitor that blocks NF-кB migration to the nucleus. ACHP treatment shows no effect on the rapid internalization of Cldn5, implying a mechanism independent of NF-кB control of gene expression. These data are shown in a new figure (Figure 6) in the revised manuscript.

      (8) Figure 2 - CD206 are quite sparse however, Iba1 would work well to look at microglial activation.

      Thank you for this suggestion, which we have followed. To assess microglial activation, we have immunostained for Iba1 and quantified the data. These are now included in Figure 2 – figure supplement 3. The data show that there is an increase in Iba1 immunostaining following E. coli infection in both WT and Cdh5-CreER; Tlr4floxed mice, with more in the former than the latter, but the difference is not statistically significant.

      (9) Suggest performing the LAMP+ co-localization experiment at <1hr, prior to NFKB nuclear localization and transcriptional changes. This would better support it, this is (or is not) independent of the NFKB. Could also test this with an NFKB inhibitor, do they still see the CLDN5 internalization when NFKB is blocked?

      Thank you for these suggestions. We have done both of these analyses, and the results are presented in Figure 6. The results show that (1) Cldn5 is internalized within 1 hour and (2) its internalization is independent of NF-кB signaling inhibition by 1 uM ACHP. Since ACHP treatment shows no effect on the rapid internalization of Cldn5, that implies a mechanism independent of NF-кB control for gene expression.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) The most important caveat is that the Cdh5-CreER model is known to recombine in leptomeningeal fibroblasts (10.1038/s41586-023-06993-7, 10.1101/2025.05.13.653681), and Cdh5 expression in these populations is now well described (10.1038/s41467-02341580-4, 10.1016/j.neuron.2023.09.002). Although the authors did not observe recombination in their reporter (details of the tamoxifen injection protocol should be provided), it is imperative to validate the specificity of their model to Tlr4 in endothelial cells, leveraging their sequencing data and providing additional IHC or ISH to confirm this. Alternatively, Tlr4 could be deleted in a more specific model, e.g., the Pdgfb-iCreERT2 or Slco1c1-CreERT2. It is also important to do the same with the LysM model, to confirm that the lack of impact of macrophage Tlr4 is not due to failure to delete the gene. This is again important to the interpretation of the study, since the authors propose that the endothelium, specifically, is the driver of the meningitis response.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclear-localised GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite right. The Results section of the revised manuscript has an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      We have also conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup>-directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text in the Results section now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. The phenotype of Cdh5-CreER; Tlr4floxed mice – a dramatically reduced infection-associated transcriptional response – argues that the floxed Tlr4 target was recombined at appreciable efficiency in those mice (Figure 1D and 1E). For Lyz2<sup>Cre</sup>; Tlr4floxed mice the principal phenotype is an up-regulation of infection-associated transcripts in a subset of dural border cells in the absence of infection; the transcriptional response to infection was largely unaffected in all leptomeningeal cell types (Figure 1D and 1E). We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4VEKO and Tlr4MKO mice.”

      (2) The authors did not examine the consequences of Tlr4 cKO on the course of meningitis or bacterial burden. Knowing the impact of this would strengthen the paper and allow us to determine if the endothelial responses are helpful or harmful in meningitis progression.

      For the revised manuscript, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) TLR4 is a known receptor for LPS. It is unsurprising (especially in the in vitro experiments) that Tlr4 knockout reduces NF-kB signalling and other downstream changes to endothelial cells. Furthermore, it is uncertain if the infection was left to continue, similar changes to the endothelium would nonetheless occur through other mediators such as IL1B and TNFa.

      We agree that it makes logical sense that Tlr4 KO decreases NF-кB signaling. The interesting next question is: what are the mechanistic underpinnings of the responses that are downstream of TLR4 and NF-кB? The cell culture experiments with WT vs. Tlr4KO bEnd.3 cells identify one set of cell biological responses related to Cldn5 and junctional integrity, and the NF-кB inhibition experiment (Figure 6) implies that rapid internalization of Cldn5 occurs in the absence of NF-кB mediated transcriptional changes. Regarding the possibility that other mediators such as IL1B or TNFα might, at least partially, make up for the lack of TLR4 signaling later in the infection, that is an open question at present.

      (3) The arachnoid fibroblast 2 cluster should be renamed to dural border cells based on their high expression of Slc4a10, Adamtsl3, Tmeff2, etc which are all highly enriched in dural border cells. I suspect this cluster is also highly enriched for Slc47a1, probably the most specific marker for these cells (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002).

      Thank you for this comment. The reviewer is correct. These are dural border cells and they express Slc47a1, as seen in a new supplemental Figure 1 – figure supplement 4, which shows UMAP plots for many leptomeningeal cell type-specific genes. We have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (4) It would be helpful to provide higher resolution images of Cldn5 in the leptomeningeal mounts. At the current resolution, it is difficult to tell if there is a similar internalisation/disruption phenotype to what is observed in vitro. Notably, this finding is similar to another recent publication on Cldn5 recycling (in the context of stroke) (10.1186/s40478-025-02125-6).

      Higher resolution images of Cldn5 in leptomeningeal vessels without or with E. coli infection are now shown in Figure 3 - figure supplement 1C. There is a visual impression of greater area occupied by Cldn5, which is confirmed by quantification (Figure 3A and B). This effect appears to be due to both an average increase in vessel diameter and a redistribution of some of the Cldn5 away from plasma membrane junctions. Thank you for pointing out the interesting and relevant Cottarelli et al (2025) paper, which we had not read. This is now referenced.

      Minor points

      (1) Typo: prominant should be spelled prominent.

      Thank you for catching that one. It is now corrected.

      (2) Strictly speaking, the arachnoid layer is not epithelial (despite Cdh1 expression). They are fibroblasts that acquire barrier-forming properties.

      Thank you for that comment. That appears to be the consensus view, and we will go along with it.

      (3) Notably, LyzM Cre will also recombine in other myeloid populations, so I wouldn't describe it as a macrophage.

      Thank you for this comment. We agree, and we have therefore changed the text and figure labels from “macrophage” to “myeloid”.

      (4) It is interesting and notable that ICAM1 expression is observed in nonendothelial populations, in the IHC, too, perhaps.

      We agree. ICAM1 may be a broader marker/mediator of inflammation than is generally recognized.

      (5) In F1B, your labels on the right image to the arachnoid barrier and pial surface are presumably meant to refer to the image on the left with DPP4 and laminin labelling? The subarachnoid should be between the laminin and DPP4 layers (although it will be collapsed in your preparations).

      Thank you for catching this error. The vertical bars were sized erroneously, and the labels were also placed erroneously. These have now been corrected.

      (6) I would reference the papers that defined leptomeningeal cell type markers (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002) when you define your cell types.

      Thank you. We have done that, and we have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (7) I would change references to the subarachnoid space in your figures to the leptomeninges (which include the SAS, but extend either side of it).

      Thank you. The labels have been changed to “leptomeninges”.

      (8) In Figure 2 - Supplement 1A, it looks like the populations are mislabelled.

      Thank you. This has been corrected to be consistent with the assignments in Figure 1B

    1. Author response:

      We are pleased that the reviewers found the repeated single-neuron sequencing and the finding of less than 1% transcriptomic variability to be original and striking, valued the single-neuron Dscam isoform repertoires and the scale of the functional screen, and judged the evidence solid to compelling. We provide below our provisional response and an outline of the revisions we plan.

      Overall plan: We intend to submit a revised version that addresses the public reviews and the recommendations to the authors. Because our conclusions rest on data already in the manuscript, the revisions are clarifications, added analysis of existing data, tempered language, and improved figures, rather than new experiments. Given the focused nature of these revisions, we would be happy for the editors to assess the revised version without re-involving the reviewers.

      One factual note for the Assessment and public reviews: The morphological RNAi screen comprised 213 cell-surface receptor genes; the figure of “140 genes” in one public review is the number that produced strong-to-severe phenotypes (Grade 3–5 at >40% penetrance), not the number screened. We will make this unambiguous in the revised text.

      Main changes in the revision:

      (1) We will explain that the RNAi screen was performed blind and independent of the RNA sequencing experiments. This was intentional, so that functional perturbation and transcriptomic identity would serve as independent lines of evidence, but could be compared with each other.

      (2) We will revise the Methods and Results to clarify how the morphological RNAi screen and behavioral subset should be interpreted, including conservative treatment of the negative-control distribution and mild-to-moderate phenotypes.

      (3) We will soften language that overstated certainty. Differentially expressed molecules are now described as prioritized candidates and convergent evidence, not definitive determinants.

      (4) We will reframe Gr59d/PXGS experiments as morphological rewiring and ectopic branching, and no longer imply a pSc-like conversion or functional rewiring.

      (5) We will add a limitations paragraph addressing RNAi off-target/background concerns, the absence of direct aPa functional testing, and the need for future mechanistic validation.

      (6) We will disclose or remove any figure panels that overlap with the companion PXGS manuscript and revise legends/labels to make it more clear.

      We hope these revisions make the logic of the study clearer and align the strength of the claims with the evidence.

      Below is our more detailed (provisional) response (not sure if this is required at this stage):

      Response to the eLife Assessment:

      Clearer articulation of the experimental logic. The Assessment’s central request (Reviewer 2) concerns the relationship between the transcriptomic experiments and the functional screen. The two were performed independently on purpose; the RNAi screen was assembled from a comprehensive survey of the literature rather than from the results of our differentially expressed genes from single-cell RNA sequencing. Thus, the RNAi screen was performed and graded blind in parallel with the single cell sequencing, with the gene identities unmasked only after both were complete. This was intentional, so that the sequencing (i.e., which molecules differ between neurons) and the screen (which molecules are functionally required) would provide mutually unbiased corroborating evidence rather than self-referential support/circular reasoning. We will state this more explicitly in the Introduction, in the Results where the screen is introduced, and in the Methods.

      Additional controls in the RNAi screen. We will treat the 39 genes that were identified in our single-cell RNA sequencing to be not expressed in the pSc neuron as a randomized negative control set in our RNAi experiments. We will state more explicitly the empirical RNAi false positive rate for a miswiring phenotype is 6/39 = 15%, likely due to RNAi off-target effects. We will also more clearly state that our claims about cell surface receptor functions are restricted to strong-to-severe phenotypes at high penetrance reproduced by at least two independent RNAi lines and corroborated independently (differential expression and/or single neuron qPCR).

      A more complete characterization of the re-wiring. We will state more clearly that mis-expressing the pSc-enriched cell surface receptors within Gr59d neurons partially shifts the arbour toward a pSc-like pattern (e.g., increased ectopic branching), and does not reproduce the full anatomical wiring, and that functional/behavioral re-wiring was not tested.

      Response to Reviewer 1:

      Reviewer 1 found the work valuable and data-rich, and the Dscam expression bias interesting. They noted over-confident language and asked how rigorously the differentially expressed genes were identified.

      Over-confident language. We will rewrite the two flagged sentences. The claim that the ~10 differentially expressed molecules are “likely the most important” will become a correlational statement, while also noting the lack of an aPa-specific Gal4 driver for direct testing. Our sentence that, “Our RNA sequencing data is biologically inadequate without a functional characterization of each molecule within the specific neuron” will be replaced with a clearer statement that gene expression data can nominate candidates, and functional perturbation of each gene is required to demonstrate necessity and sufficiency (i.e., biological function); which is exactly why we paired the RNA sequencing with an independent RNAi screen.

      Rigor of the differential-expression calls. We will more clearly state the statistical criteria in the Results (absolute log2 fold change ≥ 2 and Benjamini–Hochberg-adjusted p < 0.05). We will also note the small replicate numbers for the pooled pSc versus aPa comparisons, and emphasize that the central gene calls are independently supported by the blind RNAi screen and, for five genes, by single neuron qPCR. The full statistical workflow is in the Methods.

      Response to Reviewer 2:

      Reviewer 2 considered the findings potentially important but raised concerns about the experimental logic, the rigour of the screen, the completeness of the re-wiring, figure quality, and overlap with our PXGS companion paper. We will address each.

      Experimental logic. Beyond the design of our independent, blinded RNAi screen described above, we will add to the Introduction the rationale for sequencing single identified neurons (rather than subclasses) along with the two-pronged strategy, and add a summary paragraph at the start of the Discussion.

      Developmental stage choices. We will clarify our justification for the P14 pupal stage (the period when the mechanosensory neuron is actively elaborating its arbour while also enabling dissection). We will also clarify the rationale and caveats for comparing the pupal pSc neuron with the adult Gr59d neuron (i.e., the wiring occurs at the pupal stage, but the pupal Gr59d neurons could not be isolated at sufficient quality; the pSc pupal samples are less age-synchronized, so we simply used the comparison to identify the genes shared with the adult comparison).

      Transcriptome precision controls. We will state that the ten single pSc neurons passed the same quality controls for neuronal markers (elav, nSyb) and glial markers (Repo, moody < 20 reads) as all single-neuron libraries, which argues against any contamination by the attendant glial cell, and the aPa transcriptome is used as a different identity comparison.

      Off-target rate. As stated above, we will add the false positive rate for RNAi and restrict our confidence claims to those genes/cell surface receptors with multiple lines of evidence (e.g., strong phenotype, multiple RNAi lines, RNA sequencing, etc).

      Rewiring completeness and the behavioral prediction. As stated above, we will clarify that true re-wiring of the Gr59d neuron requires a future experiment, where a bitter tastant stimulus would elicit a grooming response.

      Response to Reviewer 3:

      We thank Reviewer 3 for judging our work to be fundamental in significance and the evidence compelling, with no major criticisms. Our clarifications above will further reinforce our hypothesis that the differential expression of specific cell surface receptors “do, in fact, control synaptic patterns,” which the reviewer highlighted.

      We are grateful for the reviewers’ time and for eLife’s model. We believe the planned revisions substantially clarify the experimental logic and tighten the claims, and we look forward to submitting the revised version.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Zonation definition under injury has been shown to be sustained broadly, but is not sufficiently validated and quantified, especially considering the resolution of the 10x Visium system and the potential variation of outcomes based on how to define zones.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g.,Alb, Mup20, Cyp2f2,Pck1,Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2,Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g.,Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) The model is built entirely in APAP injury, which specifically targets pericentral hepatocytes. It remains unclear whether the proposed mechanism applies to other liver injuries (e.g., partial hepatectomy, CCl4).

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Baseline proliferation appears higher than expected in homeostasis (Figure 1B), and fold change analysis (not absolute counts) may be needed to assess zonal proliferation suppression (Figure 1D).

      We thank the reviewer for this insightful comment. The baseline proliferation rates observed in our study are consistent with previously reported zonal distributions (PMID: 33632817; PMID: 33632818), with approximately 70% of proliferating hepatocytes located in zone 2, 20% in zone 3, and 10% in zone 1 under homeostatic conditions. To further address the reviewer’s concern, we performed a fold-change analysis of Ki-67<sup>+</sup>hepatocytes across different zones. This analysis revealed that only the mid (zone 2) and pericentral regions exhibited significant changes, whereas no statistically significant differences were observed in the other zones (as shown in Author response image 1). Importantly, when considered together with the absolute cell counts, these results indicate that the apparent suppression of proliferation is most pronounced in the mid zone, likely due to its relatively higher baseline proliferation under homeostatic conditions. In contrast, this effect is less evident in the fold-change analysis, as zones with low baseline proliferation show limited dynamic range for detecting relative changes.

      Author response image 1.

      Fold changes of Ki67-positive cells across liver zones (PC, Mid, PP) at 0, 3, 6, 12 and 24 h post-APAP. Fold change was the number of Ki67-positive cells in the three regions at each time point after APAP treatment divided by the number of positive cells in each region at 0 hour post-APAP. (a) denotes significance between PC and Mid regions, (b) denotes significance between PC and PP regions, and (c) denotes significance between Mid and PP regions.

      (4) AAV-based overexpression raises potential confounds (altered CYP activity before injury) and shows incomplete penetrance that is not quantified (Figure 5 - Figure 6).

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (5) The functional link between proliferation suppression and improved survival is inferred, but direct survival /injury readouts are limited.

      We thank the reviewer for this insightful comment. To more directly evaluate the functional link between proliferation control and liver injury, we manipulated Btg2, a downstream effector of the Atf4–Chop axis and a known inhibitor of cell proliferation. Knockdown of Btg2 using AAV8–CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial knockdown efficiency, we observed a clear exacerbation of liver injury, as evidenced by an approximately 2-fold increase in serum ALT levels and a ~1.5-fold expansion of necrotic areas. In parallel, hepatocyte proliferation was significantly increased (~1.8-fold increase in Ki67⁺ hepatocytes) compared to control mice (Revised Figure 6K–N). Conversely, Btg2 overexpression produced the opposite phenotype, markedly attenuating liver injury while suppressing hepatocyte proliferation (Revised Figure 6G–J). Together, these gain- and loss-of-function data provide direct evidence linking proliferation control to injury severity, thereby supporting a causal relationship between suppressed proliferation and improved liver outcomes (Revised manuscript, page 14, lines 399–406).

      Reviewer #2 (Public Review):

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP but pericentral hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1, but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D, but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work?

      We thank the reviewer for this important question. We fully agree that the differential susceptibility between mid‑zone and pericentral (PC) hepatocytes is likely rooted in the zonal gradient of cytochrome P450 expression. Our spatial transcriptomics data (Revised Figure 2A) show that mid‑zone hepatocytes express Cyp2e1, Cyp1a2, and Cyp3a11 at levels intermediate between PC and periportal (PP) zones. This intermediate expression may generate sufficient NAPQI to activate stress signaling but not so much as to cause immediate mitochondrial collapse, thus “buying time” for adaptive responses. We also appreciate the reviewer’s observation that Cyp2e1 mRNA levels remain highest in the PC zone even after APAP (Revised Figure 3B). However, mRNA abundance does not necessarily reflect functional protein level. In the PC zone, massive necrosis rapidly compromises cellular integrity; as shown in Revised Figure 3D, Cyp1a2 protein declines sharply around the central vein, and we observed similar degradation for Cyp2e1 (data not shown). Consequently, despite sustained Cyp2e1 transcripts, PC hepatocytes are unable to mount an effective stress response because they are already undergoing cell death. By contrast, mid‑zone hepatocytes retain sufficient metabolic capacity to activate the Atf4‑Chop axis while preserving cellular function.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall, as shown in F2E?

      We thank the reviewer for these important observations. We agree that the current spatial transcriptomics (ST) data alone do not provide sufficiently strong support for this conclusion. The limited sensitivity of ST for detecting rare proliferative events further constrains its utility in this context. At baseline, only ~1% of ST spots are Ki67-positive (Revised Figure S1I), and this fraction becomes even lower during the early phase following APAP injury. As a result, there are insufficient Ki67+ spots to robustly assess zonal distribution using ST, which precludes a reliable spatial validation of proliferation patterns at this resolution. For this reason, our primary evidence for zonal proliferation dynamics relies on Ki67 immunohistochemistry (Revised Figure 1), which provides single-cell resolution and higher sensitivity. These data show a marked reduction in Ki67+ hepatocytes specifically in the midlobular zone at 3-6 hours post-APAP, supporting a transient suppression of proliferation in this region. In addition, we agree that the transcriptional evidence for cell cycle arrest was not strong the S and G2/M scores showed no overt difference, and the gene sets used were suboptimal. We have therefore moved these analyses to the supplement and toned down the claims. We have also clarified this limitation in the manuscript (Revised manuscript, page 18, line 518-524)

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, the localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of the knockout of Btg2 is also not given.

      We thank the reviewer for this insightful comment. We have included the missing data. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup>hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N, Revised manuscript, page 14, line 399-406).

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3?

      We thank the reviewer for this insightful comment. We have included an inset of the original image to better show BTG2 staining in revised Figure 6F. During our study, we tested BTG2 expression in mice transduced with AAV‑TBG‑EGFP or AAV‑TBG‑BTG2 for three weeks without APAP challenge. We observed that BTG2 in these non‑injured livers was predominantly cytoplasmic (Author response image 2), contrasting with the nuclear localization seen after APAP treatment (Figure 6F). Regarding whether it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3, we think Ddit3 promotes BTG2 expression (as shown in revised Figure F6F), but APAP is necessary for its nuclear translocation.

      Author response image 2.

      Immunohistochemical detection of Btg2 in liver tissue from mice transduced with AAV-TBG-EGFP or AAV-TBG-Btg2 for 3 weeks without APAP treatment.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group, presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address whether it is really Atf4 that acts on Ddit3 in APAP toxicity.

      We thank the reviewer for this insightful comment. We agree that Atf4 was not among the top 30 active transcription factors in our initial analysis; however, when we extended the list to the top 50, Atf4 was included. We have therefore updated Revised Figures S4F and G to show the top 50 transcription factors. We also appreciate the reviewer’s suggestion to further investigate whether Atf4 directly acts on Ddit3 in the context of APAP toxicity. While this still shows a less pronounced midlobular enrichment for Atf4 compared with Ddit3, we sought additional evidence for a functional Atf4-Ddit3 link. In primary hepatocytes treated with APAP, we observed nuclear co‑localization of Atf4 and Ddit3 (Author response image 3A) and increased nuclear protein levels of both factors (Author response image 3B), supporting their potential cooperative role. We agree that direct analysis of AAV‑Atf4 samples generated for Figure 5 would provide more definitive evidence; unfortunately, co‑staining for Atf4 and Ddit3 on those tissue sections didn’t work well.

      Author response image 3.

      Subcellular localization of Atf4 and Chop in primary hepatocytes following APAP treatment. (A) Immunofluorescence staining of Atf4 and Chop in primary hepatocytes treated with 10 mM APAP for 6 hours or left untreated (UT). Nuclei were counterstained with DAPI. Scale bar as indicated. (B) Primary hepatocytes were treated with 0, 5, or 10 mM APAP for 6 hours. Cytoplasmic and nuclear fractions were isolated and analyzed by western blot. Lamin B1 and α-Tubulin were used as markers for the nucleus and cytoplasm, respectively

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus.

      We thank the reviewer for this helpful comment. To better demonstrate ATF4 nuclear localization, we have added enlarged insets of the original representative images in revised Figure 5A. These magnified views more clearly show nuclear ATF4 staining after APAP treatment, addressing the concern about extranuclear signal.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting the expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)?

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      (8) It is laudable that the authors tried to extend their findings to humans by using snRNA-seq data from a published study (line 391), but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion.

      We thank the reviewer for this clarification. The analysis mentioned in line 391 originally referred to spatial transcriptomics (ST) data from two ALF patients, not snRNA-seq. For the snRNA-seq dataset, we analyzed all 10 patients, but snRNA-seq lacks spatial resolution and cannot reliably assign zonal identity. We stipulate that snRNA-seq requires viable cells and thus likely excludes necrotic/peri-necrotic areas. Therefore, direct zonal comparison with our ST data was not possible. We have now clarified this in the revised manuscript (Revised manuscript, page 18, line 510-519).

      Reviewer #3 (Public Review):

      The main concern is that the overexpression of ATF4 and DDIT3 is causing reduced cell death and damage by APAP. This makes it harder to understand if these genes are truly increasing survival or if they are just reducing the injury caused by APAP. It may be better to perform overexpression immediately after, or at the same time as APAP delivery. Alternatively, loss-of-function experiments using AAV-shRNAs against these targets could be useful.

      We thank the reviewer for raising this important point. We agree that overexpression prior to APAP administration leaves open the question of whether the observed protection reflects true cytoprotection or simply reduced initiation of injury. To address this, we pursued loss‑of‑function approaches. Due to their very low basal expression, AAV‑shRNA‑mediated knockdown of endogenous Atf4 and Ddit3 proved inefficient. We therefore targeted Btg2, a downstream mediator of Ddit3 that inhibits proliferation. Knockdown of Btg2 resulted in a significant increase in APAP‑induced liver injury, as evidenced by elevated ALT levels and expanded necrotic areas (Revised Figure 6K-N). These results indicate that the ATF4‑DDIT3‑BTG2 axis limits hepatocellular damage, consistent with a protective role. We have clarified this point in the revised manuscript (page 15, line 407-414)

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Clarify how zones were defined when necrosis disrupted pericentral areas. Provide marker validation across time and whether necrotic spots are excluded or not from zonal analysis.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g., Alb, Mup20, Cyp2f2, Pck1, Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2, Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g., Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) Test whether the ISR-Btg2 program applies in other models; even targeted validation via qPCR and IF would be valuable.

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Proliferation quantification in liver sections in Figure 1: how to define the zones and why, at the basal level, there is a high proliferation rate in the mid zone? From Figure 1B-C, all three zones showed decreased hepatocyte proliferation, although the mid zone had a higher baseline. Will the mid-zone stand out by converting to the fold change of Ki-67+ hepatocytes decrease?

      We thank the reviewer for these insightful comments. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x<sup>2</sup> + z<sup>2</sup> – y<sup>2</sup>) / (2z<sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885). Consistent with previous reports (PMID: 33632817; PMID: 33632818), we observed a higher baseline proliferation rate in the mid zone, where approximately 70% of proliferating hepatocytes reside under basal conditions, compared to 10% in zone 1 and 20% in zone 3. However, when analyzing the fold change in Ki-67+ hepatocytes, only Mid and PC region showed significant difference in Ki-67+ hepatocytes, other zones showed no significant differences (as shown in the fold-change results in Author response image 1), indicating that the mid zone does not stand out in the fold change analysis. See Author response image 1.

      (4) The authors need to strengthen the causal chain with rescue experiments, e.g., Atf4/Chop overexpression and Btg2 knockdown. Link proliferation suppression to survival/ALT directly.

      We thank the reviewer for these constructive comments. Besides existing data from Figure 5 (Atf4 overexpression), we included Btg2 knockdown data in the revised Figure. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup> hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N) (Revised manuscript, page 14, lines 399–406).

      (5) Transduction efficiency, distribution, and expression levels via the AAV overexpression need to be quantified. Key CYP genes in the APAP metabolic pathway need to be assessed to exclude confounds.

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (6) The authors claim that the requirement of the Atf4/Chop at the early stage of APAP injury protects hepatocytes from proliferation for survival. What is the consequence if we remove the protective mechanism?

      We thank the reviewer for this insightful question. In our model, early induction of Atf4 and Chop functions as a cell survival checkpoint. Removal of this protective mechanism is predicted to result in two deleterious outcomes: (1) Acute exacerbation of necrosis due to the inability of hepatocytes to manage stress-induced bioenergetic demands, and (2) Impaired long-term regeneration due to depletion of the surviving cell pool. We directly tested the acute prediction (< 24 h) in Author response image 4. We deleted Ddit3 specifically in hepatocytes. Initial attempts using AAV-CasRx failed due to negligible baseline Atf4/Chop expression in healthy liver, preventing effective knockdown. We therefore generated hepatocyte-specific Ddit3 knockout mice (Alb<sup>∆Ddit3</sup>; Author response image 4B). Immunohistochemistry confirmed APAP-induced Chop induction occurs primarily in the centrilobular zone by 6 h (Author response image 4A). Following a two-dose APAP regimen (Author response image 4C), Alb<sup>∆Ddit3</sup> mice displayed significantly larger areas of centrilobular necrosis compared to Ddit3<sup>fl/fl</sup> controls (Author response image 4D; **p < 0.01). Thus, hepatocyte-intrinsic Chop limits acute APAP injury, consistent with its proposed early protective role.

      Author response image 4.

      Hepatocyte-specific deletion of Ddit3 exacerbates APAP-induced liver injury. (A) Immunohistochemical staining of Chop in liver sections at 0,3 and 6 h post-APAP. Red arrows indicate Chop-positive hepatocytes. Scale bar = 50μm. Quantification of zonal distribution of Chop-positive cells in liver sections at 6 h post-APAP is conducted . The statistic is the percentage of Chop-positive hepatocytes in each layer over the total number of Chop-positive hepatocytes. n=3 mice. (B)The construction, genotyping strategy and genotyping results of Alb<sup>∆Ddit3</sup> mice. P: positive control; WT: Wild-type; Neg: Blank control(ddH<sub>2</sub>O). (C) Schematic figure illustrating the experimental strategy for the administration of two doses of APAP to Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice. (D) H&E staining showing liver morphology from Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice at 6 h post-second dose of APAP. Injured area is outlined by black dashed lines. Scale bars = 200 μm. The percentage of injury area is quantified. n = 3- 4 mice/group. Data are represented as means ± SD; *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001; ns, not significant.

      (7) Is there any human relevance to the sensitivity of APAP injury regarding the Atf4/Chop axis?

      We thank the reviewer for this insightful comment. During our study, we analyzed a spatial transcriptomics dataset from APAP patients. In one of two analyzed patients, mid-zone hepatocytes exhibited transcriptional signatures remarkably consistent with our murine findings, including: (1) upregulation of Atf4-Chop pathways, and (2) downregulation of cell proliferation genes (Author response image 5). This suggests that this axis may also be involved in the response to APAP injury in humans. However, given the limited sample size, definitive conclusions cannot be drawn at this stage. We have now included this point in the Discussion section (Revised manuscript, page 18, line 510-519).

      Author response image 5.

      Spatial transcriptomics (GSE223561) reveals zonal gene expression changes in APAP patients. Heatmap of ISR, cell death, and cell cycle gene expression across zonal regions in healthy versus APAP‑treated human livers. 

      (8) Several IHC stainings have a weak signal and need inserts to zoom in for a clear view of the positive signals. Figure 5A, E, G, and Figure 6D, F.

      We thank the reviewer for this observation. We agree that the immunostaining signals for several target genes are relatively weak, which reflects their low endogenous expression levels. To address this, we have included higher-magnification insets in the indicated panels (Revised Figure 5A, E, G and Figure 6D, F) to show the positive signals.

      Reviewer #2 (Recommendations for the authors):

      (1) What is the functional classification of DEG in F2A based on? GO terms?

      We thank the reviewer for this constructive question. The functional classification of differentially expressed genes (DEGs) in F2A is based on Gene Ontology (GO) terms. For each DEG, we retrieved its associated GO annotations across the three main categories (biological process, cellular component, molecular function). In cases where a gene was assigned multiple GO terms, we prioritized the most representative or significantly enriched term for functional interpretation. This clarification has been incorporated into the revised figure legend and the according GO number has been included in the figure.

      (3) The rationale for focusing on CHOP is not clear because Ddit3 is not shown in the spatial transcriptomics in F2A and is not significant in F2B, contradicting what is stated in line 206.

      We thank the reviewer for raising this important point. We apologize that Ddit3 was missing from the original figure. In the revised manuscript, we have included an updated version of Figure 2A, which now shows that Ddit3 is indeed one of the differentially expressed genes (DEGs) in the Mid zone at both 3 and 6 hours post-APAP. We agree with the reviewer that, as shown in Figure S1G (previous Figure 2B), Ddit3 did not reach statistical significance, due to its relatively low expression level in that analysis. Nevertheless, when we examined transcription factor (TF) activity in the Mid zone during early AILI, Ddit3 and Atf3 ranked as the top two most highly expressed TFs among the top ten with the highest activity, whereas Atf4 ranked seventh (Revised Figure 4B and Figure S3B). Given that Ddit3 frequently co-worked with Atf4 and that the Atf4–Ddit3 axis plays a well-established role in cellular stress adaptation, we considered this pathway to be biologically relevant and worthy of further investigation.

      (3) The term "redistribution" used in line 197 to describe the expression of Cyp2e1 and other Cyps in the midlobular zone seems inappropriate, considering that they just continue to be expressed there, whereas pericentral hepatocytes are dying in F3B; the same applies to "Gene Expression Shift" in F3H.

      We thank the reviewer for this important clarification. We have revised the text (Revised manuscript, page 9, line 234-236) to state that selective loss of Cyp‑expressing pericentral hepatocytes leads to the mid‑zone becoming the primary site of residual Cyp activity. The figure label has been changed from “Gene Expression Shift” to “Peri‑necrotic Cyp retention” and the legend now explicitly notes that this is an apparent zonal shift due to necrosis, not active redistribution.

      Reviewer #3 (Recommendations for the authors):

      (1) Please do not use abbreviations like AILI. This makes the paper more difficult to read.

      We thank the reviewer for pointing this out. We have replaced AILI with the full term “APAP-induced liver injury” to ensure easiness for readers.

      (2) It will be important to clarify how pericentral, mid, and periportal were defined. In Figure 1, it appears that some of the pericentral hepatocytes that are Ki67 positive are quite mid-zonal. It would be important to have rigorous definitions for the location determination.

      We thank the reviewer for this constructive comment. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x <sup>2</sup> + z <sup>2</sup> – y <sup>2</sup>) / (2z <sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1:

      I recommend that the title be changed to not focus on sex differences to avoid misunderstanding.

      We thank Reviewer #1 for this suggestion and agree that the original title could create a misleading impression. We have updated the title from "Sex-specific behavioral and thalamo-accumbal circuit adaptations after oxycodone abstinence" to "Thalamo-accumbal circuit adaptations following extended oxycodone abstinence" to more accurately reflect the scope of the findings.

      The authors should also address the lack of difference physiologically compared to the behavior as a caveat more clearly in the discussion.

      We thank the reviewer for this important suggestion. We have revised the Discussion to explicitly address this dissociation. Specifically, we added the following to the PVT-NAcSh synaptic strength section: " The absence of sex differences in PVT-NAcSh synaptic measures suggests that this pathway, as characterized here, represents a shared neuro-adaptation to prolonged oxycodone abstinence rather than a substrate for the heightened relapse vulnerability observed in females. The mechanisms driving sex-specific relapse likely involve additional circuit elements, such as sex hormone-dependent modulation, upstream inputs, or cell-type specific plasticity." This point is also summarized in the abstract.

      Reviewer #2:

      A major weakness of this study is the disconnect between the behavioral and neurophysiological data reported. While a striking sex difference in relapse-like behavior is observed, there are no statistically significant sex differences in any of the neurophysiological data reported. Moreover, without an experiment to functionally test the role of the PVT-NAc projection in relapse-like behavior following prolonged oxycodone, these two arms of the study seem divorced.

      We respectfully disagree with the characterization that the behavioral and neurophysiological data are "divorced." The two arms of the study converge on a consistent and meaningful finding: PVT-NAcSh synaptic strength increases specifically after prolonged abstinence, this is the same time point at which enhanced cue-induced relapse is observed in both sexes. The absence of sex differences in synaptic measures does not weaken this convergence; it refines it by suggesting that circuit-level potentiation is a shared neuro-adaptation, while the sex-specific behavioral phenotype likely reflects additional modulatory mechanisms acting on this shared substrate. We have revised the Discussion to explicitly address this dissociation, as noted in our response to Reviewer #1 above. We acknowledge that functional manipulation of the PVT-NAcSh circuit would be required to establish causality, and we state this clearly in the manuscript.

      In the introduction the authors state they aim to test the hypothesis that increased synaptic strength in PVTNAcSh projections are necessary for drug-seeking. This study does not include the required experiments to test this hypothesis.

      We have revised the relevant section in the Introduction to accurately reflect the scope of our study: " We aimed to determine whether synaptic strength in PVT-NAcSh projections is affected following oxycodone abstinence and whether such changes are associated with cue-induced relapse and drug-seeking. Additionally, we examined whether there are sex-specific differences in either cue-induced relapse or PVT-NAcSh synaptic transmission after either 1 (acute) or 14 (prolonged) days of forced abstinence. Our results demonstrate that sex-specific enhancement in cue-induced relapse emerges after prolonged abstinence but not during acute abstinence from oxycodone self-administration. Although both males and females show increased cue-induced relapse after prolonged abstinence, females exhibited a greater relapse rate compared to males. Both sexes showed similar increases in PVT-NAcSh synaptic strength after prolonged abstinence, while synaptic strength was not altered after acute abstinence compared to saline controls. Together, these findings reveal a time-dependent increase in PVT-NAcSh synaptic strength and a sex-specific effect of prolonged abstinence on cue-induced relapse, while synaptic enhancements after prolonged abstinence were not sex-specific." This revision avoids implying a necessary or causal role for the circuit, which we did not test.

      Reviewer #3:

      The PVT-NAcSh synaptic strengthening after prolonged abstinence is statistically indistinguishable between sexes, while females but not males show a time-dependent escalation in oxycodone seeking from 1 to 14 days of abstinence. The Discussion proposes hormonal modulation or differences in upstream inputs as possible explanations, but none of these are tested and the gap is left unresolved.

      We agree that the mechanistic basis of the behavioral sex difference remains an open question that the current study does not resolve. As noted in our response to Reviewer #1, we have revised the Discussion to explicitly acknowledge this dissociation and to clarify that PVT-NAcSh synaptic strengthening represents a shared neuro-adaptation rather than a mechanism specific to the female behavioral phenotype. We maintain that identifying this dissociation is itself a scientifically meaningful finding.

      The intrinsic excitability recordings come from NAcSh MSNs with no confirmation that those neurons receive direct PVT input, which was raised in the original review, acknowledged in the revision, and not experimentally addressed.

      We have added the following clarification to the excitability section of the Discussion: “It should be noted that the intrinsic excitability recordings were designed to characterize general properties of NAcSh MSNs following oxycodone abstinence, independent of their synaptic inputs. As such, the excitability data should be interpreted as reflecting changes in the NAcSh MSN population broadly rather than in PVT-connected neurons specifically. The standing theory suggests that MSN excitability decreases as a homeostatic response to increased glutamatergic input [23,43,58]. Our data do not support a compensatory decrease in excitability in either sex at either abstinence time point”. These recordings were never intended to be circuit-specific; the experiment was designed to characterize NAcSh MSN excitability at the population level, which is a valid and informative question in its own right.

      The male prolonged-abstinence excitability trend has approximately 20% statistical power and is non-significant, yet the Discussion interprets it as a potential neuro-adaptation that could facilitate signal flow through the PVT-NAcSh circuit and contribute to relapse, which goes well beyond what the data support.

      We have revised the relevant Discussion text to ensure the male excitability trend is interpreted appropriately. The revised text now reads: "In males, a non-significant trend toward increased excitability was observed after prolonged abstinence, with a large effect size (Cohen's d = 1.18); however, given that this group was substantially underpowered (approximately 20% power), this finding should be interpreted with caution and cannot be taken as evidence of a neuro-adaptation. Whether this trend, if confirmed in future studies with larger cohorts, reflects a broader MSN population response or is specific to PVT-connected neurons remains an open and interesting question." The speculative mechanistic interpretation previously present in this section has been removed.

      The failure to distinguish between D1 and D2 MSNs remains a significant limitation given that cell-type specific plasticity at PVT-NAc synapses has been shown to be directly relevant to opioid seeking in prior work.

      We agree that distinguishing between D1 and D2 MSNs would provide important mechanistic insight, and we acknowledge this explicitly as a limitation and a future direction in the Discussion. The use of transgenic Cre rat lines for cell-type-specific recording in a self-administration model requires significant additional infrastructure and was beyond the scope of the present study. This is precisely the direction our laboratory is currently pursuing, and the present findings provide empirical motivation for those experiments.

      The Conclusion builds a mechanistic framework around D2 MSNs, PV interneurons, and D1 MSNs that is drawn from studies using different drugs or experimental designs, and none of these cell-type-specific mechanisms are tested in the present experiments.

      We thank the reviewer for this important critique. We have revised the opening of the Conclusion to clarify that the cell-type-specific framework is grounded in prior literature and represents a hypothesis for future investigation rather than a conclusion drawn from the present data. The revised text now reads: " When considered alongside prior work, our findings highlight the need to examine the anatomical and cell-type specific organization of PVT inputs to the NAcSh in the context of opioid relapse. Based on existing literature, PVT projections onto D2 MSNs and PV interneurons may contribute to relapse vulnerability, while adaptations involving D1 MSNs may underlie incubation of craving, though these mechanisms remain to be directly tested in the oxycodone self-administration model used here". We believe this framing accurately represents the relationship between our findings and the broader literature without overstating what the present data demonstrates.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      We thank the reviewer for this positive assessment of the study.

      Weaknesses:

      In Figure 3, the authors used C-Laurdan imaging to assess membrane fluidity and showed that knockdown of SPHK2, the human ortholog of sphk-1, led to increased membrane rigidity. However, the authors did not co-stain with a lysosomal marker, making it unclear whether the observed effect is specific to lysosomal membranes or reflects general membrane changes. Co-staining with LysoTracker or applying segmentation masks to isolate lysosomal signals would significantly improve interpretation.

      We agree with the reviewer that it is important to isolate lysosomal signals for interpreting the C-Laurdan data. We therefore repeated and extended the C-Laurdan experiments in combination with LysoTracker staining and selectively analyzed LysoTracker-positive regions. These analyses showed pronounced increases in GP values in LysoTracker-positive vesicles after SPHK2 knockdown, supporting the conclusion that SPHK2 depletion increases endolysosomal membrane rigidity. We also performed LysoTracker-based analysis in the tau-fibril and fatty-acid experiments to better assess lysosome-associated membrane properties (see new and updated Figures 3B-E, Figures 5A and B, Figures S3A, B, F, G, Figures S4A-F, and Figures S5A-D for details). The respective Results sections have been revised accordingly.

      Line 173 states that Lipofectamine 2000 increases membrane fluidity based on GP index changes, but this is incorrect. A higher GP index indicates increased membrane order (i.e., reduced fluidity), so the statement should be revised. Additionally, Lipofectamine 2000 can itself alter membrane rigidity, posing a risk of false-positive interpretations. To confirm the role of SPHK2 in this phenotype, the authors should use a CRISPR/Cas9 knockout model instead of relying solely on siRNA transfection, which may be confounded by the delivery reagent. Without lysosomal co-staining and SPHK2 KO validation, the authors cannot conclusively claim that SPHK2 loss affects endolysosomal membrane integrity.

      We thank the reviewer for pointing out the incorrect wording regarding Lipofectamine. A higher GP index indicates increased membrane order/rigidity, not increased fluidity. Since the revised main figure now includes the SH-SY5Y data (Figure 3A-D), in which Lipofectamine alone did not significantly alter GP values (see Figure S3A, B), we removed the misleading statement from the Results.

      We also agree that Lipofectamine can affect membrane properties and therefore needs to be carefully controlled. In all siRNA-mediated experiments, SPHK2 siRNA was compared to a matched control siRNA condition exposed to the same transfection reagent. We state in the Methods that cells were transfected with either SPHK2 or scrambled control siRNA and that the medium was exchanged after 6 h “to minimize lipofectamine impact on the endolysosomal system”. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      Importantly, we have now added lysosomal co-staining to address the reviewer’s concern about compartment specificity. This analysis showed that SPHK2 KD resulted in a pronounced increase in GP values within LysoTracker-positive compartments, demonstrating increased membrane rigidity at lysosomes. Thus, the revised data support the conclusion that SPHK2 KD increases lysosome-associated membrane rigidity, rather than only causing nonspecific effects on other cellular membranes.

      We also clarified the relationship between the current siRNA-based assay and our previous CRISPR inhibition-based analysis. In the revised Results, we now write: “While SPHK2 KD alone significantly increased galectin puncta above the matched control, its effect was more modest than in our previous CRISPR inhibition-based analysis [20]. This difference likely stems from the earlier readout required for the combined siRNA/tau fibril assay, when transient Lipofectamine-associated effects still increased the control background.”

      This addresses why the SPHK2 KD effect appears smaller in the current siRNA/tau-fibril assay than in our previous CRISPR inhibition-based analysis. The previous study, which is now peer-reviewed and published in the journal Autophagy, used a CRISPR inhibition-based strategy to reduce SPHK2 levels, which resulted in a highly significant increase in sfGFP-LGALS3 foci formation compared to the control [1]. Thus, the SPHK2 phenotype is not supported solely by the current siRNA experiment.

      In addition, we sought to genetically validate the RNAi phenotypes using mutant strains. However, mutant strains were not available for all sphingolipid metabolism hits analyzed in this study. We therefore used the sphk-1 mutant strain available at CGC (CZ24969; sphk-1(ju831)) to validate one of the key SL metabolism hits independently of RNAi. The revised manuscript states: “As genetic validation independent of RNAi, we tested an available sphk-1 mutant strain, which also showed a robust increase in hypodermal sfGFP::LGALS3 foci (Figure S1A).” This result supports the conclusion that genetic perturbation of sphingosine kinase activity compromises endolysosomal integrity in vivo.

      Together, the revised manuscript addresses the reviewer’s concerns by correcting the GP interpretation, controlling the siRNA experiments against matched Lipofectamine-treated controls, adding LysoTracker-based lysosome-associated C-Laurdan analysis, relating the current siRNA data to our previous CRISPR inhibition-based analysis, and providing genetic validation for the available sphk-1 mutant.

      The section titled "Fibrillar tau increases membrane rigidity and exacerbates endolysosomal damage" (lines 177-215) requires substantial revision. The narrative jumps abruptly between worms and cell models, making it hard to follow the logic. The use of the F3ΔK281::mCherry strain is introduced without explanation or context. It is unclear whether this strain is relevant to lysosomal membrane rupture, as no reference or justification is provided. The authors should clarify whether this reporter is intended to detect lysosomal membrane permeabilization (LMP). If so, it would be more appropriate to use established LMP reporters, such as lysosome-targeted fluorescent sensors, galectin-based reporters, or dextran leakage assays. Based on the current data in Figure 3G, it is difficult to draw firm conclusions regarding membrane rupture levels.

      We agree that this section required clarification, and we have substantially revised the Results to improve the logic and separation between model systems.

      First, we now introduce the C. elegans reporter strain earlier in the manuscript, in the first Results section. In the revised text, we explain both the tau construct and the actual lysosomal damage reporter: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We also clarify the principle of the Galectin reporter: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture.” Thus, F3ΔK281::mCherry is not the reporter for lysosomal membrane permeabilization; the membrane-damage readout is sfGFP::LGALS3 puncta formation.

      Second, we reorganized the manuscript to separate the human cell experiments from the C. elegans experiments more clearly. The revised section “Fibrillar tau and SPHK2 KD act in concert to exacerbate endolysosomal damage and seeded tau aggregation” now focuses on human cell data. The C. elegans experiments are now presented in a separate section, “Tau transmission sensitizes endolysosomal membranes to sphingolipid perturbations in vivo.” We believe that this revised structure now clearly distinguishes the role of the Galectin reporter from the tau transmission model, separates the human cell and C. elegans data, and avoids the abrupt transitions between model systems noted by the reviewer.

      To support the conclusion that sphingolipid metabolism gene knockdown alters membrane properties, the study would benefit from direct lipidomic analysis. Measuring changes in sphingolipid profiles in both C. elegans and cell models would provide biochemical evidence for the proposed disruption of lipid homeostasis. Given the availability of lipidomics platforms, this type of analysis should be feasible in both worms and human cells and would significantly strengthen the mechanistic claims regarding membrane fluidity and integrity.

      Because we did not perform lipidomics in the present study, we have revised the wording throughout the manuscript to avoid implying that we directly measured lipid composition. Instead, we now refer to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our experimental interventions.

      We agree that lipidomic analyses will be important in future studies to define how perturbation of sphingolipid metabolism changes lipid composition in C. elegans and human cells. However, lipidomics itself would not directly establish which lipid changes causally drive the membrane rigidification observed in our study. Membrane fluidity is a biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Thus, even if lipidomics identified changes in sphingolipid profiles, these changes could not be directly translated into a predictable effect on membrane fluidity without additional biophysical validation, using Laurdan dye imaging or FRAP. Moreover, whole-cell or whole-animal lipidomics would not resolve whether the relevant lipid changes occur specifically at endolysosomal membranes, which are the focus of our study.

      We have now clarified this point in the Discussion. Specifically, we state that “even detailed lipidomics would not by itself identify which lipid changes are responsible for the observed membrane rigidification” and that future lysosome-enriched or organelle-specific lipidomic approaches should be combined with direct manipulation of candidate lipid species, followed by measurements of membrane fluidity and rupture, to determine which lipid changes causally contribute to endolysosomal membrane rigidification. In the present study, we therefore focused on direct quantitative biophysical readouts of membrane properties in C. elegans. We used FRAP of the lysosomal membrane protein LAAT-1::mCherry to assess lateral mobility within lysosomal membranes and showed that knockdown of sphingolipid-metabolism genes increased the time to half-maximal recovery, indicating reduced lysosomal membrane fluidity. Notably, knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased membrane rigidity. This makes it unlikely that the observed rigidification is caused by accumulation or depletion of a single shared lipid species. Rather, perturbations at different steps of sphingolipid metabolism may lead to distinct lipidomic changes that nevertheless converge on a common biophysical outcome: reduced endolysosomal membrane fluidity. In parallel, we used C-Laurdan imaging to quantify membrane order and found that SPHK2 knockdown in SH-SY5Y human neuroblastoma cells increased GP values, consistent with increased membrane rigidity. Two-channel thresholding of LysoTracker-positive compartments further showed that SPHK2 knockdown increased GP values in lysosome-associated regions.

      Thus, although lipidomics will be valuable to define the underlying lipid changes in future work, the current data already provide convergent quantitative evidence from independent membrane-fluidity readouts across C. elegans and human cell models. This cross-model consistency strengthens the robustness and reproducibility of the central conclusion that perturbation of sphingolipid metabolism alters endolysosomal membrane properties and promotes membrane rupture.

      The conclusions of the study rely heavily on imaging-based assays, including FRAP, C-Laurdan, and fluorescence microscopy. While these approaches provide valuable spatial and qualitative insights, they are inherently indirect and subject to interpretive limitations. To strengthen the mechanistic claims, the authors should incorporate additional biochemical or quantitative approaches. For example, lipidomics would allow direct measurement of membrane lipid composition changes, and western blotting or quantitative proteomics could assess levels of membrane-associated proteins involved in endolysosomal function or stress responses. Including such data would significantly improve the robustness and reproducibility of the study's conclusions.

      We agree that lipidomic and proteomic analyses will be important in future studies to define which sphingolipid species and/or membrane-associated proteins contribute to the observed rigidification of endolysosomal membranes. In response to this point, we have revised the wording throughout the manuscript to more precisely distinguish our experimental interventions from inferred changes in lipid composition. Because we did not directly measure lipid composition in the present study, we now refer more specifically to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our data, rather than implying that global sphingolipid homeostasis was directly quantified. We retain “sphingolipid imbalance” only in interpretive or model-based statements where appropriate.

      However, we respectfully disagree that the current data are only qualitative. FRAP and C-Laurdan GP imaging are established quantitative biophysical approaches: FRAP provides quantitative parameters such as the time to half-maximal recovery and the mobile fraction, whereas C-Laurdan GP provides a ratiometric measurement of membrane lipid order and packing. Similarly, the Galectin puncta assay is an established quantitative readout of lysosomal membrane permeabilization. Thus, while these approaches are imaging-based, they provide quantitative readouts of membrane mobility, membrane order, and membrane rupture, respectively.

      We also note that lipidomic and proteomic profiling, although valuable, would not by itself establish which lipid or protein changes causally drive the membrane rigidification observed in our study. Membrane fluidity is an emergent biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Therefore, an increase or decrease in a given lipid or protein species cannot be directly translated into a predictable change in membrane fluidity without additional biophysical validation. This point is further supported by our observation that knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased endolysosomal membrane rigidity. These perturbations would be expected to affect lipid composition in different, possibly even opposing, ways, making it unlikely that the shared rigidification phenotype is caused by accumulation or depletion of one single lipid species. Rather, distinct lipidomic changes may converge on a common biophysical outcome: reduced endolysosomal membrane fluidity.

      We have clarified this point in the Discussion and now state that future lysosome-enriched or organelle-specific lipidomic/proteomic approaches should be combined with direct manipulation of candidate lipid or protein species, followed by measurements of membrane fluidity and rupture, to determine which changes causally contribute to endolysosomal membrane rigidification. Such experiments would address the distinct question of which molecular components mediate the effect. By contrast, the central aim of the present study was to test whether genetic perturbation of enzymes involved in sphingolipid metabolism alters membrane fluidity and thereby promotes endolysosomal rupture and tau seeding.

      For this question, direct biophysical measurements of membrane fluidity and quantitative readouts of membrane rupture are the most relevant assays. We therefore used complementary quantitative approaches in two distinct model systems: FRAP of the lysosomal membrane protein LAAT-1::mCherry in C. elegans and C-Laurdan GP imaging in human cells. The fact that perturbing sphingolipid metabolism reduced endolysosomal membrane fluidity in C. elegans and increased lysosome-associated membrane rigidity in human cells supports the robustness and reproducibility of the central conclusion across independent model systems. In the revised manuscript, we further strengthened the human-cell data by adding SH-SY5Y neuroblastoma cells as a neuronal-like model and by combining C-Laurdan imaging with LysoTracker-based analysis to assess lysosome-associated membrane properties.

      To further address causality, we manipulated membrane fluidity independently of sphingolipid metabolism enzymes using fatty acid supplementation. Increasing membrane rigidity with PA exacerbated tau-induced endolysosomal rupture and seeded aggregation, whereas increasing membrane fluidity with ALA reduced tau-induced membrane rigidification, endolysosomal rupture, and seeded aggregation. Thus, the revised manuscript combines genetic perturbation of sphingolipid metabolism, quantitative membrane-fluidity measurements, whole-cell and lysosome-associated C-Laurdan analysis, and Galectin-based rupture assays across complementary models.

      Regarding lysosomal function, we agree that functional readouts are informative, but lysosomal membrane rupture and global lysosomal degradative capacity are related but not identical readouts. This distinction is supported by Yong et al., who reported that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal functions [2]. Accordingly, the absence of overt defects in general lysosomal activity would not necessarily exclude membrane damage.

      The human cell experiments were performed exclusively in HEK293T cells, which are not physiologically relevant for modeling Alzheimer's disease or lysosomal function in neurons. Given that the study aims to draw conclusions related to tau aggregation and lysosomal membrane integrity, the use of a more disease relevant cellular model is essential. There are several established AD-relevant cell models, including iPSCderived neurons, neuroblastoma lines expressing tau, or microglial models, which would better reflect the cellular context of tauopathies. Validation of key findings in at least one of these systems would substantially enhance the biological relevance and translational impact of the study.

      We have expanded and clarified the human cell data in the revised manuscript. Specifically, we now include SH-SY5Y human neuroblastoma cells for key C-Laurdan experiments assessing membrane rigidity after SPHK2 knockdown. We also show that recombinant tau fibrils increased membrane rigidity in SHSY5Y and HEK293T cells, including in LysoTracker-positive compartments.

      Importantly, the HEK293T cells are used for specific, established quantitative assays rather than as a model of neuronal toxicity. In particular, HEK293T sfGFP-LGALS3 cells are used to quantify galectin puncta formation as a readout of endolysosomal rupture, and HEK tau-Venus biosensor cells are used to quantify seeded tau aggregation. Thus, SH-SY5Y cells and HEK293T cells are used for complementary purposes: SHSY5Y cells provide a more neuronal-like human cell context for membrane-rigidity measurements, whereas HEK293T reporter/biosensor cells provide robust quantitative assays for galectin puncta formation and tau seeding.

      In addition, tau-associated neuronal dysfunction and toxicity were assessed in vivo, in functional C. elegans touch receptor neurons. In the revised manuscript, we show that ALA supplementation mitigated the age-dependent touch-response deficit and reduced neurotoxicity in animals expressing F3ΔK281::mCherry in touch receptor neurons. We have also revised the wording throughout the manuscript to avoid implying that HEK293T cells are used to model neuronal toxicity.

      Finally, the relevance of these hits to human neuronal tau seeding is also supported by our previous study, in which conserved hits from the C. elegans screen, including sphingosine kinase perturbation, were validated in human iPSC-derived neurons for their effect on seeded tau aggregation [1]. We now cite this published study where appropriate. Together, the revised manuscript combines neuronal-like human SH-SY5Y cells, established HEK293T tau-seeding and galectin reporter assays, in vivo neuronal readouts in C. elegans, and prior validation in human iPSC-derived neurons, thereby strengthening the biological relevance of the conclusions while using each model for the assay in which it is most informative.

      The authors reported that PUFA supplementation rescues neurotoxic phenotypes by increasing membrane fluidity. However, the data supporting this claim rely entirely on confocal imaging, shown in both the main and supplemental figures. To substantiate the mechanistic link between PUFA treatment and improved lysosomal membrane properties, the authors should include functional assays demonstrating that PUFAs are indeed incorporated into lysosomal membranes. Additionally, lipidomics analysis would be valuable to identify which lipid species are altered upon supplementation and correlate these changes with the observed phenotypic rescue. Furthermore, the conclusion that PUFAs rescue "neurotoxic phenotypes" is not appropriate based on data derived solely from HEK293T cells, which are not neuronal. To make claims about tau-related neurotoxicity, the authors should validate their findings in a more relevant neuronal model, such as SH-SY5Y neuroblastoma cells expressing tau or iPSC-derived neurons. This would better reflect the cellular environment of Alzheimer's disease and provide stronger support for the proposed therapeutic potential of PUFA supplementation.

      We agree that PUFA supplementation can have effects beyond membrane fluidity and that our data do not directly demonstrate incorporation of ALA into lysosomal membranes. We have therefore revised the Discussion to explicitly acknowledge this limitation. In the revised text, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” At the same time, we note that the opposing effects of PA and ALA, together with the sphingolipid-metabolism knockdown data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models. To strengthen the link between ALA and lysosome-associated membrane properties, we combined CLaurdan imaging with LysoTracker-based analysis. In the revised Results, we show that ALA prevented tau-induced membrane rigidification and that LysoTracker-based analysis indicated effects on lysosome-associated membrane properties. ALA also reduced tau-induced endolysosomal rupture and seeded aggregation in human cell models.

      Regarding lipidomics, we refer to our response above and to the revised Discussion. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such data would not by itself establish how these changes affect membrane fluidity without additional biophysical validation.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is not based on HEK293T cells. HEK293T cells were used for established quantitative assays of Galectin puncta formation and seeded tau aggregation. The neurotoxicity experiments were performed in vivo in C. elegans touch receptor neurons, where ALA supplementation reduced galectin foci formation, mitigated age-dependent touch-response deficit and reduced neuronal toxicity.

      While the authors demonstrate that ALA supplementation mitigates neurotoxicity in C. elegans expressing aggregated tau (F3ΔK281::mCherry), the current data are not sufficient to conclude that ALA directly rescues tau aggregation toxicity via a lysosome-specific mechanism. It remains unclear how lipid composition is altered upon ALA treatment and whether these changes correlate with functional improvement of lysosomal pathways. The manuscript does not provide mechanistic insight into how ALA enhances lysosomal health or attenuates endolysosomal damage. Moreover, supplementation with PUFAs like ALA can activate a wide range of cellular processes beyond lysosomal function, including alterations in membrane fluidity, signaling cascades, and oxidative stress responses. The authors should clarify how they distinguish the lysosome-related effects from these alternative pathways. For example, did they observe specific lysosomal markers or structural improvements in lysosomes upon ALA treatment?

      Additional data or controls would be necessary to support a lysosome-specific protective mechanism and to exclude the involvement of other PUFA-responsive pathways in the observed phenotypes.

      We agree that our data do not prove that ALA acts exclusively through a lysosome-specific mechanism or that ALA is directly incorporated into lysosomal membranes. We have therefore revised the manuscript to avoid this interpretation and explicitly acknowledge alternative PUFA-responsive pathways. In the revised Discussion, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” We further conclude more cautiously that the opposing effects of PA and ALA, together with the sphingolipid metabolism perturbation data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models.

      To strengthen the lysosome-related aspect of the mechanism, we added LysoTracker-based analysis to the C-Laurdan experiments. In the revised Results, ALA prevented tau-induced membrane rigidification, and LysoTracker-based analysis indicated that ALA also affected lysosome-associated membrane properties. ALA further reduced tau-induced Galectin puncta formation and seeded tau aggregation in human cell models. These data support an effect of ALA on lysosome-associated membrane order and rupture, while not excluding additional effects through lipid signaling, oxidative stress responses, or other PUFA-responsive pathways.

      Regarding lipid composition, we refer to the revised Discussion and our response above. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such analyses would need to be organelle-specific and combined with biophysical validation to determine how candidate lipid changes affect membrane fluidity and rupture.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is based on the C. elegans experiments, not on HEK293T cells. HEK293T cells were used for quantitative Galectin puncta and tau-seeding assays, whereas neuronal dysfunction and toxicity were assessed in vivo in touch receptor neurons in C. elegans. In the revised Results, we state that ALA supplementation mitigated galectin foci formation, age-dependent touch-response deficit and reduced neuronal toxicity in animals expressing F3ΔK281::mCherry in these neurons.

      Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans, decrease the fluidity of endolysosomal membranes and promote their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Overall, the conclusions of this paper are supported by the data, with a few aspects requiring further clarification and elaboration.

      (1) A reference to Figure S2E-G, which shows that KD of SL biosynthesis genes do not affect the plasma membrane, is missing from the main text.

      We thank the reviewer for pointing this out. We have added the reference to the respective figures in the main text when discussing the plasma membrane FRAP experiments.

      (2) In Figure 3C, lipofectamine alone shows that it increases membrane rigidity (increased GP values), not membrane fluidity.

      We thank the reviewer for pointing out this incorrect wording. A higher GP index indicates increased membrane order/rigidity, not increased membrane fluidity. Since the revised main figure now includes the SH-SY5Y data, in which Lipofectamine alone did not significantly alter GP values, we removed the misleading statement from the Results. Importantly, all siRNA-mediated knockdown experiments were compared to matched control siRNA conditions exposed to the same transfection reagent. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      (3) In Figure 3F, the EV cntl condition expressing F3:mCh tau should have increased LGALS3 foci compared to the mCh EV cntl according to Ref (20) and its Figure 2G (at least for Day 5 animals), which would be indicative of the tau spreading in hypodermal tissue. What C. elegans age was examined in Figure 3F? Can the authors provide evidence of the transmission of the F3:mCh tau from the touch receptor neurons to the hypodermis in the EV [similar to Figure 2C & D from Ref (20)] and compare it to the KDs? Otherwise, it seems that KD of SL genes impacts not only endolysosomal rupture but significantly affects tau accumulation/spreading as well (e.g., shown later in HEK cells, where SPHK2 KD increases the formation of tau-Venus foci).

      We thank the reviewer for raising this important point. The analysis referred to by the reviewer has now been moved to the revised C. elegans section and is presented as Figure 4A and B. The experiments were performed in the reporter strain used in our genome-wide screen In Ref (20), now published in Autophagy [1]. We clarified the purpose of the reporter strain and the relationship between tau transmission and the galectin puncta readout. In the revised manuscript, we now state: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We further clarify that sfGFP::LGALS3 puncta formation, not F3ΔK281::mCherry, is the readout of endolysosomal rupture: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture”.

      Furthermore, we now better explain that transmitted tau sensitizes endolysosomal membranes to additional perturbations rather than necessarily inducing a strong lysosomal rupture phenotype on its own. In the experiments shown in Figure 4A and B, we compare F3ΔK281::mCherry animals with matched mCherry-only control animals that also express sfGFP::LGALS3 in the hypodermis. We now state: “We compared animals expressing F3ΔK281::mCherry in touch receptor neurons, from where it is transmitted to the hypodermis, with matched controls expressing mCherry alone in the same neurons. In both strains, sfGFP::LGALS3 is expressed in the hypodermis to monitor endolysosomal membrane damage.” We have also clarified the age of the animals in the revised figure legends.

      To experimentally address whether the enhanced rupture phenotype could be explained by altered tau transmission, we quantified hypodermal F3ΔK281::mCherry levels after sphk-1 RNAi (new Figure 4C, D). Importantly, sphk-1 RNAi did not increase hypodermal F3ΔK281::mCherry levels, arguing that the enhanced rupture phenotype is not due to increased tau transmission. Moreover, C. elegans neurons are largely refractory to systemic RNAi under the conditions used here [3]. We have added this important information to the Discussion. Specifically, the revised manuscript states that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here,” supporting the interpretation that the RNAi treatments primarily affect endolysosomal integrity in the recipient tissue rather than neuronal tau expression itself.

      Finally, we would like to clarify that the increased tau-Venus foci in HEK cells should not be interpreted as a direct induction of tau aggregation by SPHK2 KD. Only upon addition of recombinant tau fibrils did SPHK2 KD significantly increase tau-Venus foci formation (Figure 3 H, I). This is consistent with the control experiments performed in human iPSCs in our previous study and supports our interpretation that perturbation of sphingolipid metabolism increases susceptibility to seeded tau aggregation by promoting endolysosomal rupture and tau seed escape, rather than by directly increasing tau aggregation or tau transmission.

      (4) Sphingolipids are essential membrane components and signaling molecules. Does KD of SL genes in C. elegans and the subsequent endolysosomal rupture cause any major, intermediate, or minor defects/phenotypes (in non-aggregation prone models, w/t.)?

      We agree that sphingolipids are essential membrane components and signaling molecules and that perturbing sphingolipid metabolism can have broader physiological consequences. In the revised manuscript, we address this point in two ways.

      First, we directly tested whether SL gene knockdown can induce endolysosomal rupture independently of aggregation-prone tau by using matched control animals expressing mCherry alone in touch receptor neurons while also expressing sfGFP::LGALS3 in the hypodermis. In these animals, knockdown of most SL-related hits resulted in nearly all animals displaying hypodermal sfGFP::LGALS3 foci, indicating that perturbation of SL metabolism can compromise endolysosomal integrity in the absence of transmitted F3ΔK281::mCherry (Figure 4A, B).

      Second, we have added a Discussion paragraph to place these findings into a broader physiological context. We now clarify that endolysosomal membrane rupture and global lysosomal function are related but not identical readouts. In support of this distinction, we discuss work showing that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal function [2]. Thus, membrane damage can occur even when general lysosomal activity is not overtly disrupted.

      We also discuss a recent study published during the revision of this manuscript that independently identified SPHK-1 as an important regulator of lysosomal integrity in C. elegans, showing that strong sphk1 loss-of-function causes lysosomal sphingosine accumulation, membrane rupture, impaired degradative function, cargo accumulation, developmental defects, and reduced lifespan [4].

      Importantly, while that study focused on a strong loss-of-function mutation in a single SL-metabolism gene, our data show that knockdown of multiple SL-metabolism genes, including genes involved in both SL biosynthesis and degradation, converges on reduced endolysosomal membrane fluidity and increased rupture. This suggests that the observed membrane rigidification and rupture are not specific to one mutant background but represent a broader consequence of perturbing SL metabolism at multiple points. A systematic characterization of all organismal phenotypes caused by each SL gene knockdown was beyond the scope of the present study. Therefore, the revised manuscript now makes clear that the study focuses on endolysosomal membrane fluidity and rupture because these membrane-level changes are directly linked to tau seed escape and seeded tau aggregation, while broader physiological consequences may vary depending on the strength and context of the perturbation.

      Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous (yet unpublished) work of an RNA screen using a new C. elegans Tau model. They then used cell culture and C. elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and use probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds, etc.

      Weaknesses:

      The main weakness is that this work builds on not-yet-peer-reviewed manuscript that established a new C. elegans Tau model and RNAi screen that aimed to identify genes involved in the propagation of Tau.

      This reviewer misses essential information of the C. elegans Tau strain (not included in the method section): e.g., promoter used for the expression, information on the used Tau variant, expression pattern, and aggregation, etc.

      We thank the reviewer for raising this point. The related study establishing the C. elegans tau transmission model and RNAi screen has now been peer-reviewed and published in Autophagy [1]. We now cite the published article throughout the revised manuscript instead of the previous preprint.

      We also agree that the current manuscript should be understandable without requiring the reader to consult the previous paper for the basic logic of the model. We therefore added a clearer introduction of the reporter strain in the Results. Specifically, we now explain that the strain expresses the aggregation-prone tau fragment F3ΔK281::mCherry in touch receptor neurons, that this tau fragment is transmitted to the hypodermis, and that endolysosomal membrane damage is monitored in the hypodermis using sfGFP::LGALS3. We further clarify that sfGFP::LGALS3 remains diffuse under steady-state conditions and forms puncta upon endolysosomal membrane damage, when luminal β-galactosides become exposed. Thus, the revised manuscript now provides the key information needed to understand the experimental system used here, while the published Autophagy paper is cited for the full characterization of the tau transmission model, expression pattern, aggregation properties, and original genome-wide RNAi screen.

      Throughout the study, I missed data on:

      (1) Effect of the knockdown on Tau expression, localisation (with lysosomal membrane?), aggregation, and proteotoxicity. The effect of the RNAi-mediated knockdown could also simply lead to a reduced expression of Tau that, in turn, leads to suppressed propagation.

      We agree that it is important to distinguish effects on tau expression/transmission from effects on endolysosomal membrane integrity. In the C. elegans experiments, F3ΔK281::mCherry is expressed in touch receptor neurons and transmitted to the hypodermis, where sfGFP::LGALS3 reports endolysosomal membrane damage. We now describe this more clearly in the revised Results.

      The RNAi treatments target genes involved in sphingolipid metabolism under systemic RNAi conditions. Because C. elegans neurons are largely refractory to systemic RNAi in the absence of sensitizing backgrounds [3], which we did not use, a direct RNAi-mediated reduction of neuronal F3ΔK281::mCherry expression is very unlikely. We have added this point to the Discussion, stating that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here.”

      Experimentally, we also tested whether sphingolipid perturbation alters transmitted tau levels (new Figure 4C, D). Specifically, we quantified hypodermal F3ΔK281::mCherry after sphk-1 RNAi and found no increase, arguing that the enhanced rupture phenotype is not due to increased tau transmission.

      Moreover, in the cell-based tau-Venus assay, SPHK2 knockdown alone did not induce detectable tau aggregation in the absence of exogenously added tau fibrils (Figure 3H, I). Only upon addition of recombinant tau fibrils did SPHK2 knockdown significantly increase tau-Venus foci formation. We now state this explicitly in the revised Results and conclude that disruption of sphingolipid metabolism is not sufficient on its own to initiate detectable tau aggregation under the conditions tested here but rather increases cellular susceptibility to seeded tau aggregation when tau fibrils are present.

      Together, these data argue against a direct effect of sphingolipid gene knockdown on tau expression or spontaneous tau aggregation. Instead, they support our interpretation that perturbation of sphingolipid metabolism compromises endolysosomal membrane integrity, thereby facilitating tau seed escape and seeded aggregation when tau seeds are present.

      (2) A quantification of RNAi knockdown is needed to judge the efficiency of the RNAi, in particular for the combinatorial RNAi experiments involving 2 and even 4 genes in parallel. Ideally, these analyses should be validated with mutants for these genes.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      Further:

      (3) Figure 4 H, I: Would Tau also aggregate in the absence of externally added Tau?

      No. In the tau-Venus biosensor cell line, SPHK2 knockdown alone did not increase tau-Venus foci formation (now Figure 3H, I). Tau-Venus foci increased only after addition of recombinant tau fibrils and were further enhanced by SPHK2 knockdown. We now state this explicitly in the Results.

      (4) How specific is the effect for Tau? It would help if the authors could assess other amyloid proteins.

      We agree that similar membrane-level mechanisms may apply to other amyloid assemblies. We have therefore added recent literature to the Discussion supporting the broader concept that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes. The revised manuscript states: “This interpretation is consistent with recent ultrastructural studies showing that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes.” We further clarify that “whether this mechanism is specific to tau or also applies to other amyloid assemblies remains to be determined.”

      Whether perturbation of SL metabolism similarly affects endolysosomal escape and seeded aggregation of other disease-associated amyloid proteins is an important question that we plan to address in future work. However, these experiments require additional disease-specific models, aggregation assays, and validation, and are therefore beyond the scope of the present revision.

      (5) The connection between sphingolipids and AD is not new. See He et al, 2010, Neurobiol. Aging + numerous publications and also not between Tau seeding and lysosomal rupture: Rose et al., PNAS 2024 (that has been cited by the authors).

      We agree and our manuscript does not aim to establish these associations as new. We state explicitly that alterations in sphingolipid metabolism have been reported in aging and AD, and that endolysosomal rupture is increasingly recognized as a critical step in tau seed escape and propagation.

      The novelty of our study lies in mechanistically connecting these two previously established areas. Specifically, we show that genetic perturbation of enzymes involved in sphingolipid metabolism reduces endolysosomal membrane fluidity, promotes membrane rupture, and thereby increases susceptibility to tau seed escape and seeded aggregation. We have revised the Introduction and Discussion to better emphasize this mechanistic contribution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Figure formatting and annotation need improvement. Panel letters throughout the figures should be in uppercase, and gene names in pathway diagrams should be italicized for consistency. Several scale bars are missing, including in Figures 1C, 2A, and 2H, and should be clearly indicated in the figures and legends. In Figure 1C, the age of the worms used in the assay is not specified. While the Methods section mentions "age-synchronized animals," the precise age at the time of imaging or experimentation is not stated. It would strengthen the study to explore whether membrane integrity phenotypes vary between young adults (day 1) and older adults (day 7 or 10) across the different conditions. Figure 1B lacks sufficient detail describing the galectin puncta assay used. A brief explanation of the assay rationale and readout would help contextualize the findings.

      We thank the reviewer for pointing this out. We have revised the figures and figure legends accordingly by standardizing panel labels, adding scale bars where missing, and providing the age of animals used in the assays. We also expanded the description of the galectin puncta assay in the Results to explain the rationale and readout of sfGFP::LGALS3 puncta formation.

      Regarding the reviewer’s suggestion to compare young and aged animals, we agree that age-dependent changes in endolysosomal membrane integrity are an interesting question. However, the purpose of the present study was to investigate how perturbation of sphingolipid metabolism affects endolysosomal membrane fluidity and rupture under the assay conditions used in our original screen. A systematic comparison across aging is beyond the scope of the current revision. We have therefore clarified the animal ages used in the relevant figure legends and Methods.

      In Figure S1A, the authors show co-knockdown of multiple genes, including one condition with simultaneous RNAi against four targets. Because different RNAi clones can vary in knockdown efficiency, it is important to provide validation of gene knockdown levels (e.g., by qRT-PCR) shown in both panels a and b.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      In Figure 2E, the FRAP recovery curves show only ~60% recovery in controls after 25 seconds, and an even lower recovery (~40%) in hpo-8 and spp-10 RNAi conditions. The authors should discuss why the recovery is incomplete and what it implies about the mobile fraction of the protein or membrane components in these conditions.

      We agree that incomplete FRAP recovery is informative. For this reason, we report both the time to half-maximal recovery (thalf) and the maximal recoverable fluorescence signal. Increased thalf indicates reduced lateral mobility of LAAT-1::mCherry within the lysosomal membrane, consistent with reduced membrane fluidity. In addition, a reduced maximal recovery suggests that a larger fraction of the reporter is immobile or only slowly mobile during the time window analyzed. This may reflect stronger confinement of LAAT-1::mCherry within even more rigid membrane domains. However, because RNAi efficiency may differ between clones and we have not assessed their individual KD efficiency, we avoid overinterpreting differences in the absolute strength of recovery defects between individual KDs. Instead, we conclude that KD of sphingolipid-metabolism genes identified in our screen consistently reduces lysosomal membrane fluidity, as reflected by increased thalf and, in some cases, reduced maximal recovery.

      In Figure S3A, the Western blot for SPHK2 shows unequal loading between the control and siSPHK2 lanes. The blot should be normalized to a loading control and quantified to demonstrate knockdown efficiency.

      We have quantified SPHK2 levels relative to GAPDH across independent experiments and present the normalized quantification (Figure S3C-E).

      Key experimental details are missing from the manuscript. The strains of C. elegans and RNAi bacteria used were not described, and there is no information on biological replicates. The authors should clarify how many times each experiment was performed and provide more transparency on experimental reproducibility.

      We thank the reviewer for pointing this out. The C. elegans strains and RNAi bacterial clones used in this study were established and fully described in our previous study, which has now been published in Autophagy [1]. We now cite the published article throughout the revised manuscript and have added additional information in the Results section to explain the key features of the strains used here.

      We have also revised the Methods and figure legends to improve transparency regarding experimental details. The figure legends include the number of biological replicates, the number of animals or cells analyzed, and the statistical tests used for each experiment. In addition, the Statistical Analysis section in the Methods now summarizes how replicate numbers and sample sizes are reported across the study. Finally, the source details for the strains and RNAi clones used in this study are now provided in Tables S1 and S2, respectively. These revisions should improve the experimental clarity and reproducibility of the data shown.

      References:

      (1) Sandhof CA, Martin N, Tittelmeier J, Schlueter A, Pezzali M, Schoendorf DC, et al. A novel C. elegans model for MAPT/Tau spreading reveals genes critical for endolysosomal integrity and seeded MAPT/Tau aggregation. Autophagy. 2025;21(12):2963-81. Epub 20250904. doi: 10.1080/15548627.2025.2551676. PubMed PMID: 40851193; PubMed Central PMCID: PMCPMC12758218.

      (2) Yong J, Villalta JE, Vu N, Kukurugya MA, Olsson N, Lopez MP, et al. Impairment of lipid homeostasis causes lysosomal accumulation of endogenous protein aggregates through ESCRT disruption. eLife. 2024;12. Epub 20241223. doi: 10.7554/eLife.86194. PubMed PMID: 39713930; PubMed Central PMCID: PMCPMC11666243.

      (3) Calixto A, Chelur D, Topalidou I, Chen X, Chalfie M. Enhanced neuronal RNAi in C. elegans using SID-1. Nat Methods. 2010;7(7):554-9. doi: 10.1038/nmeth.1463. PubMed PMID: 20512143; PubMed Central PMCID: PMC2894993.

      (4) Li Y, Zhang J, Li M, Yang L, Wang X. Sphingosine kinase SPHK-1 maintains sphingolipid metabolism to protect lysosome membrane integrity in C. elegans. Mol Biol Cell. 2026;37(1):ar1. Epub 20251105. doi: 10.1091/mbc.E25-04-0182. PubMed PMID: 41191545; PubMed Central PMCID: PMCPMC12696880.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

      We sincerely appreciate your careful review and positive feedback. We are glad to hear that you are satisfied with the revised manuscript and acknowledge the scientific value and solid evidence of our work. Thank you again for all your efforts and valuable suggestions.

      Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The corresponding content has been revised.

      We sincerely thank you for your detailed review and valuable comments. We greatly appreciate your recognition of the cell-type-specific heterogeneity of the JAK-STAT pathway and the clinical value of JSP and STAT4 as predictive biomarkers for immunotherapy in triple-negative breast cancer. We have thoroughly revised the manuscript according to your previous suggestions, and all raised concerns have been fully addressed.

      Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

      Many thanks for your careful evaluation and valuable suggestions. We highly appreciate your affirmation of the therapeutic significance of this study. We have fully revised the manuscript to enrich the molecular mechanisms, and all your comments have been properly resolved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The most content has been revised.

      Minor corrections:

      (1) The icon about "prognosis" in graphic abstract is overly childish.

      The graphical abstract has been redrawn. The inappropriate prognosis icon is deleted accordingly.

      (2) Please double check the whole content to avoid typos. For instance, "2.2" and "2.3" have been repeated twice.

      We appreciate your reminder. The duplicate numbering of 2.2 and 2.3 resulted from Word’s automatic heading feature. We have disabled this function and fixed all repeated section numbers. In addition, we have carefully checked the full text and corrected all typos.

      (3) It will be more interesting if the oncogenic role of STAT4 could be verified via cell cloning assay.

      We appreciate your thoughtful comment. Considering the limited revision time, we cannot add the cell cloning assay in the current version. Our present data sufficiently validate the oncogenic function of STAT4, and the main conclusions remain reliable.

      Reviewer #3 (Recommendations for the authors):

      I recommend publishing the revised manuscript in eLife.

      We sincerely thank you for your positive evaluation and endorsement for the publication of our revised manuscript. We greatly appreciate your rigorous review and insightful comments that have substantially improved the quality and readability of this work.

      We sincerely appreciate all reviewers and editors for your thorough reviewing work and thoughtful feedback. Your suggestions have helped us greatly improve this manuscript. Thank you very much.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Strengths:

      Several studies in the past have predicted different functions of the anterior (INT1) vs posterior (INT2-9) epithelial cells of the C. elegans intestine based on their anatomy and ultrastructure, but detailed characterization of differences in gene expression between these cell types (and whether indeed these are different 'cell types') was lacking prior to this study. The genes and drivers identified to be exclusively expressed in the anterior vs posterior segments of the intestine will be very helpful to selectively modulate different parts of the C. elegans intestine in future studies.

      Another strength of this study is the careful experimental design to test how the anterior vs posterior cell types of the intestine respond differently to food deprivation and recovery after return to food. These comparisons between 'states' of a cell in different physiological conditions are difficult to pick up in single-cell analyses due to low sequencing depth, which can fail to identify subtle modulation of gene expression.

      The TRAP-associated bulk RNA-seq approach used in this study is more suitable for such comparisons and provides additional information on post-transcriptional regulation during metabolic stress.

      A key finding of this study is that pyruvate levels modulate the translation state of anterior intestinal cells during fasting. Characterization of pyruvate metabolism genes, especially of the enzymes involved in its mitochondrial breakdown, provides novel insights into how gut epithelial cells respond to the acute absence of food.

      Weaknesses:

      Unlike previous TRAP-seq studies (PMID: 30580965, 36044259, 36977417) that reported sequencing data for both input and IP samples, this study only reports the sequencing data for IP samples. Since biochemical pulldowns are variable across replicates, it is difficult to know if the observed differences between different conditions are due to biological factors or differences in IP efficiency. More importantly, since two different TRAP lines were utilized in this study and a large proportion of the results focus on the differences between the translational profiles of INT1 vs INT2-9 cells, it is essential to know if the IP worked with similar efficiency for both TRAP strains that likely have different expression levels of the HA-tagged ribosomal protein. One way to estimate this would be to perform qRT-PCR of genes that are known to be enriched in all intestinal cells and determine whether their fold-enrichment over housekeeping genes (normalized to input) is similar in INT1 vs INT2-9 TRAP strains and across the fed vs fasted conditions. The authors, in fact, mention variability across biological replicates, due to which certain replicates were excluded from their WGCNA analysis.

      We appreciate the reviewer's comments. We agree that the lack of matched input sequencing libraries limits our ability to directly assess IP efficiency across replicates, conditions, and TRAP strains. However, several features of the dataset support the conclusion that the major differences reported here reflect biological rather than purely technical variation. First, the RPL-22-3xHA construct was integrated into each line to improve consistency across experiments. Second, although the INT2-9 TRAP strain yielded more RNA than the INT1 strain, as expected given the larger number of labeled cells, downstream analyses were performed on normalized count data rather than raw counts. Third, principal component analysis showed robust separation by promoter identity across all conditions, and expected INT1-enriched genes such as ins-7 were recovered in the INT1 dataset. Together, these observations support the interpretation that the TRAP datasets capture reproducible, cell-type-specific differences in ribosome-associated transcripts. Nonetheless, we agree that direct input-normalized measurements would further strengthen the study, and we will explicitly note this as an important limitation.

      It appears that GFP expression is also detectable in INT2 (in addition to strong expression in INT1 in Fig.1A). Compared to INT3-9, which looks red, INT2 cells appear yellow, suggesting that the expression patterns of the two TRAP drivers are not mutually exclusive, which changes the interpretation of many of the results described in the study.

      We agree that the Pges-1ΔB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2. Because Pges-1ΔB is an engineered promoter derived from the intestine-specific Pges-11 promoter, this low-level INT2 expression is not unexpected. However, we note that the expression level in INT1 is substantially higher than in INT2. Thus, although the expression patterns of the two TRAP drivers are not completely mutually exclusive, Pges-1ΔB still provides the most selective available tool for enriching the INT1 translatome in the context of the current study.

      Some parts of the study overemphasize the differences between the INT1 vs INT2-9 cell types, which is a biased representation of the results. For example, the authors specifically point out that 270 genes are differentially expressed in opposite directions in INT1 vs INT2-9 cell types during acute (30 min) fasting without mentioning the 1,268 genes that are differentially expressed in the same direction. They also do not mention here that 96% of the genes are differentially expressed in the same direction in INT1 and INT2-9 cell types after prolonged (180 min) fasting, suggesting that the divergent translational responses of these cell types are only observed in the first 30 minutes of food deprivation. Similar results have also been reported for the effect of fasting on locomotory and feeding behaviors, where 30 min of fasting produces more variable effects, which become more consistent after longer periods of fasting (PMID: 36083280). Hence, the effects of brief food deprivation should be interpreted with caution.

      The intestine functions as a discrete and cohesive organ, so the expected result is that there would be no differences across the different cell types. For us, the surprise was that, in fact, there are differences between these cells at all. However, the point is well taken, and we have added a statement in the text to reflect that many genes change similarly in INT1 and INT2-9, while the differences reflect important functional divergence between these cell types.

      Many of the interpretations of this study primarily rely on pathway enrichment analyses, which are based on the known function of genes. The function of uncharacterized genes that were found to be differentially expressed in INT1 vs INT2-9 cell types, e.g., the ShKT proteins, was not explored in this study. In addition, overreliance on pathway enrichment tools (instead of functional validation) has resulted in several conflicting findings. For example, one of the main messages of this study is that INT1 cells specialize in immune and stress response in response to fasting, which relies on pathway analysis in Figs 5E and 5F. However, pathway analysis at a different time point (shown in Figure S5A) indicates that INT2-9 cells show a much stronger increase in translation of stress and pathogen-responsive genes compared to INT1 cells. Hence, some of the results should be interpreted as different translational effects in INT1 vs INT2-9 cells after different lengths of food deprivation, without making broad claims about selective pathways being affected only in specific cell types.

      We agree that some interpretations in the manuscript relied heavily on pathway enrichment analyses and should be stated more cautiously. In particular, we agree that the current data are most consistent with state-dependent differences in translational responses between INT1 and INT2-9 cells across different durations of food deprivation, rather than with the strongest version of a claim that specific pathways are selectively engaged only in one intestinal subset. We also agree that uncharacterized genes, including the ShKT family, were not mechanistically explored in the present study and should be presented as important candidates for future investigation.

      The authors have compared their TRAP-seq results with genes enriched in the anterior and posterior intestine clusters from a previously published whole-animal adult scRNA dataset (PMID: 37352352). They claim that their TRAP-seq results are in agreement with the findings of the scRNA study. However, among the 10 genes from the 'posterior intestine' scRNA cluster in Fig.S1E, six are downregulated in the INT1 vs INT2-9 comparison, while four are upregulated. Hence, there is no clear agreement between the two studies in terms of the top enriched genes in the anterior vs posterior intestine, which should be considered for cross-study comparisons in the future.

      We have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. We instead assessed the expression levels of our INT1 up-regulated genes in their intestinal cluster and found that they have higher expression in the anterior intestine cluster (new Figure S1C). These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9.

      The authors describe in the manuscript that they have performed INT1-specific RNAi for two C-type lectin genes that are upregulated during fasting. Due to a recent expansion of C-type lectin genes in C. elegans, there is a high chance of off-target effects of RNAi that is designed for members of this gene family. More trustworthy results could have been obtained using CRISPR-based loss-of-function alleles for these genes, one of which is publicly available. Also, the authors do not provide any explanation for why knockdown of these stress-response genes, which are activated in INT1 cells in response to food deprivation, results in improved resistance to pathogens. This, in fact, suggests a role of INT1 cells in increasing pathogen susceptibility, and not pathogen resistance, during food deprivation.

      We agree that RNAi targeting C-type lectin family members may be susceptible to off-target effects, and that validation with CRISPR null alleles, where available, would strengthen these findings. In the current study, we used INT1-specific RNAi as a cell-specific first-pass approach to test candidate gene function. We also agree that the pathogen phenotype requires cautious interpretation. Specifically, the finding that knockdown of fasting-induced INT1 lectin genes improves pathogen resistance does not support a simple protective model for these genes. Instead, it suggests that INT1-expressed stress-response genes modulate host susceptibility or host-pathogen interactions.

      Many of the studies in this field (e.g., references 2-4 in this article) have investigated the effects of food deprivation ranging from 4 hr to 24 hr, which results in activation of starvation responses in C. elegans. In contrast, the authors have used shorter time periods of fasting (30 min and 180 min), and most of their follow-up experiments have used 30 min of food deprivation. Previous work has shown that the effects of food deprivation can either accumulate over time (i.e., the effect gets stronger with longer food deprivation) or can be transient (i.e., only observed briefly after removal of food and not observed during long-term food deprivation). Starvation-induced transcription factors such as DAF-16/FoxO and HLH-30 show strong translocation to the nucleus only after 30 min of fasting. Though gene expression changes in all stages of food deprivation are of biological relevance, the authors have missed the opportunity to explore whether increased INS-7 secretion from the anterior intestine is dependent on these starvation-induced transcription factors (which can be easily tested using loss-of-function alleles) or is due to other fast-acting regulatory mechanisms induced due to the absence of food contents in the gut lumen. A previous study (PMID: 40991693) has shown that DAF-16 activation during prolonged starvation shuts down insulin peptide secretion from the intestinal epithelial cells. Hence, it is not clear if increased INS-7 secretion is only a feature of short-term food deprivation or is also a signature of long-term starvation (e.g., at 8 hr or 16 hr timepoints). Since most of the INS-7 secretion data in this study are for 30 min of fasting, it remains unknown whether the discovered regulators of INS-7 secretion can be generalized for extended food deprivation that triggers major metabolic changes, such as fat loss (e.g., conditions shown in Figure 1D).

      We agree that short-term food deprivation and prolonged starvation likely engage distinct regulatory mechanisms, and that our study primarily addresses an early phase of food deprivation rather than the full spectrum of starvation responses described in prior work. We selected the 30 min fasting condition because our previous study showed that INS-7 secretion is induced within this interval and returns to baseline upon refeeding, even before detectable intestinal fat loss. We also included a 180 min fasting condition to capture a later state associated with metabolic changes. However, we agree that the present study does not determine whether the regulators of INS-7 secretion identified here also govern secretion during more prolonged starvation (for example, 8 hr or 16 hr), nor does it test whether starvation-responsive transcription factors such as DAF-16 or HLH-30 contribute to this regulation. We appreciate that determining how this response transitions during prolonged starvation will be an important direction for future work.

      Two previous studies (PMID: 18025456, 40991693) have shown a strong reduction in the expression of ins-7 in the anterior intestine using GFP-based reporters (both promoter fusions and endogenous CRISPR-generated) and in whole-animal RNA-seq data from starved animals. These results are in contrast to the increased INS-7 secretion from INT1 cells during fasting that is reported in this study. The authors here have reported that INS-7 translation is higher in INT1 compared to INT2-9 during fed, acute fasted, and chronic fasted conditions, but they have not shown whether INS-7 translation is upregulated during acute and chronic fasting in INT1 cells in their TRAP-seq analysis. Knowing whether increased INS-7 secretion during acute fasting is due to increased transcription, translation, or secretion of INS-7 is crucial to resolve the discrepancy between these studies.

      In our dataset, INS-7 translation in INT1 tended to increase during acute fasting relative to the fed state (log<sub>2</sub>FC = 0.69), although this effect did not reach statistical significance after adjustment for multiple comparisons. Consistent with this trend, our secretion assay showed that INS-7 release from INT1 increases during fasting. However, we agree that the current data do not distinguish whether this increase in secretion is driven by enhanced synthesis, regulated release of pre-existing peptide stores, or a combination of both.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to understand whether the discrete segments of the C.elegans intestine were specialized to carry out distinct functions during an animal's exposure and adaptation to a fast-changing nutrient environment. To achieve this, the authors used a method called Translating ribosome affinity purification (TRAP), which provides a snapshot of what genes are being translated into proteins (and therefore functionally prioritized by the animal) under different fasting and re-feeding conditions. By expressing the TRAP constructs in two distinct segments of the intestine (INT1) and (INT2-9), the authors were able to identify how these segments responded to changing nutrient availability.

      Already under steady state nutrient conditions, the authors found that INT1 and INT2-9 appeared to have different 'tasks', with INT1 expressing more immune- and stress-response related genes. Exposing animals to different regimens of starvation and refeeding also showed marked differences between the intestinal segments, and the gene expression patterns in INT1 were consistent with INT1 cells playing an integrative role in linking nutrient cues to the secretion of insulin molecules that regulate fat metabolism with food intake. In summary, the data presented catalogue, for the first time, gene expression differences between two areas of the intestine, suspected to play different roles, and through clever experiments, links these gene expression changes to responses to nutrient availability.

      Strengths:

      The data presented catalogue - for the first time and in a careful manner - gene expression differences between two areas of the intestine. They strongly support the presence of intriguing differences between two areas of the intestine in immune, metabolic, and stress-response regulation, and link these gene expression changes to the responses of these regions to nutrient availability.

      Weaknesses:

      The conclusions of this paper are mostly well-supported by data, but the relevance of the changing gene expression patterns could be better clarified and extended in the discussion.

      We thank the reviewer for this constructive comment. In the revised manuscript, we have now expanded the discussion to more clearly interpret these dynamic translatomic changes in the context of intestinal subset specialization. The most pronounced difference between INT1 and INT2-9 cells is the enrichment of stress-response genes in INT1. Based on the present findings, together with our previous work identifying INS-7 as an INT1-secreted signal (PMID: 39127676), we propose that INT1 cells are sentinel enteroendocrine cells that integrate information from the luminal environment and the metabolic state of intestinal cells.

      Reviewer #3 (Public review):

      Summary:

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7), and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding).

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species with the beads during the immunoprecipitation step. This is a minor point, however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      We agree that a limitation of TRAP-seq is that it is best suited for relative comparisons across cell types or conditions, rather than for determining the absolute expression level of a given transcript in a specific cell type. As the reviewer notes, this limitation arises in part from nonspecific recovery of background or sticky RNAs during the immunoprecipitation step, complicating the interpretation of absolute expression levels. For this reason, our analysis was designed to focus primarily on differentially enriched transcripts between INT1 and INT2-9 cells and across feeding states, rather than on assigning absolute expression levels to individual genes. We appreciate the reviewer’s recognition of this point. Our study uses TRAP-seq specifically to define relative translatomic differences between intestinal subsets and physiological states, which is well aligned with the strengths of this approach.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not delve a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in the context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      We agree that the current study does not fully resolve the molecular mechanisms by which the candidate genes identified by TRAP-seq regulate the phenotypes examined here. Our primary goal was to generate a spatially resolved translatomic framework for INT1 and INT2-9 cells across feeding states, and to perform focused validation of selected candidates to establish the physiological relevance of the profiling results. We therefore view the current functional analyses as an initial validation and proof of principle, rather than a comprehensive mechanistic dissection of the molecular pathways for each candidate. We appreciate the reviewer’s recognition that these findings provide an excellent starting point for future studies.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions:

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes is enriched for expected molecular functions and forms coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      Testing additional promoter::mNeonGreen reporters across a wider range of fold-change and statistical thresholds could be somewhat useful for calibrating ranked TRAP-seq candidate genes. However, given the variation in strains bearing extrachromosomal arrays, we did not consider this a stringent enough test, given that the sensitivity and dynamic range of RNA-seq far outpaces genetic fluorescence-based reporters. For these reasons, we focused on the top differentially enriched candidates to provide not only an initial validation of the dataset, but also to determine whether these candidates regulate biological functions in INT1 cells and thus serve as potentially useful biological readouts in future efforts.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      We agree that validating the INT1-specific RNAi phenotypes with independent genetic approaches, including null-allele or degron-based alleles for essential genes, would further strengthen the conclusions. In the current study, we used INT1-specific RNAi as a rapid and spatially restricted strategy to functionally test candidates identified by TRAP-seq and to determine whether these genes contribute to the specialized physiological functions of INT1 cells. We consider these experiments an initial validation of candidate function rather than a complete genetic dissection, which could be conducted in future efforts to study other aspects of INT1 function.

      Impact:

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors need to describe the fasting method used in detail. Was fasting performed on unseeded NGM plates or in liquid (M9 buffer)? Were the animals washed with buffer prior to starvation? If yes, how many times? These details are critical for any researcher to follow up on their results.

      We have clarified the fasting/refeeding procedure in the revised Methods section. Briefly, worms were washed off OP50-seeded NGM plates with M9 buffer, washed three times in M9 buffer, and then transferred to unseeded NGM plates for fasting. For refeeding, worms were collected from the unseeded NGM plates with M9 buffer and transferred back to OP50-seeded NGM plates.

      (2) The authors claim that "INT1 and INT2-9 cells maintain fundamentally different molecular identities independent of any and all acute or chronic conditions", which they primarily based on Principal Component Analysis (PCA). The circles shown in Figure 2A are arbitrary, and many such circles can be drawn in the 2D space to separate the samples in different ways. The authors should show this comparison in a translatome-wide similarity heatmap with hierarchical clustering (similar to Fig.2C, but with all the experimental conditions and their replicates on both x- and y-axes).

      The ellipses shown in Figure 2A were generated using the stat_ellipse() function in ggplot2, which calculates the mean and covariance of the PC1 and PC2 for each line and draws ellipses corresponding to the 95% confidence level. The separation between lines is primarily driven by PC2, which accounts for 14% of the variance in the translatomic dataset. Although this difference is less pronounced when considering the full translatome, samples from the same line nevertheless cluster together, supporting line-specific differences in translatomic profile.

      (3) Figure 4 of the study shows 18 Venn diagrams for genes that are differentially expressed between INT1 and INT2-9 cell types in fed, fasted, and refed conditions. In the absence of any statistical comparisons, it is difficult to interpret whether the extents of overlap (higher or lower than expected) are significant. Ideally, P values for hypergeometric tests should be provided for the overlap regions.

      We appreciate the reviewer’s suggestion. We explored the use of hypergeometric testing, implemented through the SuperExactTest package in R, to assess the statistical significance of the overlaps shown in the Venn diagrams. However, this analysis yielded significant P values for essentially all overlap regions, including cases in which the degree of overlap was not especially informative and did not align with the interpretation presented in the text. This outcome likely reflects the dependence of the test on the size of the input gene sets and background universe, which can make statistical significance difficult to interpret meaningfully in this context.

      (4) The claims made in lines 242-244 (Figure 5D) need to be supported by P values from hypergeometric tests.

      Similar to the previous point.

      (5) The interpretation of Figures 6E and 6F described in lines 296-297 needs to be supported by statistical analyses. The authors claim that the undulating pattern of expression of the turquoise module genes is stronger in INT1 compared to INT2-9. However, based on Figures 2C and 6F, it appears that the expression change is not necessarily weaker in INT2-9, but instead is different, i.e., the expression of turquoise module genes goes up during fasting in INT1 and goes down after refeeding, while their expression goes up during fasting and stays up after refeeding in INT2-9 cells.

      We appreciate the reviewer’s point and agree that the turquoise module shows dynamic regulation in both cell populations. The key difference is not the presence versus absence of an undulating pattern, but rather the magnitude of that change, which is greater in INT1. Because of the limited number of biological replicates in some conditions, particularly the fasting group, we interpreted these results cautiously and used a nonparametric approach to assess differences in average module expression between states. This analysis indicated that the turquoise module changes significantly in both lines, but with a larger effect size in INT1. We have included the corresponding statistical analysis and effect size in the revised manuscript.

      (6) Since the INS-7 coelomocyte uptake assay was used extensively in this study, some representative microscopy images should be included to complement the quantification.

      We have added a new Figure 6G showing representative images corresponding to the quantification presented in Figure 6H.

      (7) The authors claim that INT1-specific fmo-2 RNAi results in reduced basal INS-7 secretion, but they do not have the direct statistical comparison for this. Were experiments shown in Figures 6G and 6K done on the same day?

      In the original Figure 6K (now Figure 6L), the data are presented as the percentage of normalized INS-7::mCherry fluorescence intensity relative to fed animals treated with vector RNAi. A statistical comparison between fed animals treated with INT1-specific fmo-2 RNAi and fed vector RNAi controls was performed and was significant. We have also clarified that the experiments shown in the original Figures 6G and 6I (now Figures 6H and 6J) were performed on the same day.

      (8) It is not clear why blocking the mitochondrial breakdown of pyruvate (Figures 7E and 7F) does not mimic the fasted state in terms of increased INS-7 secretion from INT1 cells. Doesn't this contradict the proposed model in this study? Can the authors speculate why this is the case?

      We do not interpret inhibition of pyruvate dehydrogenase or pyruvate carboxylase as equivalent to the fasted state. Rather, our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion. To directly examine this possibility, we performed the experiment in Figure 7G, which tests whether maintaining pyruvate levels in INT1 cells during fasting is sufficient to suppress INS-7 secretion. The results are consistent with this interpretation and further support a model in which decreased intracellular pyruvate is a key determinant of fasting-induced INS-7 secretion.

      Minor comments:

      (1) Figures 2E and 2G are very similar and represent the same result in two different ways (unbiased vs guided comparison). One of these should be moved to the supplementary figures.

      Although these figures show similar patterns, they were derived from two independent analytical approaches, WGCNA and differential expression analysis. We therefore interpret the concordance between these independent methods as strengthening the robustness of the association and increasing confidence in the biological relevance of the observed pattern.

      (2) In Figures S1C, S1D, and S1E, a more significant P-value is shown with a smaller circle, and a less significant P-value is shown with a larger circle. This is confusing to the reader and should be inverted.

      We appreciate the reviewer’s comment and have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes used in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. Our further examination of marker expression across the intestinal clusters in the Ghaddar et al. (PMID: 37352352). dataset confirmed this limitation. These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9. We also note that the spatial identities of the intestinal clusters in Ghaddar et al. (PMID: 37352352) were not clearly established in the text or by spatial transcriptomic evidence, making it difficult to assign the annotated anterior, middle, and posterior clusters to specific intestinal cells. We have revised the manuscript accordingly and replaced the original figure panels.

      (3) Line 131 mentions the comprehensive characterization of the translatomic differences between INT1 and INT2-9 cells under each acute and chronic condition. However, the paragraph only discusses the differences in the fed condition. This is confusing, and the authors should mention the comparison between these cell types under acute and chronic conditions in subsequent sections where it is described.

      In this paragraph, we indeed discuss the ‘fed’ condition, but in subsequent sections we follow with details analyses of acute versus chronic, as well as regional differences across the intestine. We have clarified this in the opening sentence of the referenced paragraph.

      (4) The Venn diagrams in Figure 4 look very similar, and it is hard to differentiate between how 4A is different from 4I, how 4B is different from 4J, etc. The authors should include the labels for 'acute' or 'chronic' above each Venn diagram to guide the reader through these panels.

      We have added labels indicating the acute and chronic conditions to the left side of each Venn diagram in the revised Figure 4.

      (5) It is not clear in the figure legends how Figure 5E is different from Figure S6A, and how Figure 5F is different from Figure S7A. This should be better described in the figure legends.

      We have added a sentence to better describe this in the figure legend.

      (6) Figure 6C: Survival parameters such as median lifespan, number of animals for each condition, etc., should be reported for the different conditions.

      We have revised Figure 6C to indicate the number of animals analyzed in each condition, and the median survival for each group is now reported in the corresponding figure legend.

      (7) The colors used for control RNAi and clec-160 RNAi are very similar in Fig.6C. Easily distinguishable colors should be used.

      We have changed the colors as suggested.

      (8) The INT1-specific RNAi strain should be first described in line 285.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (9) Line 304: 'REF' should be replaced with the reference.

      We have replaced the “REF” with the reference (PMID: 39127676)

      (10) The P value for statistical comparison between the fed and 30 min refed states should be shown in Figures 6G, 6I, and 6K.

      We have now included the p value for the comparison as suggested. Figures 6G, 6I, and 6K are now labeled as 6H, 6J, and 6L, respectively.

      (11) In Figure 7, the authors should consider replacing the 'refed' label with 'recovery' because the pyruvate treatment was done in the absence of 'feeding' (= bacteria consumption).

      We appreciate the reviewer’s point. However, we chose to retain the label “refed” in Figure 7 to maintain consistency across the set of conditions examined, including 2% glucose and OP50 supernatant, which likewise do not involve bacterial consumption despite not showing effect on the refeeding response of INS-7 secretion.

      (12) The full form of DISN should be mentioned in the figure legend of Figure 7.

      We have included the full form of D1SN in the figure legend of Figure 7A.

      (13) Line 367: 'normalization' should be replaced with 'return to basal levels'. 'Normalization of INS-7 secretion' might also mean normalization of INS-7::mCherry signal to CLM::GFP signal.

      We have revised the wording per the reviewer's suggestion.

      (14) The methods section has a quantitative RT-PCR section, but it is not clear if RT-PCR data are reported in any of the figures. Also, no qPCR primers are listed in Table S3.

      We have removed the quantitative RT-PCR part from the methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors describe that the RPL-22-3xHA constructs are not integrated, at the very end, in the section "Limitations of the data". An earlier mention of this caveat would have been useful. In addition, it would help if the authors could provide their defense (which I think is very valid) of using non-integrated strains in the results section, as they describe the experimental setup. Also, some details were missing, which left me wanting to know: Were there expression differences? How were they accounted for? Was expression normalized between these two constructs, and if so, how?

      The RPL-22-3xHA construct was integrated into each line to ensure more consistent transgene expression across experiments. Because the INT2-9 construct is expressed in a larger number of cells than the INT1 construct, the INT2–9 samples yielded greater amounts of pulled-down nascent RNA, as reflected in the supplemental table and in the higher aligned RNA counts observed for the INT2-9 samples. To account for these differences, differential expression analysis was performed using DESeq2, which corrects for library size by estimating sample-specific size factors with the median-of-ratios method. Raw counts are then normalized using these size factors, thereby accounting for differences in sequencing depth and minimizing confounding effects due to variation in library size. Such differences are common in RNA-seq experiments, particularly when comparing samples derived from distinct input populations.

      (2) The 'acute' and 'chronic' exposures are thought through and carefully defined. The question I do have is whether the 3-hour fasting can be considered chronic fasting, given how surprisingly fast the animals lose their fat content. Could these kinetics indicate that the 30-minute fasting is reflective of mechanisms during which senses change in food availability, whereas the 30 minutes represents acute fasting (with chronic fasting - meaning fasting, during which the animal activated alternative pathways - occurring later)? While this may appear to be pure semantics, it could influence how the authors interpret their results. One method to more objectively separate an 'acute' from a 'chronic' stage may be to conduct a time course of fat loss-does fat loss plateau after 3 hours? The timing when the rate of decrease levels off could be more indicative of the beginning of a chronic phase.

      We appreciate this important point and agree that it should be more clearly discussed. We interpret acute fasting as a pre-fat-loss state, since it is 30 minutes off food and no difference in fat levels are detectable at this stage (Fig 1B). The translatomic changes observed under acute fasting therefore likely reflect food-sensing mechanisms and early preparatory responses that promote subsequent fat mobilization. In contrast, chronic fasting (180 minutes off food – see Fig 1D) appears to represent a post-fat-loss state, in which fat stores have already been depleted, and the corresponding translatomic changes likely reflect the effects of sustained metabolic stress.

      (3) The age of the animals used has to be more explicitly stated. Were these animals egg-laying? Or L4/young adults? This is likely to impact the changes that the animals undergo.

      Day 1 young adults were subjected to the fasting. Great care was taken to ensure consistency across biological replicates.

      (4) What is the rationale, in the authors' view, that stress response genes are apparently more enriched than metabolic or mitochondrial enzymes, and membrane receptor changes? Are the latter mostly regulated by PTMs/localization changes, etc?

      Based on our current data, we cannot exclude the possibility that metabolic or mitochondrial enzymes, as well as membrane receptors, are regulated in INT1 and INT2–9 cells through mechanisms not captured at the translatome level, including post-translational modification or changes in subcellular localization under different fasting and refeeding conditions.

      (5) The refeeding experiment with latex beads and killed OP50 is very clever. Details on when INS-7 was evaluated in the caoelomocytes would help the reader better understand and interpret these results.

      INS-7mCherry signal was evaluated in the coelomocytes immediately after refeeding; we included this information in the methods section and referenced our previous paper.

      (6) In the Discussion, I was looking for a more detailed context for how to think about the differences and similarities in the RNA-seq data between the two segments, and perhaps a discussion of whether there were any indications that the two segments communicated with each other.

      The data show that the most pronounced difference between INT1 and the rest of the intestine at the RNAseq level, is the expression of stress response genes in INT1. Although there are some nuanced differences, the prevalence of stress response genes persists across feeding and fasting conditions. This difference, combined with the evidence that INT1 cells secrete the enteroendocrine peptide INS-7 (Fig 6 and PMID: 39127676) is strongly reminiscent of the mammalian enteroendocrine cells, which also secrete peptides and show strong expression of stress response genes (PMID: 37626258 and 27148273). We suggest that this category term reflects not only a canonical stress response, but also a broader response to shifts in the luminal environment, which INT1 cells are anatomically poised to detect well before the absorption of nutrients has begun further down the intestine (INT2-9). Thus, we believe INT1 cells are a newly defined enteroendocrine cell type within the C. elegans intestine.

      Regarding communication between INT1 and INT2-9 – this is an intriguing possibility that we have considered, given that peptide genes and receptors are found in the RNAseq datasets. The extent to which the expression of these genes leads to functional effects is the subject of future investigation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1A - It would be better to also show single fluorescent protein channels to assess the specificity of the expression patterns. A schematic or labels of where the INT1 vs. INT2-9 boundaries are located would be helpful to non-experts.

      (2) Figure 4 - At present, the Venn Diagrams are a very complicated way to visualize all of the comparisons/conditions. I would recommend that the authors consider using UpSet plots to better summarize the relevant comparisons they would like to make. The same consideration applies to Figure 5D.

      (3) Line 301 - Description of the INT1-specific RNAi strategy. I think it would be better to bring this information earlier, close to line 285, where the authors first mention performing INT1-specific RNAi experiments.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (4) Line 304 - I think the authors meant to cite a reference where the REF placeholder text is found

      We have replaced the “REF” with the reference.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data from patient iPSC-derived neurons and a mouse model were collected using solid methods, but the evidence supporting key claims is incomplete, and some technical aspects fall short of expectations. Despite these limitations, the study provides a useful foundation for exploring the relationship between mitochondrial defects and primary cilia in neural development.We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We have also:

      - Improved figure clarity and consistency

      - Corrected errors in gene annotations and normalization

      - Refined the mechanistic framework linking OCRL, mitochondria, and cilia

      New Experimental Data:

      Figure 5, Supplementary Figure 4. We confirmed mitochondrial defects by generating ocrl-KO zebrafish (Supplementary Figure 4). We first assessed mitochondrial reactive oxygen species (mitoROS) using MitoSOX staining. Next, we evaluated mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos. Finally, we assessed mitochondrial content via TOM20 staining. For all analyses, we focused on the ocular and cranial regions of the zebrafish to maintain consistency (see Author response image 1). Notably, previous studies have reported that ocrl-KO zebrafish exhibit seizures and brain developmental abnormalities, supporting their relevance as a model for Lowe syndrome-like phenotypes [1].

      Author response image 1.

      Figure 6 c. In addition to quantifying the proportion of ciliated cells in the IOB mouse brain, we measured cilia length and compared it between IOB and WT brain sections. Our results show that IOB mice exhibit elongated cilia compared to WT controls, suggesting that OCRL deficiency is associated with stress-related alterations in ciliary structure. These findings are consistent with previous studies reporting that cilia elongation can be associated with increased ROS levels and mitochondrial dysfunction [2,3].

      Public Reviews:

      Reviewer #1 (Public review):

      The preparation of the manuscript requires improvement. There are many errors in the presentation of data.

      We thank the reviewer for this important comment. We have carefully revised the manuscript to improve the clarity, accuracy, and consistency of data presentation.

      Specifically, we have corrected inconsistencies in gene nomenclature (e.g., CO2 vs COX2, DLOOP) across the text, figures, and legends. We standardized normalization methods and ensured consistency between figures and descriptions. We revised figure labels, legends, and annotations for clarity and accuracy. We corrected referencing errors and ensured appropriate citation of prior work. We improved overall figure quality and readability. In addition, we performed a thorough review of the entire manuscript to eliminate typographical errors and ensure consistency in terminology and data interpretation.

      The use of references needs to be re-considered. Sometimes a reference is used when in fact the results included in that paper are the opposite of what the authors intend.

      We thank the reviewer for this important comment. We have carefully re-evaluated all references throughout the manuscript to ensure that they accurately reflect the findings they are cited to support. In cases where the cited studies did not fully align with our interpretation or could be misleading, we have either revised the text to more accurately represent the original findings or replaced the references with more appropriate sources. We have also clarified instances where prior studies report differing or context-dependent results to avoid overinterpretation.

      The authors conclude the paper by claiming that mitochondrial dysfunction and impairments of the ciliary SHH contribute to abnormal neuronal differentiation in LS, but the mechanism by which this sequence of events might happen hasn't been shown.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct causal mechanism linking mitochondrial dysfunction, ciliary SHH signaling, and altered neuronal differentiation in Lowe syndrome. Our data demonstrate that these processes co-occur consistently across multiple model systems, supporting a potential functional relationship. However, we acknowledge that the precise sequence of events and mechanistic connections remains to be defined. To address this, we have revised the manuscript to clarify that our conclusions are based on associative findings rather than direct mechanistic evidence. We have also updated the Discussion to explicitly acknowledge this limitation and to frame our model (Figure 7) as a proposed working hypothesis. Future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary SHH signaling and how these pathways influence neuronal differentiation.

      Phenotype of increased astrocytes in both the IOB mouse brain or iPSC-derived cultures iN cells requires clarification as one of the markers used as an astrocyte marker, BRN2, is commonly used as a neuronal marker. As LS is a neurodevelopmental disorder, and the phenotype in question is related to differentiation, it is crucial to shed light on the developmental timeline in which this phenotype is seen in the mouse brain.

      We thank the reviewer for this important comment. We agree that the use of BRN2 as an astrocytic marker was inappropriate, as it is primarily recognized as a neuronal marker. Accordingly, we have revised the manuscript to remove BRN2 from the interpretation of astrocytic identity and now rely on GFAP expression as the primary astrocytic marker. We have also clarified this point in both the Results and figure legends to avoid misinterpretation. In addition, we have revised the text to more accurately describe our findings as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers.

      Regarding the developmental context, we acknowledge that Lowe syndrome is a neurodevelopmental disorder and that temporal aspects are highly relevant. In our study, the in vivo analyses were performed on adult 2-month-old IOB mouse brains, which we have now explicitly stated in the manuscript. We recognize that this limits our ability to directly assess developmental dynamics of lineage specification. We have therefore added this as a limitation in the Discussion and clarified that future studies examining earlier developmental stages will be necessary to determine when these alterations arise.

      Mitochondrial dysfunction in astrocytes has been shown to induce a ciliogenic program. However, almost the opposite is shown in this paper, with regards to ciliation. Morphology of the cilia was not assessed either, which is an important feature of ciliary homeostasis. The improper ciliary homeostasis here appears to be the improper Shh signalling, which has not been shown to be related to mitochondrial dysfunction. This leaves one wondering how exactly the different phenotypes shown in this paper are connected.

      We thank the reviewer for this important comment. We agree that the relationship between mitochondrial dysfunction, ciliogenesis, and Shh signaling is complex and not fully resolved in the current study.

      As noted by the reviewer, prior studies have reported that mitochondrial dysfunction can promote a ciliogenic program [4]. In contrast, our data show a reduced proportion of ciliated cells together with increased cilia length, indicating altered ciliary homeostasis rather than a straightforward increase in ciliogenesis. To address this point, we have revised the manuscript to describe our findings as context-dependent alterations in ciliary parameters more clearly, and we now explicitly discuss this apparent discrepancy with the literature in the Discussion. We also acknowledge the reviewer’s point regarding ciliary morphology. In the revised manuscript, we have included quantification of cilia length in addition to the proportion of ciliated cells, and we have expanded the Methods section to detail how these measurements were performed. We agree that additional ultrastructural and functional analyses would further strengthen the characterization of ciliary homeostasis, and we now include this as a limitation and future direction.

      Regarding the link between mitochondrial dysfunction, ciliary alterations, and Shh signaling, we agree that our study does not establish a direct mechanistic connection. Our data demonstrate that these phenotypes co-occur consistently across multiple models, but do not define causality. To address this concern, we have revised the manuscript to clarify that our conclusions are associative, and we now present our integrated model (Figure 7) as a working hypothesis rather than a demonstrated mechanism. We also explicitly state in the Discussion that future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary signaling and Shh pathway activity.

      This paper lacks a clear mechanistic approach. While the data validates the 3 broad phenotypes mentioned, there is a lack of connection between these phenotypes or an answer to why these phenotypes appear. While the discussion attempts to shed light on this by referencing previous studies, some of the referenced studies show contradicting results. Hence, it would be beneficial to clarify these gaps with further experiments and address the larger question of the connection between the mitochondria, Shh signalling, and astrocyte formation.

      We thank the reviewer for this important and insightful comment. We agree that the current study does not establish a direct mechanistic link connecting mitochondrial dysfunction, altered Shh signaling, and changes in neuronal versus astrocytic differentiation.

      Our primary goal in this work was to identify and validate phenotypes associated with OCRL deficiency across multiple independent model systems. We demonstrate that mitochondrial dysfunction, oxidative stress, altered ciliary/Shh signaling, and changes in neural lineage-associated markers co-occur consistently in these models. However, we acknowledge that the causal relationships between these processes remain to be defined.

      To address this concern, we have revised the manuscript to more clearly state that our conclusions are associative rather than mechanistic, and we now present our integrated model (Figure 7) as a working hypothesis that links these phenotypes through a potential mitochondria-ROS-signaling axis. We have also expanded the Discussion to explicitly acknowledge this limitation and to avoid overinterpretation of causality.

      In addition, we have carefully re-evaluated and revised the cited literature to ensure accuracy, particularly in cases where prior studies report context-dependent or seemingly contradictory effects of mitochondrial dysfunction on ciliogenesis and signaling pathways. These points are now discussed more explicitly to better position our findings within the existing literature.

      We agree that further experiments, such as targeted rescue of mitochondrial function or modulation of Shh signaling, will be necessary to establish causal relationships between these pathways. These directions are now clearly outlined in the revised Discussion as important next steps.

      Most importantly, there is no mention of how the loss of OCRL, a 5-phosphatase enzyme, results in the appearance of the mentioned phenotypes. Since there are multiple studies in the field of Lowe Syndrome that shed light on the various functions of OCRL, both catalytic and non-catalytic, it is important to address the role of OCRL in resulting in these phenotypes.

      We thank the reviewer for this important comment. We agree that the link between OCRL function and the observed phenotypes was not sufficiently developed in the original version of the manuscript.

      In the revised manuscript, we have expanded the Discussion to more clearly outline how loss of OCRL could contribute to the observed mitochondrial, ciliary, and differentiation phenotypes. OCRL encodes a PI(4,5)P₂ 5-phosphatase that regulates phosphoinositide homeostasis and membrane dynamics. Disruption of this activity is known to affect endolysosomal trafficking, actin organization, and membrane remodeling-processes that are critical for organelle maintenance and ciliary function. We now discuss how these alterations could impact mitochondrial homeostasis, for example, through defects in membrane contact sites, vesicular trafficking, or organelle quality control pathways.

      In addition, we have incorporated discussion of potential non-catalytic roles of OCRL, including protein–protein interactions and scaffolding functions, which may contribute to the coordination of intracellular trafficking and cytoskeletal organization. These aspects may provide an additional layer of regulation linking OCRL loss to both mitochondrial dysfunction and ciliary alterations.

      We emphasize that, while these mechanisms are supported by prior studies, our data do not directly test them. Therefore, we have carefully framed this section as a plausible mechanistic framework rather than a demonstrated pathway and have explicitly stated this limitation. We also outline future experiments aimed at dissecting catalytic versus non-catalytic contributions of OCRL to these phenotypes.

      There are numerous errors in the qPCR experiments performed concerning the genes that were assayed. The genes mentioned in the text section do not match those indicated in the graphs or legends. This takes away the confidence of the reader in this data.

      We thank the reviewer for this important observation. We agree that the inconsistencies between the genes described in the text and those shown in the figures and legends could reduce confidence in the data. In the revised manuscript, we have carefully rechecked all qPCR experiments and corrected the gene names across the Results, figures, and figure legends to ensure full consistency. We have also standardized the nomenclature throughout the manuscript (including consistent use of gene symbols and formatting) and verified that all plotted data correspond to the correct targets.

      In addition, we have clarified the qPCR methodology, including normalization (all data are normalized to GAPDH) and primer information, to improve transparency and reproducibility.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either an LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and a decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in an LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic Hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development, and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      We have carefully revised the manuscript to address the concerns raised, and we have revised the manuscript with additional supporting data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have checked the expression of neuronal markers NeuN and FoxG1, and apart from GFAP, they categorise Brn2 as one of the astrocytes markers that they have also checked. But Brn2 is not an astrocyte marker. It is a neuronal marker that is expressed in layer 2/3 of the cortex. In fact, Brn2 is reported to be a key driver of neurogenesis in primate telencephalon development1 and for reprogramming of astrocytes to neurons2. Hence, the only glial marker they have used here is GFAP. BRN2 is a neuronal marker. It has been used as a neuronal marker even in the reference (Zhang et al, 2013), from which the protocol for inducing iPSCs to induced neurons (iNs) was taken. Hence, the qPCR results in 1g of overexpression of BRN2 indicate an increase in expression of a neuronal marker, not an astrocyte marker.

      We thank the reviewer for this important and well-founded comment. We fully agree that BRN2 is a neuronal marker and not an astrocytic marker, and that its inclusion as an astrocyte marker in our original interpretation was incorrect.

      In the revised manuscript, we have removed BRN2 from the analysis and interpretation of astrocytic identity. We now treat BRN2 exclusively as a neuronal marker and have updated the Results, figure legends, and text accordingly. Specifically, the qPCR data previously presented in Figure 1g are now interpreted as reflecting neuronal marker expression, not astrocytic differentiation. We have also revised our conclusions to avoid overinterpretation of astrocyte abundance. Our findings are now described more accurately as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers. In this context, GFAP remains the primary astrocytic marker used in this study.

      We acknowledge the reviewer’s point that reliance on a single astrocytic marker is a limitation. This has now been explicitly stated in the Discussion, and we note that additional astrocyte markers will be required in future studies to more comprehensively define lineage-specific changes.

      Incorrect marker usage (BRN2 as astrocyte marker)

      We thank the reviewer for identifying this critical issue. We corrected the classification of BRN2 as a neuronal marker. Also, we re-analyzed the interpretation accordingly, revised all relevant text and figures. Importantly, Astrocyte conclusions are now based primarily on GFAP expression, and we explicitly acknowledge this limitation in the Discussion

      The graphs for the RT-PCR results indicate that gene expression values are normalized to actin whereas the legend mentions that they are normalized to GAPDH. This needs clarification.

      We thank the reviewer for pointing out this inconsistency. We confirm that all qPCR data were normalized to GAPDH, and the reference to actin was an error. This has now been corrected throughout the figures, legends, and text to ensure consistency.

      The use of wording to refer to the generation of induced neurons (iNs) should ideally be changed from "we developed"; as the protocol from Zhang et al, 2013 seems to have been directly adapted in this paper.

      We thank the reviewer for this helpful suggestion. We agree that the wording was inappropriate. In the revised manuscript, we have replaced “we developed” with language indicating that iNs were generated using an established protocol, and we now explicitly state that the method was adapted from Zhang et al., 2013 [5].

      OCRL KO iPSCs were obtained from Herbert Lachman's lab and not generated in this study. Hence, the Ran et al, 2013 reference is not necessary.

      We thank the reviewer for this clarification. We agree that the OCRL knockout iPSCs were obtained from Herbert Lachman’s laboratory and were not generated in this study. Accordingly, we have removed the Ran et al., 2013 reference and revised the manuscript to clearly state the origin of the OCRL KO iPSC line.

      The reference for Figure 1c is given as Ran et al, 2013 which is wrong. It should be Zhang et al, 2013.

      We thank the reviewer for noting this error. We have corrected the reference for Figure 1c from Ran et al., 2013 to Zhang et al., 2013 in the revised manuscript.Limited in vivo mitochondrial characterization

      Figure 1D, E: GFAP is a cytoskeletal marker but its expression here is very grainy and looks like an artifact. Is it possible to show the astrocyte phenotype using other astrocytes nuclei and cytosolic markers such as NFIA and S100B, respectively?

      We thank the reviewer for this important suggestion. We acknowledge that GFAP is a cytoskeletal marker and that the signal in the current images may appear granular. We have carefully re-evaluated the staining and image processing to ensure that the signal represents true GFAP expression and have improved the image quality and presentation in the revised figures. We agree that inclusion of additional astrocytic markers such as NFIA and S100B would further strengthen the characterization. While we were not able to include these additional markers in the current revision, we now explicitly acknowledge this as a limitation in the Discussion and note that future studies will incorporate a broader panel of astrocyte markers to more comprehensively define astrocytic identity.

      Since the authors have not used enough markers to understand the cell-type composition in WT and OCRL KO/mutant lines, it's not sufficient to conclude that the NPCs preferentially favour astrocytes over neuronal lineage. Any conclusive comments regarding the cell-state/cell-type specification necessitate evidence such as genetic lineage tracing using reporters for neuronal and astrocyte markers, and/or RNA/ATAC/scRNA sequencing.

      We thank the reviewer for this important point. We agree that the current marker panel is not sufficient to definitively determine cell-type composition or to conclude preferential lineage specification.

      In the revised manuscript, we have tempered our conclusions and now describe our findings as changes in neuronal versus astrocytic marker expression, rather than evidence of a shift in lineage fate. We also explicitly acknowledge this limitation in the Discussion. We agree that approaches such as genetic lineage tracing, reporter-based assays, and single-cell transcriptomic or epigenomic analyses (e.g., scRNA-seq or scATAC-seq) would be required to rigorously define cell-state transitions and lineage outcomes. These are important directions for future studies and are now highlighted in the revised manuscript.

      Figure 2:

      The mt-DNA gene CO2 was checked, not COX2. A typographical error in the written section, which does not match with the qPCR graph of the same.

      We thank the reviewer for noting this inconsistency. We confirm that the gene analyzed was CO2, and the reference to COX2 in the text was incorrect. This has now been corrected throughout the manuscript to ensure consistency between the text, figures, and qPCR data.

      The word "neurogenesis" is used very loosely throughout the paper. In the opinion of the reviewer, there is no evidence presented that there is a defect in neurogenesis in either of the models used in this paper.

      We thank the reviewer for this important comment. We agree that the term “neurogenesis” was used too broadly and is not directly supported by our data. In the revised manuscript, we have removed or replaced this term where appropriate and now refer more precisely to changes in neuronal versus astrocytic marker expression.

      Line 150: They say that they have examined the functional properties of mitochondria during neurogenesis but it would have been better to understand OXPHOS at various time points of neurogenesis to actually conclude reduced OXPHOS 'during neurogenesis'. Moreover, genes related to other pathways such as glycolysis could have been checked to understand the bioenergetics of LS patients. Also, oxidative stress could have been checked using more than one marker. Since they are trying to understand the functional role of mitochondrial defects during neurogenesis, they could have performed live imaging of mitochondrial potential during various stages of neurogenesis. Isolation of mitochondria from LS patients and transcriptomics/proteomics might provide further clues about mitochondrial defects.

      We thank the reviewer for these thoughtful suggestions. We agree that our data do not capture mitochondrial function across multiple stages of neurogenesis. In the revised manuscript, we have modified the wording to avoid implying temporal analysis “during neurogenesis” and instead describe mitochondrial parameters in differentiated cells. To strengthen the study, we have included additional in vivo validation in the zebrafish model, where we assessed multiple mitochondrial readouts, including mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos, oxidative mitochondrial stress ROS (mitoROS) using MitoSOX staining, and mitochondrial content (TOM20), supporting mitochondrial dysfunction across systems.

      Elevated astrocytic reaction during the differentiation of NSPCs in the Lowe syndrome (IOB) mouse model.

      Title: What does astrocyte reaction mean? This term should not be used without clear evidence of reactive astrocytes being present in the model.

      We thank the reviewer for this important comment. We agree that the term “astrocytic reaction” is not appropriate without specific evidence of reactive astrocytes. In the revised manuscript, we have removed this terminology and replaced it with more accurate wording, describing our findings as altered astrocytic marker expression. This change better reflects the data and avoids overinterpretation.

      In 1a, no quantification of the mouse brain size is given. From the given images alone, there appears to be no obvious decrease in brain size between the WT and IOB mice. This contradicts the text which indicates that the IOB mouse brain is smaller.

      We appreciate this important point. We have:

      Removed claims regarding reduced brain size

      Clarified that our analysis was limited to available sections and no definitive conclusion about global brain morphology can be made

      Figure 3E: Why is the astrocyte to neuron ratio measured using a cytoskeletal marker for astrocytes, GFAP but a nuclear marker for neurons, NeuN? Ratios to measure the percentage or proportion of astrocytes to neurons can only be checked by markers of the same nature such as GFAP to MAP2 (neuronal cytoskeletal marker) or NFIA (astrocyte nuclear marker) to NeuN.

      We thank the reviewer for this important point. We agree that comparing a cytoskeletal marker (GFAP) with a nuclear marker (NeuN) is not ideal for deriving cell-type ratios. In the revised manuscript, we have removed the astrocyte-to-neuron ratio analysis and now present these data as relative marker expression/signals rather than proportions. We have also clarified this limitation in the text and Discussion.

      (3) Lack of clarity in the experiments performed on the mice brains. PAX6 is used here as a neuronal marker along with NeuN, a mature neuronal marker. This is misleading as PAX6 is rather a marker for neural stem/progenitor cells and not neurons. The age of the mice has also not been mentioned, which is crucial considering the different markers used to characterize the mouse brain as well as since the authors are indicating that there is an abnormal neurodevelopment in the IOB mouse during development. Again, BRN2 is used here as an astrocyte marker. However, it is a neuronal marker. Hence, the phenotype of increased astrocytes currently is held by GFAP expression alone. Another astrocyte marker should be used.

      We thank the reviewer for these important points. We have revised the manuscript to correct marker interpretation, now describing PAX6 as a progenitor marker rather than neuronal, and BRN2 as a neuronal marker, removing it from astrocyte-related analysis. We have also explicitly stated the age of the mice (2 months) in the Methods and Results. In addition, we have tempered our conclusions, describing the data as changes in marker expression rather than definitive cell-type shifts, and we now acknowledge that reliance on GFAP as a single astrocytic marker is a limitation, which is discussed in the revised manuscript.

      Figure 6: Increase in astrocytes, mitochondrial dysfunction, and ciliary Shh signalling are 3 phenotypes discussed in this study. However, no experiments were done to shed light on the mechanistic connection between these phenotypes. This is reflected in the abstract shown in Figure 6. There is no comment on the mechanism behind these phenotypes.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct mechanistic link between mitochondrial dysfunction, altered ciliary Shh signaling, and changes in astrocytic markers. Our aim was to identify and validate these phenotypes across multiple models. In the revised manuscript, we have clarified that Figure 7 represents a proposed working model based on associative findings rather than a defined mechanism. We have also revised the Discussion to explicitly acknowledge this limitation and to outline future experiments required to establish causal relationships between these

      Reviewer #2 (Recommendations for the authors):

      The authors report interesting findings in two different experimental models but the manuscript would benefit significantly from an analysis of a potential causal relationship between different findings. They often mention neural stem cells or the neuron/glia switch but their analysis of the mouse mutant is restricted to the adult stage. A more consistent analysis of specific brain regions would also be beneficial.

      We thank the reviewer for this constructive comment. We agree that establishing causal relationships between the observed phenotypes is an important next step. In the revised manuscript, we have clarified that our conclusions are based on associative findings and have expanded the Discussion to outline experimental strategies that could address causality in future studies. We also acknowledge that our in vivo analysis is restricted to adult (2-month-old) IOB mouse brains, which limits our ability to assess developmental dynamics such as neural stem cell behavior or neuron-glial transitions. This limitation is now explicitly stated in the Discussion. Finally, we agree that region-specific analysis would strengthen the study. Due to the availability of samples, our analysis was not systematically performed across defined brain regions. We now acknowledge this limitation and note that future studies focusing on specific regions (e.g., cortex, hippocampus) will be important to better understand the spatial aspects of the phenotype.

      Figure 1: The authors only measured the expression of marker genes by qRT-PCR. This could reflect higher expression levels in individual cells rather than a change in the proportion of neurons and astrocytes. They need to determine the cell proportions of astrocytes and neurons in addition. Moreover, there is a poor marker choice. Loss of FOXG1 expression could indicate a loss of telencephalic identity. BRN2 is expressed by cortical neurons.

      We thank the reviewer for this important comment. We agree that qPCR-based marker analysis does not directly reflect cell-type proportions and may instead represent changes in gene expression per cell. Accordingly, we have revised the manuscript to avoid conclusions about cell proportions and now describe the data as changes in marker expression. We have also corrected marker interpretation, removing BRN2 from astrocyte analysis and clarifying that FOXG1 reflects telencephalic identity. These limitations and the need for more comprehensive cell-type characterization are now acknowledged in the Discussion.

      In Figure 2, the authors determine the properties of mitochondria and claim that functional mitochondrial activities are decreased during neurogenesis in mutant iN cells. They need to take into account that according to Figure 1 the proportion of neurons and astrocytes may be changed. Hence, the decreased mitochondrial activity may reflect a fundamental difference between neurons and astrocytes. The authors need to clearly distinguish between neurons and astrocytes in their analysis. Moreover, the use of the term neurogenesis is confusing. They are analysing the neuron-to-glial switch, not the formation of neurons.

      We thank the reviewer for this important comment. We agree that differences in cell-type composition may influence mitochondrial measurements. In the revised manuscript, we have tempered our interpretation, describing these data as changes in mitochondrial parameters at the population level rather than neuron-specific effects. We also acknowledge this limitation in the Discussion and note that cell-type-specific analyses will be required in future studies. In addition, we have revised the terminology throughout the manuscript, removing the term “neurogenesis” and instead referring to changes in neuronal versus glial marker expression to more accurately reflect the scope of our analysis.

      Figure 3: The authors claim that astrocyte numbers are elevated in the IOB mouse model, however, it seems as if the authors analysed late postnatal, potentially adult brains but no age of the brains is provided. Given the large time lag between the formation of astrocytes and their analysis, the increased number of astrocytes could be due to a number of processes including altered proliferation and cell death. The authors need to investigate the proportion of astrocytes and neurons closer to the neuron-to-glia switch. Cell fate experiments like the long-term application of BrdU would be much better suited and would provide mechanistic insights. Again, markers are not adequate to reach their conclusion. Pax6 is only expressed in a tiny subset of neurons, Brn2 on the other hand is not astrocyte-specific as it is expressed in cortical neurons as well. Moreover, qRT-PCR analyses were done in the cortex and hippocampus whereas the boxes in Figure 3D are located in the basal ganglia. It would be much more informative and provide better comparisons to perform gene expression analysis and cell counts in the same brain regions.

      We thank the reviewer for these important and constructive comments. We agree that our analysis is limited by the use of adult (2-month-old) IOB mouse brains, which do not allow direct assessment of developmental processes such as the neuron-to-glia transition. We have now explicitly stated the age of the animals and clarified this limitation in the Discussion, including the possibility that changes in astrocytic markers may reflect processes such as proliferation or survival rather than lineage specification.

      We also agree that our marker panel was insufficient for definitive conclusions. Accordingly, we have revised the manuscript to remove overinterpretation, corrected marker usage, and now describe the data as changes in marker expression rather than cell-type proportions. The need for more rigorous approaches, such as lineage tracing (e.g., BrdU) and expanded marker panels, is now acknowledged as a future direction. Finally, we thank the reviewer for pointing out the inconsistency in the brain regions analyzed. We have clarified the regions used for qPCR, and we now explicitly acknowledge this limitation, noting that future studies will aim to perform region-matched molecular and histological analyses for more accurate comparisons.

      Experiments in Figure 4 assess "whether changes in mitochondrial activity are involved in the altered differentiation of stem cells and progenitor cells in the LS mouse model" in 3-month-old brain sections. The murine adult brain only contains a few neural stem cells in the SVZ and in the dentate gyrus. Instead, this analysis needed to be done at late embryonic/early postnatal stages to capture the neuronal/glial switch. In addition, RT-PCR and immunostainings should be performed in the same brain region as stated above.

      We thank the reviewer for this important point. We agree that analysis in adult (2-month-old) brains does not capture developmental stages such as the neuron-glia transition. We have revised the manuscript to remove implications of developmental analysis and now describe these data as mitochondrial parameters in adult tissue, explicitly acknowledging this limitation in the Discussion. We also clarify the brain regions used for qPCR and immunostaining and note as a limitation that these were not fully matched; future studies will perform region-specific, developmentally timed analyses.

      Figure 5: The authors examine a potential link between mitochondrial defects and primary cilia. Mutant iN cell cultures contain lower levels of SHH mRNA and show concomitantly lower expression of the SHH target genes GLI1 and PTCH1. The authors link this finding with a reduced proportion of ciliated cells but the reduced SHH signalling is most likely explained by the decreased SHH expression. The authors also limit their analysis of primary cilia to one brain region, but they should also include the cortex and hippocampus as these regions were used for their qRT-PCR analysis. SHH signalling acts as a switch to stop the proteolytic processing of GLI3 and to promote the formation of the GLI3 activator form. It is therefore important to determine the ratio of GLI3 repressor and GLI3 activator using western blots. The authors claim that they found defective cilia formation, but cilia are poorly characterised. Are there differences in intraflagellar transport, the formation of the transition zone, etc? Is ciliary length altered? The authors only make a correlative link between mitochondrial defects and cilia but present no experiments to investigate causation. They should at least discuss potential mechanisms which could explain defects in cilia.

      We thank the reviewer for these insightful comments. We agree that reduced SHH pathway activity may be influenced by decreased SHH expression, and we have revised the text to avoid overattributing this effect to ciliary changes. Our conclusions are now framed as associative, not causal.

      We have expanded our cilia analysis to include quantification of both the proportion of ciliated cells and cilia length, and clarified these methods in the manuscript. We also acknowledge that additional characterization (e.g., intraflagellar transport, transition zone structure, GLI3 activator/repressor ratios) would further strengthen the analysis, and we now include this as a limitation and future direction. Regarding regional analysis, we agree that broader brain region coverage would be valuable. Due to sample availability, our analysis was limited, and this is now explicitly acknowledged as a limitation, with future studies aimed at region-matched analyses (e.g., cortex and hippocampus). Finally, we have expanded the Discussion to outline potential mechanisms linking mitochondrial dysfunction and ciliary alterations, while clearly stating that causal relationships remain to be established.

      (1) The methods section does not contain any information on how immunostainings on brain sections were performed.

      We agree that the description of immunostaining on brain sections was missing. We have now added a detailed protocol for brain section immunostaining in the Methods section to improve clarity and reproducibility.

      (2) The abbreviation "RT-PCR" is used for both, real-time PCR and reverse transcription PCR

      We also acknowledge the inconsistent use of the term “RT-PCR.” In the revised manuscript, we have standardized the terminology, using “qPCR” (quantitative real-time PCR) throughout to avoid confusion

      Conclusion

      We believe that these revisions significantly strengthen the manuscript. While the study remains primarily associative, it provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome.

      References:

      (1) Ramirez IB-R, Pietka G, Jones DR, Divecha N, Alia A, Baraban SC, et al. Impaired neural development in a zebrafish model for Lowe syndrome. Hum Mol Genet. 2012;21:1744–59. https://doi.org/10.1093/hmg/ddr608

      (2) Kim JI, Kim J, Jang H-S, Noh MR, Lipschutz JH, Park KM. Reduction of oxidative stress during recovery accelerates normalization of primary cilia length that is altered after ischemic injury in murine kidneys. Am J Physiol Renal Physiol. 2013;304:F1283-1294. https://doi.org/10.1152/ajprenal.00427.2012

      (3) Moruzzi N, Valladolid-Acebes I, Kannabiran SA, Bulgaro S, Burtscher I, Leibiger B, et al. Mitochondrial impairment and intracellular reactive oxygen species alter primary cilia morphology. Life Sci Alliance. 2022;5:e202201505. https://doi.org/10.26508/lsa.202201505

      (4) Ignatenko O, Malinen S, Rybas S, Vihinen H, Nikkanen J, Kononov A, et al. Mitochondrial dysfunction compromises ciliary homeostasis in astrocytes. J Cell Biol. 2022;222:e202203019. https://doi.org/10.1083/jcb.202203019

      (5) Zhang Y, Pak C, Han Y, Ahlenius H, Zhang Z, Chanda S, et al. Rapid Single-Step Induction of Functional Neurons from Human Pluripotent Stem Cells. Neuron. 2013;78:785–98. https://doi.org/10.1016/j.neuron.2013.05.029

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling the study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying with the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) a dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on the spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      With regards to the reviewer’s comment on the overlap between BMA and BMD trait-influencing variants, we would like to add that the correlation of effect sizes within the overlap is -0.95 as we would expect from bifurcating differentiation of mesenchymal stem cells into osteoblasts or adipocytes: the underlying common biology driving these traits results in shared traits with a negative correlation in the effects.

      The study also investigated the genetic overlap between BMA and twelve (or 13) "brain and body" traits and identified significant genetic correlations with BMI, cognitive ability, and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      There are, however, some assumptions, findings, and their interpretation, which require more critical focus.

      Sex-specificity is well described and studied here. Men have higher BMA than women, but post-menopausal women catch up in the BMA values. The authors believe that calvarial marrow has a number of features that make it particularly well-suited to the study of BMA process - which is clinically important in other bone sites. It has a simple "sandwiched" structure that they are able to model. This is true only to some extent: a condition called "Hyperostosis frontalis interna", of unknown etiology (described by Smith & Hemphill in 1956) - is characterized by irregular overgrowth of the inner table of the frontal bone (symmetric/bilateral). Although not of clinical significance, typically benign, studies report a prevalence of 12%; However, it's most common in postmenopausal women - where prevalences up to 49% in women over the age of 65 - have been reported. Thus, sexual dimorphism is obvious and the effect of estrogen is likely shared with whichever bone - and marrow - age-related pathology. So, for women not using HRT, this new layer of the bone might interfere with the calvarial BMA readings and in turn, affect the BMA-related analyses.

      Thank you for bringing the "Hyperostosis frontalis interna" condition to our attention. It is particularly interesting to hear that the etiology is unknown and one may suspect that some kind of calvarial bone marrow dysregulation may be part of the cause. Our model for bone marrow location was trained on simulated data which included variation in the thickness of all anatomical layers (including the inner table), so it will be robust to some thickening of the inner table. It might not be robust to the most extreme cases of inner table thickening (as described in some case reports), but these are rare. Further, it should also be noted that the other calvarial bones, representing a much greater fraction of the calvarial surface, remain largely unaffected by the thickening and would therefore yield correct localisation of the bone marrow layers. In summary, although the severe cases of hyperostosis frontalis interna have the potential to affect our identification of the bone marrow layer, the low frequency of such cases and the restriction of the phenotype to the frontal bone means that the potential for bias is very limited.

      It would be interesting to develop a method for detection of thickened inner bone so that the condition’s prevalence can be quantified in a large sample like the UK Biobank and its genetic architecture be determined. This could help elucidate the etiology.

      The authors suspect that the effect of BMA on BMD may be biased in women; they should comment on those "with low BMD and high BMA" given that hyperostosis frontalis might be an issue. A strong effect of SNPs in the ESR1 chromosomal region might be akin to the above concern.

      Thank you for raising this point, which we have followed up with a new analysis.

      According to ICD-10 data in UK Biobank there are only N=105 individuals with an M85.2-diagnosed disorder. Given the total sample size of N=446,814 individuals with ICD-10 data, this would translate to a prevalence of 0.02%, which speaks for an underdiagnosis in this sample such that we cannot simply remove diagnosed individuals to control for a potential diagnostic confound.

      We have therefore taken a different approach to investigate this potential issue: As you elaborated in your previous comment, the prevalence of hyperostosis frontalis increases with age in females. The literature also suggests that prevalence rates do not differ between males and females in young age / prior to menopause. Therefore, we have studied the association between BMD and BMA for males and females separately, and in two age groups based on a median split of our sample (left plot: younger than 65, right plot: subjects older than 65). In these plots, the relatively large shift in female BMA and BMD is visible with the large yellow cloud at low BMA and high BMD in the left plot disappearing in the right plot. Despite this, we observe:

      (1) Associations in both males (blue) and females (yellow), suggesting that the associations were not driven only by females.

      (2) BMA-BMD association is largely similar across the two age groups.

      If we consider that the old age group is likely to contain more cases of hyperostosis frontalis than the young group, and if we consider that old-aged females are more likely to be in this condition than men of any age, then we would expect an impact of hyperostosis frontalis on our measures to result in observable differences in Author response image 1. This is not the case. We see global age-related shifts in BMA in women, yet the association with BMD remains similar across age groups.

      The technical properties of our neural network (trained on simulated data) makes it unlikely that frontal bone will contaminate the bone marrow detection globally (description above) and these results show that hyperostosis frontalis is not a considerable issue in our analysis.

      Author response image 1.

      Then, there is a perfect overlap of the BMA SNPs that are shared with BMD (497 of 500 variants), which may prove a "face validity" of the MRI-derived BMA. However, the BMD in the study was heel-derived eBMD - which is a good proxy for osteoporosis and is mostly driven by trabecular bone. Thus, there might be a concern that the BMA metrics capture some trabecular BMD.

      The reviewer is correct in pointing out that the BMA causal variants are a near-perfect subset of the BMD causal variants. The reviewer raises the concern that the BMA measurements may capture some trabecular BMD, however it should be noted that the correlation of effect sizes for the BMA/BMD overlapping causal SNPs is negative (-0.95). If our measure of BMA had been erroneously capturing trabecular BMD then we would expect to see a positive correlation of effect sizes for the BMA/BMD overlapping causal SNPs, not a negative one.

      Next, integrating mapped genes with existing bone marrow single-cell RNA-sequencing data revealed patterns of adipogenic lineage differentiation and lipid loading. The problem here is that the scRNAseq studies of the Bone Marrow niche are overwhelmingly mouse. The authors might wish to justify why they are relevant to humans (in the absence of the human-specific scRNAseq).

      We thank the reviewer for pointing this out. We noticed that, although Figure 4 and the Methods do explicitly state that the scRNAseq data is from mouse, it is not stated in the text of the Results. This is now corrected.

      The mouse is commonly used as the model organism for in vivo investigation of human phenotypes and bone marrow adiposity is no exception because, although mice have lower bone marrow adiposity than humans, the timing and sequence in bone marrow adiposity development are similar. BMA research makes extensive use of mouse models literature as exemplified by this review of research within the field (https://www.frontiersin.org/journals/endocrinology/articles/10.3389/fendo.2016.00127/full) and this article recent article (Koh et al. 2024. “Adult skull bone marrow is an expanding and resilient haematopoietic reservoir”. https://www.nature.com/articles/s41586-024-08163-9)

      We updated the results section (line 279):

      “Mesenchymal stem cells of the BM niche commit to either the adipogenic or the osteogenic lineage (Figure 4A) and both the number committing to the adipogenic lineage and their level of lipid-loading influences the total level of BMA. This aspect of BM biology is shared between humans and mice (29), so we made use of an existing mouse scRNAseq dataset of BM mesenchymal lineage cells (30) to study variation in the expression of BMA-associated genes as cells differentiate (Figure 4B).”

      For genetic correlation analysis, the authors selected 7 body and 6 brain traits. The latter traits reflect cognition (general cognitive ability and educational attainment) and brain-related disorders. This selection might seem arbitrary. The interpretation of genetic correlation with cognitive ability, education, and Parkinson's disease was attributed to the recently discovered vascular channels that link calvarial bone marrow to the meninges. This is a fascinating hypothesis, which requires functional proof. However, there might be simpler explanations. Thus, the diploe and the inner table of the calvarium are drained by the same veins as the dura. From the anatomy textbook, we know that diploic veins connect the pericranial and endocranial venous system through the skull.

      Whilst it is true that we did not systematically compare the results of the BMA GWAS to all potentially relevant brain and body phenotypes, we did use criteria to select the phenotypes we compared to. As stated in the manuscript (line 304):

      “We selected body traits (BMD, BMI, waist-to-hip ratio, systolic and diastolic blood pressure, type-2 diabetes, coronary artery disease) that have a logical connection to BMA given the mesenchymal stem cells origin of BM adipocytes and their role in bone, fat, and vasculature (29). For the brain, we selected traits reflecting cognition (general cognitive ability and educational attainment) and disorders that are prevalent in adulthood (insomnia, multiple sclerosis, Parkinson's disease, Alzheimer's disease) since it is primarily in adulthood that the adiposity of calvarial BM experiences a substantial change”

      We entirely agree that the suggestion that the genetic correlation between BMA and cerebral traits may be mediated by the vascular channels linking calvarial bone marrow to the meninges is merely a hypothesis. We have therefore updated the text of the Discussion (line 470):

      “We tentatively speculate that calvarial MALPs may be involved in sensing perivascular flows of CSF from the meninges to the BM and in influencing the BM’s hematopoietic response, and that this might be the basis of the observed genetic overlap between BMA and some cerebral traits. However, more conventional anatomical pathways may also be relevant, as the diploë and inner table communicate with meningeal and dural venous systems through diploic veins.”

      Reviewer #2 (Public review):

      Summary:

      This study develops a new artificial intelligence method for high-throughput analysis of skull bone marrow from MRI data, which may be useful for large-scale biological analyses. Using this method, the authors then attempt to estimate skull bone marrow adiposity (BMA) using T1-weighted signal intensity from MRI scans of ~33,000 people, followed by genome-wide association analysis; however, the approach is inadequate because T1-weighted signal intensity is not validated for measurement of bone marrow adiposity. If it could be validated, the study would be an important advance in understanding of bone marrow adiposity and skeletal biology.

      Strengths:

      This paper is well-written, and the figures are nicely presented. The neural network method used for analysing skull bone marrow is innovative, and the authors validate this through several approaches. Therefore, the authors have achieved the aim of developing a method for large-scale analysis of skull bone marrow from MRI data.

      The GWAS is reasonably well-powered and addresses potential ethnicity differences, with one GWAS done across white males and females, and a separate GWAS in non-white participants. The methodology also conforms to common GWAS standards, including for mapping genetic variants to candidate genes. Moreover, the study further investigates the biological roles of these genes by analysing their expression in single-cell RNA sequencing data.

      Weaknesses:

      The fundamental weakness is that T1-weighted MRI signal intensity (T1W) is used as an estimate of BMA, but it has never been validated for this. The authors show that this T1W parameter measures something that is heritable and can be compared between subjects, but they don't show that it actually measures (or even estimates) calvarial BMA. There is an attempt to do so by comparing the T1W parameter with data from quantitative T1 images: the authors show a reasonable correlation with some of the quantitative T1 image data. However, this still does not show that the parameter is measuring BMA; it could be measuring some other biological characteristic, but this remains unclear. So, there is a need to validate the T1W parameter against an established measure of BMA, such as the bone marrow fat-fraction or proton density fat fraction measured from multi-echo MRI analysis.

      Without validating this BMA measurement method, it is not possible to interpret the GWAS or other findings reported in the study.

      We reject this criticism.

      Although T1-weighted has not been validated as a quantitative measure of fat-fraction, there are several studies showing that it is a semi-quantitative measure of fat content (e.g. Loevner et al 2002, Shen et al 2013, Zhang et al 2020) and we also provide data that support this (figures S9-11).

      Semi-quantitative measures are used in many biomedical GWASes for instance even highly heritable neuropsychiatric disorders (such as schizophrenia and bipolar disorder) involve assessment by clinicians where the test-retest kappas are in the range 0.4-0.6.

      Further, we would suggest that the shortcoming of the imperfect correlation of T1w signal intensity with fat content is more than outweighed by our precision in identifying the calvarial BM cavity and the fact that the flat calvarial bone marrow has a wide range of adiposity in middle-aged and elderly individuals (compared to other bones). This lies at the root of why:

      We clearly recapitulate the known sex and age profiles, as well as the effect of HRT.

      We estimate high BMA heritabilities (43% in males and 23% in females)

      We find clear sex differences (which is a known feature of BMA biology)

      We identify a large number of the genes already known to affect BMA from earlier animal and cell work

      A noisy measurement of an entity with strong biological signal (a well-defined bone marrow cavity with variation in BMA across subjects) will often be more informative than a highly precise measurement of a poorly defined entity with little signal.

      A less critical weakness is that the GWAS has been done only on a single cohort, without replicating the findings in a follow-up cohort. For example, the authors could repeat their analysis on the remaining ~50,000 UK Biobank imaging participants for whom MRI data is now available. However, this would be pointless without knowing what biological characteristic(s) the T1W parameter is actually reflecting.

      We disagree with this comment. We separated the UKB data into discovery and replication sets prior to running the GWAS, so these datasets are independent:

      (1) Further, we ran the discovery (white british individuals) GWAS separately for males and females (prior to combining) and reported in the results section: “We found them to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      (2) We performed our replication GWAS in non-white British males and females. As noted in the results section: “Out of the 168 significant discovery SNPs, 62% replicated at P<.05, and 39% of the 41 lead SNPs replicated at P<.05 (Table S6). One locus replicated at genome-wide significance (P<5e-8). Furthermore, 92.7% of the lead SNPs of the discovery sample showed same effect direction in the replication sample (Table S5).”

      Reviewer #3 (Public review):

      Summary:

      This manuscript, "Estimating bone marrow adiposity from head MRI and identifying its genetic 2 architecture", brings together the groups of Drs. Kaufmann and Hughes in a tour de force work to develop an artificial neural network that localizes calvaria bone marrow in T1-weighted MRI head scans, with the goal of studying its composition in several large MRI datasets, and to model sex-dimorphic age trajectories, including the effect of menopause.

      Strengths:

      Bone marrow adiposity is a very active tissue with far-reaching implications for tissue crosstalk and human health than we had initially recognized. Although MRI has been used to measure BM, studies such as the one by these two groups are still lacking whereas very large datasets are analyzed using advanced AI machine learning tools coupled with genetic studies and a specific pathology. The groups had to develop new methods and new AI machine-learning tools for the imaging analyses.

      Weaknesses:

      Some aspects of the work that authors could add additional clarification.

      (1) Imaging Limitations: The authors provide an excellent overview and references supporting the use of MRI as a method for assessing marrow fat, particularly with some specific modifications. However, MRI images can be affected by various factors, including the presence of other tissues as well as specific MRI settings, which are much harder to precisely control when using different datasets.

      We thank the reviewer for his positive assessment of our review of methods.

      Regarding MRI settings: We agree with the reviewer that differences in scan protocols can create substantial differences in the resulting images between samples. Different tools exist for harmonization of imaging data across sites, but they usually operate on tabulated data and there is no one-size-fits-all approach yet [1]. Here, we took a different approach to prevent confounding bias: We generated a large set of simulated data for training of the neural network. The simulations circumvented potential issues emerging from confound biases in training sets that we might have seen had we had combined multiple samples with different scan protocols. Nevertheless, applied to real data the models may still face confound issues, such as better BMA estimates for some scan protocols over others. We have addressed these issues as follows: (1) Validation analyses (10 repeat scans of 30 individuals and twin pairs, figure 1d and 1e) are based fully on data that was acquired on the same scanner with the same protocol. (2) Analysis in UK Biobank included data from different scan sites albeit harmonized protocols. Here we accounted for scan site in all statistical models (including GWAS).

      (1) Dominik Kraft, Gloria Matte Bon, Édith Breton, Philipp Seidel, Tobias Kaufmann; Removing scanner effects with a multivariate latent approach: A RELIEF for the ABCD imaging data?. Imaging Neuroscience 2024; 2 1–7. doi: https://doi.org/10.1162/imag_a_00157

      Regarding the presence of other tissues: We recognise in the existing text of the results section that sometimes inner or outer table voxels are wrongly identified as bone marrow, but we show that this does not have a major impact on the correct identification of the bone marrow cavity. The existing text reads:

      “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap. However, this did not result in a corresponding fall in the ratio of the predicted intensity of BM to its true intensity, because the typical BM intensity was only marginally higher than the neighbouring bone intensity. This property of the typical relative intensities of these anatomic structures also explains why the intensity ratio at high overlap is not centred on 1: any misidentification of cortical bone as BM, will typically result in an underestimate of true BM intensity (Figure S2). The neural network performed well and intensity ratios were in the range 0.9-1.1 for the vast majority of head types (Figure S3).”

      Also note that we implement a number of QC measures to exclude scans where there is evidence that we may have failed to correctly identify the bone marrow cavity. The existing text reads:

      “We used two additional QC metrics to filter out calvaria where BM location was likely to have failed. First, we set an upper limit of 30 on the standard deviation of the intensity of the outer table as scans with higher values were clear outliers and were probably cases where the location of both outer table and BM has failed (Figure S6). Second, for each calvarium, we computed the Mahalonobis distance for all vertices in the two dimensions “first layer of the BM” and “BM intensity” (Figure S7). By manual inspection we found that data points with MD > 25 often had errors in BM layer identification, typically where the network had erroneously predicted a higher and more intense layer to be the BM. We considered a calvarium as failing this QC criterium if more than 0.5% of vertices have MD > 25. This criterium is very strict as errors on only 0.5% of data points in a calvarium would not significantly affect the average BM intensity for a calvarium.”

      (2) The specific density of cranial bones as it relates to the types of bone marrow: Cranial bones are extremely dense structures, which naturally interfere with MRI imaging. While it is thought that cranial bones have mostly "red bone marrow", this is only true for a short time in humans. How sensitive is their system in differentiating between red and yellow BM?

      We implemented several measures to ensure that our method would be robust to anatomical variation between individuals. As noted in the current version of the Methods section: “In order to train the neural network model, we generated a large synthetic dataset of intensity arrays, with known boundaries between anatomical structures, by simulating the thickness and intensity of the different structures located between the outer skin and the subarachnoid space. The simulation incorporated the following real-world complexities:

      Different anatomical architectures (skin, subcutaneous fat, aponeurosis, outer table, BM, inner table, dura mater, arachnoid space), including when a structure is not present throughout the calvarium

      Variation in thickness and intensity between vertices (on the same calvarium)

      A wide variety of different calvarium types with different combinations of levels of BM adiposity, bone thickness, and subcutaneous adiposity.”

      Further, as noted in the Results section:

      “We evaluated the performance of the neural network on simulated data using two metrics (Figure 1B): 1. the overlap between the predicted and the true BM location, and 2. the ratio between the predicted intensity of the BM and the true intensity of the BM”. The accuracy in localising the bone marrow layers was good, with the only exception being: “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap”

      We also validated our procedure on real data (see Results section, subsection “Procedure validation on real data and heritability estimate”). Briefly, we checked the accuracy of our method using a dataset from the Consortium for Reliability and Reproducibility, a twin dataset from the Human Connectome Project and by manually checking many hundreds of UKBiobank scans.

      We are thus confident that we accurately identify the bone marrow cavity irrespective of whether the bone marrow is red (low adiposity) or yellow (high adiposity).

      (3) Both items above are further complicated by aging, but aging is not a linear event as we have learned. There are specific bursts of aging in humans around the age of 45 and early 60s. How do the system and model predict or incorporate these peaks of aging? It seems from the data shown that aging is reflected more as a linear phenomenon. Is this because additional aging datasets are needed?

      We agree with the reviewer that ageing probably occurs in bursts rather than being a linear process. We do see a non-linear relationship between age and BMA in our data (see figure 2B), with a more rapid rise in BMA between the ages of 45 and 65, than later in life (in women). As a result of this, when we model BMA using regression, we use orthogonal polynomials of degree 2 which allows for a non-linear relationship. However, we cannot observe bursts of BMA increase in our data because it is cross-sectional. Longitudinal data would be required to obtain information on the nature and timing of any bursts in bone marrow adiposity.

      (4) The authors describe in richness of detail their AI learning programming and how it extracted the data from datasets. The authors also show some important correlations with specific genes, SNPs. What is not clear is how conditions such as anemia for example. An expected finding would be that patients with chronic anemia have lower bone marrow (BM) signal intensity on MRI scans than healthy people. This is because the signal intensity of BM depends on the fat-to-cell ratio in the tissue.

      We agree with the reviewer that conditions affecting the bone marrow niche have a potential to affect and be affected by bone marrow adiposity, with leukemia being a known example. This is why we believe that a method, such as the one we present here, has the potential to be useful in several biomedical fields (hematology and osteology).

      Furthermore, patients with a host of musculoskeletal disorders ranging from osteopenia to osteoporosis, sarcopenia, and osteosarcopenia will also have altered MRI scans. When using such large datasets how did the authors control or exclude these pathological conditions, or were all these conditions likely present?

      We did not exclude specific pathologies. We were careful to train our NN model on a wide variety of skull thicknesses, bone marrow adiposity, and subcutaneous adiposity and to evaluate the performance of the model on simulated and real datasets (see answer to your point 2).

      Reviewer 1 raised the issue of individuals displaying Hyperostosis frontalis interna (thickening of the inner table) and we recognize that in extreme case of this condition, where there is a major change in the anatomy of the calvarial bone, our method would probably not correctly localise the bone marrow. However, such extreme cases are rare and thus would not have a major impact on our results derived from over thirty thousand individuals. We demonstrate this with an extra analysis performed in response to the point about hyperostosis frontalis interna made by reviewer 1.

      (5) Some of the genes and SNPs although significant showed very small correlations. What is their likely physiological significance?

      Bone marrow adiposity is a polygenic trait and we have identified 41 statistically significant loci. We had a discovery sample of approximately 30k individuals which is modest for a GWAS study, so these 41 loci are a lower bound on the number of genes influencing the BMA trait. When a large number of genes influence a trait, the effect size of an individual gene is typically relatively small. However, the SNP heritability estimates of 31.5% indicates that we are able to explain approximately one third of the phenotypic variation with the effect sizes estimated by our GWAS: this is quite a high fraction relative to many other GWASs of biomedical traits.

      (6) The authors could use this excellent manuscript to expand their discussion to include the need for studies like theirs to be also complemented by multi-OMICS studies that will include proteomics and lipidomics of BM, bones, and muscles.

      We agree with the reviewer and hope that such studies will be undertaken in the future. We attempted to point in this direction in the last sentence of the Discussion (line 517): “Future studies can build on our developments to further validate the proposed measure of bone marrow composition and to study its effect on bone, blood, and brain”. Word count limits prevented us from further expanding on the specific kinds of studies that should be performed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      More moderate concerns include:

      (1) In the "Genetic correlation and overlap" part, it is unclear why a high effect correlation with high SNP overlap is suggestive of vertical pleiotropy, while "moderate overlap and a low correlation of effect sizes ... are more indicative of horizontal pleiotropy". This is not intuitive.

      In vertical pleiotropy (genetics > phenotype A > phenotype B): genetics drive phenotype A and phenotype A drives phenotype B (B has few direct genetic drivers of its own). If we perform a GWAS of phenotype A and a GWAS of phenotype B, one would expect to see a high overlap in the causal variants (because phenotype B is largely indirectly determined by the genetics of phenotype A) and the correlation should be high (either negative or positive) because there is a cause-effect relationship between A and B (whether phenotype A has a positive or negative effect on phenotype B).

      In horizontal pleiotropy: the A and B phenotypes share the same genetic loci (but have no phenotypic influence on each other). In this case, we would observe high overlap in the associated loci, but one would not expect to see a high correlation in effect sizes (across loci) because there is no a priori reason to expect that genes associated with both phenotype A and phenotype B would have a consistent (negative or positive) effect across loci. For example, gene X may increase A and B, whereas gene Y may increase A but decrease B.

      We have made a small update to the relevant part of the Results section and otherwise rely on the explanation above (which will be publicly available along with the manuscript):

      Line 325: “These patterns of high overlap and high effect correlation within this overlap are suggestive of vertical pleiotropy i.e. a molecular mechanism influencing one trait, that in turn influences a second trait, such that most of the variants driving the first trait either have the same or the opposite direction of effect on the second trait.”

      (2) "possible causal effect of BMA on cognition" asks for a formal analysis, like Mendelian randomization.

      Given the current wording, the reviewer is justified in asking for a formal analysis. Since we did not perform this analysis, we have changed the wording:

      Line 434: “Since these are two highly correlated traits [36], a high overlap and correlation of genetic effects for both traits with BMA may be consistent with the hypothesis that BMA could have a causal effect on cognition (Table 1).”

      (3) ll. 132-134: please reword this sentence for clarity: "Using the network-estimated location of the BM within ... averaged these across all datapoints...". Please define threshold of desirable overlap between the predicted BM and the true BM (=0.7?).

      Background: The model predicts the BM localisation for a datapoint (which interval of layers of the 50 layers is bone marrow). We tested the model on a wide variety of simulated data and aim for the overlap to be as close to 1 as possible, but some error is inevitable. We found that the average overlap between the true and predicted bone marrow was only below 0.7 when the bone layer is only 4 mm thick (meaning that the bone marrow is only 1-2 mm thick). This demonstrates the high accuracy of our method in identifying a very small anatomical feature.

      When applying our method to real data, we do not know the truth and therefore cannot compute the overlap between the predicted and true value. It is therefore not possible to identify datapoints where the overlap is poor (e.g. lower than 0.7) and filter them out.

      Given the above, we struggle to understand in what way an overlap threshold is relevant to how we compute the signal intensity for a datapoint. Nevertheless, we recognize that the sentence pointed to by the reviewer is poorly formulated and have tried to make it clearer:

      Line 130-133: “To obtain the BM signal intensity for an individual datapoint of the calvarium, we used the network model to estimate the location of the BM within the datapoint’s intensity array and averaged these BM intensities to get the BM intensity for that datapoint. Then, we averaged these datapoint intensities across the calvarium to produce the global BMA measure for the scan.”

      In the GWAS Results, please clarify the phrases - what was "significantly genetically correlated (Rg=.94)" (also, l. 402, "genetic correlation between the sexes" - in what?).

      Genetic correlation is a statistical measure that quantifies the extent to which two traits (or the same trait in two different cohorts) are influenced by the same genetic factors. Simply put, it is the effect sizes of the SNPs in the two GWASs of interest that are correlated (after correcting for confounding effects, such as linkage desequilibrium). When comparing two GWASs, the standard formulation is to refer to their “genetic correlation”. We made a modification to the text to clarify this:

      Line 239: “We found the male and female GWASs to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      "a more than two-fold difference between the sexes" - in which metric?

      We feel that what is being compared is stated clearly in the original sentence:

      Line 262: “A comparison of the male and female effect sizes of the top lead SNPs of each locus revealed 6 loci in which there is a more than two-fold difference between the sexes (loci 10, 18, 26, 30, 32, 37 in Table S5)”.

      (4) Also In GWAS Results, a locus Dlx5 is called "SHFM" in the Supplementary Table.

      Background:

      We identified 41 genome-wide significant loci and named the locus after the gene closest to the top lead SNP (bold in Figure 3C). Other genes in each locus for which genome-wide significant SNPs were eQTLs, are listed below the closest gene in normal font (Figure 3C).

      In table S5, we report details of the top lead SNP for all 41 loci. We report only the nearest gene to the top lead SNP.

      For locus 14, SHFM1 is the closest gene to the top lead SNP whereas DLX5 and DLX6 are genes in the locus for which genome-wide significant SNPs were eQTLs. This explains why DLX5 appears under SHFM1 in Figure 3C, but does not appear in Table S5.

      (5) Please reword MRI jargon - "Dixon method", vertix - should be introduced, as well as abbreviation "KDE".

      We had recognised that the word “vertex” would be confusing and had replaced it by datapoint, but had unfortunately missed one occurrence in the text. This is now corrected.

      Thank you for pointing out the lack of introduction of the term “KDE”. This was only explained in the supplementary materials, but has now been added to the main text:

      Line 500-508: “To ensure between-subject comparability, we used the intensity normalised nu.mgz volume output by FreeSurfer. We validated this approach through comparison with well-established intensity normalization methods; Kernel Density Estimation (KDE), WhiteStripe (WS), Gaussian Mixture Model (GMM), Fuzzy C-Means (FCM), and Z-score normalization (ZS). We found the highest test-retest reliability with our approach (Figure S9), and, together with KDE (based on reference signal intensity in WM), the highest correlation with quantitative T1 relaxation maps (Figure S10).”

      Reviewer #2 (Recommendations for the authors):

      (1) This would be an extremely useful advance for the bone and BMA fields if only it could be confirmed that the T1W signal intensity is actually measuring BMA in some meaningful way. Or, even if not BMA, to confirm what other biological characteristic(s) it is in fact capturing. This is essential for interpreting the findings.

      (2) I note that you have compared the normalized T1W parameter with quantitative T1 data (e.g. Figure S10). However, this doesn't address the fundamental issue, because even these quantitative T1 data (e.g. from MP2RAGE) may not be measuring calvarial BMA. T1W sequences have been used to estimate BM cellularity (if not BMA directly) but are not nearly as precise as water-fat imaging. For example, one study found a reasonable correlation (0.71) between T1 relaxation times and BM fat (https://www.nature.com/articles/s41598-019-57030-5). So, if your normalized T1W parameter shows a correlation of -0.44 with the T1 MP2RAGE MRI signal (Figure S10), what does this mean in terms of how well your parameter reflects the actual BMA adiposity? We can't know this, because we also don't know if the T1 MP2RAGE signal reflects calvarial BMA.

      (3) I think my recommendations are clear from the public review. Ideally, you would be able to compare the skull BM normalized T1W parameter with PDFF data that have T2* correction (since the skull BM cavity is quite small and so may suffer from T2* effects relating to tissue inhomogeneity). But even if you had only dual-echo BMFF data, this would still be much more informative than relying only on T1 data. I hope this can be done so that the findings of the study can be properly interpreted.

      As explained above, we reject this reviewer’s claim that T1-weighted signal intensity cannot be used to perform a GWAS of BMA: other studies have shown that T1-weighted signal intensity is a semi-quantitative measure of fat fraction, we have performed extra analyses that confirm this, and our results further demonstrate this. For further detail on why we reject this criticism, see our response to this reviewer’s comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism, showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine, and a mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines, and their intermediates and formate. Single-cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions, such as respiratory chain complex assembly, and mitochondrial functions, with downregulation of genes related to amino acid transport, cellular growth, and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells, highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids are a marvel of nature. It is fascinating that nature decided to make two versions of the same molecule, in this case, an amino acid. While the L-stereoisomer plays well-known roles in biology, the D-stereoisomer seems to function in obscurity. Research into these novel signaling molecules is gathering momentum, with newer stereoisomers being discovered. D-serine has been the most well-studied among the different stereoisomers, and we still continue to learn about this novel neurotransmitter. The roles of these molecules in the context of metabolism is not well studied. The authors aim to elucidate the metabolic role of D-serine in the context of neuronal maturation with implications for 1C metabolism and in cell proliferation. The metabolic role of these molecules is just beginning to be uncovered, especially in the context of mammalian biology. This is the strength of the manuscript. The authors have done important work in prior publications elucidating the role of D-amino acids. The advancement of the field of D-amino acids in mammalian biology is significant, as not much is known. The presentation of RNA seq data is a valuable resource to the community, however, with caveats as mentioned below.

      Weaknesses:

      The following are some of the issues that come out in a critical reading of the manuscript. Addressing these would only strengthen and clarify the work.

      (1) Kinetic assessment of D-serine versus L-serine: While the authors mention that D-serine is not a good substrate for SHMT2 compared to L-serine, the kinetic data are presented for only D-serine. In a substrate comparison with an enzyme, data must be presented for L-serine as well to make the conclusion about substrate specificity and affinity. Since the authors talk about one versus another substrate, there needs to be a kinetic comparison of both with Km (affinity). (Ref Figure 2 panel).

      We agree with the reviewer that kinetic parameters for l-serine are important for evaluating the substrate specificity of SHMT2. The hydroxymethyl-transferase activity of SHMT2 toward l-serine has been previously characterized by our co-author Tetsuya Miyamoto (Miyamoto et al., FEBS Journal, 2024; PMID: 37700610), which is appropriately cited in the manuscript (line 143). In that study, the kinetic parameters for l-serine were determined, with a Km of 0.07 ± 0.009 mM and a Kcat of 33.5 ± 0.9 min<sup>-1</sup>. The strong chiral selectivity of SHMT2 for the l-enantiomer in the hydroxymethyl-transferase reaction was also demonstrated in that work. On the other hand, the primary aim of the present study is different from characterizing d-serine as a catalytic substrate for SHMT2. Rather, our goal was to determine whether d-serine interferes with the hydroxymethyl-transferase reaction of SHMT2 by interacting with the l-serine binding site. Accordingly, the analyses shown in Fig. 2B-D were designed to evaluate whether d-serine could structurally occupy or interfere with the l-serine binding pocket of SHMT2. Therefore, our experiments focused on assessing the potential inhibitory effect of d-serine rather than performing a full kinetic comparison of d-serine and l-serine as substrates.

      (2) Molecular Dynamics simulations, while a good first step in modeling interactions at the active site, rely on force fields. These force fields are approximations and do not represent all interactions occurring in the natural world. Setting up the initial conditions in the simulations can impact the final results in non-equilibrium scenarios. The basic question here is this: Is the simulated trajectory long enough so that the system reaches thermodynamic equilibrium and the measured properties converge? Prior studies have shown mixed results with the conclusion that properties of biological systems tend to converge in multi-second trajectories (not nanosecond scales as reported by the authors) and transition rates to low probability conformations require more time. (Ref Figure 2C).

      We thank the reviewer for raising the important point regarding the limitations of molecular dynamics (MD) simulations, including the dependence on force fields and the potential effects of simulation length and initial conditions. We agree that MD simulations represent approximations of molecular behavior and that longer trajectories may be required to fully explore rare conformational states in biological systems.

      In the present study, however, the MD simulations were not intended to provide a comprehensive thermodynamic description of SHMT2 conformational dynamics. Rather, they were used as a structural assessment to evaluate whether d-serine could plausibly occupy the canonical l-serine binding site of SHMT2. As shown in Fig. 2BC and supplementary movie 1, the simulations did not support stable occupation of the l-serine binding pocket by d-serine. Importantly, this structural observation is consistent with our biochemical data showing that d-serine does not inhibit the hydroxymethyl-transferase activity of SHMT2 when l-serine is used as the substrate (Fig. 2D). Together, these results indicate that the inhibitory effect of d-serine on one-carbon metabolism is unlikely to be mediated through direct inhibition of SHMT2.

      As the reviewer correctly notes, it remains possible that d-serine interacts with SHMT2 at sites distinct from the canonical l-serine binding pocket and could exert potential allosteric effects. Indeed, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that SHMT2 exhibits dehydratase activity toward d-serine. However, this reaction is not directly linked to mitochondrial one-carbon metabolism. Therefore, further extensive simulations exploring alternative conformational states or potential allosteric interactions would extend beyond the scope of the present study.

      Importantly, the key conclusions of this study do not rely solely on MD simulations but are supported by multiple independent experimental approaches, including metabolomics, enzymatic assays, and mitochondrial transport analyses.

      (3) The authors use N2a cell line to demonstrate D-serine burden on primary cortical neurons. N2a is an immortalized cell line, and its properties are very different from primary neurons. The authors need to mention a rationale for the use of an immortalized cell line versus primary neurons. The transcriptomic profile of an immortalized cell line is different compared to a primary cell. Hence, the response to D-serine may vary between the two different cell types.

      We thank the reviewer for raising this important point regarding the differences between immortalized cell lines and primary neurons. As the reviewer notes, N2a cells and primary cortical neurons (PCNs) differ in several aspects, including their degree of differentiation and proliferative capacity. We appreciate the opportunity to clarify the rationale for using both systems in this study.

      One-carbon metabolism is known to be particularly active in highly proliferative or relatively undifferentiated cells. Primary cortical neurons are initially obtained as immature neuronal populations and gradually undergo maturation during culture (Fig. S6). In our experiments, we observed that sensitivity to d-serine and dependence on one-carbon metabolism were primarily evident in immature neuronal states rather than in fully mature neurons (Fig. 4EF).

      In this context, immature PCNs share certain metabolic characteristics with proliferative neural cell lines such as N2a cells. Consistent with this idea, the inhibitory effects of d-serine on one-carbon metabolism and cell proliferation were observed in both immature PCNs and N2a cells (Fig. 1E–G, Fig. 2E, Fig. 3AB, and Fig. 4F). Thus, the use of N2a cells provides a complementary experimental model for studying the metabolic vulnerability of immature neural cells that depend on one-carbon metabolism. Importantly, the key findings were consistently reproduced in primary cortical neurons, supporting the physiological relevance of the observations made in N2a cells. For clarity, we added descriptions in lines 107-108 and 167-168 in our revised manuscript.

      (4) In Figure 4D, the authors mention that D-serine activates the cleavage of caspase 3. Figure 4D shows only cleaved caspase 3 as a single band. They need to show the full blot that contains the cleaved fragments along with the major caspase 3 band.

      In our experiments, we used an antibody that specifically recognizes cleaved caspase-3 and does not recognize full-length caspase-3 (Cell Signaling Technology, anti-cleaved caspase-3 antibody, clone 5A1E). Therefore, the Western blot detects only the cleaved caspase-3 fragment at approximately 17 kDa, which appears as a single band in the blot (please see Author response image 1). For clarity, we have also included the antibody information in the Western blot section in the revised Materials and Methods (lines 501-502).

      Author response image 1.

      An original image of western blot for Fig. 4D. An arrow indicates the bands of cleaved caspase-3 (17 kDa).

      (5) In Figure panel 4, the authors use neural progenitor cells (NPCs). They need to demonstrate that the population they are working with is NPCs and not primary neurons. There must be a figure panel staining for NPC markers like SOX2 and PAX6. Also, Figure S5 needs to be properly labeled. It is confusing from the legend what panels B-E refer to? Also, scale bars are not indicated.

      We thank the reviewer for pointing out these issues. First, we have relabeled and rearranged the figure panels in revised Figure S6B-E, and revised the figure legend to provide a clearer description of each panel. We have also added scale bars to the revised images.

      To confirm the identity of neural progenitor cells (NPCs), we performed immunostaining for Nestin, a well-established marker for NPCs, instead of Sox2 and Pax6 (Fig. S6B). Nestin is widely used as a marker for NPCs(Bernal and Arranz, 2018; Bott et al., 2019; Lendahl et al., 1990), and is known to be co-expressed with Sox2 in the mouse embryonic brain(Graham et al., 2003). In addition, Nestin has been identified as a Pax6-bound gene associated with the transcriptional program regulating neural progenitor identity (Thakurela et al., 2016). Consistent with these observations, analysis of published scRNA-seq data from the developing mouse brain (Bella et al., 2021) shows that the RNA expression profile of Nestin (Nes) closely parallels those of Sox2 and Pax6 (new Fig. S9C). Based on these lines of evidence, we consider Nestin-positive cells in our cultures to represent NPCs. Importantly, Nestin-positive cells were co-stained with cleaved caspase-3, suggesting that the apoptotic population corresponds to NPCs (Fig. S6B). Together, these observations support that the apoptotic cells observe in our culture correspond to NPCs rather than differentiated neurons.

      (6) In Supplementary Figure panel 7F, the authors mention phosphatidyl L-serine and phosphatidyl D-serine. A chromatogram of the two species would clarify their presence as they used 2D-HPLC. On an MS platform, these 2 species are not distinguishable. Including a chromatogram of the 2 species would be helpful to the readers.

      We thank the reviewer for this helpful suggestion. As the reviewer correctly noted, mass spectrometry alone cannot distinguish phosphatidyl-d-serine and phosphatidyl-l-serine. To quantify these species separately, lipids were first extracted from cells using the Bligh and Dyer method, followed by phospholipase D treatment to cleave the serine moiety from phosphatidylserine (a schematic of the procedure is shown in Fig. S8D). The released serine was then derivatized with NBD-F and analyzed by 2D-HPLC for enantioselective separation and quantification of d- and l-serine, as described in the Materials and Methods section (“Quantification of glycine and serine enantiomers”). Phosphatidyl-l-serine (Sigma-Aldrich: P0474) was used to generate a standard curve for quantification (Fig. S8E). Because phosphatidyl-d-serine is not commercially available, and because the peak height of free d-serine is equivalent to that of free l-serine in the chromatograms of our 2D-HPLC system, the same standard curve was used to estimate phosphatidyl-d-serine levels.

      As requested by the reviewer, we have now added representative chromatograms of d- and l-serine derived from phosphatidyl-serine in NPC samples (new Fig. S8F).

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development, which appears to be a discussion of prior published data from Hubbard et al, 2013, Burk et al, 2020, and Bella et al, 2021, in Supplementary Figure panels 8 A-E. This should not be presented as a figure panel, as it gives the false impression that the authors have performed the experiment, which is clearly not the case. However, its discussion can well serve as part of the manuscript in the discussion section.

      We appreciate this helpful suggestion. The datasets used in Fig. 4K and Fig. S9 were derived from previously published transcriptomic studies (Hubbard et al, 2013; Burk et al, 2020; Bella et al, 2021). While these studies reported transcriptomic profiles during neuronal development in vitro or in vivo, they did not specifically analyze serine metabolism, one-carbon metabolism, or d-serine biosynthesis, which are the focus of the present study. Therefore, we reanalyzed these publicly available datasets from a metabolic perspective, focusing on genes involved in serine metabolism and one-carbon metabolism. This re-analysis allowed us to examine a developmental shift in serine enantiomer metabolism, we presented the results as figure panels rather than simply citing the datasets in the Discussion. Re-analysis of publicly available transcriptomic datasets to address new biological questions has become a common approach in genomics and transcriptomics studies.

      To avoid the impression that these experiments were performed in this study, we have clearly indicated the original references and clarified this point in the figure legends (Fig. 4K and Fig. S9). These panels are intended to provide supportive evidence for the developmental shift in serine enantiomer metabolism discussed in this study.

      (8) The entire presentation of the section on enantiomeric shift of serine metabolism during neural development (lines 274-312) is a discussion and should be part of the discussion section and not in the results section. This is misleading.

      We thank the reviewer for this comment. This point overlaps with the concern raised in comment 7. Please see our response to comment 7 for a detailed explanation and the revisions made in the manuscript.

      (9) The discussion section is not well written. There is no mention of recent work related to D-serine that has a direct bearing on its metabolic properties. In the discussion section, paragraph 1, the authors mention that their work demonstrates the selective synthesis of D-serine in mature neurons as opposed to neural progenitor cells. This concept has been referred to in prior publications:

      (a) Spatiotemporal relationships among D-serine, serine racemase, and D-amino acid oxidase during mouse postnatal development. PMID:14531937.

      (b) D-cysteine is an endogenous regulator of neural progenitor cell dynamics in the mammalian brain. PMID:34556581.

      We thank the reviewer for this helpful suggestion and for drawing our attention to these studies. We have revised the Discussion to better place our findings in the context of previous work related to d-serine.

      Specifically, we have added references describing the spatiotemporal relationship between d-serine and serine racemase (Srr) during brain development (PMIDs 14531937 and 33592203) (line 333). These studies highlighted the tissue-level (PMID 14531937) and cellular-level (PMID: 33592203) relationship between Srr expression and development. These studies highlighted the role of d-serine in supporting the functional maturation of neurons during postnatal development. In contrast, our study addresses a complementary question of why d-serine is NOT present during embryonic and early postnatal stages of brain development, when proliferative metabolic activity is high. This question is fundamentally different from the previous reports. Our point is that d-serine is not favorable because it interferes with one-carbon metabolism, which is essential for cell proliferation. Therefore, this concept has not been referred to in prior publications. To clarify this point, we added the following sentence to the Discussion (line 325-328): “In addition to the known role of d-serine in the functional maturation of differentiated neurons, our findings highlight a previously unrecognized, stereoselective regulation of cellular metabolism by d-serine, and provide a rationale for its selective synthesis in mature neurons where proliferative metabolic activity is no longer required’.

      We also appreciate the reviewer bringing our attention to the study describing d-cysteine as a regulator of neural progenitor cell dynamics (PMID: 34556581). We have added the description regarding the overlapping and distinct functions of d-cysteine and d-serine to the Discussion (lines 418-436).

      (10) In the abstract, in lines 101 and 102, the authors mention "how d-serine contributes to cellular metabolism beyond neurotransmission remains largely unknown". In 2023, a paper in Stem Cell Reports by Roychaudhuri et al (PMID:37352848) showed that d and l-serine availability impacts lipid metabolism in the subventricular zone in mice, affecting proliferative properties of stem-cell derived neurons using a comprehensive lipidomics approach. There is no mention of this work even in the discussion section, as it bears directly on l and d-serine availability in neurons, which the authors are investigating. In the discussion section in lines 410-411, the authors mention the role of d-serine in neurogenesis, but surprisingly don't refer to the above reference. The role of d-serine in neurogenesis has been demonstrated in the Sultan et al (lines 855-857) and Roychaudhuri et al references.

      We thank the reviewer for highlighting these relevant studies. We have revised the statement in line 99 to avoid overgeneralization and to reflect that the metabolic roles of d-serine are incompletely understood rather than largely unknown. In addition, we have incorporated discussion of previous works, including the work by Roychaudhuri et al. (2023), into the Discussion section (lines 418-436) of our revised manuscript.

      (11) Both D-serine and the structurally similar stereoisomer D-cysteine (sulfur versus oxygen atom) have a bearing on 1C metabolism and the folate cycle. With reference to the folate cycle, Roychaudhuri et al in 2024 (PMID:39368613) have shown in rescue experiments in mice that supplementing a higher methionine diet provides folate cycle precursors to rescue the high insulin phenotype in SR-deficient mice. Since 1C metabolism is being discussed in this manuscript, the authors seem to overlook prior work in the field and not include it in their discussion, even when it is the same enzyme (SR) that synthesizes both serine and cysteine. Since the field of D-amino acid research is in its infancy, the authors must make it a point to include prior work related to D-serine at least, and not claim that it is not known. The known D-stereoisomers are not many, hence any progress in the area must include at least a discussion of the other structurally related stereoisomers.

      We are grateful to the reviewer for drawing our attention to this relevant study.

      We have now incorporated the findings from Roychaudhuri et al. 2024 into the Discussion. In that study, Srr-/- mice exhibited reduced levels of DNMT1 and DNMT3A, resulting in reduced DNA methylation activity. Notably, supplementation with a methyl-donor diet (containing choline, betaine, and methionine) restored the aberrant insulin phenotype in Srr-/- mice. These findings are relevant to our study, as DNA methylation depends on S-adenosyl-methionine (SAM), which is generated through one-carbon metabolism.

      We have expanded the Discussion to include the relationship between d-serine and d-cysteine, both of which are synthesized by Srr in the revised manuscript (lines 418-436).

      (12) Racemases (serine and aspartate) in general are promiscuous enzymes and known to synthesize other stereoisomers in addition to D-serine, D-cysteine, and D-aspartate. A few controls, like D-aspartate, D-cysteine, or even D-alanine must be included in their study to demonstrate the specific actions of D-serine, especially in the N2a cell treatment experiments. Cysteine and Serine are almost identical in structure (sulfur versus oxygen atom), and both are synthesized by serine racemase (published). Cysteine has also been very recently shown to inhibit tumor growth and neural progenitor cell proliferation. (PMIDs: 40797101 and 34556581). How the authors' work relates to the existing findings must be discussed, and this would put things in perspective for the reader.

      We thank the reviewer for this comment regarding the need to demonstrate the specificity of d-serine. As shown in Fig. S7A, d-serine, but not other d-amino acids commonly detected in mammals (Gonda et al., 2023), including d-aspartate, d-alanine, and d-proline, induced the cleavage of caspase-3 under the same experimental conditions, supporting the specific effect of d-serine in our system. We agree that d-cysteine shares structural similarity with d-serine and is also synthesized by Srr, suggesting potential functional overlap with d-serine. To place our findings in this context, we have added a paragraph in the Discussion (lines 418-436) describing the similarities of d-serine and d-cysteine. While both molecules may exert anti-proliferative effects, their underlying mechanisms appear to differ. Notably, supplementation with SAM, methionine, or glutathione did not rescue d-serine-induced growth inhibition in our system (Fig. S5J), suggesting that its effects are not primarily mediated through methylation or sulfur metabolic pathways. Instead, d-serine suppresses cellular proliferation by limiting mitochondrial l-serine availability and one-carbon metabolism. These observations highlight mechanistic divergence between d-serine and d-cysteine.

      Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition, and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in the central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine not only serves as a traditional neurotransmitter in the central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on the one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Weaknesses:

      (1) The detailed mechanism by which D-serine competes with L-serine for its mitochondrial transport is not investigated. For example, although the authors made some discussion, they did not provide direct genetic or biochemical evidence linking these effects to the specific transporters, such as SFXN1.

      We thank the reviewer for this important comment regarding the mitochondrial l-serine transport mechanism. To address this point, we performed additional experiments using N2a cells in which Sfxn1 was knocked down by siRNA. Under semi-permeabilized cell conditions, we newly examined the effect of d-serine on mitochondrial L-serine transport.

      Interestingly, even under conditions where Sfxn1 expression was markedly suppressed, d-serine still inhibited mitochondrial d-serine transport. Given that SFXN1 is known to function redundantly with its paralogs (SFXN2–SFXN5) in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only through SFXN1 but potentially through multiple members of the SFXN transporter family. These new data have been added as Fig. S3, and the corresponding results and discussion have been incorporated into the revised manuscript (lines 158–165 and 379-385).

      (2) Unlike tumor cells, where SHMT2 usually plays a predominant role in catalyzing serine/THF-derived one-carbon metabolism, normal cells may employ both SHMT1 and SHMT2 to do the work. Even under certain conditions that SHMT2-mediated one-carbon metabolism is suppressed, the activity of SHMT1 could be elevated for compensation. Thus, it is important to investigate whether D-serine affects SHMT1 activity or changes the balance between SHMT1- and SHMT2-mediated one-carbon metabolism. To this aim, the authors are strongly encouraged to perform a metabolic flux assay (MFA) by using 13C-labeled L-serine in the model cells in the presence and absence of D-serine.

      We thank the reviewer for this thoughtful comment regarding the potential contribution of cytosolic SHMT1 to one-carbon metabolism. As the reviewer notes, while mitochondrial SHMT2 is generally considered the predominant enzyme supporting one-carbon metabolism in proliferating or tumor cells, SHMT1 in the cytosol may function in a complementary manner in normal cells. In our experiments using primary cortical neurons, we indeed observed that the sensitivity to d-serine differed between immature and mature neuronal states (Fig. 4E). This observation suggests that the relative contribution of SHMT1- and SHMT2-mediated one-carbon metabolism may vary depending on the differentiation status of the cells.

      However, the primary focus of the present study was to investigate the mechanism underlying the anti-proliferative effect of d-serine in proliferative or undifferentiated neural cells. Our data demonstrate that d-serine inhibits one-carbon metabolism primarily by limiting mitochondrial l-serine availability through inhibition of mitochondrial l-serine transport. Therefore, a detailed analysis of the compensatory balance between SHMT1 and SHMT2 after disruption of mitochondrial serine transport falls beyond the central scope of the present study. Importantly, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that both SHMT1 and SHMT2 exhibit strong stereoselectivity for l-serine in their hydroxymethyl-transferase activity and do not utilize d-serine as a substrate. These findings make a direct effect of d-serine on hydroxymethyl-transferase activity of SHMT1 unlikely.

      Nevertheless, we agree that differences in the anti-proliferative effects of d-serine across cell types or differentiation states could reflect variations in cellular dependence on one-carbon metabolism or potential compensation by SHMT1. To address this point, we have expanded the Discussion section to clarify the possible contribution of SHMT1 and the limitations of the present study (lines 367–376). While isotope tracing analysis would be valuable for further dissecting compartmentalized one-carbon fluxes, such analyses would primarily address the relative contributions of SHMT1 and SHMT2 rather than the mitochondrial serine transport step that constitutes the central mechanism identified in this study.

      (3) A defect in serine-derived one-carbon metabolism may cause multiple cellular stress responses. It is valuable to detect whether cellular NADPH/NADH, GSH, or ROS is altered before and after D-serine treatment.

      We appreciate this insightful comment regarding potential cellular stress responses associated with impaired one-carbon metabolism. Consistent with the reviewer’s suggestion, we examined markers related to redox status. Transcriptomic analysis revealed compensatory changes in genes involved in mitochondrial metabolic function, including components of the NADH dehydrogenase complex (Fig. 2IJ and Fig. S4A), suggesting metabolic adaptation to d-serine treatment. In response to the reviewer’s comment, we measured GSH levels in NPCs and found that d-serine treatment led to a reduction of GSH (new Fig. S7F), indicating altered redox balance. However, supplementation with exogenous GSH did not rescue d-serine-induced cell death (Fig. S7E). These results suggest that while d-serine induces changes in cellular redox status, including GSH depletion, redox imbalance alone is unlikely to be the primary driver of cell death in this context.

      (4) The physiological relevance between D-serine and neural cell maturation/death should be further tested and discussed, since the dosage of D-serine used in the in vitro assay is much higher than that in physiological conditions.

      We thank the reviewer for this comment regarding the physiological relevance of the d-serine concentrations used in our study. We agree that the concentrations of d-serine required to compete with l-serine for mitochondrial transport are higher than those typically observed under physiological conditions, which we acknowledged in the Discussion of our manuscript (lines 386-388). Importantly, however, in vivo, d-serine levels are tightly regulated in a spatiotemporal manner during brain development, with low levels during embryonic and early postnatal stages and increased levels upon neural maturation. This temporal regulation coincides with the transition from proliferative neural progenitor states to differentiated neurons, thereby limiting the potential for d-serine to interfere with one-carbon metabolism during periods of active cell proliferation. Thus, while the concentrations used in vitro may exceed physiological levels, they allow us to uncover a latent metabolic effect of d-serine that may become relevant under specific cellular or developmental contexts. To clarify these points, we have revised the Discussion (lines 388-395) in the revised manuscript to more explicitly address the relationship between d-serine dosage and physiological relevance.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      We thank the reviewer for this insightful comment regarding the identity of the mitochondrial l-serine transporter. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed. Briefly, we conducted new experiments using siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Notably, even with marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that SFXN family members (SFXN1–SFXN5) are reported to function redundantly in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only via SFXN1 but potentially across multiple SFXN paralogs. These results have been incorporated into the revised manuscript and are presented in Fig. S3, with the corresponding discussion added to the Results and Discussion sections (lines 158–164 and 379-385).

      (2) The effective concentrations of D-serine used in vitro (IC<sub>50</sub> ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments - such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools - and the mechanisms governing their regulation - may further influence its metabolic impact in vivo.

      We appreciate this insightful comment regarding the physiological relevance of the d-serine concentrations used in vitro. We agree that the effective concentrations observed in our assays exceed typical bulk brain levels. To address this point, we have expanded the Discussion (lines 386-399) to consider conditions under which locally elevated d-serine concentrations may arise in vivo. These additions provide a more nuanced interpretation of the relationship between the concentrations used in vitro and the potential physiological contexts in which d-serine may exert metabolic effects.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

      We thank the reviewer for this comment regarding the potential generality of the anti-proliferative effects of d-serine. We agree that it is important to clarify whether the observed effects are specific to neural lineages or reflect a broader vulnerability of proliferative cells. In our study, we focused on neural progenitor cells and neural tumour models. However, the mechanism identified here, namely limitation of mitochondrial L-serine availability, targets a fundamental metabolic pathway that supports cell proliferation. Therefore, it is possible that similar effects may extend to other highly proliferative cell types beyond the neural lineage. To address this point, we have expanded the Discussion (lines 408-410) to clarify that the observed effects of d-serine may reflect a general metabolic vulnerability associated with proliferative states, while also noting that the degree of sensitivity is likely to depend on cell-type-specific reliance on mitochondrial one-carbon metabolism.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor issues:

      The authors mention terms like neural tissue (line 109) and neural tumor cells (line 182). Neural tissue can mean anything under the sun. They need to mention the specific tissue being studied or investigated.

      We have changed the terms.

      Reviewer #2 (Recommendations for the authors):

      (1) The figure items were not well organized, and the legends were difficult to read since they apparently lacked key information. For example, in Figure 2D, did the author intend to show SHMT2 activity? It is not clear how this assay was performed (in vitro or in vivo experiment?). Additionally, more detailed information should be provided in the Methods section.

      We have improved the organization of figures and revised figure legends for better readability.

      (2) The writing of the manuscript should be significantly improved, using a professional editing service, in order to increase the readability.

      We appreciate the reviewer’s comment. The manuscript was professionally edited prior to submission. Nevertheless, we have made targeted revisions to the Introduction and Discussion sections to improve overall readability. We hope that these revisions have improved the clarity of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The core mechanism centers on competition for mitochondrial L-serine transport, yet the identity of the transporter(s) involved remains speculative. While Kory et al. (2018) identified SFXN1 as a mitochondrial L-serine transporter, this connection is not directly addressed in the current study. It would strengthen the manuscript to clarify whether SFXN1 or related isoforms are expressed in the neural cell models used and whether their expression patterns correspond with the observed D-serine sensitivity. Even if functional validation is beyond the current scope, a more detailed discussion of potential transporter candidates and the limitations of existing data would provide important mechanistic context and help frame future directions.

      We thank the reviewer for this insightful comment regarding the mitochondrial l-serine transporter and the potential involvement of SFXN family members. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed.

      Briefly, we performed siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Even under conditions of marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that members of the SFXN family have been reported to function redundantly in mitochondrial serine transport (e.g., Kory et al., 2018), these findings suggest that d-serine may affect l-serine transport not only through SFXN1 but potentially across multiple SFXN paralogs.

      These new results have been incorporated into the revised manuscript (Fig. S3), and the relevant discussion has been expanded to clarify the potential roles of SFXN family transporters and the current limitations in defining the exact transporter responsible (lines 158–165 and 379-385).

      (2) The authors use ex vivo brain slice cultures with tumor xenografts to demonstrate tissue-level relevance, which is a valuable strength of the study. However, additional context would enhance its translational significance. It would be helpful to discuss whether in vivo D-serine administration (e.g., ICV or systemic) is feasible and safe, especially given the high concentrations required in vitro. Briefly addressing whether genetic models, such as Srr knockout mice, support a role for D-serine in tumor progression or neurodevelopment would also strengthen the interpretation.

      We thank the reviewer for this important comment regarding the translational relevance of d-serine administration in vivo. To address this point, we performed additional exploratory experiments using a subcutaneous tumour xenograft model in nude mice. Because the inhibitory effect of d-serine on one-carbon metabolism becomes evident under l-serine–limited conditions, tumor-bearing mice were fed an l-serine/glycine–deficient diet and administered d-serine in drinking water.

      However, when d-serine was provided at high concentrations (≥ 500 mM) in drinking water, the mice exhibited marked behavioral abnormalities, including increased aggression and other abnormal behaviors. Due to these adverse effects, the experiment was ethically terminated. While our study focuses on the inhibitory effect of d-serine on one-carbon metabolism, d-serine is also a physiological co-agonist of the NMDA receptors, and therefore high systemic concentrations may influence neuronal excitability in vivo. These observations suggest that the concentrations required to achieve antitumor effects may be associated with significant neurological side effects.

      Based on these findings, we consider that direct administration of d-serine itself may have limited therapeutic applicability as an antitumor reagent. Instead, future development of derivatives or strategies that retain the metabolic inhibitory effect while minimizing NMDA receptor–mediated effects may be required.

      Regarding serine racemase (Srr) knockout models, xenograft tumour experiments would require an immunodeficient background, which makes the generation and use of such compound models technically challenging. Therefore, we did not pursue this approach in the present study. We have incorporated these considerations into the revised Discussion (lines 413-417) to clarify the translational implications and current limitations of our findings.

      Bella DJD, Habibi E, Stickels RR, Scalia G, Brown J, Yadollahpour P, Yang SM, Abbate C, Biancalani T, Macosko EZ, Chen F, Regev A, Arlotta P. 2021. Molecular logic of cellular diversification in the mouse cerebral cortex. Nature 595:554–559. DOI: https://doi.org/10.1038/s41586-021-03670-5, PMID: 34163074

      Bernal A, Arranz L. 2018. Nestin-expressing progenitor cells: function, identity and therapeutic implications. Cellular and Molecular Life Sciences 75:2177–2195. DOI: https://doi.org/10.1007/s00018-018-2794-z, PMID: 29541793

      Bott CJ, Johnson CG, Yap CC, Dwyer ND, Litwa KA, Winckler B. 2019. Nestin in immature embryonic neurons affects axon growth cone morphology and Semaphorin3a sensitivity. Molecular Biology of the Cell 30:1214–1229. DOI: https://doi.org/10.1091/mbc.e18-06-0361, PMID: 30840538

      Graham V, Khudyakov J, Ellis P, Pevny L. 2003. SOX2 Functions to Maintain Neural Progenitor Identity. Neuron 39:749–765. DOI: https://doi.org/10.1016/s0896-6273(03)00497-5, PMID: 12948443

      Lendahl U, Zimmerman LB, McKay RDG. 1990. CNS stem cells express a new class of intermediate filament protein. Cell 60:585–595. DOI: https://doi.org/10.1016/0092-8674(90)90662-x, PMID: 1689217

      Thakurela S, Tiwari N, Schick S, Garding A, Ivanek R, Berninger B, Tiwari VK. 2016. Mapping gene regulatory circuitry of Pax6 during neurogenesis. Cell Discovery 2:15045. DOI: https://doi.org/10.1038/celldisc.2015.45, PMID: 27462442

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs), providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      We thank the reviewer for this evaluation

      Weaknesses:

      (1) Reliance on GAL induction introduces metabolic shifts (raffinose -> galactose -> glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. Alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be discussed as a way to avoid carbon-source changes.

      Indeed, this could be an improvement, and we mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3.

      (2) While proteomics is solid, a comprehensive supplementary table listing all identified proteins (with enrichment and statistics) would enhance transparency.

      Indeed, we now provide source data showing LFQ intensities, fold-enrichment and statistics for all detected proteins.

      (3) Importantly, the authors note that the method is particularly useful "in conditions where direct tagging of Nup84 interferes with its function, while sub-stoichiometric nanobody binding does not." After this sentence, it would be valuable to add concrete examples, such as experiments examining NPC integrity in aging or stress conditions where epitope tags can exacerbate phenotypes. These examples will help readers identify situations in which this approach offers clear advantages.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants, and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks, and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ. While useful applications including timelapse imaging, affinity purification, or proximity labeling are envisioned, addressing some outstanding technical questions would give a clearer picture of the sensitivity and temporal resolution of this approach.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

      We thank the reviewer for this evaluation

      Weaknesses:

      The decrease in nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over during this time. However, it is also possible that the nanobody: Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Figure 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      We thank the reviewer for this evaluation

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful for the authors to verify that their pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs.

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure but consider such studies to be beyond the scope of this study.

      Overall, the data clearly show that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      We thank the reviewer again for the constructive feedback and thoughts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 1A, and although it is partially mentioned in the legend, it would be helpful to indicate precisely when cells are grown in raffinose, when galactose is added for induction, and when glucose is used to terminate expression.

      We included “galactose” and “glucose” to Panel A to indicate induction and termination of expression, respectively.

      (2) Related to the previous point, consider mentioning the GAL4-ER-VP16 (ADGEV) estradiol-inducible system as an optional strategy to avoid carbon shifts and potentially reduce cell-to-cell variability.

      We mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3

      (3) Add a brief sentence explaining that the ZZ tag is derived from Protein A and binds IgG Fc.

      This information is now added on p.2

      (4) The statement "all Nups significantly coenriched with VHH[Nup84]-ZZ..." is likely inaccurate, since not all Nups are labeled in panels F-H, and some basket components are missing in panel I (particularly basket components such as Nup60, Nup1). Consider revising to "most Nups significantly coenriched...". In panel I, please include a clearly non-enriched protein as a visual reference for the color scale.

      We are very grateful to the reviewer for pointing this out. We accidentally used a faulty filtering on the dataset to generate figure panel I, omitting several Nups that were reproducibly found in all replicas. All Nups, except for Gle1 and Pom33, were detected reproducibly.

      We have made the following adjustments to the figure panel and accompanying text:

      In Fig. 1I, we included the missing Nups and 5 proteins that co-purified with VHH[Nup84] but not specifically enriched, as the reviewer suggested. They cluster in a separate group and their abundance is not going up in time. We randomly selected these 5 proteins from the list of genes that were reproducibly found in all four timepoints.

      For clarity, we removed the NTRs

      We changed the text to “we found that all Nups, except Gle1 and Pom33, significantly coenriched with VHH[Nup84]-ZZ” on p.2.

      We updated the methods section, describing the clustering method and how we selected the 5 random proteins

      (5) Provide a supplementary spreadsheet with LFQ intensities, fold-enrichment, and statistics for all detected proteins. This will address questions about missing Nups and support transparency.

      This information is now added as Source data Figure 1.

      (6) Directly after the statement "Amongst others this is useful in conditions where direct tagging of Nup84 interferes with its function, while sub stoichiometric nanobody binding does not," it would be useful to include concrete instances, such as stress or aging conditions, where Nup84 tagging may sensitize NPC integrity.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      (7) In panels K and L, since individual points correspond to biological replicates, overlaying a box plot obscures much of the data. Consider overlaying the means per replicate instead of box plots: see the "SuperPlots" approach for a clear explanation of how to present this (PMID: 32346721).

      We thank the reviewer for the “SuperPlots” suggestion, and we agree that representing the data in this way improves the visualization of individual biological replicates. We have updated the summarizing overlay in figures in panel K and L to represent the means per replicate instead of boxplots.

      (8) I spotted a few typos ("Lasty" ? "Lastly"; "in maintained" vs. "is maintained").

      Thank you, these are corrected

      Overall, this is a neat, well-executed methodological advance with clear value to the NPC field and potentially other complex assemblies. I look forward to seeing a revised version.

      Thank you!

      Reviewer #2 (Recommendations for the authors):

      Based on the recent structural analyses and NPC modeling using this nanobody, how accessible is the Nup84 epitope expected to be within the fully assembled NPC? While the data shown indicate that nanobody labeling of NPCs is readily detectable, stating this clearly would help motivate the approach and interpret the resulting data.

      We now included such a statement in the introduction on p.1.

      The decrease of nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over due to cell division during this time window. However, it is also possible that nanobody:Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      Reviewer #3 (Recommendations for the authors):

      (1) As mentioned above, to assess the general relevance of this tool, it would be informative to verify whether the VHH[Nup84] nanobody can access and detect NPC species under conditions that challenge their structural organization or biogenesis, for example, in nucleoporin mutants or under stress. The authors could, for instance, analyze the localization of VHH[Nup84] in yeast strains harboring clustered NPCs (nup133Δ), or following stresses known to impact NPC organization (e.g., osmotic stress or energy depletion; PMID: 34762489).

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure and tried to include such data. Unfortunately, this was not successful, and further efforts are beyond the scope of his study. Following the reviewer’s suggestion, we expressed VHH[Nup84] in nup133∆N (nup133∆2-300) (Doye, Wepf, and Hurt 1994) following the experimental set-up in panel A and examined its localization. However, at t=2hrs hardly any nanobody signal was detectable in nup133∆N (see Author response image 1, upper panel A) and only after overnight expression nanobody-labelled NPC clusters are detectable (bottom panel A). Considering that expression levels of free mNG are also lower at t=2hrs in nup133∆N cells compared to WT cells (Author response image 1, panel B), it appears that protein expression under the Gal system is generally reduced in a nup133∆N background. These expression level differences between nup133∆N and WT preclude statements about the accessibility of the Nup84 epitope in nup133∆N. We note that nup133∆N cells do not have general mRNA export defects (Doye, Wepf, and Hurt 1994), so other inducible systems may be better suited for such analysis.

      Author response image 1.

      Expression level differences in WT and Nup133∆N cells. Left: localization of VHH[Nup84]-mNG in Nup133∆N cells at t=2hr following a 20-minute induction pulse and after overnight 0.5% galactose (ON) induction. Right: mNG levels in WT and Nup133∆N cells at t=2hr following a 20-minute induction pulse. Brightness/contrast settings are identical between the two panels. All panels are sum slices projections from 30 z-slices of 0.1µm. Scale bar = 5 µm.

      (2) Since outer rings are found on both sides of NPCs (i.e., the cytoplasmic and nuclear faces), could the authors indicate whether the VHH[Nup84] nanobody can enter the nucleus and probe the nuclear outer rings? Along these lines, it would be useful to provide a summary of the structural organization of NPCs in the introduction.

      Thank you, we have added a sentence on the localization of Nup84 in NPCs in the introduction. Based on what is known about influx (nuclear transport receptor-independent nuclear entry) of proteins with similar size and surface properties (Popken et al. 2015; Timney et al. 2016), the nanobody can rapidly enter the nucleus and hence bind Nup84 on both the nuclear and cytoplasmic side. We have no data to answer if binding might initially be biased towards cytosolic VHH[Nup84] binding the cytoplasmic outer rings.

      (3) The authors state that VHH[Nup84] and direct Nup84 detection are indistinguishable (p. 2). Could they provide images of the endogenously tagged Nup84-GFP strain for comparison?

      We have now included a pairwise comparison in a Figure 1 – supplement 1.

      Minor corrections:

      (1) There are a few typos that need correcting: 'Nup84Δ' (p. 1; should read 'nup84Δ') and 'promotor' (p. 2; should read 'promoter').

      Thank you, these are corrected

      (2) The reference 'Veldsink et al. 2025' (quoted in the PunctaFinder analysis description on page 8) does not appear in the References section.

      Thank you, these are corrected.

      We thank the reviewer again for the constructive feedback and thoughts.

      References

      Doye, V., R. Wepf, and E. C. Hurt. 1994. 'A novel nuclear pore protein Nup133p with distinct roles in poly(A)+ RNA transport and nuclear pore distribution', EMBO J, 13: 6062-75.

      Khmelinskii, Anton, Philipp J. Keller, Holger Lorenz, Elmar Schiebel, and Michael Knop. 2010. 'Segregation of yeast nuclear pores', Nature, 466: E1-E1.

      Popken, Petra, Ali Ghavami, Patrick R. Onck, Bert Poolman, and Liesbeth M. Veenhoff. 2015. 'Size-dependent leak of soluble and membrane proteins through the yeast nuclear pore complex', Molecular Biology of the Cell, 26: 1386-94.

      Timney, Benjamin L., Barak Raveh, Roxana Mironska, Jill M. Trivedi, Seung Joong Kim, Daniel Russel, Susan R. Wente, Andrej Sali, and Michael P. Rout. 2016. 'Simple rules for passive diffusion through the nuclear pore complex', Journal of Cell Biology, 215: 57-76.

      Zsok, J., F. Simon, G. Bayrak, L. Isaki, N. Kerff, Y. Kicheva, A. Wolstenholme, L. E. Weiss, and E. Dultz. 2024. 'Nuclear basket proteins regulate the distribution and mobility of nuclear pore complexes in budding yeast', Mol Biol Cell, 35: ar143.

    1. Author response:

      Reviewer #1:

      Major comments

      (1) Although I myself believe that the datasets in this study should be more consistent and comprehensive, the authors should perform a data mining analysis of previously reported transcriptomic changes of these mutants or similar mutants in the same longevity pathway and compare the reported changes with their findings to highlight the necessity and advances of this study.

      According to this suggestion, we have compared the differentially expressed genes identified in this study to previous gene expression studies involving these long-lived mutant strains. To our knowledge no previous studies have examined gene expression in sod-2 or ife-2 mutants, and at the time that we performed the RNA sequencing gene expression in osm-5 worms had not been examined (it took us a long time to complete this paper). We have included weighted Venn diagrams to illustrate the overlap and supplemental tables to list the overlapping di erentially expressed genes. For our current study, we felt it was important to compare RNA-seq data generated under exactly the same experimental and analysis paradigms in order to best compare across the nine long-lived mutants. These new analyses are included in Figures S19 – S25 and Table S2 . Please see lines 111-114, Figure S19-25, and Table S2.  

      (2) This manuscript does not perform any regulon or transcription factor (TF) analyses. TFs are the drivers of the transcriptomic changes and multiple conserved TFs (e.g., daf-16) have already been identified in these pathways. Therefore, it is necessary to examine and compare the regulons/TFs in these new datasets by bioinformatics. Such analyses can: a) provide more information of the driving force of these transcriptomic changes; b) show the role of these known longevity TFs; c) propose new TFs driving longevity; d) support the findings of 'longevity strategies' and 'longevity groups' from the perspective of TFs.

      According to this suggestion, we have now performed transcription factor analysis on the RNA-seq data to determine which transcription factors might be driving the longevity-associated transcriptional changes. To do this we used two complementary approaches: (1) transcription factor inference, which is based on the coordinated expression changes of known transcription factors; and (2) motif enrichment analysis, which is based on identifying transcription factor binding motifs in the promoters of di erentially expressed genes. After identifying which transcription factors were identified for each individual mutant, we then compared the identified transcription factors across all nine mutants. Interestingly, while 33 of the same transcription factors were implicated in group 1 and group 2 longevity mutants, 25 are modulated in different directions (activated in group 1, repressed in group 2 or vice versa) while only 5 are modulated in the same direction. This indicates that although group 1 and group 2 longevity mutants may modulate overlapping pathways to achieve long lifespan, in most cases these pathways are modulated in opposite directions. These new analyses are included in Figure S31 and Table S5. Please see lines 194-208, Figure S31, and Table S5.  

      (3) osm-5 and daf-2 are categorized into two different groups in this study. Since the longevity of cilia (-) mutants is through daf-16, the same master TF driving daf-2 longevity, please perform further analyses or discussion to clarify this issue.

      Loss of daf-16 is generally detrimental to lifespan. Disruption of daf-16 decreases the lifespan of all nine long-lived mutants that we examined (see supplemental table in our review paper PMID:37127095). However, loss of daf-16 also decreases wild-type lifespan. Thus, without further evidence it is hard to distinguish between the loss of daf-16 non-specifically decreasing lifespan verse activation of DAF-16 actually contributing to lifespan extension. In daf-2 mutants and the long-lived mitochondrial mutants there is increased nuclear localization of DAF-16 and upregulation of DAF-16 target genes. The differentially expressed genes in the long-lived mitochondrial mutants exhibit about a 50% overlap with the differentially expressed genes in daf-2 mutants (see Author response image 1). In contrast, osm-5 mutants show upregulation of some DAF-16 upregulated genes, no change in some DAF-16 upregulated genes and downregulation of other DAF-16 upregulated genes (see Author response image 1). Only about 10% of the differentially expressed genes in osm-5 mutants overlap with differentially expressed genes in daf-2 mutants. We believe that these results are consistent with loss of DAF-16 causing a general decrease in lifespan and not specifically contributing to osm-5 longevity. These comparisons will be included in a manuscript that we are currently preparing on osm-5 mutant longevity.

      Author response image 1.

      (4) This manuscript focused on genes whose RNAi suppressed the mutants longevity. Please also use bioinformatics to analyze the functions of those whose RNAi extends the mutants longevity, because these genes could tell the health price these mutants pay and help improve ageing interventions by reducing side effects.

      We perform enrichment analysis for both genes upregulated and downregulated in the long-lived mutant strains. The downregulated genes are involved in translation, ribosome biogenesis and gene expression. For the RNAi screen, we aimed to identify genes that are contributing to longevity and so we looked for a decrease in the lifespan of long-lived mutants when treated with RNAi. We did not screen for genes that extend the long-lived mutants longevity. While we did, nonetheless, identify multiple RNAi clones that increased either daf-2 or nuo-6 lifespan, there were not enough genes to identify any patterns of enrichment.

      (5) (OPTIONAL) I strongly suggest a comprehensive comparison of these transcriptomic changes in long-lived mutants with published age-related transcriptomic changes in wild type worms.

      According to this suggestion, we have now compared the differentially expressed genes that we identified in the nine long-lived mutants with genes that were found to be differentially expressed with aging. Interestingly, the group 2 long-lived mutants show a larger overlap for genes modulated in the opposite direction as aging (genes downregulated during aging are upregulated in eat-2 and osm-5 mutants). We have added this new analysis to our manuscript. Please see lines 210-223, Figure S32 and Table S6.

      Minor comments

      (1) Please further clarify the analysis of DEGs correlated with lifespan extension in Fig. 2 by a depiction. In Fig. 2C and D, please label data dots from different strains with different colors.

      According to this suggestion, each strain has been labelled a different colour.

      (2) In Fig. 3 and S20, please label the percentage of overlapping genes on top of each bars.

      We have now labelled the percentage of overlapping genes in Figure 3 and S20 (now S27).

      Reviewer #2:

      Major comments

      (1) While the authors identified a set of 196 upregulated genes, the rationale for narrowing these down to the three final candidates (C08F11.7, ugt-62, and K05C4.9) is not clearly described. The authors show that genetic inhibition of several genes, including DC2.5, C05B5.5, T07C4.5, and W03B1.7, decreases lifespan in both nuo-6 mutants and wild-type animals. However, the authors did not describe why these additional validated candidates, which also showed significant effects on longevity, were not pursued for further

      characterization. The authors should explicitly state the criteria used to prioritize these three genes over the other validated genes.

      Due to the costs and time involved in generating and characterizing new strains, we decided that we would select three strains to study further as a proof-of-principle. When deciding which genes to study further, we considered several approaches. In the end, we chose to use the strength/reproducibility of the increase in weighted mortality to identify genes with a clear, consistent impact. C05B5.5 and T07C4.5 were ruled out because they had an inconsistent impact on weighted mortality (Figure S28). W03B1.7 was ruled out because it did not have a strong enough e ect on weighted mortality (Figure S28). That narrowed it down to C08F11.7, ugt-62, DC2.5, and K05C4.9. Of those 4, C08F11.7, ugt-62, and K05C4.9 have the greatest consistent impact on weighted mortality (Figure S28) and so these genes were chosen. We have updated the manuscript to include this justification for focussing on C08F11.7, ugt-62, and K05C4.9. Please see lines 273-278.

      (2) The authors conclude that longevity can be mediated by multiple molecular pathways. However, it remains unclear whether these distinct strategies can operate simultaneously or are mutually exclusive. The authors need to test whether lifespan extension in a Group 1 mutant is further enhanced or suppressed by the knockdown of a key Group 2-specific genes. These experiments would help determine these pathways act additively, antagonistically, or as partially redundant survival programs.

      This is an excellent suggestion. While our data identify several genes that are regulated in opposite directions in group 1 and group 2 longevity mutants, we do not yet know the extent to which each of these genes contribute to the longevity of group 1 and group 2 mutants. The three genes that we focused on for further characterization (C08F11.7, ugt-62 and K05C4.9) are upregulated in group 1 longevity mutants but not group 2 mutants. Contrary to what might be expected, RNAi knockdown of these genes does not decrease the lifespan of the group 1 longevity mutant daf-2 but does decrease the lifespan of the group 2 longevity mutant eat-2. We recently reviewed the e ect of di erent resilience pathways on the lifespan of long-lived genetic mutants. Disruption of daf-16, sek-1, skn-1, hsf-1, ire-1 and trx-1 can decrease lifespan in both group 1 and group 2 longevity mutants, but also decreases lifespan in wild-type worms suggesting that at least in some mutants the e ect on longevity may be non-specific. Disruption of hif-1 does not a ect the longevity of group 2 mutants, but does a ect the lifespan of some group 1 mutants (clk-1, isp-1, nuo-6) but not others (daf-2, glp-1). To more definitively answer the question, it would be interesting to cross different combinations of group 1 and group 2 longevity mutants to see the extent to which different longevity groups synergize. This is something we are currently working on for a separate manuscript. We have added these points to the revised manuscript. Please see lines 363-381.

      (3) The authors provide interesting data on overexpression of the three candidate genes. However, whereas C08F11.7 clearly demonstrates both necessity and sufficiency for lifespan extension, overexpression of ugt-62 and K05C4.9 does not independently extend lifespan. To strengthen the manuscript, the authors should expand the discussion of these divergent results and clarify possible explanations.

      According to this suggestion, we have expanded our discussion to discuss possibilities of why these genes might be having different effects on lifespan. Please see lines 411-423.

      (4) Key citations are missing and the authors should add multiple citations including the following ones. Please cite the following paper and discuss the authors' finding with respect to the related work (Lee et al PMID: 40814218). Add citations in the sentence describing changes in the transcriptome of C. elegans associated with age (Lee et al., PMID: 38508494). Furthermore, please cite papers describing the overviews of survival assay using C. elegans (Kwon et al., PMID: 40436148, Hwang et al., PMID: 40436147).

      We have added the suggested citations to the revised manuscript. Please see lines 211 (Ref #39), 307 (Ref #41), 423 (Ref #54) and 436 (Ref #55).

      Minor comments

      (1) To improve readability, please provide the full names for all abbreviations at their first appearance in the manuscript.

      We have added the full names for each abbreviation on first appearance.

      (2) Please ensure that the labels in the figures match the text exactly. For instance, if different promoters are used for generating overexpression animals, it may be helpful to indicate the specific promoter in the figure panel or legend for clarity.

      We have ensured that the nomenclature in the text and the figures is the same. We have noted the promoter used for the overexpression strains in the figure legend.

      (3) For all lifespan and stress resistance assays, please include the total number of animals (n) and the number of independent biological replicates (N) in the figure legends to confirm statistical reliability.

      We have added the number of animals and independent biological replicates to the figures and figure legends.

      (4) Please clearly specify the exact developmental stage of the animals used for the survival assays in the Materials and Methods section.

      We have updated the methods to describe the developmental stages used for the survival assays.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA, and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples, and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time the conservative function of MATR3 has been revealed in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. It also provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine, and human) verification, combined with scRNA-seq, LACE-seq, CO-IP, and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

      We greatly appreciate positive comments and constructive feedback on our manuscript.

      We are encouraged that you recognize the novelty, rigour, and clinical relevance of our study on MATR3 in oocyte development and OMA. We have carefully considered your comments and revised the manuscript accordingly. We will further expand the OMA patient cohort in future studies to verify the causal relationship between MATR3 and OMA.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      We greatly appreciate your constructive comments and suggestions. We are grateful for the recognition of our conditional knockout mouse model and experimental design. We have carefully addressed all the weaknesses mentioned by the reviewer, including the validation of key experiments, quantitative and statistical analysis, consistency between text and figures, and detailed methodological descriptions. Details are described point-by-point below.

      Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9's central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

      We greatly appreciate your positive and insightful comments on our study. We are grateful for the recognition of our novel dual-mechanism model, comprehensive multi-omics analysis, and translational potential from basic research to clinical OMA. We have carefully addressed the weaknesses raised by the reviewer, including in-depth discussion of the GDF9-centered regulatory network and clarification of the origin of granulosa cell abnormalities. More details point-by-point responses are provided below.

      Recommendations for the authors:

      Point-by-point responses to reviewers’ comments

      We thank the reviewer very much for his/her reviewing of our work, and we appreciate the constructive comments and suggestions that have helped us to prepare an improved revision. Based on the comments of the reviewer, we have carefully revised the manuscript by performing some new experiments.

      Reviewer #1 (Recommendations for the authors):

      (1) Did most of the follicles cultured in vitro reach the antral follicle stage after 6 days?

      We greatly appreciate your insightful question. We statistically analyzed the survival rate and antral follicle ratio of in vitro-cultured follicles after 6 days of culture. Due to differences in culture systems and protocols, the follicle survival rate in our study (57.43 ± 3.11%) was different from that reported in previous literature (92 ± 10%). However, the proportion of antral follicles among surviving follicles was highly consistent between our results and published data (83 ± 13% vs 80.87 ± 3.27%) (Cortvrindt and Smitz 2002).

      Author response image 1.

      Ratio and survival rate of antral follicles after 6 days of culture. n = 3. Data are represented as mean ± SD.

      (2) In Figure 2F, at which stage did MATR3 begin to affect oocyte diameter?

      Thank you for your careful observation. Our morphological analysis of oocytes collected from PD14 and PD23 mice showed no significant difference in oocyte diameter between the cKO and Ctrl groups at the GO stage (Fig. S3D, E). However, oocytes in the cKO group became significantly smaller than those in the Ctrl group once they reached the FGO stage (Fig. 2E, F). Taken together, these results indicate that the growth defect caused by MATR3 deletion begins to manifest during the transition from the GO to FGO stage, with significant reduction in oocyte diameter clearly observed at the FGO stage as shown in Figure 2F.

      (3) What was the developmental potential of oocytes in Matr3-knockout mice?

      Thank you for this important question. Compared with the Ctrl group, oocytes derived from cKO mice showed a drastically reduced fertilization rate (91.55 ± 1.96% vs 10.55 ± 4.78%) and almost completely failed to develop to the blastocyst stage (76.62 ± 7.56% vs 3.67 ± 3.38%). These results clearly demonstrate that maternal deletion of Matr3 severely compromises the developmental potential of mouse oocytes, including fertilization capacity and subsequent early embryonic development.

      Author response image 2.

      Results of in vitro fertilization of oocytes. 2-cell: 2 days after fertilization; blastocyst: 4 days after fertilization. Data are represented as mean ± SD. ***P < 0.001.

      (4) The legend labels in the figures should not be bold.

      Thank you for your valuable suggestion. We have revised all the figures accordingly.

      Reviewer #2 (Recommendations for the authors):

      This manuscript investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout (cKO) mouse models combined with in vitro follicle culture approaches. The topic is relevant to the field of reproductive biology and provides potentially important insights into the molecular mechanisms regulating oocyte maturation and follicle development.

      While the study presents interesting observations and utilizes both in vivo and in vitro experimental systems, several issues need to be addressed before the manuscript can meet the expected standards. These include concerns related to data interpretation, validation of experimental approaches, completeness of methodological descriptions, and clarity in data presentation. In addition, the manuscript requires substantial language editing to improve clarity and readability.

      The comments below outline major issues that should be addressed to strengthen the manuscript, as well as specific minor points regarding presentation and clarity.

      (1) The manuscript requires substantial revision to improve the written language and grammar. Numerous sentences are unclear or awkwardly phrased, which makes interpretation of the results difficult in several sections. The authors are strongly encouraged to have the manuscript professionally edited or thoroughly revised for language and clarity before making a resubmission.

      We sincerely appreciate the careful and constructive comments on the language quality and clarity of the manuscript. We fully agree that the written language, grammar, and sentence structure need substantial improvement to ensure the results are presented clearly and accurately.

      To address these concerns thoroughly, we have carefully revised the entire manuscript, including correcting grammatical errors, refining awkward phrasing, and restructuring unclear sentences to enhance readability and logical flow. In addition, we have sought professional language editing support to further polish the English expression and ensure the manuscript meets the linguistic standards of the journal.

      All revisions related to language and clarity have been completed, and we believe the revised version is significantly improved in terms of readability and precision.

      (2) Interpretation of oocyte maturation results (Line 140; Figure 2E, H). The manuscript states: "During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig. 2E, H)." However, Figure 2H appears to show that a small proportion of knockout oocytes do reach the MII stage. Therefore, the description in the text seems inconsistent with the data presented. The authors should clarify the exact maturation rates in both groups, revise the text to accurately reflect the data, and provide statistical analysis to support the stated conclusions.

      Thank you for your valuable comment. A small proportion of knockout oocytes from PD23 cKO mice can indeed develop to the MII stage. We have revised the corresponding description and supplemented the statistical analysis of maturation rates to support our conclusion.

      Line 140: “During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig.2E, H).” have been replaced by “During in vitro maturation, the proportion of oocytes from PD23 cKO mice developing to metaphase II stage was significantly reduced (Fig.2E, H, 54.9±2.08% vs 9.57±1.11%).”

      (3) Human oocyte sample size: In Figure 1D, it is unclear how many human oocytes were analyzed. It is important to specify the sample size (n) for all experiments. The authors should clearly indicate the number of oocytes analyzed in this experiment. Provide this information either in the figure legend or in the main text.

      Thank you for this important comment. We agree that the altered subcellular localization of MATR3 in human OMA oocytes is of great physiological significance for understanding the functional role of MATR3 during oocyte development.

      Unfortunately, during a 3‑month period of sample collection, we examined MATR3 localization in immature oocytes that failed to reach the MII stage, obtained from 11 women undergoing IVF treatment. Among these samples, only one donor’s oocytes exhibited the NSN chromatin configuration. Excitingly, these NSN‑stage oocytes from this donor clearly showed the loss of MATR3 nuclear localization, which strongly supports the critical role of MATR3 during oocyte growth and maturation. We have now clearly stated the sample size (n = 11) in the figure legend and main text as suggested. In future studies, we will continue to collect more human oocyte samples to further validate these observations with an expanded sample size.

      (4) Figure annotation issue: The figure legend for Figure 1F refers to an arrow, but no arrow is visible in the figure panel. Please correct this inconsistency by either adding the appropriate arrow to the figure or revising the legend accordingly.

      Thank you for pointing out this error. We have revised the figure legend for Figure 1F accordingly to correct this inconsistency.

      (5) Description of follicle analysis (Line 146): The sentence: "This was reinforced by the data of available follicles within the follicles of mice on PD35 (Fig. 2I, J)." is incorrect or poorly phrased. It should likely read: "...available follicles within the ovaries of mice at PD35...".

      We really appreciate your constructive suggestion on the phrasing. We have revised this sentence in the revised manuscript accordingly.

      (6) Quantification of proliferating cells: Figure 2K shows Ki-positive cells, but quantitative analysis is not provided. The authors should quantify the number or proportion of Ki-positive cells in both control and cKO groups and include statistical analysis to support any claims regarding differences in proliferation.

      Thank you for your valuable suggestion. We have quantified the number of Ki‑67‑positive granulosa cells in both control and cKO groups and performed the corresponding statistical analysis.

      The quantitative results have been added to Fig. S3G, and the relevant description has been supplemented in the main text at Line 150 to support our conclusion regarding cell proliferation differences.

      Line 150: “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K) in cKO mice were lower than those found in the Ctrl.” have been replaced by “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K, Fig. S3G) in cKO mice were lower than those found in the Ctrl (73.55±13.29% vs 24.65±7.80%).”

      (7) Validation of findings in the in vivo cKO model (Figure 3): The development of an in vitro follicle culture system is an interesting and valuable component of the study. However, several key analyses performed in vitro (e.g., transcription assays and analysis of epigenetic markers) should ideally also be validated in oocytes derived from the in vivo cKO model.

      Thanks for the valuable concern. We fully agree with you that the in vivo cKO model should be used to validate several key analyses performed in vitro. We collected growing oocytes from Ctrl and cKO mice and conducted transcription assays as well as analysis of epigenetic markers. The results showed that Matr3 knockout significantly downregulated transcriptional activity in GO and increased H3K9me2 levels (Author response image 3), which is consistent with our in vitro findings (Fig 3D E I J). These in vivo results confirm that MATR3 plays a critical role in regulating GO transcriptional activity and H3K9me2 levels.

      Author response image 3.

      Matr3 knockout results in the reduction of transcriptional activity. A EU staining (green) in GO collected from Ctrl and cKO. n = 15. B Quantification of the mean fluorescence intensity of EU in oocytes. C H3K9me2 staining (red) in GO collected from Ctrl and cKO. n = 15. D Quantification of the mean fluorescence intensity of H3K9me2 in oocytes. Scale bar: 20 μm. Data are represented as mean ± S.D. ***P < 0.001.

      (8) To strengthen the conclusions, the authors should consider repeating key experiments using oocytes directly isolated from the cKO mice. This would help confirm that the observed effects are not artifacts of the in vitro culture system.

      Thank you for this valuable and constructive suggestion. We fully agree that the conditional knockout mouse model is essential for verifying the physiological significance of MATR3 in vivo.

      To address this point, we have validated multiple key in vitro findings using oocytes directly isolated from cKO mice. For instances, the changes in oocyte transcriptional activity (EU staining) (in Comments 7), H3K9me2 levels (in Comments 7), GDF9 levels (Fig 4.B C E F), and OO-Mvi (Fig 6.A B, Author response image 4) all showed consistent trends with our in vitro knockdown results. In addition, the complete infertility phenotype of cKO female mice further demonstrates that MATR3 is indispensable for oocyte growth and meiotic maturation. We have also provided supplemental data from GDF9 rescue experiments and sequencing analysis performed in the mouse model.

      Author response image 4.

      Matr3 knockdown impairs the structural integrity of oocyte OO-MVi. A p-ERM staining (green) showing the OO-Mvi in oocyte from NC and si-Matr3. B Quantification of the number of Oo-Mvi vesicles (n = 6). Scale bar: 20 μm. Data are represented as mean ± SD. ***P < 0.001.

      In conclusion, the core conclusions of this study are supported by the mutual validation of key experimental results from MATR3-specific knockdown in vitro and Matr3 conditional knockout mouse models in vivo.

      (9) (1) Validation of MATR3 knockdown: The in vitro MATR3 knockdown experiment presented in Figure 4G raises an important concern: it is unclear whether Matr3 knockdown was effectively achieved in the oocytes analyzed. The authors should provide direct evidence of knockdown efficiency, for example, immunostaining for MATR3 protein on the oocytes. Without such validation, it is difficult to interpret the functional outcomes observed.

      As requested, we have provided direct evidences of the knockdown efficiency via immunostaining, which is now presented in Fig.S2D.

      In this experiment, oocytes from early growing follicles (approximately 150 μm in diameter) were microinjected with Matr3 siRNA. Following 5 days of continuous in vitro culture, oocytes from both the NC and si-Matr3 groups were isolated and subjected to immunofluorescence staining to assess protein levels. As shown in the figure, the oocytes at this stage exhibited the characteristic non-surrounded nucleolus (NSN) chromatin configuration. We observed robust MATR3 protein expression within the nucleus of NC oocytes, whereas the MATR3 protein levels were markedly reduced in the si-Matr3 group. These results confirm the successful construction of the Matr3 knockdown model in early growing follicle oocytes.

      (9) (2) Furthermore, it would be more convincing if the authors could perform the Gdf9 supplementation experiments using follicles isolated from the cKO mice, rather than relying solely on siRNA knockdown in vitro. Such experiments would provide clearer and more physiologically relevant evidence. If these experiments were attempted but did not produce similar results, this should be discussed.

      Thank you for your valuable and insightful suggestion. We fully agree that performing GDF9 supplementation experiments using follicles isolated from cKO mice would provide more direct and physiologically relevant evidence to strengthen our conclusions.

      Unfortunately, when we attempted to conduct GDF9 rescue experiments on follicles from cKO mice, neither the control nor cKO follicles were able to develop to the antral follicle stage (n=3). We speculate that this was caused by insufficient bioactivity of the veterinary-grade FSH been used, as compared to the imported FSH been provided by NHPP. Unfortunately, this particular FSH product has been discontinued. We are currently actively seeking and attempting to purchase new, qualified FSH reagents to repeat these experiments and further validate our findings in future work.

      (10) Figure citation order: Figures are not cited sequentially in the text. For example, Figure 6J is described first (line 250), followed by Figure 6A. Figures should be discussed in logical order, typically starting from panel A. Please revise the text to ensure that figure panels are introduced sequentially.

      Thank you for this careful and important comment.

      We have carefully revised the citation order of all figure panels in the main text, especially for Figure 6, to ensure they are introduced sequentially from panel A to the last panel in logical and numerical order, rather than being cited out of sequence.

      The corresponding adjustments have been made in the revised version of the manuscript.

      (11) Figure 6J interpretation: The purpose of the images shown in this figure is unclear. The authors should provide higher magnification images to clearly visualize the Oo-Mvi structures and include quantification of the observed phenotype to support the interpretation. It is important because the main findings of the paper heavily rely on these results.

      Thank you for this valuable and constructive suggestion. We fully agree that higher‑magnification images and quantitative analysis are essential to clearly demonstrate the Oo‑Mvi structures and reliably support our conclusions, especially given the importance of these results to the main findings of this study.

      Accordingly, we have replaced the original panels in Figure 6A with higher‑magnification images to better visualize Oo‑Mvi structures. In addition, we have supplemented the corresponding quantitative analysis of the observed phenotype to strengthen the interpretation of this figure (Fig 6B). All revisions have been incorporated into the revised manuscript.

      (12) Incomplete Materials and Methods section: The Materials and Methods section lacks important experimental details required for reproducibility. Specifically, the Matr3 flox mouse model. Either provide the appropriate reference describing the Matr3 floxed mice or include details on how the floxed allele was generated.

      Thank you for your valuable and careful comment. We fully agree that detailed experimental information in the Materials and Methods section is crucial for ensuring the reproducibility of the study. And we apologize for the omission of key details regarding the Matr3 flox mouse model.

      In response to your suggestion, we have thoroughly supplemented the relevant experimental details in the Materials and Methods section of the revised manuscript, including the specific construction strategy of the Matr3 floxed allele. These detailed descriptions will enable other researchers to reproduce our mouse model and verify the experimental results.

      All supplementary information has been integrated into the revised manuscript to meet the requirements of experimental reproducibility. We greatly appreciate your guidance in helping us improve the completeness and rigour of our study.

      (13) Follicle isolation: The manuscript does not describe how growing follicles were isolated. Please specify whether follicles were isolated using enzymatic digestion or mechanical dissection and provide sufficient methodological detail so that other researchers can reproduce the experiments.

      Thank you for this valuable comment. We agree that adding this information is essential for ensuring the reproducibility of our experiments. We have supplemented the corresponding description in the Materials and Methods section. Briefly, growing follicles were isolated by mechanical dissection using insulin syringes under a stereomicroscope, without any enzymatic digestion.

      Reviewer #3 (Recommendations for the authors):

      (1) Since KDM3B and MATR3 interact in cell lines, does this relationship affect the functional localization of KDM3B within oocytes? Specifically, does the localization of KDM3B change in cKO mice (e.g., nuclear export or aggregation)?

      Thank you for this insightful and constructive question. To address whether the interaction between KDM3B and MATR3 influences the functional localization of KDM3B in oocytes, we performed immunofluorescence staining to examine the subcellular distribution of KDM3B in cKO oocytes. Our results demonstrated that the nuclear localization of KDM3B remained unaltered; no obvious nuclear export or abnormal aggregation was observed in MATR3-deficient oocytes (Author response image 5).

      Based on these observations combined with our other experimental data, we propose that MATR3 regulates oocyte transcriptional activity through its physical interaction with KDM3B, rather than by controlling the nuclear targeting of KDM3B. Notably, despite unchanged nuclear localization of KDM3B in MATR3 cKO oocytes, we detected significantly elevated global levels of H3K9me2 and markedly reduced transcriptional activity (in Comments 7). These findings indicate that KDM3B loses its physiological function of demethylating H3K9me2 and promoting transcription in the absence of MATR3.

      In line with this mechanism, previous results showed that KDM3B knockout in female mice leads to follicle arrest at the secondary follicle stage and consequent infertility (Liu et al. 2015). Collectively, we conclude that in MATR3 cKO oocytes, although KDM3B is properly localized in the nucleus, it fails to execute its H3K9me2 demethylase activity, thereby impairing normal transcriptional regulation during oocyte development.

      Author response image 5.

      Matr3 knockout has no effect on KDM3B localization. KDM3B staining (red) in GO collected from Ctrl and cKO. n   = 50. Scale bar: 40 μm.

      (2) As the GDF9 rescue experiment only partially restores the phenotype, it is suggested to select 2-3 novel targets with high binding intensity and significant expression changes from the LACE-seq data (e.g., Igf2bp2 or Ccnb1 mentioned in the text) to further illustrate that MATR3 regulates a network.

      Thanks for this meaningful and constructive suggestion.

      We have supplemented the relevant data in Fig. S9 and further elaborated on these findings in the Discussion section in our revised manuscript, following your advice. Briefly, we collected growing oocytes from Ctrl and cKO mice and performed RT‑qPCR analysis to verify the expression of Igf2bp2 and Ccnb1 - two representative novel targets with strong binding intensity and significant expression changes identified from our LACE‑seq bioinformatics analysis. The results showed that both genes were significantly downregulated in cKO oocytes compared with controls, supporting the notion that MATR3 regulates a functional RNA network during oocyte development.

      (3) Are the granulosa cell defects primary or secondary? It is recommended to collect ovaries from earlier-stage cKO mice (e.g., PD7 or PD10) to examine the levels of FOXL2 and PCNA in granulosa cells.

      Thank you for your valuable comment. To clarify whether the granulosa cell defects are primary or secondary, we further investigated the temporal effect of MATR3 deficiency on granulosa cells by collecting ovaries from PD7, which is a critical period for primordial follicle activation and early follicular development.

      To evaluate the status of granulosa cells, we performed immunofluorescence staining on PD7 ovarian sections using FOXL2 (a specific marker for granulosa cells) to quantify the number of granulosa cells, and Ki67 (a proliferation-related marker) to assess the proliferative capacity of granulosa cells. The results, as shown in Author response image 6, demonstrated that there were no significant differences in either the number of granulosa cells or their proliferation levels in primary follicles between cKO and Ctrl.

      These findings are consistent with the data in our supplementary Fig S3F, where we observed no significant differences in the number of primordial follicles and growing follicles at PD7 between the two groups. Collectively, these results indicate that the activation of primordial follicles is not affected by MATR3 deficiency in oocytes, and the impairment of granulosa cells caused by oocyte-specific Matr3 knockout occurs at the secondary follicle stage rather than the early stages of follicular development. Therefore, we conclude that the granulosa cell defects in cKO mice are secondary to the oocyte dysfunction induced by MATR3 deficiency.

      Author response image 6.

      Loss of MATR3 in oocytes does not affect the number and proliferation of granulosa cells in primordial follicles. A Immunohistochemistry results showing granulosa cells in PD7 ovaries from Ctrl and cKO. B Quantification of granulosa cell number in the largest cross-section of primary follicles. C Quantification of the proliferation rate of granulosa cells in the largest cross-section of primary follicles. n = 15. Data are represented as mean ± SD. n.s., not significant.

      References:

      (1) Cortvrindt RG, Smitz JE. 2002. Follicle culture in reproductive toxicology: a tool for in-vitro testing of ovarian function? Human reproduction update 8: 243-254.

      (2) Liu Z, Chen X, Zhou S, Liao L, Jiang R, Xu J. 2015. The histone H3K9 demethylase Kdm3b is required for somatic growth and female reproductive function. International journal of biological sciences 11: 494-507.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting

      Strengths:

      (1) Novel evolutionary insight - reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

      Weaknesses:

      (1) The study infers 1C-THF depletion mostly genetically and indirectly (growth rescue with adenine) without direct quantification of folate intermediates or fluxes. Biochemical confirmation, LC-MS-based metabolomics of folates/1C donors, or isotopic tracing would strengthen mechanistic claims.

      We agree with the reviewer that quantification of 1C-THF intermediates would strengthen our conclusions. However, the chemical methodologies to extract folates from L. monocytogenes are not established in our lab. Moreover, quantification of C1-substituted folates requires comprehensive biochemical and analytical expertise that we also do not have and which we cannot cover though co-operations. However, to further strengthen our arguments, we have introduced an experiment in the updated manuscript that demonstrates synthetic lethality of a ΔgcvPAB ΔglyA mutant with a deletion of fold (Fig. 7A). This gene encodes 5,10-methylene-tetrahydrofolate dehydrogenase/ 5,10-methylene-tetrahydrofolate cyclohydrolase, which is the third enzyme involved in N5, N10-methylene-THF generation next to GlyA and GcvPBA. Synthetic lethality of a ΔgcvPAB ΔglyA double mutant with a fold deletion is best explained by GcvPAB and GlyA also feeding the N5,N10-methylene-THF pool.

      (2) In multiple result sections, the authors report data from technical triplicates but do not mention independent biological replicates (e.g., Figure 2C, Figure 4A-B, Figure 6D). In addition, some results mention statistical significance but without a detailed description of the specific statistical tests used or replicates, such as Figure 2A-C, Figure 2E, and Figure 2G-I.

      We thank the reviewer for this helpful comment. Experiments were usually repeated three independent times, with each repetition including three technical replicates. Mean values and standard deviations were usually calculated from the technical replicates of a representative run. Statistical significance was calculated using t-tests for pairwise comparisons or t-tests using Bonferroni-Holm correction for multiple comparisons. We made sure that this is explicitly explained for each experiment in the figure legends. Wherever other calculations were used, we also clarified this in the legends.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Freier et al examines the impact of deletion of the glycine cleavage system (GCS) GcvPAB enzyme complex in the facultative intracellular bacterial pathogen Listeria monocytogenes. GcvPAB mediates the oxidative decarboxylation of glycine as a first step in a pathway that leads to the generation of N5, N10-methylene-Tetrahydrofolate (THF) to replenish the 1-carbon THF (1C-THF) pool. 1C-THF species are important for the biosynthesis of purines and pyrimidines as well as for the formation of serine, methionine, and N-formylmethionine, and the authors have previously demonstrated that gcvPAB is important for bacterial replication within macrophages. A significant defect for growth is observed for the gcvPAB deletion mutant in defined media, and this growth defect appears to stem from the sensitivity of the mutant strain to excess glycine, which is hypothesized to further deplete the 1C-THF pool. Selection of suppressor mutations that restored growth of gcvPAB deletion mutants in synthetic media with high glycine yielded mutants that reversed stop codon inactivation of the formatetetrahydrofolate ligase (fhs) gene, supporting the premise that generation of N10-formyl-THF can restore growth. Mutations within the folk, codY, and glyA genes, encoding serine hydroxymethyltransferase, were also identified, although the functional impact of these mutations is somewhat less clear. Overall, the authors report that their work identifies three pathways that feed the 1C-THF pool to support the growth and virulence of L. monocytogenes and that this work represents the first example of the spontaneous reactivation of a L. monocytogenes gene that is inactivated by a premature stop codon.

      Strengths:

      This is an interesting study that takes advantage of a naturally existing fhs mutant Listeria strain to reveal the contributions of different pathways leading to 1C-THF synthesis. The defects observed for the gcvPAB mutant in terms of intracellular growth and virulence are somewhat subtle, indicating that bacteria must be able to access host sources (such as adenine?) to compensate for the loss of purine and fMet synthesis. Overall, the authors do a nice job of assessing the importance of the pathways identified for 1C-THF synthesis.

      Weaknesses:

      (1) Line 114 and Figure 1: The authors indicate that the gcvPAB deletion forms significantly fewer plaques in addition to forming smaller plaques (although this is a bit hard to see in the plaque images). A reduction in the overall number of plaques sounds like a bacterial invasion defect - has this been carefully assessed? The smaller plaque size makes sense with reduced bacterial replication, but I'm not sure I understand the reduction in plaque number.

      The observation that the ΔgcvPAB mutant forms fewer plaques was not our claim, and we have already addressed the possibility of an invasion defect by quantifying bacterial numbers during infection of 3T3 cells. As shown in Fig. 2A, the ΔgcvPAB mutant invades 3T3 cells (the same cells used in the plaque formation assays) as efficiently as the wild type but exhibits reduced intracellular growth. Therefore, the plaquing defect is not due to impaired invasion. Furthermore, we also have analyzed the intracellular dissemination of the ΔgcvPAB mutant in 3T3 fibroblasts compared to a ΔactA mutant by microscopy. This shows that the ΔgcvPAB is evenly distributed throughout the infected host cells as the wild type and unlike the ΔactA mutant, which forms clusters (Fig. S1). Both experiments indicate that the reduced plaque area results from impaired intracellular growth rather than defects in invasion or cell-to-cell spread. The apparent reduction in plaque numbers in ΔgcvPAB-infected 3T3 cells is likely due to a strong reduction in plaque size, with only the largest plaques remaining visible. We have rephrased the relevant section to clarify this point and avoid any potential confusion:

      “In agreement with our previous results, only small plaques were formed in 3T3 cells upon infection with the ΔgcvPAB mutant (plaque area: 15±19% of wild-type level) and small plaques were also formed by the complemented strain in the absence of IPTG (50±19%).”

      (2) Do other Listeria strains contain the stop codon in fhs? How common is this mutation? That would be interesting to know.

      We determined the frequency of inactivated fhs genes among 30,000 publicly available L. monocytogenes genomes. The analysis identified only 10 isolates carrying truncated fhs alleles. These isolates fell into two groups: (i) EGD-e and its descendants, and (ii) a cluster of five ST2 food isolates. These findings have been added as a separate results section.

      (3) Based on the observation that fhs+ ΔgcvPAB ΔglyA mutant is only possible to isolate in complex media, and fhs is responsible for converting formate to 1C-THF with the addition of FolD, have the authors thought of supplementing synthetic media with formate and assessing mutant growth?

      No, we did not test formate supplementation. However, we included additional experiments testing adenine and thymine supplementation (Fig. 6E and 7B). These results show that purine and thymine become limiting in mutants lacking 1C-THF synthesizing pathways.

      Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al. demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1CTHF synthesis in L. monocytogenes, and the lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays, and murine infection models to demonstrate this.

      Weaknesses:

      (1) The primary finding of the study is that the perturbation of any of the 3 metabolic pathways important for the synthesis of 1C-THF results in reduced growth and virulence of L. monocytogenes. However, there is no evidence demonstrating the levels of 1C-THF in the various knockouts and suppressor mutants used in this study. It is important to measure the levels of this metabolite (ideally using mass spectrometry) in the various knockouts and suppressor mutants, to provide strong causality.

      As already outlined above, we do not have any experimental possibilities to measure 1C-substituted THF in L. monocytogenes extracts directly. However, to provide additional evidence for our interpretation that “Three pathways feed the 1C-THF pool…”, we included additional genetic experiments.

      The first experiment demonstrates that the growth defect of the fhs- ΔglyA igcvPAB strain in synthetic medium lacking IPTG—which reflects the synthetic lethality of fhs with glyA and gcvPAB—can be rescued by the addition of adenine (Fig. 6E). This indicates that the fhs/fold pathway, GlyA, and the glycine cleavage system are essential due to their combined contribution to purine biosynthesis. This result confirms the metabolic model presented in Fig. 1A and thus supports the hypothesis that all three pathways contribute to 1C-THF biosynthesis.

      The second experiment additionally demonstrates synthetic lethality of gcvPAB and glyA with the fold gene. FolD acts downstream of Fhs and is one of the three enzymes synthesizing N5, N10methylene-THF shown in Fig. 1A. The fold gene is essential in EGD-e (PMID: 36114002), most likely explained by fhs inactivation. However, we were able to delete fold in an EGD-e background carrying a reconstituted fhs gene and the resulting fhs<sup>+</sup> Δfold strain was as viable as a fhs<sup>+</sup> ΔglyA ΔgcvPAB strain (Fig. 7A). However, a fhs<sup>+</sup> Δfold ΔglyA igcvPAB strain required IPTG for growth in BHI medium (Fig. 7A), indicating that the simultaneous deletion of glyA and gcvPAB is not tolerated in the absence of fold, similar to what is observed in the absence of fhs. Notably, this growth defect was not rescued by adenine supplementation (Fig. 7B), but was alleviated by thymine addition, which is also consistent with the metabolic model shown in Fig. 1A.

      Even though we are unable to directly demonstrate reduced 1C-THF levels, we hope that these two genetic approaches together with the revised title and heading of the relevant paragraph will, in the reviewers’ eyes, support our hypotheses.

      (2) The story becomes a little hard to follow since macrophage-CFU assays and murine infection model data precede the in vitro growth assays. The manuscript would benefit from a reorganization of Figures 2,3, and 4 for better readability and flow of data.

      We respectfully disagree with the reviewer. The attenuation of the ΔgcvPAB mutant in macrophages and fibroblasts was the primary motivation for further investigating its phenotype. Therefore, we chose to begin the manuscript with results from various virulence studies before presenting the sections that provide mechanistic explanations. In our view, this sequence represents a more logical and coherent way to present the data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Synthetic medium assumptions: LSM "mimics" intracellular limitation but isn't chemically validated against host cytosolic composition. Nutrient availability conclusions could be biased.

      This is correct. We added this information:

      “LSM broth is a chemically defined medium that contains all components required for growth at defined concentrations, but it has not been chemically validated to reflect host cytosolic conditions…”

      (2) Glycine toxicity: The paper introduces the concept of glycine toxicity in ΔgcvPAB mutants. However, the conditions under which glycine becomes toxic could be further elucidated. Why glycine causes toxicity despite being an essential metabolite in other contexts requires a more in-depth mechanistic explanation.

      The concept of glycine toxicity in GCS mutants has been described previously by other researchers. We have added further details to better explain glycine toxicity and how it accounts for the growth phenotype of the ΔgcvPAB mutant:

      “If glycine cannot be catabolized (and 1C-THF cannot be generated) by the GCS due to deletion of gcvPAB, glycine might be re-routed to the serine hydroxymethyl transferase GlyA for serine formation, even though this would consume 1C-THF and therefore even further deplete the cell for 1C-THF.” and

      “In the complete absence of glycine, growth of the ΔgcvPAB mutant was largely unaffected, presumably because glycine cannot be converted to serine by GlyA anymore, thereby conserving the 1C-THF pool.”

      (3) Figures:

      (a) Figure 1B, scale bar?

      A scale bar was added.

      (b) Figure 1C, the standard error bars for igcvPAB (both with and without IPTG) are relatively wide, indicating high variability in the data. This suggests that the results in the igcvPAB group are not as consistent as the wild-type (wt) or ΔgcvPAB groups. Please show the original data and perform a statistical test (e.g., t-test or ANOVA).

      The original data have been added to Fig. 1B. t-test results (with Bonferroni-Holm correction) are now included for all samples.

      (c) Figure 2D, scale bar?

      These are sections of agar plates. From our point of view, a scale bar does not add relevant information.

      (d) Figure 3, please check the labels of the Figure 3 legend. (D) and (E) or A-B?

      Thanks, corrected.

      (e) Figure 5E, quantification of plaque areas?

      The plaque areas were quantified. A blot showing these quantitative data is now presented in Fig. 5F.

      Reviewer #2 (Recommendations for the authors):

      Line 58: There are published studies that indicate that syncytiotrophoblasts are actually resistant to Listeria infection and that it is extravillous trophoblasts that are likely to serve as entry points for Listeria into the placenta (see, for example, Lowe et al, Infect Immun. 2018 Volume 86 Issue 6 e00801-17).

      We thank the reviewer for this comment and have removed our statement claiming that syncytiothrophoblasts are the entry point as this is not relevant to the understanding of the work presented here.

      Reviewer #3 (Recommendations for the authors):

      (1) Please mention the number of times experiments were performed as independent biological replicates, wherever applicable.

      We added this information to the figure legends wherever this was necessary.

      (2) Please provide the details of the type of statistical analysis used for the various graphs, either in the figure legends or in the materials & methods section.

      This information was also added to the figure legends wherever it still was missing.

      (3) Can the authors comment on how the weight-loss phenotype of animals and the variation in the size of the spleen between animals infected with wild-type and mutants in Figure 2 can be explained without any significant changes in the CFU? Additionally, I did not see details regarding the number of animals used in the murine infection model experiments. Please mention this along with the type of statistical analyses used.

      We do not see a contradiction here, as the apparent differences are explained by the distinct time points at which CFU numbers (day 3 post-infection) and spleen size (day 9 postinfection) were measured. Starting from day 8, the difference in weight between animals infected with the wild-type strain and those infected with the ΔgcvPAB mutant becomes clear for the first time. At day 3, no significant differences are detected in either CFU numbers or weight. By day 9, when the weight difference has become apparent, differences in spleen size are also observed. To improve clarity, the time points at which each analysis was performed have been added to Fig. 2G and Fig. 2I. The number of infected animals and the type of statistical analysis used are now specified in the figure legend.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback.

      We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them.

      Reviewer #1 raised two important concerns regarding our methodology. First, the determination of allele dosage is insufficiently explained, which is central to our ploidy assignment and downstream analyses. Second, the setup and sample sizes of the common garden experiments are unclear, raising questions about the robustness of our conclusions. We accept these criticisms and will address them as follows.

      Regarding allele dosage, we will add a detailed step-by-step description of our calling pipeline in the Methods section, including the criteria for peak height ratios and thresholds used to assign copy numbers. We will also clarify a crucial biological detail: the common reed (Phragmites australis) is an allotetraploid in its origin. As a consequence, many molecular markers, including the widely used SSR markers in previous studies, behave as disomic markers (i.e., two homeologous copies inherited in a Mendelian manner). Therefore, observing more than two alleles at a locus is indeed indicative of higher-level ploidy (hexaploidy or octoploidy) in this system. We will explicitly state this to resolve any confusion about why tetraploids in our dataset are treated as having a maximum of two alleles, while hexaploids and octoploids can carry more.

      Regarding the common garden experiment, we will explicitly report the replication number for each lineage-by-treatment combination and clarify the experimental design. We will also discuss the statistical approaches used given the sample sizes, while acknowledging that the consistency between experimental results and distributional patterns lends additional support to our conclusions.

      Reviewer #2 raised three substantive framing issues. First, ploidy is completely confounded with genetic background, yet our manuscript places undue emphasis on polyploidy as a causal factor. Second, our species distribution models treat each lineage as a homogeneous entity, failing to capture within-lineage variation and thus repeating the oversimplification we criticize. Third, we insufficiently explore the evolutionary significance of asymmetric introgression, gene flow, and the novelty of combining SDM with experiments. We fully agree with these points and will revise accordingly.

      To address the confounding issue, we will substantially reframe the manuscript to de-emphasize claims about polyploidy as a causal driver, and instead focus on the adaptive differentiation among distinct genetic lineages that happen to differ in ploidy. The Discussion will explicitly state that dissecting ploidy effects from background genetic effects will require future experimental approaches.

      To address the simplification in SDMs, we will add a clear acknowledgment of this limitation, discussing how it may affect predictive accuracy and suggesting that future studies incorporating population-level genomic data could more directly assess evolutionary potential.

      To address the insufficient exploration of introgression and the novelty of our approach, we will expand the Introduction to better highlight the value of coupling controlled experiments with SDMs at the intraspecific level. In the Discussion, we will elaborate on the evolutionary significance of asymmetric introgression, including testable hypotheses about how gene flow might mediate the spread of heat-tolerance alleles and influence lineage geographical limits under climate change.<br /> We also thank the reviewer for the suggestion to emphasize the experimental validation of SDM efforts, which we will incorporate into a revised Introduction.

      Looking beyond the present study, we envision three complementary directions that build upon our current findings. Expanding common garden experiments to include admixed individuals would test whether introgressed genomic blocks confer fitness advantages under thermal stress. Leveraging the population genomic framework established here, we will transition to whole-genome resequencing for selection scans and genotype-environment association analyses to pinpoint adaptive loci and reveal whether heat-tolerance alleles are preferentially transferred via asymmetric introgression. We will also integrate transcriptomic profiling with phenotypic measurements to identify candidate genes whose expression correlates with thermal performance and introgressed ancestry, helping to disentangle ploidy effects from genetic background. Together, these directions span expanded phenotyping, whole-genome resequencing, and transcriptome-guided discovery, forming an integrated framework that moves from the correlative patterns reported here toward mechanistic understanding. These perspectives are briefly outlined in our Discussion, and we hope the present study will serve as a foundation for these future investigations, which we plan to pursue in subsequent work.

      We believe these revisions will substantially strengthen the manuscript.

    1. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. Author response:

      The following is the authors’ response to the original reviews.

      This is a summary of the changes that have been made to the Reviewed Preprint:

      (1) The data from RettBASE which was analysed in the manuscript has been added in the form of four supplementary tables. Supplementary Table 1 contains the download of all MeCP2 mutations contained in RettBASE. Supplementary Tables 2-4 contain subsets of this data which were used in Figure 2B and Figure 3 SF2. Supplementary Table 4 also has the HGVS nomenclature for both e1 and e2 isoforms and the ClinVar Variation ID for each allele. Wording has been changed to clarify that analysis in the manuscript used this data from RettBASE and not the information that was deposited in ClinVar.

      (2) Similarly, Supplementary Tables 5-8 contain the gnomAD data that was used in the preparation of Figure 2, Figure 2 SF1 and Figure 3 SF1. The “high confidence” alleles in Supplementary Table 8 have been annotated with their HGVS names.

      (3) The criteria for selecting “high confidence” RettBASE and gnomAD alleles have been more explicitly stated in both the Results and Materials and Methods sections.

      (4) An additional “high confidence” RTT mutation (c.1152_1195) has been added to Figures 2B and 3B.

      (5) Figure 3 Supplementary Figure 2 has been added to show reading frame data for all frameshifting deletions in the C-terminal deletion-prone region (CT-DPR), showing that +2 frameshifts predominate in this larger data set, not just in the “high confidence” set. This has necessitated changing the previous Fig. 3 SF2 to Fig. 3 SF3.

      (6) A summary of the genetic alterations described in the manuscript, and their outcomes, has been added as Figure 7.

      (7) A simple flow chart which assists in the classification of human CTDs as “likely benign” or “likely pathogenic” has been added as Figure 8. This will aid future assessment of novel mutations in this region.

      (8) Additions have been made to the Materials and Methods section to comply with reporting guidelines.

      (9) Minor changes have been made to the text to correct typographical errors and to clarify meaning.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors scrutinized differences in C-terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implications in Rett syndrome diagnosis and proposes a therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome, and carries the potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments translate genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences that give rise to dramatic phenotypic consequences.

      Weaknesses:

      There are many base-editing and protein-expression changes throughout the manuscript, and they cause confusion. It would be helpful to readers if authors could provide a simple summary diagram at the end of the paper.

      We have added a summary diagram, as suggested (Figure 7). We have also provided a flowchart which shows how to classify human CTDs as “likely benign” or “likely pathogenic” based on their location.

      Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown to be feasible in generated cell lines, but has yet to be shown in vivo in animal models.

      This work is the obvious next step and is in progress. However, with the rise in pre- and neonatal genetic testing we felt it was important to disseminate our findings as soon as possible. The family pedigree in Figure 3C is a clear illustration of this need

      Reviewer #3 (Public review):

      Summary:

      Guy et al. explored the variation in the pathogenicity of carboxy-terminal frameshift deletions in the X-linked MECP2 gene. Loss-of-function variants in MECP2 are associated with Rett syndrome, a severe neurodevelopmental disorder. Although 100's of pathogenic MECP2 variants have been found in people with Rett syndrome, 8 recurrent point mutations are found in ~65% of disease cases, and frameshift insertions/deletions (indels) variants resulting in production of carboxy-terminal truncated (CTT) MeCP2 protein account for ~10% of cases. Many of these occur in a "deletion prone region" (DPR) between c.1110-1210, with common recurrent deletions c.1157-1197del (CTD1) and c.1164_1207del (CTD2). While two major protein functional domains have been defined in MeCP2, the methyl-binding domain (MBD) and the NCoR interacting domain (NID), the functional role of the carboxy-terminal domain (CTD, beyond the NID, predicted to have a disordered protein structure) has not been identified, and previous work by this group and others demonstrated that a Mecp2 "minigene" lacking the CTD retains MeCP2 function suggesting that the CTD is dispensable. This raises an important question: If the CTD is dispensable, what is the pathogenic basis of the various CTT frameshift variants? Prior work from this group demonstrated that genetically engineered mice expressing the CTD1 variant had decreased expression of Mecp2 RNA and MeCP2 protein and decreased survival, but those expressing the CTD2 variant had normal Mecp2 RNA and protein and survival. However, they noted that differences between the mouse and human coding sequences resulted in different terminal sequences between the two common CTD, with CTD1 ending in -PPX in both mouse and human, but CTD2 ending in -PPC in human but -SPX in mouse, and in the previous paper they demonstrated in humanized mouse ES cells (edited to have the same -PPX termination) containing the CTD2 deletion resulted in decreased Mecp2 RNA and protein levels. This previous work provides the underlying hypotheses that they sought to explore, which is that the pathological basis of disease causing CTD relates to the formation of truncated proteins that end with a specific amino acid sequence (-PPX), which leads to decreased mRNA and protein levels, whereas tolerated, non-pathogenic CTD do not lead to production of truncated proteins ending in this sequence and retain normal mRNA/protein expression.

      In this manuscript, they evaluate missense variants, in-frame deletions, and frame shift deletions within the DPR from the aggregated Genome Aggregated Database (gnomAD) and find that the "apparently" normal individuals within gnomAD had numerous tolerated missense variants and in-frame deletions within this region, as well as frameshift deletions (in hemizygous males) in the defined region. All of the gnomAD deletions within this region resulted in terminal amino acid sequences -SPRTX (due to +1 frameshift), whereas nearly all deletion variants in this region from people with Rett syndrome (from the Clinvar copy of the former RettBase database) had a terminal -PPX sequence, due to a +2 frameshift. They hypothesized that terminal proline codons causing ribosomal stalling and "nonsense mediated decay like" degradation of mRNA (with subsequent decreased protein expression) was the basis of the specific pathogenicity of the +2 frameshift variants, and that utilizing adenine base editors (ABE) to convert the termination codon to a tryptophan could correct this issue. They demonstrate this by engineering the change into mouse embryonic stem cell lines and mouse lines containing the CTD1 deletion and show that this change normalized Mecp2 mRNA and protein levels and mouse phenotypes. Finally, they performed an initial proof-of-concept in an inducible HEK cell line and showed the ability of targeted ABE to edit the correct adenine and cause production of the expected larger truncated Mecp2 protein from CTD1 constructs.

      The findings of this manuscript provide a level of support for their hypothesis about the pathogenicity versus non-pathogenicity of some MECP2 CTT intragenic deletions and provide preliminary evidence for a novel therapeutic approach for Rett syndrome; however, limitations in their analysis do not fully support the broader conclusions presented.

      Strengths:

      (1) Utilization of publicly available databases containing aggregated genetic sequencing data from adult cohorts (gnomAD) and people with Rett syndrome (Clinvar copy of RettBase) to compare differences in the composition of the resulting terminal amino acid sequences resulting from deletions presumed to be pathogenic (n+2) versus presumed to be tolerated (n+1).

      (2) Evaluation of a unique human pedigree containing an n+1 deletion in this region that was reported as pathogenic, with demonstration of inheritance of this from the unaffected father and presence within other unaffected family members.

      (3) Development of a novel engineered mouse model of a previously assumed n+1 pathogenic variant to demonstrate lack of detrimental effect, supporting that this is likely a benign variant and not causative of Rett syndrome.

      (4) Creation and evaluation of novel cell lines and mouse models to test the hypothesis that the pathogenicity of the n+2 deletion variants could be altered by a single base change in the frameshifted stop codon.

      (5) Initial proof-of-concept experiments demonstrating the potential of ABE to correct the pathogenicity of these n+2 deletion variants.

      Weaknesses:

      (1) While the use of the large aggregated gnomAD genetic data benefits from the overall size of the data, the presence of genetic variants within this collection does not inherently mean that they are "neutral" or benign. While gnomAD does not include children, it does include aggregated data from a variety of projects targeting neuropsychiatric (and other conditions), so there is information in gnomAD from people with various medical/neuropsychiatric conditions. The authors do make some acknowledgement of this and argue that the presence of intragenic deletion variants in their region of interest in hemizygous males indicates that it is highly likely that these are tolerated, non-pathogenic variants. Broadly, it is likely true that gnomAD MECP2 variants found in hemizygous males are unlikely to cause Rett syndrome in heterozygous females, it does not necessarily mean that these variants have no potential to cause other, milder, neuropsychiatric disorders. As a clear example, within gnomAD, there is a hemizygous male with the rs28934908 C>T variant that results in p.A140V (p.A152V in e1 transcript numbering convention). This pathogenic variant has been found in a number of pedigrees with an X-linked intellectual disability pattern, in which males have a clear neurodevelopmental disorder and heterozygous females have mild intellectual disability (see PMIDs 12325019, 24328834 as representative examples of a large number of publications describing this). Thus, while their claim that hemizygous deletion variants in gnomAD are unlikely to cause Rett syndrome, that cannot make the definitive statement that they are not pathogenic and completely benign, especially when only found in a very small number of individuals in gnomAD.

      We have included the possibility that mutations found in gnomAD may give rise to less severe neurological conditions in the discussion.

      (2) The authors focus exclusively on deletions within the "DPR", they define as between c.1110-1210 and say that these deletions account for 10% of Rett syndrome cases. However, the published studies that are the basis for this 10% estimate include all genetic variants (frameshift deletions, insertions, complex insertion/deletions, nonsense variants) resulting in truncations beyond the NID. For example, Bebbington 2010 (PMID: 19914908), which includes frameshift indels as early as c.905 and beyond c.1210. Further specific examples from RettBase are described below, but the important point is that their evaluation of only frameshift variants within c.1110-1210 is not truly representative of the totality of genetic variants that collectively are considered CTT and account for 10% of Rett cases.

      The vast majority of C-terminal truncating mutations do occur within the “CT-DPR”, likely due to its C-rich nucleotide sequence and the presence of microhomologies within the region. Looking at frameshifting deletions in RettBASE that start after the NID, a large proportion of these end within the CT-DPR and result in a -PPX ending. We decided to restrict our analysis to the c.1110-1210 region to avoid including the rarer examples that may have a different reason for their pathogenicity. We do not assert that all C-terminal truncations are pathogenic due to this mechanism, but current evidence suggests that most are.

      (3) The authors say that they evaluated the putative pathogenic variants contained within RettBase (which is no longer available, but the data were transferred to Clinvar) for all cases with Classic Rett syndrome and de novo deletion variants within their defined DPR domain. Looking at the data from the Clinvar copy of RettBase, there are a number (n=143) of c-terminal truncating variants (either frameshift or nonsense) present beyond the NID, but the authors only discuss 14 deletion frameshift variants in this manuscript. A number of these variants have molecular features that do not fall into the pathogenic classification proposed by the authors and are not addressed in the manuscript and do not support the generalization of the conclusions presented in this manuscript, especially the conclusion that the determination of pathogenicity of all c-terminal truncating variants can be determined according to their proposed n+2 rule, or that all of the 10% of people with Rett syndrome and c-terminal truncating variants could be treated by using a base editor to correct the -PPX termination codon.

      It is important to state here that we did not use the data in ClinVar for our analysis, but the original information that was held in RettBASE. We have clarified this in the manuscript and have now included supplementary tables containing the data we downloaded from both RettBASE and gnomAD. Table 1 contains all the RettBASE entries with MECP2 mutations, while Tables 2-4 contain subsets of this data pertaining to CTDs. We have extended our “high confidence” set of RTT alleles to contain one more that was previously overlooked (c.1152_1195) to bring the total number of alleles to 15. Taken together these alleles account for 158 individual entries in RettBASE. We have now included an analysis of all frameshifting mutations in the CT-DPR (Supplementary Table 3, Fig. 3 SF2) which covers 69 different mutations and 260 individual entries. Of these, +2 frameshifts make up the large majority, in contrast to the gnomAD data shown in Supplementary Table 8 and Fig. 3 SF1.

      (4) The HEK-based system utilized is convenient for doing the initial experiments testing ABE; however, it represents an artificial system expressing cDNA without splicing. Canonical NMD is dependent on splicing, and while non-canonical "NMD-like" processes are less well understood, a concern is whether the artificial system used can adequately predict efficacy in a native setting that includes introns and splicing.

      We disagree with this opinion. We show that the loss of protein and mRNA seen with knock in mouse and human alleles is recapitulated when using a cDNA-based transgene in the HEK system, demonstrating that the mechanism of loss does not involve factors bound at splice junctions etc. We also demonstrate the effect of the A to G change at the stop codon is the same whether we do this by base editing our cDNA transgene in T-REx cells or by making the CTD1 X>W knock-in mouse. Both result in increased levels of a slightly extended but still truncated protein.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The phrase in the title, "an alternative therapeutic approach" is only insinuated in the manuscript, making it rather inappropriate to be in the title.

      In this study we use adenine base editing to modify RTT-causing CTD mutations in T-REx cells which is clearly a precursor to developing a therapy, utilising the new findings in this study. We therefore feel that the use of “an alternative therapeutic approach” can be justified.

      Reviewer #2 (Recommendations for the authors):

      I have a few minor comments for the authors to consider:

      (1) Please double-check Figure 2 Supplementary Figure 1, as the allele count for E394K does not appear to be in the thousands; rather, E397K seems to be the variant shown in the graph.

      Yes, this was an error and has been corrected to E397K in the text. Thank you for spotting it.

      (2) On page 12, the phrase “common DNA sequence features shared by all CTDs that give rise to RTT” might be better described as “amino acid sequence features.”

      This has been altered in the text as suggested.

      (3) On page 3, the sentence "analysis of patient mutations and experimental data from mouse models support a role in transcriptional repression" cites Gabel et al. 2015 and Kinde et al. 2016, which focus on null alleles but not patient mutations. It would be appropriate to also cite Johnson et al. 2017, which analyzed MeCP2 T158M and R106W patient mutations.

      This has now been cited.

      (4) In the same sentence, Bajikar et al. 2025 are described as studying the "acute loss of MeCP2," but gene expression was analyzed after a week or longer period of time, not minutes to hours as in degron-mediated degradation systems.

      The word “acute” has been removed from the text.

      (5) On page 7, the statement that "E394K is common... who were later found to have additional pathogenic MECP2 mutations" should include a supporting reference.

      We have now included the references Moncla et al (2002) and Wan et al (1999) to address this.

      (6) Similarly, on page 10, the sentence "these findings question the validity of two cases where individuals presented with classical Rett..." is missing a reference to the case report mentioned.

      The references Bienvenu et al (2000) and Philippe et al (2006) have been added.

      (7) While the manuscript is well written and full of detail, if space is an issue, the authors might consider tightening sections that reiterate findings from their 2018 HMG paper.

      A section discussing the CTD2 allele from the 2018 HMG paper has been removed from the results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Overall, the manuscript is rather dense and potentially challenging to follow easily, especially for a non-expert reader.

      Minor edits have been made to the text which will hopefully make it easier to follow. We have also added two new figures (Figures 7 and 8) to summarize the different alleles and edits which appear in the paper, and to show how to determine whether a CTD in the region is likely to be benign or pathogenic.

      (2) The introduction of data presented in Figure 1 within the manuscript introduction seems inappropriate and should be moved to the results section.

      We would say that this is unconventional rather than inappropriate, and is referred to in the introduction, so we would prefer to leave it as it is.

      (3) Providing specific, common nomenclature for genetic repository variants (rs numbers, gnomAD IDs, etc) somewhere would be beneficial. This is an issue because of the complexity of numbering (either coding or protein) for MECP2 due to the different transcript-based numbering systems.

      This nomenclature is now included in the supplementary tables of data from RettBASE and gnomAD.

      (4) As described in the public comments, there are a number of MECP2 genetic variants listed in the Clinvar copy of RettBase, resulting in c-terminal truncations that are not mentioned or discussed within the manuscript. Without the level of detail present in the original form of RettBase (number of events, de novo, etc) in the currently available Clinvar iteration, it is unclear why a number of variants, even within the limited DPR region, were not mentioned. A supplementary file including the more complete information from the RettBase version, with a complete listing of all c-terminal truncating variants, and an explicit rationale for the exclusion of variants would be helpful.

      We have now included supplementary tables with our download of all MECP2 mutations which were held in RettBASE. We have further added tables with the subset of mutations that we have analysed and have more explicitly stated our criteria for defining the “high confidence” sets of mutations. We did not download the data relating to “evidence of pathogenicity” (ie de novo?, absent from parents etc) from RettBASE, but annotated our list of CTDs with this information while RettBASE was still available. This was used in Supplementary Table 3.

      (5) A discussion of the limitations, notably that the fact that the focus exclusively on deletion variants within a restricted region (c.1110-1210) does not truly represent all genetic variants that cumulatively account for 10% of Rett cases, is needed. Furthermore, as pointed out, not all frameshift variants, even those that are n+2, result in the -PPX termination that is presented as the pathogenic basis of c-terminal truncations and amenable to correction by ABE. This should be noted in the discussion, as well as consideration of the late nonsense variants that cause c-terminal truncations (some of which would be very similar to the deletion variants discussed but without the proposed primary pathogenic driver, -PPX).

      We do not claim to explain the pathogenic mechanism of all C-terminal frameshift mutations found in cases of Rett syndrome. There will certainly be some that do not fit our explanation. However, we believe we have shown evidence that a large proportion of CTDs in RTT will be amenable to the therapy we propose.

      (6) Regarding point 3 in the public review, specifically:

      (a) n=7 nonsense variants (S360X, K363X, E397X, R453X, E455X) that do not carry the destabilizing -PPX sequence.

      (b) n=136 frameshift indel variants beyond the NID.

      (i) n=11 that have indels that extend past the native stop codon, n=4 of which start within the DPR domain (c.1110-1210) but would have a different terminal sequence than their proposed pathogenic -PPX sequence.

      (ii) n=125 frameshift indels with terminal breakpoint before the native stop codon

      (c) n=89 that have start or stop points within c.1110-1210

      (d) n=72 not mentioned within the manuscript.

      (e) n=27 are n+1, with 26/27 having what the authors term as the "tolerated" -SPRTX ending, but 1/27 having a frameshift beyond this region (c.1133_1361)

      (f) n=45 are n+2, with 32/45 ending in -PPX (supporting authors conclusion), but 10/45 will use the frameshift stop codon preceding the -PPX and have a different terminal sequence, and 3/45 result in a frameshift termination beyond the -PPX sequence.

      (g) n=36 have breakpoints either before c.1110 or after c.1210

      (h) n=16 start before c.1110, with 11/16 ending before c.1110. 4 of these 11 are n+2, but would use the earlier frameshift stop codon and not have -PPX terminal sequence. For 5/16, the indel extends past c.1210, with the n+2 leading to frameshift termination codons beyond the -PPX sequence.

      (i) n=20 indels start beyond c.1210, with 8/20 being n+1 and 12/20 being n+2, with neither leading to the -SPRTX or --PPX termination sequences characterized in this manuscript.

      As mentioned previously, we have used data taken from RettBASE, not from ClinVar. Both RettBASE and ClinVar will contain MECP2 mutations found in cases of RTT which are not the causative mutation. Databases of this kind contain sequencing errors and mutations that have been mistakenly assigned as causative. It is therefore imperative that the publicly available information is screened to only include mutations that meet stringent criteria. This is why we chose to start by looking at high confidence sets of mutations, with our conclusions supported by analysis of all such mutations in our region of study.

      As mentioned in response to point 5 above, we do not claim to explain every mutation in the region, but believe this study reveals an important disease mechanism for a large proportion of CTDs, leading to a potential therapy. It also contains significant information for predicting the likely prognosis of individuals with CTDs, who may remain healthy but are currently informed that their mutation is likely pathogenic. At present it seems that this is often based solely on the presence of a frameshifting mutation with similarities to bona fide RTT CTDs, without strong evidence of pathogenicity.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. putida bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

      We thank the reviewer for raising this important comment. We agree that receptor-mediated salt taxis is a well-established mechanism in some bacteria, including the classic study by Qi and Adler (PNAS, 1989). We note, however, that the original manuscript did discuss this possibility and cited Qi and Adler in line 117: “We also note that we did not observe any significant difference in the tumble rates between the control and NaCl gradient cases (Figure 3h; Figure S2, SM), suggesting that NaCl gradients do not interfere with chemoreceptors (Qi and Adler, 1989).” In Figure S2, the run-time distributions show no significant difference between the no-gradient condition and the NaCl-gradient condition, with fitted tumble rates of λ = 0.41 s<sup>−1</sup> and λ = 0.38 s<sup>−1</sup>, respectively. We also do not observe a directional bias in run duration, namely longer runs up the salt gradient and shorter runs down the gradient, which would be expected for canonical chemoreceptor-mediated taxis.

      The basis for assigning the observed migration to diffusiophoresis is that the NaCl gradient produces a directional drift of the bacterial body without a measurable change in the run-and-tumble statistics. This behavior is consistent with our previous work [1], where non-motile bacteria were shown to undergo diffusiophoretic migration toward higher salt concentration. Because that migration occurred in non-motile cells and across different bacterial types and morphologies, it supports the interpretation that native bacterial surface charge can drive a non-specific diffusiophoretic response in salt gradients.

      That said, we agree with the reviewer that a genetic control would provide a stronger test against receptor-mediated chemotaxis. We will therefore perform additional experiments using a ∆cheA strain. Because CheA is required for canonical chemotactic signal transduction, observing the same directional migration in the ∆cheA mutant would directly test whether the NaCl gradient response persists in the absence of receptor-mediated chemotaxis. We will include these new data and revise the manuscript to more explicitly distinguish diffusiophoretic drift from chemoreceptormediated salt taxis.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      We thank the reviewer for the constructive comments. We agree that the proposed steering mechanism should be supported by evidence that directly connects the salt gradient response to bacterial orientation dynamics, not only to trajectory-level transport statistics.

      We would like to clarify that the original manuscript already includes an orientation-based analysis in Figure 4f,g in the main text. The corresponding methodology and results are described in lines 173–183 and in the Supporting Information. Specifically, we quantified cell steering by measuring the change in body angle, ∆θ, along individual run trajectories as a function of arc length, s, using the orientation correlation ⟨cos(∆θ)⟩<sub>s</sub>. In the absence of salt gradients, the orientation correlation decays slowly with arc length, indicating persistent swimming along the initial

      Author response image 1.

      Instantaneous angular velocity as a function of heading angle relative to the salt gradient orientation. (a) Experimental and (b) simulated mean angular velocity of cells as a function of heading angle θ (measured relative to the gradient direction; θ = 0° points toward the gel/high-salt side) in the absence (blue) and presence (red) of a NaCl gradient. Positive and negative values indicate counterclockwise and clockwise rotation, respectively, with arrows showing rotation direction. Under the gradient, both experiments and simulations show a signed, angle-dependent rotation rate that is largest near θ = ±90° and approaches zero near θ = 0° and 180°, consistent with a restoring torque that steers cell heading toward the gradient direction. Simulations reproduce this behavior with comparable magnitude to the experimental measurements, and in the absence of a gradient, angular velocity remains relatively small in the no-gradient case with no consistent directional bias across heading angles.

      run direction. Under a salt gradient, the correlation decays more rapidly, indicating stronger directional reorientation during runs. Because the tumble statistics do not change significantly between the no-gradient and salt-gradient conditions, this enhanced orientational decorrelation is not attributed to increased tumbling or rotational noise. Instead, it is consistent with continuous deterministic steering during runs, as expected from a diffusiophoretic torque acting on the cell body–flagellar bundle system.

      To further address the reviewer’s concern, we performed an additional orientation-dynamics analysis using the same dataset shown in Figure 4. Following the approach used by Stehnach et al. [2], we calculated the instantaneous angular velocity during individual runs as a function of the cell heading angle relative to the salt gradient direction. The cell orientation was obtained from the run trajectories, and tumble events were excluded because they produce large transient angular velocity spikes that are not representative of continuous steering during runs.

      The new analysis is shown in Author response image 1. Under the no-gradient condition, the angular velocity remains small and nearly independent of heading angle. In contrast, under the salt gradient condition, the angular velocity becomes strongly heading-dependent. The angular velocity is largest when cells swim nearly perpendicular to the salt gradient, where a steering torque is expected to be maximal. Moreover, the sign of the angular velocity indicates rotation toward alignment with the gradient direction. This behavior is consistent with the proposed diffusiophoretic torque mechanism and provides a direct link between the observed transport behavior and salt-gradient-induced reorientation dynamics.

      We plan to add this angular velocity analysis to the revised manuscript and revise the relevant text to make the connection between trajectory statistics, orientation dynamics, and the proposed torque mechanism clearer.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      We thank the reviewer for raising this important point. We agree that alternative physical mechanisms associated with the imposed salt gradient should be considered explicitly. In the revised manuscript, we will add a more detailed discussion explaining why flow-mediated or other hydrodynamic mechanisms are unlikely to account for the observed steering behavior.

      First, the characteristic diffusiophoretic drift velocity in our experiments is approximately u<sub>d</sub> ≈ 1 µm/s, corresponding to only about 1–5% of the typical swimming speed of P. putida. Thus, the proposed mechanism does not require externally driven advection of the cells. Instead, the salt gradient produces a weak but persistent differential diffusiophoretic slip on the cell body and flagellar bundle, which can generate a reorienting torque during active swimming.

      We also considered whether diffusio-osmotic flow along the channel walls could generate sufficient shear to induce rheotaxis. In a dead-end channel, the diffusio-osmotic velocity profile can be estimated as [3]

      which gives a wall shear rate (at z = h) of

      This value is below the shear rate threshold reported by Marcos et al. [2], where rheotactic drift becomes negligible for S < 0.1 s<sup>−1</sup>. Therefore, the shear generated by diffusio-osmotic wall flow in our experiments is too weak to explain the observed directional reorientation. This effect would be even smaller for P. putida, whose thin flagellar bundle is expected to experience weaker shear-induced alignment than organisms with larger flagellar structures.

      We further considered viscotaxis as a possible mechanism. However, the viscosity difference between 1 mM and 100 mM NaCl solutions is marginal. This contrast is far smaller than the approximately 4–5-fold viscosity difference reported to induce strong viscophobic turning in bacteria [4]. Thus, salt-gradient-induced viscosity variations are insufficient to account for the measured steering response.

      Taken together, these estimates indicate that hydrodynamic shear, rheotaxis, and viscotaxis are too weak under our experimental conditions to explain the observed migration and orientation dynamics. In contrast, the proposed nonuniform diffusiophoretic mechanism naturally accounts for the key observations: directional migration up the salt gradient, enhanced curvature during runs, heading-dependent angular velocity, and the absence of significant changes in tumble statistics. We will include this analysis in the revised manuscript to clarify why diffusiophoretic steering is the dominant mechanism under the present conditions.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

      We thank the reviewer for this helpful suggestion. We agree that the manuscript should more clearly position the proposed mechanism within the broader literature on physically induced microbial transport and swimmer reorientation.

      In the revised manuscript, we have added a discussion comparing our results with prior studies on rheotaxis and viscosity-gradient-induced steering. Specifically, we now discuss the work of Marcos et al. [4], which showed that shear flow can generate a torque on the helical flagellum of Bacillus subtilis, producing rheotactic alignment independent of chemical sensing. We also discuss the work of Stehnach et al. [2], which showed that viscosity gradients can steer Chlamydomonas reinhardtii through asymmetric viscous drag on its two flagella, producing viscophobic turning down the viscosity gradient.

      The mechanism proposed in the present work is distinct from both of these cases. Unlike rheotaxis, it does not require externally imposed shear flow. Unlike viscophobic turning, it does not rely on a substantial viscosity contrast. Instead, we propose that a salt concentration gradient generates differential diffusiophoretic motion of the cell body and flagellar bundle, producing a torque that continuously reorients swimming cells along the gradient. This mechanism therefore identifies salt gradients as a distinct physical cue capable of steering bacteria through surface-mediated transport rather than through flow, viscosity contrast, or canonical chemoreceptor signaling.

      We have added the following text to the revised manuscript: ”Our findings also relate to other physical mechanisms of microbial reorientation. Bacterial rheotaxis, arising from a torque generated by shear flow acting on the helical flagellum, steers Bacillus subtilis independent of any chemical gradient [4]. Viscosity gradients similarly drive a viscophobic turning in the alga Chlamydomonas reinhardtii, where uneven viscous drag on its two flagella produces a torque that reorients cells down the gradient [2], a behavior confirmed by measuring angular velocity as a function of heading angle and revealing a sinusoidal form, ω(θ) = −ω <sub>visc</sub> sin(θ), which resembles what we report here for diffusiophoresis in (Figure S4). This identifies salt gradients, independent of flow or viscosity, as a distinct physical route by which swimming cells can be steered and guided, broadening the set of known nonchemoreceptor mechanisms for directed microbial transport.” References

      (1) V. S. Doan, P. Saingam, T. Yan, S. Shin, A trace amount of surfactants enables diffusiophoretic swimming of bacteria, ACS Nano 14 (10) (2020) 14219–14227.

      (2) M. R. Stehnach, N. Waisbord, D. M. Walkama, J. S. Guasto, Viscophobic turning dictates microalgae transport in viscosity gradients, Nature Physics 17 (8) (2021) 926–930.

      (3) S. Shin, E. Um, B. Sabass, J. T. Ault, M. Rahimi, P. B. Warren, H. A. Stone, Size-dependent control of colloid transport via solute gradients in dead-end channels, Proc. Natl. Acad. Sci. 113 (2) (2016) 257–261.

      (4) Marcos, H. C. Fu, T. R. Powers, R. Stocker, Bacterial rheotaxis, Proceedings of the National Academy of Sciences 109 (13) (2012) 4780–4785.

    1. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists. However, the evaluation is inadequate due to four major weaknesses described below. Despite these limitations, the authors conclude that current general-purpose LLMs lack adequate accuracy, which is already widely recognized. Its key contribution should instead be to provide concrete recommendations for the development of specialized LLMs for this task, which is completely absent. Developing such specific LLMs would be highly valuable, as they could substantially reduce the time required by researchers to analyze signaling networks.

      Strengths

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

      Weaknesses:

      (1) The authors evaluate LLM performance using only three signaling networks: "hypertrophy", "fibroblast", and "mechanosignaling". Given the large number of well-established signaling pathways available, this is not a comprehensive assessment. Moreover, the analysis need not be restricted to signaling networks. Other network types, including metabolic and transcriptional regulatory networks, are already accessible in well-known databases such as KEGG, Reactome, BioCyc, WikiPathways, and Pathway Commons. Including these additional networks would substantially strengthen the evaluation.

      We agree with the reviewer that our evaluation of LLM performance is not comprehensive of all signaling networks, and that the benchmarking was previously limited to signaling networks. The purpose of this study is to benchmark LLMs against peer-reviewed computational models that make testable predictions and are highly validated experimentally, of which these three signaling networks are strong examples. KEGG, Reactome, Biocyc, WikiPathways, and Pathway Commons are databases that contain collections of individual reactions or pathways but are not computational models themselves and have not been experimentally validated in that sense.

      While this study focuses primarily on signaling networks, we agree that it would be useful to evaluate how well LLMs perform in generating networks of another type, for which predictive and validated computational models are available. Therefore, in new Figure 3 we now test the ability of LLMs to reconstruct the E. Coli core metabolic network, as well as test its ability to predict growth on metabolic substrates using flux balance analysis. We find that Claude Opus 4.6, GPT 5.2 Pro, and Gemini 3 Pro Preview perform well at reconstructing reactions from E. Coli core metabolism, but these reactions are not sufficient to predict growth on a variety of substrates. Given this expansion of scope, we replaced “signaling networks” in the title with “biochemical networks”.

      (2) In LLM evaluation, the authors use the gene lists that exactly match those in their "ground truth" networks, thereby fixing the set of nodes and evaluating only the predicted edges. However, in practical research, the relevant genes or nodes are not fully known. A more realistic assessment would therefore include gene lists with both genes present in the ground-truth network and additional genes absent from it, to evaluate the ability of the LLM to exclude irrelevant genes.

      We agree with the reviewer that evaluating the capacity of these LLMs to exclude additional genes is interesting. But because biological networks are always incompletely known, there is no “ground truth” of genes absent from a given network. Therefore, for the most rigorous benchmarking against a “ground truth”, we examine the positive predictive ability of LLMs. However, in response to this comment and point 3 below, we further examine additional measures of performance that include “false positives”.

      (3) The authors report only the recall/sensitivity of the LLM, without assessing specificity. In practical applications, if an LLM generates a large number of incorrect interactions that greatly exceed the correct ones, researchers may be misled or may lose confidence in the LLM output. Therefore, a comprehensive evaluation must include both sensitivity and specificity. Furthermore, it would be informative to check whether some of the "false positives" might in fact represent biologically plausible interactions that are absent from the manually curated "ground truth". Manually generated "ground truth" can overlook genuine interactions, and the ability of LLMs to recover such missing edges could be particularly valuable. This may even represent one of the most important potential contributions of LLMs.

      We agree with the reviewer that additional metrics could inform the evaluation of LLM performance. Therefore, as recommended, we calculated sensitivity, specificity, precision, negative predictive value (NPV), accuracy, and F1 score for each of the network models (hypertrophy, fibroblast, and mechanosignaling). These new results are summarized in confusion matrices shown in a new Supplementary Figures 3, 4, and 5. We performed this additional benchmarking across the 10 replicates for each LLM.

      One limitation of this approach is the substantial class imbalance within these confusion matrices. Because we interpret actual negatives as connections that are not found between any nodes of the ground truth models, there will be >10x more true negatives than any of the other classes. This makes specificity, accuracy, and the negative predictive value less informative.

      To illustrate this point, consider the ground-truth hypertrophy network which contains 191 connections between 106 nodes. The total number of possible connections between any two nodes is 106<sup>2</sup> = 11,236. Given that there are 191 actual positives, that leaves 11,045, actual negatives (as illustrated in the null predictor of Supplementary Figure 3A). As the number of node-to-node connections predicted by LLMs is on the order of a few hundred, the number of true negatives is always in the thousands, often outweighing the true positives in the specificity or accuracy calculations.

      The effects of this class imbalance are illustrated with the results of the “null predictors” in Supplementary Figure 4A, which are hypothetical models that fail to predict any connection between nodes (have predicted positive values of 0). These null predictors have sensitivities of 0 and specificities of 1 and high accuracies and NPVs because of the high true negative rates.

      The precision and F1 scores calculated using these confusion matrices are robust to these class imbalances. Indeed, there is substantial heterogeneity in the number of “false positives” connections generated by the LLMs as illustrated by the precision and F1 scores. We are hesitant to unequivocally label these novel, predicted connections as false positives because, as the reviewer points out, these connections could represent true molecular interactions that were undiscovered at the time of the ground truth models’ conception but have since been demonstrated experimentally and published.

      (4) It is widely known that applying differential equation models to highly complex biological networks, such as the three networks in the manuscript, is meaningless, because these systems involve a large number of parameters whose values can drastically alter the results. As Richard Feynman once said: "with four parameters I can fit an elephant, and with five I can make him wiggle his trunk." Thus, the evaluation of LLMs on "logic-based differential equation models" does not make much sense.

      Differential equation models have been the primary mathematical framework for studying complex systems for decades. We refer readers unfamiliar with differential equations to the Nobel prize-winning work of Hodgkin/Huxley (Physiology or Medicine1963, action potential of neurons), Prigogine (Chemistry 1977, non-equilibrium thermodynamics and pattern formation), John Nash (Economics 1994, game theory and Nash Equilibrium), Merton/Scholes (Economics 1997, dynamics of financial derivatives), and Manabe/Hasselmann (Physics 2021, dynamic modeling of atmosphere and oceans).

      We are confused by the quote of a joke by Richard Feynman about fitting equations to data in the shape of an elephant. While this famous joke is amusing, it is both misattributed (it was a recollection by Enrico Fermi in 1953 of a joke once made by John von Neumann) and deliberately hyperbolic (see https://en.wikipedia.org/wiki/Von_Neumann%27s_elephant). Regardless, the relevance of this joke to our study is unclear, because we are not fitting equations to data. As described in the text, in previous studies we validated the predictions of these three logic-based network models with experimental data that was not used to develop the models.

      Reviewer #1 (Recommendations for the authors):

      (1) All figures are in very poor resolution.

      Thank you for identifying this. We have fixed this issue, which was due to SVG embedding. We now embed as higher resolution PNG and provide full resolution files separately.

      (2) The manuscript does not include data availability or code availability.

      As described in the Methods, all code and data is now available via GitHub (https://github.com/saucermanlab/LLM-network-generation).

      Reviewer #2 (Public review):

      (1) Information on the accuracy of directionality of interaction would help understand if there is a bias towards either a positive or negative association.

      To assess if LLMs are biased in predicting either stimulatory or inhibitory connections, we examined the proportion of stimulatory and inhibitory connections for each ground truth model along with the prediction sets from the different LLMs (Author response table 1). These findings suggest that any directional bias is minimal and not conserved across the different ground truth models. These tables were not included in the revised manuscript.

      Author response table 1.

      Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).

      The primary subset of reaction types that the LLM’s tend to miss are often downstream, cell-type specific, and/or gene regulatory connection (e.g. Figure 1B). This observation is consistent with the fact that the ground truth models were constructed using experimental evidence from specific publications involving defined experimental models and cell types whereas the LLM’s presumably draw from the entire corpus of published literature.

      (2) Do all LLMs capture similar information, or are some LLMs better at capturing certain information than others? Further to this, it would be interesting to look into whether amalgamating information across all three LLMs results in a more accurate network.

      The reviewer asks interesting questions that can be qualitatively answered in Figure 1B and in the network visualizations (Supplementary Figures 1 and 2). These diagrams illustrate predicted connections that are shared between the nodes. To include a more quantitative, comprehensive assessment of this overlap, we have included Venn diagrams showing the extent to which LLMs capture shared information (Supplementary Figures 3-5). One such Venn diagram for the hypertrophy model is included in Supplementary Figure 3C. Indeed, it seems that there is substantial overlap in the “false positive” connections predicted by Claude and GPT. It could be interesting to evaluate if these connections are reflective of newly discovered molecular interactions, as referenced in our response to reviewer one point 3.

      While amalgamating information across all three LLMs might generate a more accurate network, these Venn diagrams illustrate that there remain connections within the ground truth models that are not represented in any of the prediction sets from the LLMs. Indeed, the union of all predicted connections for the hypertrophy network made by any of the 10 replicates from the different LLMs would still lack 33 ground truth connections (Supplemental Figure 3C).

      (3) Would it be possible to retrieve a confidence value of the interactions from the LLMs and conduct Precision, recall rate, AUPR and calibration analyses? These metrics would also help with the benchmarking process.

      We agree with the reviewer that including additional evaluative metrics is instructive. Therefore, for each reference network, we have included precision, specificity, recall, accuracy, and F1 score calculations (Supplementary Figures 3-5). Additionally, we conducted calibration analyses to illustrate the LLM’s reliability. While this latter analysis is interesting, we note that the calibration curves were generated using the frequency of predictions among the 10 replicates as a proxy for confidence. This frequency is highly sensitive to the temperature parameter of the LLMs. Different temperature settings could substantially influence the trajectories of these confidence curves and the distributions of the associated histograms. See Supplemental Figure 3B for calibration analyses of the hypertrophy model.

      We considered performing analyses resembling AUPR as suggested by the reviewer, but we did see a defensible way to vary a “threshold” for the classification of a connection as positive or negative. In this study, a connection predicted by the LLM either does or does not exist within the ground truth mode.

      Typos:

      (1) Generated networks to predict THE CLASSIC "fetal gene program" gene expression.

      Thank you for catching this error. We have corrected it.

      (2) A manually curated network has A functional accuracy of

      Thank you for catching this error. We have corrected it.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. Author response:

      Reviewer #1 (Public review):

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      We agree that the model description should be expanded and that quantitative agreement with the data should be assessed more thoroughly, and we plan to address this during the revision. In the initial version of the manuscript, we aimed to highlight the qualitative agreement of the data with the novel and counterintuitive predictions by the model. While the reviewer is correct that the model "generates different predictions in different regimes," the particular predictions we test (the interaction between the target noise level and the parity of the target and non-target noise levels in Experiments 1-3, and the effects of non-target noise when the target noise is held constant in Experiment 4) hold across regimes (Figure S1 shows this for the former prediction). We aim to further expand on this point in the revision.

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      We will provide more details about the model and how it was fitted to the data. Please note that the model parameters were fitted per subject and condition, not per trial: 2 hue noise parameters,  and , corresponding to the noise of the target and non-target item across 4 noise combination conditions, plus a shared identifiability noise, , across conditions, determining the discriminability of the items along the identifying dimension. This strongly limits model flexibility as only 3 parameters (including the shared  across conditions) are used to create the bias curve for each subject in each condition.

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      We are grateful to the reviewer for raising this point, and we will provide more details on it in the revision. In brief, we believe that it is the standard ideal observer assumption of one sample per trial that is unrealistic and works only in cases when there is a single signal source, so that the samples can be simplified to a single average. Consider that determining a stimulus value is a similar problem for an ideal observer to the one that a researcher who aims to decode neural data from populations of neurons (or fMRI voxels) has to solve. Different populations of neurons would provide responses that match different stimuli – in essence, creating different samples in an ideal observer framework. Thus, even without external noise, the demixing problem would be present when there is more than one stimulus, but internal noise is much more difficult to control, so in our experiments, we used multi-colored stimuli.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      This is an interesting point that we aim to address in the revision.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      We are grateful for the suggestion to expand the discussion of this point and will do so in the revision. The reviewer is correct that for very large d-prime in the identifying dimension, the demixing problem solution is trivial. However, the 2D case does not resolve to the 1D case when d-prime reaches zero. This is because the identifying dimension is still used to identify which item to report—unlike in the 1D case, when the reported dimension is the same as the identifying one. Consider what happens if the observer in our task does not remember at all which stimulus was left and which was right. It would report the other item in 50% of cases, leading to a strong attractive bias.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      We apologize for the confusion. We constrain the discriminability of items in the "identifying" dimension using the  parameter that determines the noise in that dimension for both items. The means in this dimension are fixed at an arbitrary value, as means and noise are interchangeable when considering discriminability. We will revise the description of the fitting procedure accordingly. Regarding the use of the same values in predictions, while possible, we prefer to keep predictions separate from fitting to avoid them becoming postdictions. The curves for the fitted model in Figure 2 already illustrate what the model predicts under the fitted parameter values.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      We are grateful to the reviewer for raising this point, which we will address in the revised discussion. Briefly, we believe that the number of samples in the ideal observer model does not correspond directly to the working memory “slots”.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      We are grateful to the reviewer for raising this point. We both agree and disagree with the reviewer here. The predictions reported come from an earlier paper describing the model (ref. 22). In our opinion, this represents a pure hypothesis-driven approach, where a prediction is formulated first and then tested with subsequently collected data. The model we test is normative, not descriptive; its goal is not to fit the data as closely as possible, but rather to make predictions about internal brain mechanisms. We do not suggest, for example, that demixing is the sole source of biases, so the resulting bias pattern might differ significantly from the predictions. That the model fits the data is, therefore, an additional bonus. At the same time, we agree that it is interesting to test whether the model can explain other parameters of the data. Note that our current fitting procedure was not geared toward this; we optimized the model to explain only the bias curve. In the revision, we aim to test whether the model can also explain the error variability.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      External noise accounts for only a portion of the total noise, as evidenced by behavioral errors. Even for a single item, the total noise consists of the amount of information the observer samples from the stimulus, the variability of these samples (external noise), and early (applied to each sample) and late (applied after integration) internal noise. Therefore, external noise alone might not constrain the model in the right regime. Regarding the identifying-dimension noise, as noted above, we do constrain it with the data. However, we aim to explore these points further in the revision.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      We apologize for the confusion. We will further clarify and perhaps reassess the treatment of swap errors in the revision. The model itself produces swap errors when the stimuli sources are misidentified.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

      Indeed, this is a valid point, and it's why we use a multi-initialization strategy. For each simulation of a single trial sample set (e.g., 100 random samples), we use a large number of initial points (50 in the initial submitted manuscript) to ensure the obtained EM solution is truly optimal. Additionally, we conduct a large number of simulated trials (10,000 for each parameter combination) to ensure the accuracy of the bias distribution we obtain.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

      We agree and we aim to provide a comparison with alternative models and an expanded description of the fitting results in the revision. Note, however, that the majority of existing models are descriptive, while we believe that as a normative model, the Demixing Model should be compared with other normative models, thus limiting the selection of competitors significantly.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.