10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This is a valuable study of the effects of selective or broadband amyloid deposition in medial septum (MS) cholinergic neurons on neuropathology, cognition, sleep, and hyperexcitability in aging mice carrying risk factors for Alzheimer's disease (AD). The investigation is still incomplete, with some weaknesses related to conceptualization and methodology that need to be addressed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

    4. Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ⁻ is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ⁻ and AAV‑Aβ⁺ is marked by asterisks (***) placed above AAV‑Aβ⁺. In addition, the single asterisk above KI-Aβ⁺ does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ⁻ vs WT‑Aβ⁻ or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ⁺ mice lack a preference for the novel object compared with AAV‑Aβ⁻ controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ⁻, and the notation over AAV‑Aβ⁺ is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ⁺, KI-Aβ⁺⁺, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Thank you for the positive assessment and for recognizing the novelty of our findings, the complementarity of our experimental approaches, and the breadth of our multi-modal readouts. We will address the weaknesses raised below point by point.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      This is an insightful comment. We fully agree that early functional changes preceding structural degeneration represent an important and exciting avenue, and we will explicitly acknowledge this in the revised manuscript, including the relevant literature on early cholinergic hyperactivity. Examining earlier time points in our model to capture such changes is a compelling perspective that we will discuss as a key direction for future work.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      We acknowledge this limitation and will increase the sample size for the tracing experiments in the revised manuscript. However, we wish to clarify that the tracing was performed in a separate cohort of young animals specifically to characterize baseline MS cholinergic projections independently of any amyloid-related disturbances. While we agree that replicating these experiments in APP mice would be of great interest, this falls outside the scope of the current study and will be highlighted as an important direction for future work.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      We agree with this pertinent suggestion. We will perform and report region-specific analyses of the hippocampal formation in the revised manuscript, in line with our discussion of specific amyloid accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      We respectfully maintain our analytical approach. As the test phase was terminated upon reaching a predefined cumulative exploration time of 20 seconds rather than using a fixed trial duration, computing a discrimination index is not appropriate in this context, as total exploration time is constrained by design. We instead followed the validated protocol described by Leger et al. (2013, Nature Protocols; PMID: 24263092), which controls for inter-individual differences in exploratory motivation by fixing cumulative exploration time, ensuring equivalent sampling conditions across animals. We will clarify this methodological choice in the revised manuscript.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

      We appreciate this constructive suggestion. We will provide additional methodological detail on interictal spike detection, include representative examples, and perform a vigilance state-specific analysis in the revised manuscript. We agree this will provide valuable mechanistic insight and will update the reference accordingly.

      Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Thank you for this positive assessment and for recognizing the rigor of our experimental design, the value of our multi-modal approach, and the mechanistic significance of our cholinergic lesion comparison. We will address the weaknesses below point by point.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

      Regarding amyloid broadcasting, we fully acknowledge that the precise mechanism remains to be elucidated; while this was not a primary objective of the study, it represents a fascinating and unexpected finding that we will discuss more carefully as an open question for future investigation. Regarding the causal relationship between MS<sup>ChAT</sup> cell loss and the observed phenotypes, we agree that the specific circuits involved remain to be fully resolved; however, we would like to emphasize that the convergent evidence from our three complementary models (and in particular the recapitulation of cognitive and REM sleep deficits by selective cholinergic ablation) provides strong causal support for MS<sup>ChAT</sup> neuronal loss as a key driver of these phenotypes, independent of amyloid deposition per se.

      Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Thank you for acknowledging the compelling conceptual premise of our study, the thoughtful and broad experimental framework, and the value of our longitudinal sleep analysis and lesion comparison. We will address all concerns raised below point by point.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      This is a valid point. We agree that the term “prodromal phase” requires more careful framing in the context of animal models, and we will revise the title and relevant statements accordingly. We would like to note, however, that the REM sleep disturbances we report seem to emerge prior to overt cognitive decline in our longitudinal analysis, which we consider to reflect a prodromal-like feature of the model. Nevertheless, we will adopt more precise terminology throughout the manuscript to avoid conflation with clinical staging criteria.

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      We appreciate this comment and will carefully nuance our opening statement to better reflect the current state of evidence. However, we respectfully note that several reports, including Pase et al. (Neurology, 2017; PMID: 28835407) and Ibrahim et al. (Sleep, 2024; PMID: 38001022), have demonstrated that REM sleep loss is associated with increased risk of incident neurodegenerative disorders, particularly Alzheimer's disease, supporting the broader validity of our framing. We will revise the statement to more accurately capture the complexity of this relationship while preserving its scientific relevance.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      We agree and will revise this sentence to better reflect the multi-system complexity of sleep/wake regulation, acknowledging the interplay between multiple neurotransmitters and neuromodulatory systems.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      We agree and will revise this sentence to clarify that ACh is a crucial component of a broader REM sleep-generating circuitry, rather than a sole requirement, while better contextualizing its specific contribution to REM sleep regulation.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      We will revise this statement to more accurately reflect the evidence. We would like to note, however, that while the Pase et al. (2017) cohort included a mixed dementia population, 75% of incident dementia cases (24 out of 32) were consistent with Alzheimer's disease, lending meaningful support to the relevance of REM sleep alterations specifically in the context of AD. We will ensure this nuance is clearly conveyed in the revised manuscript.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      We agree and will revise this section to acknowledge the well-established contribution of tau pathology to BF cholinergic neuronal vulnerability in human AD, including its early accumulation at Braak stages I-II. We wish to clarify, however, that the present study focuses exclusively on amyloid-driven mechanisms, and the introduction will be revised to ensure readers do not infer a purely amyloid-dependent mechanism in the broader human disease context.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      As noted in our response to weakness (1), we will systematically revise the manuscript to replace “prodromal phase” with more precise terminology that clearly distinguishes our animal model findings from clinical disease staging.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      Thank you for raising these important points. Regarding age matching, we acknowledge that chronological ages are not perfectly aligned across groups; however, we wish to emphasize that the duration of amyloid pathology is carefully matched across models. Indeed, AAV injection in MS<sup>ChAT</sup>-AppNL-G-F mice marks the onset of amyloid expression, directly paralleling the onset of pathology from birth in App<sup>NL-G-F/NL-G-F</sup> knock-in mice. We believe pathology duration represents the most biologically relevant variable for comparison in this context, and we will clarify this in the revised manuscript. Regarding the MS<sup>ΔChAT</sup> cohort, animals were culled upon reaching a comparable degree of REM sleep loss, providing a functionally meaningful matching criterion. Finally, regarding sham-operated controls, we acknowledge this limitation; however, based on our experience, surgical procedure alone has negligible effects on the cellular populations under study, and the inclusion of an additional sham group across all experimental cohorts would have required a prohibitive number of animals, raising significant ethical concerns under the 3R principles. We will address these points more explicitly in the revised manuscript.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      Thank you for this observation, we will correct the figure citation accordingly in the revised manuscript.

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      We appreciate this important observation. We will carefully re-examine our tracing data to determine whether MS<sup>ChAT</sup> axons specifically innervate the LHA, and whether amyloid deposition was detectable in this region in our MS<sup>ChAT</sup>-App<sup>NL-G-F</sup> model. We agree that clarifying the potential anatomical relationship between MS<sup>ChAT</sup> projections and LHA-MCH neurons is important to properly position our proposed circuit mechanism relative to this well-established REM sleep-promoting node, and we will address this in the revised manuscript.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      Thank you for catching this inconsistency. We confirm that this is an error in the manuscript: line 125 should read “15-16 month-old” referring to the chronological age of the animals at the time of perfusion, rather than “13-14 months,” which corresponds to the duration of AAV expression. We will correct this in the revised manuscript.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      Thank you for flagging this ambiguity. We will revise the sentence to clearly distinguish between AAV-injected and global knock-in cohorts in the revised manuscript.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      The supplementary information referenced in the main text will be provided in full in the revised manuscript, together with analyzed datasets and analysis scripts, in accordance with eLife's data sharing policy.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      We agree that a comprehensive description of our histological quantification pipeline is essential. We will provide full methodological details in the revised manuscript, including: the metric used for Aβ load quantification, ROI definitions, Fiji/ImageJ thresholding algorithms (accounting for batch effects and autofluorescence normalization), as well as the strategy for identifying and counting ChAT-positive neurons, section spacing, number of sections per animal, hemisphere coverage, normalization, and blinding procedures. All ImageJ scripts will be made available to ensure full transparency and reproducibility.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ<sup>-</sup> is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ<sup>-</sup> and AAV‑Aβ<sup>+</sup> is marked by asterisks (***) placed above AAV‑Aβ<sup>+</sup>. In addition, the single asterisk above KI-Aβ<sup>+</sup> does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ<sup>-</sup> vs WT‑Aβ<sup>-</sup> or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      We acknowledge that the current statistical annotations can be difficult to interpret when multiple experimental groups are displayed within a single plot. While the notation is consistent across figures, we agree that clarity can be improved, and we will revise all figure annotations to explicitly link significance indicators to their corresponding pairwise comparisons, using standardized brackets throughout, with all contrasts clearly specified in the figure legends.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ<sup>+</sup> mice lack a preference for the novel object compared with AAV‑Aβ<sup>-</sup> controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ<sup>-</sup>, and the notation over AAV‑Aβ<sup>+</sup> is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ<sup>+</sup>, KI-Aβ<sup>++</sup>, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      We wish to clarify that in Figure 5D, the absence of novel object preference manifests as equal exploration of both objects (approximately 10 seconds each), such that comparisons against the 10-second chance level reflect this lack of preference. We will revise the annotations and figure legend to make the statistical comparisons explicit and fully consistent with the results text.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      This is a constructive feedback, and we agree that the Discussion would benefit from pruning and recalibration. We will streamline it to focus on the most evidentially supported points, particularly those linking cholinergic loss to REM sleep and cognitive outcomes, while toning down or removing speculative statements regarding disease-stage equivalence, prion-like spreading mechanisms, and direct therapeutic implications.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      Thank you for raising this important point. We will expand the Discussion to address this apparent contradiction with the human literature. Indeed, while increased anxiety is reported in early AD stages, it tends to normalize or decrease at later stages (Botto et al., 2022; PMID: 35461471), which may partly reconcile our findings. In addition, anxiety-like phenotypes are highly inconsistent across AD mouse models, varying with model type, age, sex, and behavioral assay (Pentkowski et al., 2021; PMID: 33979573). Anxiety-related changes in human AD may reflect damage to brain regions beyond the MS cholinergic system, involving additional circuits and mechanisms not captured by our model.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

      We will nuance our wording at line 654 by replacing “comparable durations of amyloid pathology” with “comparable amyloid burden,” acknowledging that while both cohorts were aged for matched durations, the kinetics of amyloid progression may inherently differ between an AAV-driven focal model and a germline knock-in model.

    1. eLife Assessment

      This important study investigates how viral macrodomains in dual-host viruses functionally contribute to their infection in the vector, which was previously understudied. The findings that chikungunya virus (CHIKV) macrodomain mutants have unique impacts on virus replication in mammalian cells and in live mosquitoes are convincing and significant; however, the study can be strengthened by further investigating the role of ADP-ribose binding and potential other compensatory mutations. The work will be of interest to virologists and biologists who study macrodomains and ADP-ribosylation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

    3. Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      We thank the reviewer for raising this important point. The emergence of second-site mutations at residue D31 in Vero cells was actually observed in two independent experiments, each initiated by transfection of viral RNAs encoding either the N24A or N24D mutant. In the second experiment, additional timepoints were sampled for sequencing. Comparison of the two experiments revealed that, in the second replicate, only the D31N variant was recovered, whereas D31H was not detected. We will include the results of both independent experiments in the updated Figure 1 and will modify the text accordingly to improve clarity.

      In addition, we will present new experiments performed in other cell types that further confirm the reproducibility of D31 mutation emergence across different cellular contexts. Specifically, we will include results from new viral RNA transfections in BHK-21 mammalian cells and C6/36 mosquito cells, which will be included in the updated Supplementary Figure 1.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      We thank the reviewer for this correction. The sequencing was indeed focused on the region of the macrodomain surrounding the introduced mutations, covering the first 119 amino acids of nsP3. This region was selected to confirm the stability of the introduced mutations at position 24 and to monitor the emergence of potential second-site mutations in its immediate vicinity. We acknowledge that this approach does not rule out the emergence of compensatory mutations elsewhere in the macrodomain or in the full-length nsP3 protein. We will correct the text accordingly to accurately reflect the region that was analyzed.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      This is an excellent point. To address it, we introduced N24A and N24D mutations into an infectious clone of the Indian Ocean strain, a representative of the ECSA lineage, and assessed the emergence of D31 second-site mutations during viral stock production. As observed with the Caribbean strain, the D31N secondary mutation consistently emerged in both N24A and N24D Indian Ocean mutant viruses. These results will be included in the updated Supplementary Figure 1 and in the main text.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      We thank the reviewer for raising this point. We acknowledge that previous studies have reported limited infection of A549 cells by certain CHIKV strains. However, in our hands, with our viral stocks and under our experimental conditions, we observe an increase in viral titers following infection of A549 cells with WT Caribbean strain virus, indicating productive viral replication. We also confirmed that the Indian Ocean strain replicates in A549 cells under the same conditions, and we will include growth curve data for this strain in the updated version of the manuscript. Importantly, we did not observe detectable cytopathic effects in A549 cells infected with either WT or N24 mutant viruses, which is consistent with the notion that CHIKV replicates less efficiently in this cell line compared to other mammalian cell lines. We will include a sentence acknowledging this in the discussion of the revised manuscript.

      Regarding the suggestion to use human fibroblasts, we respectfully note that the primary focus of this study is the role of macrodomain catalytic activity in the mosquito host, and the experiments in A549 cells were performed to confirm the known importance of the macrodomain in interferon-competent mammalian cells. We therefore consider the current data in A549 cells sufficient to support this conclusion within the scope of the manuscript.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      We acknowledge that the full viral genome was not sequenced, and we cannot rule out the presence of additional mutations elsewhere in the genome that could contribute to the observed phenotype. However, we note that the enhanced dissemination phenotype was also observed in independent experiments performed with the Indian Ocean strain mutant viruses in Ae. albopictus, which were generated independently from the Caribbean strain stocks. The consistency of the phenotype across two independently generated sets of mutant viruses from different CHIKV lineages strongly supports the conclusion that the enhanced dissemination phenotype is linked to the macrodomain mutations. These new data will be included in the updated manuscript.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

      We thank the reviewer for pointing this out. The text will be modified accordingly to accurately reflect that we assessed transmission potential, based on viral dissemination to heads, rather than actual transmission.

      Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      We agree that the mutants may indeed have stronger binding for modified substrates than the WT protein; however, given that DSF is not a quantitative measure of binding affinity and that free ADPr is not the relevant ligand (in fact we do not know the relevant ADPr-modified molecule), we have refrained from speculating further than saying in the discussion:

      “The progressive selection of D31H over D31N in the mosquito host further suggests that subtle differences in ADP-ribose substrate recognition may influence viral fitness in the mosquito environment in ways that are not yet understood.” (Lines 368-370)

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

      Regarding the title, we agree with the reviewer that it could be misleading, as both mutants lack catalytic activity yet show distinct phenotypes in mosquitoes. We will therefore modify the title to: "Loss of macrodomain catalytic activity modulates Chikungunya virus dissemination and transmission potential in Aedes mosquitoes", which more accurately reflects that the observed phenotypes arise as a consequence of the loss of catalytic activity at position N24.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      We thank the reviewer for this important comment. Regarding the generation of viral stocks in A549 cells, transfection of N24A and N24D viral RNAs into A549 cells did not yield sufficient viral titers to produce usable stocks. However, as described in our response to Reviewer #1, we confirmed the reproducible emergence of D31 second-site mutations in C6/36 mosquito cells and BHK-21 mammalian cells, which will be included in the updated Supplementary Figure 1. These results demonstrate that D31 mutations spontaneously emerge in stocks generated directly in mosquito cells, addressing the reviewer's concern.

      Regarding the question about insect-specific alphaviruses, we examined the sequence at position 31 using a multiple sequence alignment of 14 alphaviruses with diverse host ranges, including dual-host, insect-specific, and aquatic alphaviruses. We observed that Yada Yada virus (GenBank: QGR15362.1) has an asparagine (N), Tai Forest alphavirus (GenBank: YP_009333615) has an aspartic acid (D), Mwinilunga alphavirus (GenBank: BBC45634.1) has an aspartic acid (D), Eilat virus (GenBank: QBG67155.1) has an aspartic acid (D), and Agua Salud alphavirus (GenBank: QEV83787.1) has a lysine (K) at this position. These results suggest that insect-specific alphaviruses do not share a conserved residue at position 31, and therefore no clear conclusion can be drawn regarding a direct correspondence with the compensatory mutations observed in our study. This sequence alignment with the corresponding text will be included as supplementary data in the revised manuscript.

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      We thank the reviewer for this observation. We acknowledge that the prevalence of WT virus infection shows variability between experiments and timepoints. Based on our experience with infectious blood meal experiments, this variability is sometimes observed between independent experiments even under identical experimental conditions and with the same virus. In this particular case, each set of experiments was performed with independently produced viral stocks. Specifically, for the experiments shown in Figure 3, WT, N24A-D31N, and N24D-D31H/N viral stocks were produced in parallel from transfection of viral RNAs, and the same stocks were used for sequencing, growth curves, and mosquito infections. Subsequently, when we generated single D31H and D31N mutant viruses, a new WT viral stock was produced in parallel with the D31 mutant stocks, and these independently produced stocks were used for the experiments shown in Supplementary Figure 2. Differences in absolute infection rates between experiments are therefore expected, as they reflect both the use of independently produced viral stocks and the inherent variability in the efficiency of midgut infection and dissemination between mosquito cohorts. Importantly, the comparisons between WT and mutant viruses are always made within the same experiment, using stocks produced in parallel.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    1. eLife Assessment

      This important study identifies and characterizes a set of amino acid states that can rescue protein function in the presence of substantially deleterious mutations. Some of these super-compensatory substitutions also confer substantial mutational robustness, with broader implications for understanding epistasis, protein evolution, and protein engineering. The evidence is convincing, supported by a creative reanalysis of a large deep-mutational-scanning dataset, statistically rigorous treatment of anticipated error rates, experimental validation, and analyses of additional proteins, although clearer presentation and deeper investigation of the evolutionary implications and structural mechanisms would further strengthen the study.

    2. Reviewer #1 (Public review):

      Summary:

      The study identifies and characterizes a set of amino acid states that make the protein robust to other mutations, to the point of being able to compensate mutations that render wildtype proteins entirely non-functional. The study uses a previously published dataset and uses it to find and study such super-compensators. It then analyzes the biophysics and fitness landscape structure of what may be behind the compensation, identifying stability as an important parameter that, nevertheless, is not sufficient to explain all of the compensatory effect. These findings have important implications for our understanding of protein evolution, with these super-compensators possibly acting in a role of "permissive mutations" and opening up evolutionary trajectories that may be closed without them. Perhaps the identification of such super-compensator substitutions can be incorporated into various protein design approaches.

      Strengths:

      The paper presents a compelling case with a rigorous analysis of the expected error rates of observation. While not unique, the current state-of-the-art in the field typically does include experimental error rate estimation like this work. The paper also does a good job in exploring the issue, including looking at plausible biophysical basis of super-compensators.

      Weaknesses:

      The paper lacks rigor in talking about evolutionary-related issues of the state of the fitness landscape. As an example, the paper mentions that these super-compensators flatten the landscape. While I understand where this is coming from, I think that the fitness landscape in this context is a static entity and cannot be flattened or otherwise altered. A much more accurate description is that a sequence with a super-compensator is located in a flatter-than-expected segment of the fitness landscape, or on a flat fitness ridge. These issues are more semantic in nature, and while the manuscript would benefit from it being shown to an expert in molecular evolution or fitness landscapes, this issue does not take away from the importance of the results.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents an interesting and conceptually valuable analysis of compensatory evolution using a large combinatorial deep-mutational-scanning dataset for yeast His3p.

      Strengths:

      I particularly like the identification of "super compensatory" substitutions that improve fitness across diverse genetic backgrounds and apparently reduce the sensitivity of the local fitness landscape to subsequent mutations. The work connects epistasis, protein stability, mutational robustness, and evolvability in a clear and potentially broadly relevant manner.<br /> The authors provide several complementary lines of evidence in support of this central conclusion. In particular, the new experimental validation of S189A is an important strength because it directly demonstrates that a predicted super compensator can buffer the effects of diverse deleterious substitutions, while analyses of additional DMS datasets from other proteins and assay systems suggest that the phenomenon is not restricted to the original His3p landscape.

      Weaknesses:

      The structural analysis currently relies primarily on correlations with RSA, weighted contact number, conservation, and Rosetta-predicted changes in folding or binding energy. For super compensators, the mechanistic evidence is largely limited to predicted stabilization and individual examples, such as the proposed salt bridge between 110D and R112. I believe that the newly developed structure-aware deep-learning approaches could provide useful information on the mechanism of super compensators. For example, an inverse-folding model such as ESM-IF1 could score complete multi-mutant sequences conditioned on the His3p backbone and test whether adding a super compensator restores sequence-structure compatibility across backgrounds. More recent multimodal mutation-effect or stability models could similarly be used to cross-check the Rosetta results, including models that explicitly support combinatorial mutations. I would not recommend simply comparing AlphaFold confidence scores between mutants, because current structure predictors are not necessarily sensitive to subtle mutation-induced energetic or conformational changes.

      The manuscript states that the pipeline was applied to 217 ProteinGym datasets and concludes that super compensators are broadly distributed across proteins and assays. However, this central generalization is described in only a few sentences and is largely relegated to Figure S7. The Methods do not explain which datasets contained sufficient combinatorial mutants to calculate compensatory ability or buffering, how many genotype pairs or quadruplets were available per substitution, or how differences in assay scale and library design were handled. This point requires clarification because supercompensation is inherently a background-dependent property and cannot be established from single-mutant measurements alone. ProteinGym is widely used as a substitution-effect benchmark, and many of its constituent assays primarily contain single substitutions; for example, an analysis of an earlier ProteinGym collection reported that 76 of 87 assays contained only single substitutions. It is therefore unclear how the same compensatory-interaction pipeline could be applied uniformly to all 217 datasets.

      The analysis of 335 His3p orthologs in Discussion is potentially very interesting, but co-occurrence between super compensators and putatively deleterious amino-acid states does not by itself demonstrate evolutionary compensation. Closely related species share substitutions through common ancestry, and both states could be associated with a particular lineage or ecological context. A tree-aware analysis would considerably strengthen this result. The authors could reconstruct ancestral states and ask whether acquisition of a super compensator tends to precede or accompany otherwise deleterious substitutions. Alternatively, they could use phylogenetically informed permutations that preserve substitution frequencies and shared ancestry.

    4. Reviewer #3 (Public review):

      The manuscript by Jiang and co-authors presents an analysis of experimental measurements (about 400k variants) from a deep mutational scan of the HIS3 enzyme. The authors assess the ability of a genotype to be "rescued" and show that this depends on mutation sites (in particular their solvent accessibility) and mutation effects (should be mild on folding stability or binding affinity). They further identify a set of super-compensatory mutations, and their results suggest that these mutations flatten the fitness landscape.

      This finding is interesting and likely of interest to a broad community. The analysis seems sound.

      However, I have a number of major concerns regarding the presentation and positioning of the work.

      (1) It would improve the manuscript to clarify the present contribution with respect to a previous study by the same authors, namely Pokusaeva et al. 2019. Did the authors apply the same protocol to generate a new library of mutants, or did they re-analyse an already published library? If the library is not new, ambiguous sentences like "Nevertheless, to our knowledge, the His3p library remains one of the largest and most comprehensive resources that contains multi-site mutants" should be reformulated.

      (2) Pokusaeva et al. 2019 is cited for the library and also for the deep neural network. It would be beneficial to briefly describe the architecture, the inputs and outputs, and the training procedure. Was the network trained on the current library? What is the purpose of this network? It looks more like an additive linear model (except for the global sigmoid) than a deep neural network. How does it relate to global epistasis models? The sigmoid function is designed to capture plateauing effects; doesn't that introduce some circularity issue in the reasoning?

      (3) Are the super-compensatory mutations observed (conserved) across evolution? Beyond the fact that they are accompanied by mildly deleterious mutations in natural sequences. Can we predict them with variant effect predictors?

      (4) The AAindex mention should be accompanied by a citation.

      (5) Equations should be numbered. WCN formula seems to contain misformatting issues.

      (6) A more explicit description of the structural data analysed (which PDB entry?) should be provided.

      (7) I believe the citation Van Cleve and Weissman 2015 for the ProteinGym benchmark is incorrect. Additionally, is the Rosetta citation adequate?

      (8) How is the definition of rescueability sensitive to the threshold choice?

    1. eLife Assessment

      This important study investigated whether the adoption of different explicit strategies influences implicit recalibration during visuomotor adaptation. Through a series of increasingly controlled experiments, the authors demonstrated that implicit recalibration is relatively insensitive to the specific type of explicit strategy employed but is influenced by the variability of strategic motor plans. However, the evidence supporting these conclusions is currently incomplete. The findings will be of interest to cognitive scientists studying motor control and its relationship to higher-level cognitive processes, provided that the key claims are supported by more rigorous analyses and experimental approaches.

    2. Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

    3. Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback).

      Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

    5. Author response:

      Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      We appreciate the reviewer’s thoughtful and comprehensive summary of our study.  One thing we would like to clarify is that the non-critical targets used in Experiment 1 received the same type of online continuous feedback as the critical target.Therefore the extent of implicit recalibration should be comparable between the critical and non-critical targets in Experiment 1. In contrast, for Experiment 2, we tightened the control for implicit recalibration at the non-critical targets by: 1) delivering delayed endpoint feedback for all non-critical targets while keeping the online continuous feedback for the critical target, which is known to suppress implicit recalibration, and 2) we widened the spatial gap between the critical and non-critical targets, to minimize spillover 2) As a result, the implicit recalibration was diminished at the non-critical targets, contributing to the shrinkage of the generalization curve around the critical target. 

      One other issue that we would like to make clear is that Experiment 1 was not a straight replication of a previous study, at least to our knowledge. We believe that the reviewer is referring to our previous study (McDougle and Taylor 2019), which found broader generalization for algorithmic strategies (Experiment 4). However, in that study cursor feedback was always delayed. As such, the observed broader generalization was most likely due to the strategy itself and not implicit recalibration. 

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      We are glad to know that our primary finding and conclusion appears reasonable. While we acknowledge that the finding doesn’t appear to be particularly exciting at face value, it does speak to larger questions regarding the independence of different learning systems and how just statistical or surface-level differences in training can result in relatively large differences in apparent behavior that could be easily misinterpreted as the result of system interactions.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      This is a fair concern, as RT is an indirect marker of strategy use and, by itself, cannot establish that participants used different strategies at the Critical target. Prior work, however, provided guidance for the design of our experimental manipulations. Algorithmic strategies, in which an aiming solution is computed online, are associated with longer RTs, whereas retrieval of a previously cached stimulus–response association produces substantially shorter RTs (McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Moreover, caching becomes increasingly difficult as the number of target-specific solutions increases, particularly beyond approximately four targets (Velazquez-Vargas and Taylor, 2024; Bejjanki and Taylor 2026). Our manipulation was designed around these findings: participants in the Algorithmic condition learned the 45° rotation across 10 targets, whereas participants in the Retrieval condition repeatedly encountered the 45° rotation only at the Critical target. As expected with this experimental design, RTs at the Critical target were significantly longer in the Algorithmic condition across all three experiments.

      We agree with the reviewer, however, that this RT difference could in principle reflect a more general carryover of cognitive demands from the Non-Critical targets rather than online computation at the Critical target itself. We can address this possibility more directly by asking whether RT at the Critical target exhibits the parametric signature expected of an algorithmic process. A defining feature of mental rotation is that RT scales with the magnitude of the computed aiming solution (Georgopoulos and Massey, 1987; Bhat and Sanes, 1998; McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Although rotation magnitude was fixed in the present experiments, participants’ actual reach angles varied naturally from trial to trial. Indeed, McDougle and Taylor (2019) originally demonstrated this relationship using actual reach angle rather than imposed rotation magnitude. We can therefore test whether trial-by-trial RT covaries with reach angle at the Critical target in the Algorithmic condition but not in the Retrieval condition. Such a relationship would be difficult to explain as a nonspecific carryover of cognitive load and would instead provide direct evidence that preparation time at the Critical target reflects the computation of the aiming solution.

      We also observe a second, independent difference at the Critical target: reach angles are consistently more variable in the Algorithmic condition than in the Retrieval condition across all three experiments (Figure S6). This pattern is consistent with repeated online computation producing variability in the selected aiming solution, whereas retrieval of a cached stimulus–response association produces a more stable response. Importantly, a general carryover account based solely on increased cognitive demands does not readily explain why movements to the Critical target should also be systematically more variable. Nor is this pattern easily explained by a speed–accuracy tradeoff: the Algorithmic group had more preparation time yet nevertheless exhibited greater variability. Consistent with this notion, Velázquez-Vargas and Taylor (2024) found that retrieving cached solutions produced less variable and more precise movements than movements that are not cached in the memory trace. 

      Third, we can conduct additional analysis to compare RT variability between algorithmic and retrieval groups at the critical target location. According to Logan instance theory (Logan, 1988), the retrieval of cached stimulus-response associations produces a stable RT profile, in contrast, trial-by-trial algorithmic computation can result in more variable trial-by-trial RT differences. 

      While we agree that RT differences alone should not be taken as definitive evidence of distinct strategies, our specific experimental design, the longer and more variable RTs at the Critical target, the greater trial-by-trial variability at that same target, and, if confirmed, a parametric relationship between RT and reach angle provide converging evidence that participants in the two conditions relied on different strategy implementations when preparing movements to the Critical target.

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

      We appreciate the reviewer’s thoughtful comment and agree that this is an important distinction. First, we would like to clarify the relationship among reaching variability, movement-plan variability, and the width of the implicit generalization function. Previous work has shown that when implicit recalibration occurs at a particular target location without an explicit aiming strategy, its generalization across the workspace can be described by a Gaussian-shaped function centered near the trained target location (Morehead et al., 2017). When an explicit aiming strategy is involved, however, implicit recalibration is centered closer to the planned aiming location rather than the visual target itself (McDougle et al., 2017). Thus, implicit recalibration is greatest near the direction in which the movement is planned, and a broader spatial distribution of movement plans can, in principle, produce a broader aggregate implicit generalization function. In the present study, we therefore use trial-to-trial variability in endpoint hand angle as a behavioral proxy for variability in planned movement direction, while recognizing that endpoint variability may also contain contributions from execution-related noise.

      This framework motivated the progression from Experiments 1 to 3. In Experiment 1, participants in the algorithmic condition exhibited substantially greater reaching variability, consistent with the idea that they sampled a wider range of movement plans across trials. Because error feedback associated with these different movement plans can induce implicit recalibration around each planned direction, greater variability in strategy use could contribute to the broader implicit generalization observed in the algorithmic group. In Experiment 2, we imposed stricter controls on spillover from noncritical targets, which reduced the overall breadth of generalization; nevertheless, model fits still suggested a modestly broader implicit generalization function in the algorithmic group, consistent with the remaining difference in reaching variability.

      Experiment 3 was designed to further reduce the direct influence of strategic variability on the induction of implicit recalibration by using a modified error-clamp paradigm. Error-clamp feedback has been shown to elicit implicit recalibration independently of task success and the participant’s explicit strategy (Morehead et al., 2017). We therefore used error-clamp feedback as an incidental signal to induce implicit recalibration while participants implemented either algorithmic or retrieval-based strategies. Importantly, however, Experiment 3 did not completely eliminate between-group differences in reaching variability: the algorithmic group continued to show greater variability around the critical 45-degree location than the retrieval group. We agree with the reviewer that, in principle, comparable generalization widths could arise if this broader distribution of movement plans were offset by a narrower local generalization of implicit recalibration in the algorithmic group. 

      However, not all variability is equivalent—or well described by a Gaussian distribution. Depending on the direction of the skew, variability in aiming can produce different effects on the implicit recalibration function, appearing as either broader generalization or greater amplitude. These effects are difficult to appreciate in Figures 3 and 5 for Experiments 2 and 3, respectively. We therefore sought to illustrate this more clearly in Figure 6, which shows how differences in the underlying aim distributions can shape the resulting implicit recalibration function in directionally complex ways. What complicates matters further is that implicit recalibration can asymptote (Morehead et al., 2017; Kim et al., 2018; Wilterson & Taylor, 2021). As a result, plan-based generalization can distort the implicit recalibration function in different ways depending on which side of the aim the error falls. The block-by-block analysis, suggested by reviewer 2, may shed light on this issue because we can get a sense if implicit recalibration has reached asymptote. 

      At a minimum, in the revised manuscript, we work to make clearer that subtle changes in the reach distribution may have a corresponding impact on the shape of implicit recalibration’s generalization function.  

      Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      We appreciate the reviewer raising this potential confound. We agree that the five pre-exposure trials in the Retrieval condition introduce a small difference in initial task familiarity. However, this account makes a straightforward prediction: if the shorter RTs in the Retrieval condition simply reflect five additional trials of general task experience, then the RT difference should disappear once the Algorithmic group has received a comparable amount of practice.

      We can test this directly by comparing the five pre-exposure trials in the Retrieval condition with the first five trials of the Algorithmic condition. We will also compare these pre-exposure trials with a later five-trial window from the Algorithmic condition to determine whether additional practice substantially reduces Algorithmic RTs. Assuming the observed pattern is as expected, RTs in the Algorithmic condition remain substantially longer even after considerably more than five trials of practice. Thus, the group difference cannot be explained simply by the Retrieval group having five additional trials of task familiarity. This persistent RT difference, together with our prior work showing characteristic RT differences between algorithmic computation and retrieval of cached aiming solutions, supports our interpretation that the groups relied on different strategy implementations.

      We are less certain what the reviewer means by the “early training advantage.” If this refers to angular error, we agree that the Retrieval group shows somewhat better performance very early in training, but this difference is not a central focus of the present study and largely disappears by the second or third training block. We will clarify the text so that we do not overinterpret this transient difference.

      If instead the reviewer is referring to RT, then the matched-trial analysis directly addresses the concern. Five additional familiarization trials could plausibly produce a brief initial advantage, but such an effect should dissipate within a small number of subsequent trials. In contrast, the RT difference between the Algorithmic and Retrieval conditions remains robust throughout training. We therefore do not think that general task familiarity provides a sufficient explanation for the observed RT differences.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      The reviewer raises an interesting possibility. In McDougle and Taylor (2019), however, the transition from algorithmic computation to retrieval occurred in a condition with only two targets in the task set. With repeated practice, participants needed to retain only two target-specific aiming solutions, making it feasible to replace online computation with retrieval of cached stimulus–response associations. By contrast, the Algorithmic condition in the present study contained 10 target locations. Our prior work suggests that caching becomes increasingly difficult once the number of target-specific solutions exceeds approximately four, at least over the timescale of several hundred trials (Velázquez-Vargas and Taylor, 2024; Bejjanki and Taylor, 2026). Thus, our task was designed to maintain pressure toward an algorithmic strategy throughout training.

      The reviewer nevertheless raises a more specific possibility that is not ruled out simply by the size of the target set: participants might selectively cache the aiming solution for the frequently sampled Critical target while continuing to use an algorithmic strategy at the remaining targets. We think the existing behavioral data argue against such a clear transition.

      First, reaction times at the Critical target in the Algorithmic condition remained substantially longer than those in the Retrieval condition throughout training. If participants had cached the aiming solution for the Critical target, we would expect preparation times at that location to be the same as the Retrieval condition. However, the RTs for the Algorithmic and Retrieval conditions are significantly different in the last block of training. 

      Second, within the Algorithmic condition, reaction times at the Critical target remained similar to those at the Non-Critical targets. Selective caching of the Critical target predicts a different pattern: preparation should become faster at the Critical target than at the surrounding locations, where participants would still need to compute the appropriate aiming solution. However, we do not observe a significant difference between RTs at Critical and Non-critical targets for the Algorithmic conditions at the end of the training block. Taken together, these two observations suggest that the Critical target continued to be treated similarly to the other members of the 10-target set rather than becoming a privileged, cached stimulus–response association.

      Third, participants would have to single out the Critical target as being distinct. All targets had the same visual appearance, the Critical target was never presented on consecutive trials, and participants were not informed that it had a special role in the experiment. However, we acknowledge that its higher sampling frequency, its somewhat greater separation from neighboring targets, and the location of the subsequent exclusion trials could nevertheless have made it more salient. Thus, we cannot rule out selective caching solely from the task structure.

      For this reason, we agree that the reviewer’s proposed analysis provides a useful additional test. If the Critical-target strategy progressively transitioned from algorithmic computation to retrieval, one prediction is that the generalization function in the Algorithmic condition should become narrower across successive exclusion blocks and increasingly resemble that of the Retrieval condition. We will therefore attempt to estimate the width of the generalization functions as a function of the training block between the Algorithmic and Retrieval conditions.

      There is, however, an important limitation to interpreting block-by-block generalization functions in this experiment. Implicit recalibration is both plan-based and temporally labile. Generalization is centered around the planned aiming direction (McDougle et al., 2017), and recent work indicates that implicit adaptation can decay over relatively short intervals (Zhou et al 2017; Hadjiosif et al 2023). Consequently, the first trial of an exclusion block provides the cleanest sample of the current state of implicit recalibration. Across later trials in the block, the measured response can be influenced both by temporal decay and by where the sampled target falls relative to the participant’s current aiming direction.

      To minimize systematic sampling bias, the starting exclusion target was randomized across participants. This means that these effects should average out at the group level, but individual exclusion blocks do not provide equally precise samples of the entire generalization function. A fully balanced estimate of every position within each exclusion block would require substantially more participants than were included in the present experiments. We will therefore present the blockwise analysis while interpreting changes in the estimated breadth cautiously.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback). Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

      The reviewer raises an important point. By design, the error-clamp manipulation in Experiment 3 decouples the participant’s planned movement from the visual consequence of that movement. While this gives us precise control over the error driving implicit recalibration, it could reduce agency over the cursor and thereby alter the interaction between explicit strategy and implicit learning in ways that are difficult to rule out completely. In particular, as the reviewer notes, reduced agency could potentially weaken either the influence of the explicit plan on implicit recalibration or the strategies themselves. Because Experiment 3 was intended to test for the absence of a strategy-dependent difference in implicit recalibration, we acknowledge that higher-order interactions of this kind represent an inherent limitation of our study if Experiment 3 is taken in isolation. 

      There are nevertheless several observations that make us think that reduced agency is unlikely to provide the primary explanation for the convergence between groups. First, the progression across Experiments 1–3 is informative. In Experiment 1, where participants retained normal control over the cursor, the broader generalization function in the Algorithmic condition closely mirrored the broader distribution of reach directions. This relationship suggests that the apparent difference in implicit generalization could arise from differences in where participants planned their movements rather than from a direct effect of strategy type on the implicit system. In Experiment 2, we sought to reduce the difference in the distribution of planned movements while preserving normal action–outcome contingencies and, importantly, the generalization functions became correspondingly more similar. Experiment 3 then controlled the error signal itself and again produced similar generalization across strategy conditions. Taken together, this progression favors the interpretation that strategy affects the measured generalization function indirectly, through differences in the distribution of movement plans, rather than directly altering the underlying implicit recalibration process.

      We nevertheless agree that these experiments cannot exclude all possible interactions between explicit and implicit learning systems. Indeed, whether these systems interact directly has been an important and persistent question in the sensorimotor adaptation literature. Several studies have reported evidence consistent with direct interactions (e.g., Albert et al., 2022; Maresch and Donchin 2021; t’Hart and Henriques 2024), whereas our own work has generally pointed toward indirect interactions mediated by factors such as movement planning and the current state of implicit adaptation (Taylor et al 2010; Taylor and Ivry 2011; McDougle et al 2017). Indeed, our recent study was designed specifically to distinguish these possibilities under tighter experimental control (Chen and Taylor, 2026), yet we found that the interaction between explicit and implicit processes is more complex than a simple independent-versus-interacting dichotomy. Going forward, we think it is more cautious to first rule out low-level statistical or distributional differences that could account for apparent effects before invoking higher-order interactions between learning systems.

      Finally, one motivation for the present study was that much of the literature on implicit generalization trains participants at a single target location before measuring generalization across the workspace. Under such conditions, participants have ample opportunity to retrieve a stable target-specific aiming solution. If algorithmic and retrieval strategies fundamentally alter implicit generalization, then many existing estimates of generalization may characterize implicit learning under retrieval-like conditions rather than providing a strategy-independent property of the implicit system. Across the present experiments, we find little evidence for such a fundamental difference once the distribution of movement plans and the experienced error are better controlled. We therefore think the most parsimonious interpretation of the current results is that algorithmic and retrieval strategies primarily influence implicit generalization indirectly through how movements are planned. 

      We plan to revise the manuscript to acknowledge that Experiment 3 cannot completely rule out higher-order effects associated with reduced agency under error-clamp feedback. We will also provide additional validation that participants were implementing distinct strategies, beyond the group-level RT differences, by testing whether RT in the Algorithmic condition scales with the instructed rotation magnitude, and whether RT variability shows group-level difference (Logan, 1988). If present, this relationship would provide stronger evidence that participants continued to engage the intended strategy under the clamp. We agree, however, that confirming distinct strategy use would not by itself rule out the possibility that reduced agency altered how those strategies interacted with implicit recalibration.

      Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      We thank the reviewer for this thoughtful and constructive assessment of the manuscript. We especially appreciate their recognition of the progressive experimental logic across the three experiments and of the broader theoretical motivation for distinguishing algorithmic and retrieval-based strategies. We are also grateful for the reviewer’s comments on the clarity and transparency of the work.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      We agree with the reviewer that our original interpretation of several non-significant effects was too strong, especially without providing some form of equivalence test. Our central hypothesis predicts little or no difference between algorithmic and retrieval-based strategies under conditions in which movement plans and error exposure are controlled, and we therefore face the inherent difficulty of drawing conclusions from an expected null effect. As the reviewer notes, a non-significant conventional hypothesis test does not by itself provide evidence that two conditions are equivalent.

      We therefore plan to supplement the existing analyses with quantitative assessments of the strength of evidence for the null/equivalence, using Bayesian factor analyses to confirm whether two conditions are equivalent. These analyses will allow us to distinguish between effects that are sufficiently small to support our theoretical interpretation and effects for which the data are simply inconclusive. We will also revise the manuscript throughout to avoid treating p > .05 as evidence of equivalence in the absence of such supporting analyses.

      We also now appreciate that the magnitude of implicit recalibration decreases progressively across experiments. This reduction could diminish our sensitivity to differences between the Algorithmic and Retrieval conditions and therefore represents an important qualification on the null result. Because the experiments used similar trial structures and were conducted with the same experimental equipment, the source of this reduction is not immediately clear. The block-by-block analysis suggested by Reviewer 2 may provide useful insight into how implicit recalibration evolves over the course of training and whether this attenuation emerges gradually within experiments.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Based on prior work characterizing implicit generalization in relative isolation from explicit strategy (Morehead et al., 2017; Poh and Taylor, 2019), we expected a relatively narrow generalization function, with a full width at half maximum of approximately 30°. We therefore expected probes spanning −45° to +45° around the Critical target to capture most of the function. At the same time, prior work on plan-based generalization predicts that the function should shift toward the participant’s aiming direction (Day et al., 2016; McDougle et al., 2017; Chen and Taylor, 2026), which complicates the choice of where to center the probes. Expanding the range and density of probe locations is also not cost-free, because additional exclusion trials increase temporal decay (Hajiosif et al., 2023) and begin to overlap with trained locations.

      We nevertheless agree with the reviewer that, when the fitted center approaches the edge of the sampled range, estimates of Gaussian width can become poorly constrained and may partly reflect parameter trade-offs or extrapolation beyond the observed data. We therefore plan to test whether the group differences persist when the Gaussian fits are constrained so that their centers fall within the sampled range. We will also examine complementary nonparametric measures of generalization breadth, such as the area between the group generalization curves across the sampled probe locations. Convergence across these approaches would provide stronger evidence that the reported differences reflect the observed shape of the generalization functions rather than instability in the Gaussian parameter estimates. If the results are not robust across approaches, we will revise the manuscript to qualify the interpretation of the fitted width estimates accordingly.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      This concern is closely related to that raised by Reviewer 2. We are fairly confident that participants in Experiment 3 were nevertheless engaging in the intended strategy manipulation. Participants in the Algorithmic condition showed substantially longer RTs than those in the Retrieval condition, and they were able to accurately generate the instructed angular reach directions across trials. The two conditions also differed in the variability of both RT and reach direction, consistent with online computation of an aiming solution in the Algorithmic condition and retrieval of a more stable cached response in the Retrieval condition. We can provide an additional validation by examining the relationship between RT and reach angle and comparing RT variability across two conditions. If RT scales parametrically with instructed reach angle in the Algorithmic condition but not in the Retrieval condition, this would provide stronger evidence that participants were engaging an online mental-rotation-like computation rather than simply following arbitrary spatial instructions. Moreover, RT yielded by the retrieval strategy would tend to be less variable than the algorithmic strategy.

      We agree, however, that this does not fully address the reviewer’s broader concern. Experiment 3 necessarily changed the context in which the strategy was implemented. In Experiments 1 and 2, mental rotation was used to counteract a visuomotor perturbation and was therefore directly tied to successful control of the cursor. In Experiment 3, the error clamp removed this instrumental relationship: participants still had to compute and execute different angular reach directions, but those computations no longer determined the visual cursor outcome. In that sense, the algorithmic process in Experiment 3 was less naturally embedded in the task and could reasonably be viewed as a somewhat different instantiation of the task.

      We therefore acknowledge that Experiment 3 cannot establish with certainty that the same higher-order cognitive process was engaged in exactly the same way as in Experiments 1 and 2, nor can it rule out the possibility that this change in task relevance altered how explicit strategy interacted with implicit recalibration. At the same time, when considered together with Experiments 1 and 2, we think the overall pattern remains informative. The apparent strategy-dependent difference in generalization was largest when movement plans and error exposure differed most, became smaller when these factors were better controlled while normal action–outcome contingencies were preserved, and was eliminated when the error signal itself was experimentally controlled. This progression is more consistent with an indirect influence of strategy through differences in movement planning and error exposure than with a robust direct effect of strategy type on implicit recalibration.

      Nonetheless, we agree that Experiment 3 should not be interpreted as a definitive test of whether algorithmic strategy, in its more relevant visuomotor adaptation context, can lead to different interactions with implicit recalibration compared to a retrieval strategy. We plan to revise the manuscript to make this limitation explicit.  

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      We agree that the RT difference in Experiment 3, by itself, cannot rule out the possibility that the Algorithmic condition imposed greater spatial precision demands because participants were reaching to invisible locations specified by angular instructions. We can, however, test for a more diagnostic signature of algorithmic computation by examining whether RT scales parametrically with the instructed reach angle. A general cost associated with reaching to invisible targets could increase overall RT, but it would not necessarily predict the characteristic increase in preparation time with the magnitude of the required angular transformation.

      We will therefore examine the relationship between RT and instructed reach angle in Experiment 3. If RT increases systematically with angular displacement in the Algorithmic condition, this would provide additional evidence that the longer RTs reflect online computation of the instructed aiming direction rather than simply the greater difficulty of reaching invisible targets.

      We can also compare this RT–angle relationship across experiments. If participants are engaging the same underlying algorithmic computation in Experiments 1–3, we would expect the slope relating RT to angular displacement to be similar across experiments, even if the overall intercept differs because of differences in task structure and spatial demands. A comparable slope would therefore provide converging evidence that the same computational process was engaged despite the altered task context in Experiment 3.

      We acknowledge, however, that similarity of the slopes would itself require yet another inference from a null difference and should therefore be interpreted cautiously. As with the generalization functions, we will use a Bayes factor analysis to quantify the strength of evidence for the null.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

      Based on the reviewers’ comments and the additional analyses they have suggested, we think we will be able to place our conclusions on a firmer empirical footing while also tightening and narrowing them. In the revised manuscript, we will strengthen the generalization analyses, use more appropriate inferential tools for interpreting null effects, and temper our broader claims about the insensitivity of implicit recalibration to strategy type. We will also more explicitly acknowledge the limitations of the present experiments, especially Experiment 3.

    1. eLife Assessment

      This study presents valuable findings implicating Xkr and its newly identified binding partners in regulating the exposure of apoptotic signal for phagocytosis, suggesting that Xkr might facilitate phosphatidylserine transfer at the ER-plasma membrane contact sites. However, several conclusions, including caspase-independent activity of Xkr and its role at the ER-PM contact sites, are not well supported by experimental data. Furthermore, some key controls are missing, and in numerous places the authors' discussion has extended far beyond what the data support. The findings presented are currently incomplete and do not fully support key mechanistic claims of the study.

    2. Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

    3. Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

    1. eLife Assessment

      Shin et al present significant new observations highlighting a novel oscillatory window during REM sleep that impacts neuronal dynamics and cross-regional communication and could have implications for our understanding of sleep's role in memory consolidation. This important study identifies and characterizes high-frequency oscillations (HFOs) in the prefrontal cortex (PFC) during REM sleep, reporting their temporal dynamics and coordination with hippocampal area CA1, and an intriguing dissociation between activation of CA1 neurons during REM HFOs and non-REM ripples that may impact sleep-dependent firing rate decreases. The main claims are supported with convincing evidence including an impressive range of analyses.

    2. Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high frequency events during nonREM or in hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from down regulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will open become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique, I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this, comparison it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs? What percentage of theta waves contain HFOs and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods for often each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and grossly what calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of overall calculations being done for each panel.

      The authors mention in discussion that they see increased functional connectivity between mPFC and CA1, but most data suggesting that seems to be based on LFP rather than spiking. Functional connectivity is defined best by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions. 


      Comments on revised version.

      Previously raised concerns are addressed.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Comments on revised version.

      I do have one remaining concern, which is about their continued use of the term "reactivation" during REM sleep, whereas it still seems "activation" is more appropriate. The only place they show more Post vs. Pre activation is in Figure 6F/6G which includes NREM sleep where indeed reactivation is robust (but not the main focus of this paper). There is no evidence offered that the REM ensembles are not already "pre-configured" and active at similar levels (with similar activation patterns) during Pre sleep. Notably Louie and Wilson 2001 found greater "replay" during Pre than Post during REM. Also, the first half vs. second half comparisons (e.g. Fig 6C) could be more effectively performed in Figure 6A, showing that the same ordering persists across the periods. If this point were addressed, the significance of the findings could potentially increase.

    4. Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      Distinction from wake HFOs<br /> The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent more general theta-associated phenomenon.

      Link to memory consolidation<br /> The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      Comments on revised version.

      The authors have since addressed these weaknesses. In supplementary figure S11 the authors now show that while HFOs were detectable during wake, they were not associated with gamma/theta oscillations or theta modulation of unit activity. This suggests that HFOs during REM are a distinct feature of REM sleep and not comparable to HFOs during NREM or wake. It would be interesting for future work to identify the significance of wake PFC HFOs, whether there are differences between HFOs during running compared to stationary behaviour, and their relationship to hippocampal sharp-wave ripples and memory consolidation.

      Regarding the link between REM HFOs and memory consolidation, the authors have further acknowledged this as a limitation of the study and requirement for a more specific experimental design to test related hypotheses. Nevertheless, they do show a clear trajectory of learning in the rats and corresponding increase in reactivation of task-related activity which could be associated with REM sleep HFOs. This study paves the way for future experiments to more directly test this link.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high-frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high-frequency events during non-REM or in the hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from downregulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique: I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this comparison, it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs?

      Here, we provide additional analyses to demonstrate that the phasic, theta-modulated PFC activity that we observe during HFOs is specifically tied to the occurrence of HFOs and not a strong phenomenon during non-HFO-associated theta periods. In Figure S5 M (middle and right), we find that aligning PFC multiunit activity to theta periods in phasic REM but temporally distant from HFOs does not elicit the same degree of theta-modulated activity as aligning to HFOs (as in Figure 2A and Figure S5 M, left).

      Additionally, we provide analyses of the theta periods adjacent to HFOs (at different temporal thresholds) and demonstrate that this theta-modulated spiking activity is largely absent (Figure S5; compare to Figure 2A). Unlike Figure S5, this analysis was not restricted to putative phasic REM.

      We have now added Supplementary Figures S5L-N to the revised manuscript.

      What percentage of theta waves contain HFOs, and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      Although theta oscillations are continuously expressed during REM sleep, HFOs occur only intermittently, such that only a small subset of theta cycles contain HFOs. Across all animals and epochs included for analysis, we found that ~7.4% of theta cycles contain HFOs. We present an epoch-level quantification of this in Author response image 1, where the proportions were calculated across all, tonic, and phasic theta cycles. As expected, a higher proportion of putative phasic theta cycles contain HFOs.

      Regarding differential firing rate modulation, we refer reviewer to normalized MUA plots in the manuscript (Figures 2A and 4F). We would like the emphasize that what we show is PFC multiunit activity that is normalized by the mean population firing rate during REM sleep. Thus, these figures, specifically the chain HFO aligned figure, indicate that there are peaks in activity around baseline level in the background of an overall decrease in population activity relative to baseline (i.e. the troughs surrounding the peaks have lower activity compared to baseline, as in Figure 9C). While this suggests that there is an overall decrease in firing rates during HFOs as compared to baseline theta periods without HFOs, this simply provides a qualitative account of this difference. In Figure S5N, we present firing rate comparisons during HFO chains versus theta periods during putative phasic REM bouts at least 4 theta cycles away from HFOs. We find that HFO chain-associated PFC neuron firing rates are lower compared to non-HFO theta periods, supporting our finding of activity suppression during HFOs.

      Author response image 1.

      Proportion of theta cycles with HFOs. (A) Proportion of theta cycles with HFOs across all, tonic, and phasic cycles (***p = 4.90e-05, rank sum test).

      Lastly, we appreciate the reviewer's suggestion that REM HFO quantifications could be framed relative to phasic theta cycles without HFOs. We agree that such comparisons are informative and ensure that the results we present are specific to periods with detected HFOs. In response, we have added additional analyses of theta periods outside of HFOs in Figure S5. Furthermore, while we found that a larger proportion of chain HFOs occurred during bouts of putative phasic REM compared to isolated events (Figure S5 and Figure S3C), most of the analyses that we performed were on events pooled across putative tonic and phasic, since putative phasic REM is relatively scarce (<10%).

      We also refer the Reviewer to our response to Reviewer 2’s major comment #1 below where we reiterate several of our findings that demonstrate the HFO-specificity of the reported dynamics, as well as our extended response to Reviewer 3’s Public Review comment #1 where we show that the dynamics associated with REM HFOs are absent during HFOs detected during awake behavior (comparable theta state) on the W-Track (Figure S11). We hope that the additional control analyses we present as well as our expanded explanations adequately address the Reviewer’s concerns.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods, often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods, often for each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and what gross calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of the overall calculations being done for each panel.

      We have now split and updated the figures in a more logical progression of ideas:

      Figure 1: Prefrontal cortical HFOs in REM sleep using spectral analyses.

      Figure 2: Characteristic spiking modulation in PFC during REM HFOs.

      Figure 3: HFO and gamma distinction in theta cycles, and PFC-CA1 coherence (including chain/ isolated HFOs in phasic and tonic REM stages).

      Figure 4: Differential modulation of PFC spiking activity during REM PFC HFOs vs. NREM PFC ripples.

      Figure 5: Characteristic PFC population activity during REM HFO chains.

      Figure 6: Comparison of PFC reactivation during REM PFC HFOs vs. NREM PFC ripples.

      Figure 7: Differential engagement and excitability modulation of CA1 neurons by REM HFOs.

      Figure 8: REM theta phase shifting CA1 neurons preferentially engaged by REM HFOs.

      Figure 9: (Model) Network model with ACh reproduces spiking modulation during REM PFC HFOs vs NREM PFC ripples.

      Figure 10: (Model) Model reproduces restricted REM coactivity vs. widespread NREM coactivity.

      We have also rearranged the figures to parallel the main figures and results section.

      The authors mention in the discussion section that they see increased functional connectivity between mPFC and CA1, but most data suggesting this seems to be based on LFP rather than spiking. Functional connectivity is best defined by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions.

      We have updated the manuscript accordingly. Specifically, we have modified the text on Page 13, Lines 23-24:

      “These chains are associated with increased measures of oscillatory coupling between PFC and CA1…”

      Reviewer #1 (Recommendations for the authors):

      (1) Please ensure that analytical methods are presented in the same order in the methods section as they are in the results section - panel by panel. That said, the methods section is well-written and presented.

      We apologize for any confusion this may have caused. We have now reorganized the methods section to ensure they presented in the same order as the panels in the main figures.

      (2) Please specify whether the recordings/behaviors occur during the animal's light circadian phase.

      We have added a statement in the methods under the “Behavior” section on Page 34, Lines 9-12 indicating that the experiments took place during the light phase:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Please specify how many tetrodes are in mPFC and how many in CA1?

      We have added this clarification on Page 33, Lines 27-28:

      “Tetrodes were split equally between PFC and CA1 (15, 16, or 32 tetrodes in each region).”

      (4) All mentions of "coherence" should have frequency bands specified. "Theta coherence", for example.

      We have now added this information to all relevant mentions of “coherence”.

      (5) The intuitive logic of the phase slope index (2K) should be briefly explained for maybe half a sentence in the results section. The intuition should be explained better in the methods section devoted to it.

      We have added more clarification on the phase slope index method, both in the legend of Figure 3 on Page 21, Lines 22-23:

      “Phase slope index (PSI), which is a measure of phase lag consistency across different frequencies…” and in the methods section under “Phase slope index” on Page 39, Lines 21-26:

      “In practice, PSI is used to assess the consistency of phase lag relationships between two signals across different frequency bands and is a measure that is weighted by oscillatory coherence. We opted to use PSI to estimate the directional flow of information instead of other methods, such as Granger causality, since it has been demonstrated that PSI is less prone to false positives.”

      (6) "Cofiring" and "coactivity" should be defined clearly as measures - preferably in the results section if possible. They sound similar and are somewhat jargon-y without self-explanatory meaning (or difference from each other). How should readers understand and interpret them?

      We apologize for the confusion regarding these two terms, which are both used throughout the manuscript. Here, we specifically used to term “cofiring” to specify the explicit quantification of coincident activity between pairs of neurons (e.g. Figure 4E, left) as described in the Methods section under “Ripple and HFO co-firing”. There were instances where “coactivity” was used to refer to this quantification, and they have been changed to “cofiring”. We have now added a statement to the manuscript to clarify that “cofiring” is a quantification of coincident activity between neurons on Page 7, Lines 34-35:

      “Overall cortical co-firing, which is a measure of coincident activity between neuron pairs during discrete events”

      Additionally, the term “coactivity” in the manuscript is used when describing neuronal activity in the model or when there are mentions of coincident activity other than the explicit quantification described above.

      (7) The temporal threshold for "cofiring" should be stated in the results to enable interpretation of the results.

      We apologize for the lack of clarity regarding the cofiring metric. Here, the temporal threshold that we are imposing is determined by the ripple/HFO event times (see Methods section “LFP event detection”). If both neurons in a pair emit spikes within the defined window of an event, they are considered to have “cofired” (Cheng and Frank 2008, Singer and Frank 2009, Sosa, Joo et al. 2020).

      We have now clarified this in the results section on Page 7, Lines 34-35:

      “Overall cortical cofiring, which is a measure of coincident activity between neuron pairs during discreet events…”

      We have also added a clarifying statement in the Methods section under “Ripple and HFO cofiring” on Page 40, Lines 24-25:

      “Here, cofiring assesses coincident activity between neuron pairs within the start and end times of events, and thus, no explicit temporal threshold was implemented.”

      (8) Please clarify more systematically in which region the NREM ripples were detected. The natural assumption is the hippocampus, but at times it is mentioned that they are detected in the cortex. Are they cortical in all analyses? Readers could easily get confused about this and misinterpret "ripples". To clarify further, if these are always cortical events, I suggest renaming "ripples" in this text to "cortical NREM HFOs".

      We apologize for the confusion. For all mentions of “ripples” in the text, we are referring to PFC ripples specifically in NREM sleep. When hippocampal ripples are mentioned, we differentiate them by explicitly using “sharp-wave ripples” or “SWRs”. Our decision to use “ripples” for NREM sleep was based on previous studies that investigated NREM high-frequency events in cortex (Khodagholy, Gelinas et al. 2017, Helfrich, Lendner et al. 2019, Vaz, Inati et al. 2019, Aleman-Zapata, Morris et al. 2022, Ghosh, Yang et al. 2022, Shin and Jadhav 2024). Furthermore, we elected to use the “HFO” nomenclature for high-frequency events in REM sleep, since it has been used in previous studies to describe these events (Tort, Scheffer-Teixeira et al. 2013, Bueno-Junior, Ruckstuhl et al. 2023). Thus, to remain consistent with the literature, we decided to use these terms to describe these sleep-state-specific events throughout the manuscript:

      Hippocampal sharp-wave ripples (SWRs) in NREM sleep

      PFC ripples in NREM sleep

      PFC HFOs in REM sleep

      We have now added the following statement on Page 5, Lines 1-3:

      “However, to avoid ambiguity, and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout.”

      (9) The analysis performed for 4B is not explained clearly. It is somewhat better explained in the methods. I believe the reader should understand that each HFO is treated as an event, and spiking participation per unit was measured, and then the similarity of that pattern was assessed between HFOs. I also find the x-axis being quartiles makes understanding this graph particularly difficult.

      Why not label by raw lag and show a correlation plot rather than break down by quartiles? Alternatively, labeling the millisecond values of these quartiles on the x-axis labels may make comprehension much easier.

      We apologize for the lack of clarity regarding this method. We have added points to clarify the analytical procedure used for Figure 4B (Now Figure 5B) on Page 8, Lines 25-28:

      “We represented each HFO as a binary vector of PFC neurons active during the event, and computed the Pearson correlation between the vectors of every pair of consecutive HFOs” As well as in the Figure 5 legend on Page 24, Lines 7-10:

      “Here, the PFC spiking activity during each HFO was binarized across all neurons and the Pearson correlation coefficient was calculated between adjacent events as a measure of pattern similarity. Then the relationship between pattern similarity and IEI was reported.”

      In addition to the quartile labels, we have now added the average inter-event interval (IEI) of each quartile to the x-axis labels of Figure 5B to improve comprehension as well as a statement in the figure legend on Page 24, Line 11:

      “Below each quartile is the average IEI of that quartile in milliseconds.”

      (10) "Rank order analysis" should be again defined in the results in a manner that the reader can follow the point of the analysis and figure. For example, the goal is to assess the spike sequence across HFO events by looking at the regularity of spike timing rank for each neuron in each HFO event. Also, could this correlation be more simply calculated and presented as just a standard deviation around the mean of that cell's rank?

      We apologize for not including an explanation of this analysis in the results section. We have now added more detail about the rank order analysis to the Figure 5 legend on Page 24, Lines 18-21:

      “The mean rank-order correlation from the leave-one-out cross-validation procedure. Each event’s rank was correlated with the averaged rank across all other events. The average across all events compared to a distribution of means generated by jittering (n = 1000) spike times is shown.”

      We also provided a short description of the procedure in the results section on Page 8, Lines 32-36:

      “Furthermore, we examined whether the order in which individual PFC neurons fired during HFOs within a chain was preserved across chains. For each chain, we extracted each cell's first-spike rank order and compared it to a leave-one-out template constructed from the average normalized rank across all other chains.”

      Regarding the Reviewer’s second comment about the correlation, if we presented the result as the standard deviation around the mean of the cell’s rank, the result would be similar to the example rank order in Figure 5C, which we provided as a visualization of the rank-order template procedure we used. However, this would not necessarily demonstrate the consistency of population activity across HFO chains. To demonstrate consistency in activity across HFO chains, we used a leave-one-out cross-validated approach where each chain event was assessed separately. For each event, the neurons firing during that event was ranked and normalized 0-1. Then, the average rank of each neuron across all other events was calculated. Lastly, the correlation between the ranks of the left-out event and the average ranks of the template was taken to assess similarity in sequential activity. We opted to use this method since it has been demonstrated to be effective for evaluating similarities in sequential across events (Stark, Roux et al. 2015, Valero, Viney et al. 2021).

      (11) In Figure 5B and the related results section, I gather that "spatial" relates to place field location rather than anatomical/tetrode location of the neuron? This was not my original understanding and should be stated clearly.

      We apologize for the lack of clarity regarding this result. Yes, the term “spatial” refers to spatial rate map correlation between CA1 neuron pairs, specifically during the behavioral W-Track session prior to the sleep session where cofiring was assessed. We have updated the y-axis label of Figure 5B (Now Figure 7B) to include “rate map” and have updated the results section on Page 9, Lines 32-33:

      “…we observed a higher degree of spatial rate map correlation for high cofiring pairs…”

      (12) Figure 5E and its caption are almost totally unable to be explained since axes aren't explained well in either. It is only by inference from the results text that meaning can be assumed.

      We apologize for the lack of clarity regarding Figure 5E (Now Figure 7E). We have now updated the y-axis labels on both updated figures (Figures 7E and 8B) and have reworded and added more detail to the figure legend on Page 28, Lines 1-5:

      “High REM PFC HFO cofiring CA1 neurons exhibited a greater degree of suppression during NREM PFC ripples. For a description of modulation index, see Methods section Ripple/HFO aligned modulation. Here, since CA1 neurons exhibit a robust decrease in activity in response to NREM PFC ripples, we refer to the modulation as suppressive.”

      (13) Figure 5F is interesting, supporting the concept that HFOs "protect" neurons from downscaling. However, how can a "neuron" be cofiring? Would cofiring not be defined in a pairwise manner, and so each unit of the cofiring measure would be a pair of neurons? This question applies to other panels in this figure. Please clarify this.

      We apologize for the lack of clarity regarding the exact cofiring metric that we use. For Figures 7D-F and 8A, since we wanted to relate cofiring to changes in firing rate and modulation state during NREM PFC ripples, we calculated single cofiring values for each CA1 neurons by averaging across all pairings with PFC neurons. Thus, this metric gives us an estimate of the overall cofiring strength of each CA1 neuron. We have now added a new section in the Methods under “Calculation of a single cofiring metric and separation into populations of high and low cofiring neurons” on Page 44, Lines 34-43:

      “Since we wanted to relate the above cofiring metric to other measures, we needed to obtain a single-value cofiring metric for each neuron. To do this, we averaged the cofiring values across all pairings for a neuron (e.g. 1 CA1 neuron paired with all PFC neurons) and reported it as the cell’s cofiring. Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0 (High cofiring > 0; low cofiring < 0). Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions (High cofiring > mean; low cofiring < mean).”

      We have also clarified this in the Figure 7 legend on Page 27, Lines 25-28:

      “For comparisons between cofiring and other metrics (e.g. firing rate), a single cofiring value was calculated for each neuron by averaging the cofiring metric across all neuron pairings. Additionally, high and low cofiring CA1 neurons were split based on average cofiring values > 0 and < 0, respectively.”

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Weaknesses:

      The primary weakness of the study lies in the lack of a clear distinction between global brain states and the specific events being analyzed. Because the authors compare HFOs across different sleep stages (NREM, tonic REM, and phasic REM) without sufficient controls, it is difficult to determine if the observed differences are intrinsic to the HFOs themselves or simply a reflection of the different physiological states in which they occur.

      We would first like to note in case it was unclear – as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly spiking activity patterns observed in prefrontal cortical-hippocampal circuits during these event-types that occur in the two sleep stages. As noted in response to Reviewer 1’s Comment #8 and Reviewer 3’s Comment #1, to avoid ambiguity and to remain consistent with existing literature, we use the following terms to describe these sleep-state specific events throughout the manuscript: Hippocampal sharp-wave ripples (SWRs) in NREM sleep, PFC ripples in NREM sleep, and PFC HFOs in REM sleep.

      We refer the Reviewer to our Response to Reviewer 1’s first comment where we provide additional analyses of theta periods outside of HFO times (i.e. baseline REM periods), including Figure S5. We also further address these comments as a response to Reviewer 2’s major comment #1 below.

      Furthermore, the evidence for "structured reactivation" is not yet convincing. The temporal alignment of these reactivation events appears inconsistent, with peaks occurring well before the HFO itself, and the analysis does not sufficiently control for pre-existing cellular assembly strengths.

      We have now addressed this as a response to Reviewer 2’s major comments #3 and #8 below, including Figure 6.

      Additionally, some of the sleep architecture presented appears atypical, such as very short REM bouts and direct NREM-to-REM transitions that bypass standard progression, raising questions about the consistency of the sleep detection across animals.

      We have now addressed this as a response to Reviewer 2’s major comment #2 below, including Figure S1 and Figure S3. We expect that addition of these figures will mitigate concerns about our sleep staging procedures and will clarify certain points raised by the Reviewer.

      Finally, the study does not account for potential confounds like baseline firing rates when interpreting the behavior of "high-cofiring" neurons, which may simply be the most active cells in the population.

      We have now addressed this as a response to Reviewer 2’s major comment #6 below.

      Reviewer #2 (Recommendations for the authors):

      In this study, the authors detect activity periods during REM sleep that feature high-frequency oscillations in the prefrontal cortex. They term these REM HFOs and report that they occur in "chains" phase-locked to theta oscillations. They contrast these with HFOs that appear in isolation and with ripples observed during NREM sleep, either in the PFC or in the hippocampus. The data presented is very interesting in places. The authors show that these HFOs are observed primarily during "phasic REM" periods that have higher theta power. It appears that overall firing is substantially lower surrounding these events. Intriguingly, it appears that CA1 cells that co-fire with the PFC during HFOs increase their firing rates over the course of sleep, whereas other neurons show a firing decrease. This seems to be the most striking finding of this study.

      Major:

      (1) The findings are generally intriguing, but the study makes some choices that are hard to understand. They begin comparing HFOs that occur in chains to HFOs that occur in isolation, even though it appears that chained and isolated HFOs likely occur at different times, some during tonic REM, some during phasic REM, and others during NREM. The lack of control for sleep state makes it difficult to determine if the reported differences are intrinsic to the HFO patterns or merely reflections of the underlying global brain state. Overall, I just didn't quite understand the motivation for comparing isolated and chain HFOs, as it seems natural that chains would occur during periods of greater synchrony.

      We appreciate the Reviewer for raising these points and agree that the properties of ripples and HFOs cannot be interpreted independently of the underlying brain state. As we mentioned in the initial response, we do expect that the generation of these ripples/HFOs in NREM and REM sleep are inextricably linked to global brain state (ex., cholinergic tone, as shown in the model in Figures 9-10), which results in differing patterns of activity across sleep states. We noted in the response to the public review comment #1 above that as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly, spiking activity patterns observed in cortical-hippocampal circuits during these event types that occur in the two sleep stages.

      Sleep state: Our primary goal in comparing isolated and chain HFOs (chain vs. isolated HFOs are only compared in REM periods, not NREM periods) was not to suggest that the differences that we observe are state-independent, but rather to provide evidence that temporal clustering of HFOs underlies distinct PFC dynamics, as well as enhanced CA1 engagement. Rather than attempting to dissociate ripple/HFO occurrence and the differential physiology that underlie NREM and REM sleep states, we view them as complementary—both sleep states are permissive to generation of high-frequency oscillations. Similarly, putative tonic and phasic REM substates (low vs high theta power) are differentially permissive for isolated and chain HFOs (Figure S3C). We add additional detail on tonic vs. phasic REM substages in Supplementary Figure S3, showing that the rate of HFOs and HFO chains is significantly elevated in putative phasic REM. While we would have liked to analyze REM HFOs during tonic and phasic states separately to be able to make more concrete conclusions, the scarcity of phasic REM sleep made it difficult to make accurate comparisons between the two, especially for spiking modulation. We have therefore removed the qualifying term “phasic REM” from the abstract.

      Also, we would like to clarify that while we investigated chains of events (ripples) in NREM sleep as a comparison to REM HFOs, we do not refer to them as HFOs in NREM anywhere in the manuscript. In NREM, while there are PFC ripples that are clustered into chains based on our definition (separation of <200 ms), we do not observe a prominent peak in the IEI distribution that would suggest entrainment by other oscillations (e.g. theta or spindles). This analysis was only to show that associated results are specific to REM HFO chains, and not seen during comparable chains of NREM cortical ripple events.

      Chain vs isolated HFOs: Regarding the Reviewers comment about why we chose to compare isolated and chain HFOs, the motivation is clearly demonstrated by differences in these events in spectral properties (Figure 3H-I) and spiking modulation (Figure 4F-G). Our initial motivation to investigate these chains of events in REM sleep came from our observation that PFC population activity aligned to REM HFOs was theta-modulated (multipeaked, suggesting multiple high-frequency events over a short duration). This led us to hypothesize that there may be chaining of events in REM sleep, which in line with previous studies demonstrating the clustering of events, such as spindles in NREM sleep (Darevsky, Kim et al. 2024). In that study (Darevsky, Kim et al. 2024), the authors showed that reactivation of motor patterns was more persistent during trains of spindle events as compared to isolated spindles. Furthermore, trains/chains of multiple hippocampal SWRs have been shown to underlie the replay of extended experience (Davidson, Kloosterman et al. 2009). Thus, the separation of high-frequency events into isolated and chained events has precedence and may have functional significance. Furthermore, since we see that isolated and chain events have a bias for occurring during putative bouts of tonic and phasic REM sleep, respectively (Figure S3C), characterization of both event types is an important step for understanding potential differences in interregional interactions during REM sleep. Furthermore, a recent appreciation for the role of sleep stage sub-states (Chang, Tang et al. 2025) further emphasizes the importance of investigating these events separately. We expect future studies to further dissect the roles of tonic and phasic REM states in memory and cognition, and our finding provides an account of the existence of different events for future reference.

      HFO chains during periods of greater synchrony: While it may seem natural that chaining would occur during periods of high synchrony (theta synchrony here), we show that our result is HFO-specific, especially for spiking activity modulation (Figures 4F-H and Supplementary Figures S5K-N), which makes it novel and important. Previous studies have focused on theta-gamma cross-frequency phase amplitude coupling, which has been proposed as a mechanism where slower theta oscillations temporally organize faster local population activity, thereby synchronizing neural ensembles within and across brain regions (Belluscio, Mizuseki et al. 2012). Here, we present a comparison of HFOs and gamma events and show that while gamma events can also occur in chains (added to Supplementary Figure S4H), possibly due to increased synchrony as the Reviewer stated above, we do not observe theta-modulated population activity aligned to these events (Supplementary Figure S4J), which is a defining property of HFOs that we propose underlies our results. This indicates that our reported results are unique to periods of synchrony associated with HFOs.

      Control analyses for theta periods: Regarding the comment about lack of control for sleep state, we refer the Reviewer to our response to Reviewer 1’s comments. We have performed additional control analyses comparing PFC activity during REM theta periods adjacent to detected HFOs and demonstrate that phasic PFC population activity is largely absent (Figure S5).

      We are aware that it is difficult to dissociate the generation of ripples and HFOs from the underlying brain state, since they are so tightly linked. However, we do expect that our clarifying points in addition to the REM theta state controls provide strong evidence that the results that we present are intrinsic to HFOs and not general reflections of activity during baseline theta activity in REM. Here, we further reiterate a few of the main results that demonstrate this:

      (1) Phasic spiking modulation of PFC population activity associated with HFOs is not present during baseline theta periods (Supplementary Figure S5K).

      (2) Phasic spiking modulation during HFOs is not linked to extracted gamma events (Supplementary Figure S4J).

      (3) Phasic spiking modulation, as well as activity suppression, is strongest during chains of HFOs (Figure 4F, Supplementary Figure S4J).

      (4) PFC-CA1 theta coherence surrounding HFOs increases relative to baseline (Figure 3E, z-scored relative to baseline coherence).

      (5) Assembly peaks are sequentially organized surrounding HFO chains but not during isolated or shuffled (baseline) HFO times (Figure 6C, Supplementary Figure S7E, and Author response image 2).

      (6) A higher proportion of CA1-CA1 pairs are high cofiring during HFOs as compared to baseline periods (Figure 7B), indicating specific CA1 engagement during HFOs (in addition to coherence).

      Lastly, we show that HFOs detected during periods of active behavior (high theta) on the W-Track are not associated with many of the defining features of HFOs in REM sleep, thus demonstrating the specificity of REM HFOs despite similar background theta activity (Figure S11, in response to Reviewer 3 comment #1).

      (2) It is crucial that the study provides REM-specific sleep examples for each of the data sessions, marking tonic and phasic REM and indicating when isolated and chain HFOs are observed. It remains unclear how interspersed these events are. Do some REM episodes just have isolated HFOs and others chains?

      We thank the Reviewer for raising this point and apologize for the lack of clarity. For additional transparency, we now provide example sleep state plots for each animal (Supplementary Figure S1H-I).

      We also provide hypnograms that show bouts of putative phasic REM on top of REM periods (Supplementary Figure S3). Additionally, as a compact way of demonstrating the validity of our separation of putative tonic and phasic bouts, we provide a plot showing the average velocity and spectrogram surrounding putative phasic REM bouts across all putative phasic REM transitions (Supplementary Figure S3A). Regarding the incidence of HFOs in tonic and phasic REM, we refer the Reviewer to Supplementary Figure S3D, where we show that HFO rate is significantly higher during putative bouts of phasic REM, during which they tend to be organized in chains as compared to putative tonic bouts (Supplementary Figure S3E). In line with this, we additionally report that although both isolated and chain HFOs occur in putative tonic and phasic bouts of REM sleep, the proportion of chain HFOs during phasic REM is significantly higher than that of isolated HFOs (Figure S3), indicating a bias for chains to occur in phasic REM sleep, potentially due to stronger theta input. However, since putative phasic REM accounts for <10% of REM sleep in our dataset, consistent with previous reports using similar methods (Mizuseki, Diba et al. 2011), we were unable to restrict spiking analysis to HFOs in phasic REM. We found that chain HFOs during both putative tonic and phasic REM sleep elicited theta modulated population activity in PFC, thus we pooled HFOs across REM states for analysis.

      While a larger proportion of chain events occur in putative bouts of phasic REM sleep as compared to isolated events (Figure S3), chain and isolated events occur in both putative tonic and phasic substates (Proportion of events in putative tonic REM is (1 - proportion in phasic) shown in Figure S3C, right). Thus, the majority of analyses comparing isolated and chain HFOs, especially for spiking data, were pooled across states. Instead of showing example plots for all 22 sleep epochs (22 out of 36 epochs with >5 s of putative phasic REM), we expect that this analysis will be sufficient to illustrate the distributions of isolated and chain HFOs across putative REM sleep substates. We have also removed the qualifying term “phasic REM” from the abstract.

      To clarify this, we have added a statement to the new “Limitations” section of the manuscript on Page 16, Lines 10-20:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states. (Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult. Finally, although our model predicts that distinct cell-type activity profiles shape REM sleep HFO dynamics, we did not record a suDicient number of interneurons to test these predictions directly. Future studies using appropriate behavioral tasks, longitudinal sleep recordings, and cell-type specific opto-tagging will be able to resolve these limitations and further clarify the roles of high-frequency oscillations in REM sleep.”

      We have now added Supplementary Figure S3.

      This is also important because some of the REM sleeps detected and shown in Figure S1H seem unusual. For example, in Animal 1, there are multiple bursts of REM that seem very irregular. In some other sessions, animals occasionally appear to enter REM sleep with very little preceding NREM, which goes against existing literature. It's not clear which ones of these meet the 30 s duration threshold. Are the chain events occurring in these periods?

      In reference to the hypnograms shown in Supplementary Figure S1H (Now Supplementary Figure S1I), these were generated by concatenating all 9 sleep epochs regardless of whether they passed the inclusion criterion of > 30 s of total REM sleep. Furthermore, only a subset of the epochs shown were included for analysis based on a secondary, manual inspection that is performed to confirm inclusion. Sleep state plots (e.g. Supplementary Figure S1H) were visually inspected to further confirm the transition into REM sleep, ensuring absence of noisy T/D ratio or spurious detection due to noisy signals – epochs where microarousals or persistent subthreshold fluctuations in animal movement induced noisy TD ratio increases, and thus inaccurate REM designation, were excluded. We thus used a total of 36 sleep sessions from a possible 90 sleep sessions.

      We apologize for not specifying what portions of data represented by the hypnograms were included. We have now provided updated hypnograms only illustrating the sleep epochs included for analysis (Figure S1I).

      We have now added Supplementary Figure S1.

      Regarding the Reviewer’s point “Are the chain events occurring in these periods?”: Yes, all of the chain events (and isolated events) come from these updated epochs that are now shown. No events in the excluded epochs were included, since we wanted to only analyze the data that came from curated REM epochs that we were confident in.

      (3) The analysis supporting structured reactivation was not generally convincing. Figure 4E does not provide convincing evidence of this. Indeed, reactivation strength is lowest around the time of the event, and appears highest 0.5s before. The example panel seems rather anecdotal. It's also not clear why REM reactivations should be compared to NREM ones here. I could not follow what was done in Figures 4F-H. Why should the first half and second half of an HFO event be correlated?

      We appreciate the Reviewer for raising these concerns. We agree that this section is somewhat dense and at time hard to follow, so we will clarify with further explanations and analyses (see also response to Comment #8 with new reactivation figures).

      In Figure 6B (originally Figure 4E), our intention was to show that if you simply align assembly activation to REM HFOs, a relatively flat response is observed when averaged across all assemblies, in stark contrast to assembly reactivation during NREM cortical ripples. This could suggest a couple of things: 1) There is no real assembly activation in response to REM HFOs or 2) Assembly activation is organized differentially (compared to NREM) surrounding HFOs. What we hypothesized, and then quantified based on observations, is that assembly activity is sequentially organized around HFOs. If this were true, it would suggest a consistent temporal relationship between REM HFOs and assemblies (for example, assembly 1 tends to be active 115 ms after the onset of chains, assembly 2 is active 230 ms after, etc.). Of course, the sequential pattern that we show in Figure 6A may arise trivially, especially since these types of sequential visualization plots can simply arise from noise. Thus, a cross-validation method must be utilized to ensure the sequences are indeed reflective of an underlying computation.

      In the methods section “Assembly sequence detection surrounding HFOs” we explain the splitripple/HFO procedure that we used to compare assembly sequences across two halves of the data, similar to methods used in hippocampal place cell sequence cross-validation (Plitt and Giocomo 2021, Sosa, Plitt et al. 2025). For every split and assembly alignment, a sequence similar to Figure 6A, right is generated based on the first half of aligned data. Then the second half of aligned data is sorted based on the peak indices of the first half of data and the correlation between the peak reactivation indices across all assemblies for the two datasets is calculated. A high correlation indicates high sequence similarity across the two halves of data (not two halves of an HFO as the Reviewer mentioned), suggesting temporal consistency of assembly reactivation. By utilizing this method, we show that assembly reactivation sequences across randomly chosen halves of data are most similar for chained HFOs (Figure 6C and Supplementary Figures S7C-F). We have now provided an additional control analysis that investigates this sequential assembly activity for time-shifted chain events (Author response image 2). As in Figure 6, assembly activity was aligned to the first event in each time-shifted chain. In addition to the analyses in Supplementary Figure S7, this control further indicates that sequential assembly activity is preferentially restricted to HFO chains.

      Author response image 2.

      Assembly sequences surrounding time-shifted HFO chains. (A) Distributions of r values calculated from the Pearson correlation between peak reactivation bins across all assemblies for two randomly chosen halves of the shifted HFO-aligned data. There was no difference between r values for time-shifted chains and shuffled data, indicating no structured assembly activity during periods outside of real HFOs chains.

      To further quantify this, we calculated two additional metrics: 1) the slope difference between fitted lines for the two halves of data for each split and 2) the absolute peak difference between assemblies in the two halves of data. First, for the slope difference metric, the slope of the best-fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data in Figure 6D indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. Secondly, Figure 6E is a quantification of the peak reactivation displacement between the two halves of data. If there is a high probability of small peak differences, as in the real data, this indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data.

      We apologize for the omission of the description of the slope and peak difference metrics that we used in Figures 6D,E. We have added information in the Methods section under “Assembly sequence detection surrounding HFOs” on Page 43, Lines 23-33:

      “Furthermore, the slope and peak differences were calculated as additional metrics of sequence and temporal reactivation consistency between the two halves of data, respectively. For the slope difference metric, the slope of the best fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. For peak reactivation difference, the temporal displacement of the peak reactivation index between the two halves of data was calculated and compared to shuffle. A high probability of small peak differences indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data. Shuffling of assembly strength was carried out as above.”

      (4) The authors argue that the occurrence of HFOs, rather than theta power, is the reason for lower MUA activity, but the analysis for this (Figure S4I) is quite confusing. The left panel actually seems to indicate that MUA is indeed lower when theta power is high.

      We apologize for the confusion regarding this figure and appreciate the Reviewer’s point that periods of high theta power can also appear to be associated with reduced MUA. We agree that the original presentation may not have clearly separated the contributions of theta and HFOs. What we convey with Supplementary Figure S5K (originally Supplementary Figure S4I) and new Supplementary Figures S5L-M is that the observed theta-modulated PFC spiking response (Figure 2A) cannot be solely explained by baseline theta periods (outside of HFOs) in REM sleep. Since theta oscillations are ubiquitous during REM sleep, an important control is to demonstrate that the theta-modulated population activity is not simply a consequence of ongoing theta activity. Thus, we aligned PFC activity to theta oscillations of varied power and show that the fluctuating, theta-modulated population activity is absent, indicating that HFOs associated with the theta oscillation are driving this phasic response.

      We are not solely arguing that the presence of HFOs is the driver of decreased PFC multiunit activity. In both the data and model, we show that the magnitude of theta power detected in PFC (possibly input to PFC) is inversely related to multiunit activity (Figure 9E). Our interpretation is not that theta is unrelated to MUA, but rather that HFO occurrence provides additional explanatory power beyond theta alone. Accordingly, we show directly in Supplementary Figures S5L-M, in response to Reviewer 1’s comment #1, that theta periods in phasic REM not associated with HFOs do not elicit the MUA activity suppression similar to HFOs.

      The phase-alignment performed is hard to follow and is not being applied to HFO periods. If the study is trying to argue that high-theta periods without HFOs in the same recording sessions show lower MUA, then perhaps some sort of shuffle or jitter would be more suitable. For example, in Figure 1B, it seems there are some high-theta periods that don't have HFOs and appear to have higher MUA.

      We expect that the new analyses where we provide additional baseline theta controls for periods adjacent to HFOs in Supplementary Figures S5L-M now clarifies this point. Briefly, alignment to theta phases at different temporal distances from detected HFOs does not exhibit the same fluctuating PFC activity as HFO alignment.

      Regarding the phase alignment procedure that we used for the control analysis, since REM sleep is characterized almost entirely by ongoing theta activity, the control condition was not a separate brain state but rather theta periods outside of HFOs. We therefore needed a systematic way to select comparable theta cycles and phase bins in order to make a valid comparison with HFO-aligned PFC population activity. Since we demonstrated that there is significant phase amplitude coupling between theta and HFOs, thus a theta phase preference of HFOs, we used that specific phase bin across multiple theta cycles to align PFC activity. This phase bin selection was performed separately for each epoch to account for inter epoch and animal variability in phase preference. We reasoned that this procedure would allow for a valid comparison as compared to random alignment, since activity was aligned to similar phases in the baseline theta vs HFO conditions.

      (5) As far as I could tell, the study does not distinguish between putative excitatory and inhibitory neurons in the PFC, but only in the CA1, even though these play very different roles in the model. What is the rationale for not separating these? How are reactivations to be interpreted among interneurons?

      We apologize for the lack of emphasis on this point, which was originally highlighted in Supplementary Figure S2G (now in Supplementary Figure S2F), and for omitting the explanation as to why we did not separately analyze putative excitatory and inhibitory neurons. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with recording primarily from pyramidal neurons (Supplementary Figure S2F, left). We therefore decided to pool and not separate the populations into putative pyramidal cells and interneurons for the spiking analyses presented. To further validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO-aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript, as we did not have enough interneurons to investigate them separately. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      Regarding the interpretation of reactivation in the context of interneurons, we expect reactivation reflects coordinated ensemble activity with excitatory neurons encoding task-relevant information, and inhibitory interneurons shaping timing and neural synchronization. We however did not record enough distinct interneurons to test the predictions of the model, which is now noted in the Limitations on Page 16.

      Relatedly, we find that PFC assemblies detected from pooled data have task relevant representations (Figures 5F-G).

      (6) Are the high-cofiring CA1 neurons generally higher-firing than the other cells? Could this perhaps explain why they behave differently?

      We apologize that this information was not more evident in the manuscript, as it is an important control. In Supplementary Figure S9A, we show that there was no difference in baseline firing rate between low and high cofiring CA1 neurons.

      We have now explicitly referenced this figure in the main text on Page 9, Lines 42-45:

      “Analysis of low and high cofiring CA1 neurons during REM HFOs showed that high cofiring neurons exhibited elevated activity during chained events as compared to low cofiring neurons (Figure 7D), independent of baseline firing rates (Figure S9A).”

      (7) It appears that the decreased firing around HFO's could be a consequence of the stronger firing modulation around these periods, related to time averaging, rather than suppression per se. How does the firing rate compare to other periods with similar modulation that might not have HFOs?

      We thank the Reviewer for raising this important point. We expect that the new analyses, where we provide additional baseline non-HFO-associated theta periods, and theta periods adjacent to HFOs at different temporal distance as controls in also Supplementary Figure S5L-N now clarifies this point. These figures show that the decreased firing rate is specific to HFO chains (Supplementary Figure S5N). Indeed, if the suppression that we observe is related to time averaging, or another analytical artifact, our claims of suppression during HFOs would not be valid. We present a number of results and provide further explanations to support the accuracy of our characterization of PFC population suppression.

      First, event-aligned multiunit activity was quantified as baseline-normalized population firing relative to detected events. For both NREM and REM, activity was normalized by the mean population firing rate during a baseline period within the same sleep state in which events were detected. Values are therefore expressed as deviations from baseline. This normalization allows comparison of relative changes in firing around events within each state. Importantly, values below baseline reflect reductions relative to the state-matched baseline period.

      Second, we show that smoothing activity with a larger gaussian kernel preserves the dip in population activity, consistent with suppression of activity surrounding HFOs (Figure 9C, note that this is a data figure presented in the context of the model). However, as raised by the Reviewer, this normalized measure does not on its own distinguish sustained suppression from transient deviations introduced by event-locked temporal structure in firing.

      Third, to address this, we employed an alternative method to demonstrate that HFO chains tend to occur during periods of PFC suppression (Figure 4G). Briefly, we detected events in PFC where activity fell below a threshold and calculated the probability of HFOs surrounding these “suppression” events. We refer the Reviewer to the Methods section under “Detection of population suppression” where we explain this procedure in more detail. We found that chain ripples, during which the strongest suppression is observed (Figure 4F), are associated with decreases in PFC activity (Figure 4G and Supplementary Figure S6).

      Fourth, we refer the reviewer to Figure S5N, where we compare the firing rates of PFC neurons during HFO chains and theta periods outside of HFOs during putative phasic REM bouts. The observed reduction in firing during true HFO chains compared to non-HFO periods therefore reflects HFO event-specific activity suppression rather than a common occurrence during baseline REM periods.

      Lastly, we performed a control analysis complimentary to Supplementary Figure S5K where we aligned PFC activity to the preferred theta phase of HFOs and investigated how distance from detected HFOs modulates PFC activity (Figure S5). We found that there was no consistent theta-modulated activity aligned to HFO-adjacent theta phases.

      Regarding the final comment, if the Reviewer meant “modulation” as in the theta modulation or suppression observed when PFC population activity is aligned to HFOs (Figures 2A and 4F), we are not aware of any other REM periods where this strong theta modulation or suppression of PFC population activity is present. To our knowledge, we are the first to demonstrate such a modulation of PFC population activity in REM sleep. The closest comparison that we can make is PFC activity aligned to gamma events that are coordinated with HFOs (suppression of PFC), but this is explained only with association with HFOs, as shown in Supplementary Figure S4J.

      (8) The assembly reactivation measure does not control for pre-existing assemblies. The term "activation strength" would therefore be more appropriate.

      We thank the reviewer for this important methodological point. The concern that ICA-based reactivation strength does not, by itself, distinguish behavior-induced reactivation from pre-existing assembly activity/structure is well-taken, and we have implemented several complementary analyses that directly address these concerns.

      First, the interleaved structure of our recordings (8 run epochs interleaved with 9 sleep epochs) allows us to investigate the within-session pre/post assembly strength differences (i.e. each W-Track run session has a preceding (pre) and following (post) sleep session). An increase in assembly strength from pre to post is a hallmark of behaviorally relevant assembly reactivation (Kudrimoti, Barnes et al. 1999, Peyrache, Khamassi et al. 2009). For each run epoch, the same run-derived templates were projected onto the preceding and following sleep epochs, and the pre and post strengths were compared. The distribution of post-minus-pre reactivation differences across all epoch pairs is significantly skewed toward positive values (Figure 6F), indicating that templates from the run epochs are more strongly expressed in the post-sleep epochs. This asymmetry cannot be explained by pre-existing assembly structure, which would predict similar assembly strengths.

      Second, reactivation strength in post-experience sleep increases across the experiment, with templates from later running epochs producing the strongest reactivation in the following post-sleep (Figure 6G). This increase in reactivation strength over time cannot be explained by preexisting assembly structure, which predicts similar assembly expression strength independent of experience.

      Third, the detected assemblies carry behaviorally meaningful structure. Assembly activation maps computed during running exhibit spatially organized "assembly fields" similar to single-cell place fields (Figure 6H), demonstrating that the detected assemblies represent specific spatial locations or task variables rather than behavior-independent states. Pre-existing co-firing structure unrelated to ongoing experience would not be expected to produce spatially tuned, task-locked assembly activation. Furthermore, this spatial tuning of assemblies was verified by comparison with surrogates, where assembly maps were generated using circularly shuffled activation times (1000 shuffles). Assemblies with p < 0.05 (z > 1.65) were considered to have significant spatial structure (Figure 6I).

      These results establish that what we measure is the selective re-expression of behaviorally relevant assemblies in subsequent sleep epochs, consistent with the use of "reactivation" in the established literature (Peyrache, Khamassi et al. 2009, Lopes-dos-Santos, Ribeiro et al. 2013). We have therefore retained the term "reactivation strength" and have added text to the manuscript noting these new results that justify the use of “reactivation” strength.

      We have now added Figure 6.

      We have also added the procedure for the calculation of spatial information to the Methods section under “Spatial information of assembly fields” on Page 44, Lines 8-23.

      (9) Can the study rule out that the rank-ordering in Fig 4C is related to firing rates? Higher-firing rates tend to fire earlier, and lower-firing cells later.

      We thank the Reviewer for raising this interesting point. Here, we assume that the Reviewer meant the baseline firing rates of the neurons, not the intra-HFO firing rates of the neurons. Indeed, when we look at baseline REM firing rates of these PFC neurons, we do find that neurons with higher firing rates tend to fire earlier than low-firing-rate neurons (Author response image 3). This is also true when PFC rank and firing rate are assessed for isolated REM HFOs and NREM PFC ripples (Author response image 3). Similarly, we also observe this relationship in CA1 during SWRs, during which rank order correlation is typically assessed as a method for replay detection. In line with this, a previous study has shown that CA1 neurons with high excitability at animals’ current location tend to initiate replay events (Karlsson and Frank 2009). Furthermore, high-firing-rate, rigid CA1 neurons are more active during SWRs than low-rate, plastic neurons (Grosmark and Buzsaki 2016), and there are distinct populations of neurons in both hippocampus and PFC that are preferentially active during immobility in sleep epochs (Jarosiewicz, McNaughton et al. 2002, Kay, Sosa et al. 2016, Tang, Shin et al. 2017), potentially biasing replay activity during high-frequency events. Similar dynamics may underlie activity during PFC ripples and HFOs in NREM and REM sleep, respectively. The critical point here is in the leave-one-out cross-validation that we implemented to determine sequence similarity—each left out event’s cell rank was correlated with the averaged rank of the template that was generated from all other events. This analysis provides a basis for our claim that there is preserved sequential PFC activity across HFO chains. We did not observe neuron firing consistency during isolated HFOs or during pseudo-HFO chains (coherently shifted chain HFO times), which indicates that REM HFO chains are unique temporal windows during which PFC activity proceeds in a more structured manner.

      Author response image 3.

      Firing rate difference of low and high rank neurons (A) Comparison of baseline firing rates of PFC and CA1 neurons split by average rank across all PFC REM HFOs, PFC NREM ripples, or CA1 SWRs. Baseline rates were calculated separately for NREM and REM sleep.

      Minor:

      (1) It gets confusing that the authors sometimes (but not always) refer to HFOs during NREM as "ripples" but not if they occur during REM. The terminology is inconsistent. When they refer to HFO chains, it seems they now pool between REM and NREM periods, as well as across phasic and tonic REM periods, which is confusing.

      We apologize for the confusion regarding the terminology. In the revised manuscript, we now use NREM ripples exclusively for NREM events and REM HFOs exclusively for REM events. We have removed mixed labels such as “ripple/HFO” except where a collective term is explicitly defined. We also clarified that HFO chains refer to REM events only and revised the relevant text/figure legends to avoid any implication that chain analyses pool NREM and REM events.

      (2) P7 L8: It might be helpful to emphasize "broader temporal distribution".

      We thank the Reviewer for the suggestion. We have updated the text on Page 8, Line 18:

      “Since we observed a broader temporal distribution of activity…”

      (3) P8 L9: What do they mean by spatial? Do they mean the place-fields of these same neurons during a previous task period?

      We apologize for the confusion. The Reviewer is correct. Here, we calculated the spatial rate map correlation between CA1 neurons as a measure of place field similarity during the W-Track session prior to the sleep epoch being assessed.

      For clarification, we have added “rate map” to the text on Page 9, Line 33:

      “…spatial rate map correlation…”

      We have also updated the y-axis label for Figure 7B for clarity.

      (4) P8 L27: What do they mean by "coordinated SWRs"? As opposed to what?

      Here, we are referring to our previous study where we investigated ripples in NREM sleep and showed that ripples and SWRs in PFC and CA1, respectively could either be independent from or coordinated with events in the other region (Shin and Jadhav 2024). A main result in the study showed that CA1 neurons are strongly suppressed during independent PFC ripples and that there was a relationship between activity suppression and reactivation during coordinated SWRs (CA1 SWR-PFC ripple coordination in NREM). We specifically mentioned “coordinated” since these are SWRs that are also coupled with SOs and spindles as compared to SWRs that are independent from PFC ripples (Shin and Jadhav 2024). Overall, we wanted to frame this result in the context of oscillatory coupling and mechanisms of memory consolidation.

      (5) P37 L28 says "we observed a bimodal distribution" but L31 says "unimodal". Which is it?

      We apologize for the confusion. We observed a bimodal distribution for CA1-PFC cofiring in REM sleep, but a unimodal distribution in NREM sleep. Because of these two observations, we decided to split the CA1 population into high and low cofiring neurons based on two different thresholds:

      (1) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) 0.

      (2) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) the mean of the distribution of averaged cofiring values.

      Using two separate thresholds to split high and low cofiring CA1 neurons demonstrates the robustness of the firing rate change result in Figures 7F and Supplementary Figures S9B-D.

      We added a statement that clarifies that the bimodal distribution was seen in REM sleep only on Page 44, Lines 37-43:

      “Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0. Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions.”

      Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which, if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      (1) Distinction from wake HFOs

      The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent a more general theta-associated phenomenon.

      To investigate PFC high-frequency oscillations during running behavior on the W-Track, events were extracted in the same manner as NREM and REM events (Methods). Events during wake were subset by periods where the animals’ velocity was >4 cm/s to provide a comparison of events during periods of high theta. While we were able to detect HFOs during wake that exhibited a similar spectral profile in the high frequency band, we did not observe 1) strong association with gamma or theta oscillations, 2) prominent HFO chaining, 3) HFO aligned theta modulated PFC activity, 4) comparable levels of theta phase amplitude coupling, 5) association with population suppression, or 6) a relationship between peri-event theta power and multiunit activity (Supplementary Figure S11). Many of the defining features of PFC REM HFOs are absent during wake, indicating REM specificity of the results we present.

      (2) Link to memory consolidation

      The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      To address these concerns, we have now added an explicit “Limitations” section in the main text of the manuscript that includes a statement about learning. We have also added Supplementary Figure S1 in the manuscript, which illustrates the performance of all 10 animals on the W-Track task. Finally, we have also included Figures 6F G in the manuscript, showing that PFC assembly reactivation strength during sleep epochs increases during learning.

      Reviewer #3 (Recommendations for the authors):

      Most of my specific comments were related to further clarification that will help the reader's understanding.

      (1) I'd recommend simplifying terminology. Open to debate, but would it not be simpler and clearer to just say NREM HFO vs REM HFO? Obviously, there is a need to mention how NREM HFOs have previously been referred to as cortical ripples, but I'm not sure it is such a helpful terminology to continue for the field, given, as you state, how different cortical ripples are from hippocampal SWRs. If not, I'd at least provide a clearer explanation early in the manuscript, distinguishing NREM ripples from REM HFOs but collectively still calling them 'cortical events'.

      We appreciate the Reviewer’s suggestion regarding terminology and agree that it would be simpler and clearer to use a single term, ripple or HFO, to describe these events. Initially, we had used a unified term (ripples across both states) but ultimately decided to switch to state-specific terminology due to previous comments we received on the manuscript and to emphasize the distinctions between NREM and REM events. Ultimately, we decided on calling them ripples in NREM and HFOs in REM since there is precedence for both terms in each respective sleep state (Khodagholy, Gelinas et al. 2017, Vaz, Inati et al. 2019, Bueno-Junior, Ruckstuhl et al. 2023, Shin and Jadhav 2024), but we do agree that this distinction can be confusing if not clearly stated. Thus, we have added an additional statement in the manuscript on Page 4, Line 44 to Page 5, Lines 1-3 for clarity:

      “Similar criteria were used to detect cortical high-frequency events in NREM and REM states; however, to avoid ambiguity and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout”

      (2) I assume experiments occurred during the light phase, but it would be good if this could be stated explicitly.

      We have now added text specifying that these experiments took place in the light phase on Page 34, Lines 9-12 of the Methods section under “Behavior”:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Page 2 Line 14: Rephrase to make clearer, e.g. 'that have a shift in...'.

      We have rephrased the sentence for clarification on Page 2, Lines 13-15:

      “REM HFO chains also preferentially engage CA1 neuronal populations that demonstrate a shift in their preferred theta-phase from behavior to REM sleep.”

      (4) Page 4 Line 36 - 'and find coherent shifts in TD', this is self-fulfilling. I would rephrase to something like 'resulting in...'.

      We have rephrased the sentence on Page 4, Lines 36-38:

      “We separated NREM and REM sleep stages based on theta-to-delta (TD) ratio in CA1, which revealed coherent shifts in TD ratio across CA1 and PFC at the onset and offset of REM sleep…”

      (5) Figure 1H left - clarify how many animals or multiunits this is based on.

      We have now updated Figure 1 legend to specify the number of animals and epochs included on Page 19, Line 4:

      “REM HFO aligned multiunit activity (MUA) in PFC (n = 10 animals, 36 epochs)…”

      (6) Figure 1I (right), it would be useful to see the x-axis frequency start from 0, since you are cutting the peak in power.

      We thank the Reviewer for this suggestion. We had initially set the frequency limits to 4 and 12 to specifically illustrate the absence of theta-modulated activity during NREM ripples. However, as suggested by the Reviewer, it is informative to expand the frequency range to ascertain the location of the peak frequency for NREM. Indeed, the peak frequency of NREM ripple-aligned PFC activity tends to be lower than 4 Hz, which is consistent with a single peak of activity that lasts <1 s.

      We have now updated Figure 2C with these new panels.

      (7) Figure 5 E/I - use of ** is confusing, it looks like a significance comparing e.g. quartile 2 to 1, but I think this is the correlation significance. I'd move ** to the top right corner and ideally include rho values.

      We thank the Reviewer for this suggestion. We have now added the r values for the correlation to Figures 7E and 8B to resolve any ambiguity. In addition, we updated Figure 7F, right and Figure 5B to maintain consistency across main figures.

      (8) Figure 7D - Why are stimulated and non-stimulated cells so different at baseline? This is not the case for the NREM results.

      The y-axes in Figures 10C,D (Formerly Figure 7) show the fraction of cells with at least one spike per 10ms bin, either for all pyramidal cells or all interneurons. We report in the figure legend that the stimulated cells represent only 30% of the network for any given stimulus (Page 32, Line 18), so at baseline in both the NREM and REM simulations, there are roughly 3x as many non-stimulated cells with a spike per 10ms bin compared to stimulated cells, as these populations are active at roughly equal firing rates per neuron outside of stimulation periods.

      (9) Methods - there is limited info on spike sorting procedure - can you provide a reference with further details (I couldn't find)? Is this method equally valid for identifying PFC units?

      Matclust is a MATLAB-based spike sorting graphical user interface that allows for manual curation of neuron clusters through the visualization of spike waveform amplitude, peak-to-trough, and principal components. Polygons or boxes are drawn around spike data points, and single unit clusters are resolved through refinement in multiple dimensions. It was developed by Mattias Karlsson and was first used in a publication reporting replay of remote experiences in the hippocampus (Karlsson and Frank 2009). Although there is no formal reference, it can be found at https://bitbucket.org/mkarlsso/matclust/src/master/. Other labs have used the software for clustering neurons from cortical areas (Yu, Liu et al. 2018, Proskurin, Manakov et al. 2023), which demonstrates its robustness across multiple brain areas.

      Additionally, examples of clustered neurons in PFC over the course of the experimental paradigm used here can be found in our previous publication (Shin, Tang et al. 2019). In addition to the aforementioned references, we show that PFC neurons can be accurately clustered and that neurons are stable over time, according to a number of cluster metrics.

      (10) Why were there no further analyses of pyramidal cells and interneurons beyond Figure S2?

      We thank the Reviewer for bringing up this important point, which is similar to Reviewer 2, comment #5 above. We repeat our response here. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with primarily recording from pyramidal neurons (Supplementary Figure S2F, left). All results are similar if we exclude putative interneurons. To validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      (11) Looking at Figure S1B, the second to last main block of NREM sleep shown has a clear peak passing TD threshold, but oddly not classed as REM - I can only assume this is due to the duration limits on your classification?

      Yes, this is due to the REM duration threshold that we implement in our sleep scoring algorithm. We used a minimum REM bout threshold criterion of 10 s for inclusion, as in previous reports (Rothschild, Eban et al. 2017, Zhang, Zhang et al. 2020). Additionally, we have included example sleep plots for all 10 animals in Supplementary Figure S1.

      (12) Include details of how head speed was calculated - just based on the 30fps video?

      Yes, the head speed of the animal was determined by tracking the animals’ position and calculating the speed based on cm/pixel values. We have now added more detail on this in the Methods section under “Surgical implant and electrophysiology” on Page 33, Line 44 to Page 34, Lines 1-2:

      “Additionally, the animals’ speed was calculated based on predetermined cm/pixel values and the position displacement between frames captured at 30 fps.”

      (13) Page 31 line 7 - With the reference you cite, they didn't really show tonic and phasic REM can be segregated based on theta frequency - they just defined it as such. You've done it for some, but I'd ensure all references to phasic REM are defined as putative - mostly missed within the discussion - I'd also add this as a brief limitation (without recording of eye movements or PGO waves).

      We thank the Reviewer for these suggestions. We have now ensured that all mentions of “phasic REM” are qualified with “putative” and have added the requested limitation in the Limitations section on Page 16, Lines 10-15:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states.(Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult.”

      We have also modified the wording in the Methods section under “Theta inter-peak intervals during bouts of high and low theta power” on Page 37, Lines 14-15 to indicate that the referenced article simply used the theta frequency-based method to define tonic and phasic REM – not to explicitly separate the two states:

      “Since previous studies have segregated putative tonic and phasic substages of REM sleep based on CA1 theta frequency…”

      References:

      Abdou, K., M. Nomoto, M. H. Aly, A. Z. Ibrahim, K. Choko, R. Okubo-Suzuki, S. I. Muramatsu and K. Inokuchi (2024). "Prefrontal coding of learned and inferred knowledge during REM and NREM sleep." Nat Commun 15(1): 4566.

      Aleman-Zapata, A., R. G. M. Morris and L. Genzel (2022). "Sleep deprivation and hippocampal ripple disruption after one-session learning eliminate memory expression the next day." Proc Natl Acad Sci U S A 119(44): e2123424119.

      Belluscio, M. A., K. Mizuseki, R. Schmidt, R. Kempter and G. Buzsaki (2012). "Cross-frequency phase coupling between theta and gamma oscillations in the hippocampus." J Neurosci 32(2): 423–435.

      Bueno-Junior, L. S., M. S. Ruckstuhl, M. M. Lim and B. O. Watson (2023). "The temporal structure of REM sleep shows minute-scale fluctuations across brain and body in mice and humans." Proc Natl Acad Sci U S A 120(18): e2213438120.

      Cairney, S. A., S. J. Durrant, R. Power and P. A. Lewis (2015). "Complementary roles of slow-wave sleep and rapid eye movement sleep in emotional memory consolidation." Cereb Cortex 25(6): 1565–1575. Chang, H., W. Tang, A. M. Wulf, T. Nyasulu, M. E. Wolf, A. Fernandez-Ruiz and A. Oliva (2025). "Sleep microstructure organizes memory replay." Nature 637(8048): 1161–1169.

      Cheng, S. and L. M. Frank (2008). "New experiences enhance coordinated neural activity in the hippocampus." Neuron 57(2): 303–313.

      Darevsky, D., J. Kim and K. Ganguly (2024). "Coupling of Slow Oscillations in the Prefrontal and Motor Cortex Predicts Onset of Spindle Trains and Persistent Memory Reactivations." J Neurosci 44(43). Davidson, T. J., F. Kloosterman and M. A. Wilson (2009). "Hippocampal replay of extended experience." Neuron 63(4): 497–507.

      Ellenbogen, J. M., P. T. Hu, J. D. Payne, D. Titone and M. P. Walker (2007). "Human relational memory requires time and sleep." Proc Natl Acad Sci U S A 104(18): 7723–7728.

      Foster, D. J. and M. A. Wilson (2006). "Reverse replay of behavioural sequences in hippocampal place cells during the awake state." Nature 440(7084): 680–683.

      Ghosh, M., F. C. Yang, S. P. Rice, V. Hetrick, A. L. Gonzalez, D. Siu, E. K. W. Brennan, T. T. John, A. M. Ahrens and O. J. Ahmed (2022). "Running speed and REM sleep control two distinct modes of rapid interhemispheric communication." Cell Rep 40(1): 111028.

      Grosmark, A. D. and G. Buzsaki (2016). "Diversity in neural firing dynamics supports both rigid and learned hippocampal sequences." Science 351(6280): 1440–1443.

      Helfrich, R. F., J. D. Lendner, B. A. Mander, H. Guillen, M. PaD, L. Mnatsakanyan, S. Vadera, M. P. Walker, J. J. Lin and R. T. Knight (2019). "Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans." Nat Commun 10(1): 3572.

      Jarosiewicz, B., B. L. McNaughton and W. E. Skaggs (2002). "Hippocampal population activity during the small-amplitude irregular activity state in the rat." J Neurosci 22(4): 1373–1384.

      Ji, D. and M. A. Wilson (2007). "Coordinated memory replay in the visual cortex and hippocampus during sleep." Nat Neurosci 10(1): 100–107.

      Karlsson, M. P. and L. M. Frank (2009). "Awake replay of remote experiences in the hippocampus." Nat Neurosci 12(7): 913–918.

      Kay, K., M. Sosa, J. E. Chung, M. P. Karlsson, M. C. Larkin and L. M. Frank (2016). "A hippocampal network for spatial coding during immobility and sleep." Nature 531(7593): 185–190.

      Khodagholy, D., J. N. Gelinas and G. Buzsaki (2017). "Learning-enhanced coupling between ripple oscillations in association cortices and hippocampus." Science 358(6361): 369–372.

      Kudrimoti, H. S., C. A. Barnes and B. L. McNaughton (1999). "Reactivation of hippocampal cell assemblies: effects of behavioral state, experience, and EEG dynamics." J Neurosci 19(10): 4090–4101. Lee, A. K. and M. A. Wilson (2002). "Memory of sequential experience in the hippocampus during slow wave sleep." Neuron 36(6): 1183–1194.

      Leemburg, S., V. V. Vyazovskiy, U. Olcese, C. L. Bassetti, G. Tononi and C. Cirelli (2010). "Sleep homeostasis in the rat is preserved during chronic sleep restriction." Proc Natl Acad Sci U S A 107(36): 15939–15944.

      Lopes-dos-Santos, V., S. Ribeiro and A. B. Tort (2013). "Detecting cell assemblies in large neuronal populations." J Neurosci Methods 220(2): 149–166.

      Mizuseki, K., K. Diba, E. Pastalkova and G. Buzsaki (2011). "Hippocampal CA1 pyramidal cells form functionally distinct sublayers." Nat Neurosci 14(9): 1174–1181.

      Nitsche, M. A., M. Jakoubkova, N. Thirugnanasambandam, L. Schmalfuss, S. Hullemann, K. Sonka, W. Paulus, C. Trenkwalder and S. Happe (2010). "Contribution of the premotor cortex to consolidation of motor sequence learning in humans during sleep." J Neurophysiol 104(5): 2603–2614.

      Peyrache, A., M. Khamassi, K. Benchenane, S. I. Wiener and F. P. Battaglia (2009). "Replay of rule-learning related neural patterns in the prefrontal cortex during sleep." Nat Neurosci 12(7): 919–926.

      Plitt, M. H. and L. M. Giocomo (2021). "Experience-dependent contextual codes in the hippocampus." Nat Neurosci 24(5): 705–714.

      Proskurin, M., M. Manakov and A. Karpova (2023). "ACC neural ensemble dynamics are structured by strategy prevalence." Elife 12.

      Rothschild, G., E. Eban and L. M. Frank (2017). "A cortical-hippocampal-cortical loop of information processing during memory consolidation." Nat Neurosci 20(2): 251–259.

      Shin, J. D. and S. P. Jadhav (2024). "Prefrontal cortical ripples mediate top-down suppression of hippocampal reactivation during sleep memory consolidation." Curr Biol 34(13): 2801–2811 e2809.

      Shin, J. D., W. Tang and S. P. Jadhav (2019). "Dynamics of Awake Hippocampal-Prefrontal Replay for Spatial Learning and Memory-Guided Decision Making." Neuron 104(6): 1110–1125 e1117.

      Siapas, A. G. and M. A. Wilson (1998). "Coordinated interactions between hippocampal ripples and cortical spindles during slow-wave sleep." Neuron 21(5): 1123–1128.

      Simor, P., G. van der Wijk, L. Nobili and P. Peigneux (2020). "The microstructure of REM sleep: Why phasic and tonic?" Sleep Med Rev 52: 101305.

      Singer, A. C. and L. M. Frank (2009). "Rewarded outcomes enhance reactivation of experience in the hippocampus." Neuron 64(6): 910–921.

      Sosa, M., H. R. Joo and L. M. Frank (2020). "Dorsal and Ventral Hippocampal Sharp-Wave Ripples Activate Distinct Nucleus Accumbens Networks." Neuron 105(4): 725–741 e728.

      Sosa, M., M. H. Plitt and L. M. Giocomo (2025). "A flexible hippocampal population code for experience relative to reward." Nat Neurosci 28(7): 1497–1509.

      Stark, E., L. Roux, R. Eichler and G. Buzsaki (2015). "Local generation of multineuronal spike sequences in the hippocampal CA1 region." Proc Natl Acad Sci U S A 112(33): 10521–10526.

      Tang, W., J. D. Shin, L. M. Frank and S. P. Jadhav (2017). "Hippocampal-Prefrontal Reactivation during Learning Is Stronger in Awake Compared with Sleep States." J Neurosci 37(49): 11789–11805.

      Tort, A. B., R. Scheffer-Teixeira, B. C. Souza, A. Draguhn and J. Brankack (2013). "Theta-associated high-frequency oscillations (110-160Hz) in the hippocampus and neocortex." Prog Neurobiol 100: 1–14.

      Valero, M., T. J. Viney, R. Machold, S. Mederos, I. Zutshi, B. Schuman, Y. Senzai, B. Rudy and G. Buzsaki (2021). "Sleep down state-active ID2/Nkx2.1 interneurons in the neocortex." Nat Neurosci 24(3): 401–411.

      van de Ven, G. M., S. Trouche, C. G. McNamara, K. Allen and D. Dupret (2016). "Hippocampal Offline Reactivation Consolidates Recently Formed Cell Assembly Patterns during Sharp Wave-Ripples." Neuron 92(5): 968–974.

      van der Helm, E. and M. P. Walker (2011). "Sleep and Emotional Memory Processing." Sleep Med Clin 6(1): 31–43.

      Vaz, A. P., S. K. Inati, N. Brunel and K. A. Zaghloul (2019). "Coupled ripple oscillations between the medial temporal lobe and neocortex retrieve human memory." Science 363(6430): 975–978.

      Wilson, M. A. and B. L. McNaughton (1994). "Reactivation of hippocampal ensemble memories during sleep." Science 265(5172): 676–679.

      Yang, S. R., H. Sun, Z. L. Huang, M. H. Yao and W. M. Qu (2012). "Repeated sleep restriction in adolescent rats altered sleep patterns and impaired spatial learning/memory ability." Sleep 35(6): 849–859.

      Yu, J. Y., D. F. Liu, A. Loback, I. Grossrubatscher and L. M. Frank (2018). "Specific hippocampal representations are linked to generalized cortical representations in memory." Nat Commun 9(1): 2209.

      Zhang, L. B., J. Zhang, M. J. Sun, H. Chen, J. Yan, F. L. Luo, Z. X. Yao, Y. M. Wu and B. Hu (2020). "Neuronal Activity in the Cerebellum During the Sleep-Wakefulness Transition in Mice." Neurosci Bull 36(8): 919– 931.

    1. eLife Assessment

      This study presents a useful combined analysis of human behavior, pupillometry and EEG measures to probe differences in ambiguity assessment between individuals during value-based decision-making. Using computational models, the authors aim to test for distinct groups of participants that differ in the relationship between evidence accumulation, arousal, and neural decision-making computations. However, at present, the evidence for group differences in the physiological, neural, and behavioral profile of decision-making under uncertainty is incomplete, lacking an appropriate statistical test for group differences, and a more complete assessment of the modeling results on evidence accumulation in decision-making in this study is required to robustly characterize and interpret the results.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript investigates value-based decision-making under risk and ambiguity using a combination of behavioral, pupillometric, and EEG data. Participants are stratified into three "decision styles" (ideal, aggressive, conservative) based on how their choices under known risk align with expected-value optimality. The central claim is that ambiguity aversion is not a uniform bias but reflects heterogeneous internal belief models, and that physiology tracks subjective belief rather than objective task structure. While this is an interesting conceptual question, the evidence is underwhelming given that differences between groups are not tested statistically (but just described), there are clear problems with how the computational models are implemented, and there are serious issues with sampling of participants.

      Strengths:

      The multimodal design (behavior, pupillometry, EEG) and the attempt to link a latent belief parameter to physiological signatures address a question of clear interest.

      Weaknesses:

      Framing and motivation

      (1) The framing conflates two questions that appear distinct. The motivation centers on "ambiguity aversion," but the study's actual aim - how individuals internally represent ambiguous outcomes - seems like a different question. The relationship between these two framings needs to be made explicit, because as written the motivating phenomenon and the studied phenomenon are not obviously the same thing.

      (2) Several of the contrasts the paper sets up against prior literature read as strawmen. The claim that ambiguity aversion is treated as "a single bias or fixed trait that applies uniformly" is presented as the view being overturned, but it is not clear this is a position the field actually holds - it reads as a strawman. Relatedly, the central objective-versus-subjective valuation distinction that the results are built around also reads as a strawman dichotomy rather than a genuine competing account.

      (3) The motivation for the physiological measures is overly broad. The statement linking EEG to control, attention, valuation, uncertainty, conflict, effort, and engagement is so general as to be uninformative - EEG signals have been linked to essentially everything, so this does not constrain the hypotheses or predictions. A more specific, falsifiable rationale is needed.<br /> Design, sample, and grouping.

      (4) The inclusion of the collaborative spacecraft/Apollo task is unclear. It is not explained why this task is included, and its role relative to the core ambiguity question needs justification (this also bears on the leadership analyses; see below).

      (5) The participant numbers do not add up and must be reconciled. The text reports 57 participants, yet the analyses describe three groups of roughly 32 + 32 + 31. The relationship between participants, sessions, and group n's needs to be stated clearly and consistently, because at present the sample description is internally contradictory.

      (6) The rationale for categorizing participants into three discrete groups is not established, and the approach is statistically questionable. Decision tendency appears to be a continuous variable; dichotomizing/trichotomizing a continuous measure is generally discouraged and can manufacture or distort group differences. The authors should justify why discrete groups are needed at all, and ideally show whether there are genuine group differences (e.g., evidence of discontinuity/clustering) rather than an arbitrary split of a continuum.

      Statistics

      (7) Key claims about how ambiguity affects groups differently are made without the appropriate test. To support a claim that the effect of ambiguity differs across groups, the interaction (group × ambiguity) must be shown - group-wise effects reported separately are not sufficient. This is really a key limitation of the current work.

      (8) The methods mentioned that some participants performed multiple sessions, but their data were treated as if coming from separate participants. This is incorrect for several reasons, particularly given the focus on individual differences.

      Belief parameter and terminology

      (9) The term "ideal" is not justified. It is unclear why this group is labelled "ideal" - are they Bayes-optimal, or optimal in some defined sense? If the label implies normativity, that needs to be demonstrated; otherwise it should be renamed.

      Drift-diffusion modelling

      (10) The boundary parameter is fixed (a detail which is hidden in the methods), but this is highly problematic. By enforcing the same boundary value for all participants, the model is forced to capture any variation as drift rate effects. As such, all conclusions about drift rate are not interpretable as they might reflect boundary effects in disguise.

      (11) The DDMs are fit separately per group of participants, which again precludes testing interactions. As with the behavioral analyses, fitting separate models means group differences cannot be properly compared within a single statistical framework, and interactions cannot be assessed. The paper does mention some comparison between groups, but comparing DDM parameter estimates across separately fit models is not valid.

      (12) Overall, the DDM is very complex, and the manuscript does not yet provide enough validation to make the model trustworthy. Given the number of trial-wise covariates entering the drift rate and the per-participant fitting, stronger evidence that the model is identifiable and that its parameters are recoverable/reliable is needed before the conclusions drawn from it can be accepted.

      Methods - EEG and analysis details

      (13) The high-pass filter setting appears very aggressive. The authors should confirm whether this risks removing genuine low-frequency signal of interest, particularly given that delta-band effects are later interpreted.

      (14) There is an apparent inconsistency in the epoching/time-locking. The time-frequency analysis appears to be computed on choice-locked data, yet elsewhere the epochs are described as stimulus-locked. This needs to be clarified and made consistent, as it affects interpretation of the pre- versus post-decision EEG clusters.

      (15) The mixed-effects modelling appears to omit random slopes. The authors should justify the random-effects structure (e.g., why only random intercepts), as this affects the validity of the inference.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Qin and colleagues entitled "Pupil and Neural Dynamics Reveal Belief-Dependent Decision Making Under Ambiguity" examines decision-making under risk and ambiguity using pupillometry and EEG. The study employs a lottery choice task with three levels of ambiguity (zero, low, high). Participants were classified into three groups based on their choice behavior in a condition with risk and no ambiguity: ideal (choosing in line with objective expected values), aggressive (preference for investments), and conservative (preference against investments). The authors then compared behavior, pupil, and EEG results across these groups. The study concludes that individual beliefs about ambiguity are reflected in different behavioral strategies and neural correlates.

      Strengths:

      The combination of behavior, computational modeling, pupillometry, and EEG.

      Weaknesses:

      (1) It is unclear whether group definition is theoretically justified.

      One general concern is that the strategy to form three distinct groups is not clearly motivated. The authors created the three groups, "aggressive", "ideal", and "conservative", based on the zero-ambiguity trials. However, as the authors state: "Ambiguity differs fundamentally from risk at both the physiological level (34; 6) and the behavioral level" (page 4). Under this assumption, it is questionable whether forming groups based on risk preferences is a useful strategy for studying ambiguity. What do we learn about ambiguity processing when group differences are primarily based on risk preferences? Might the present results partly be driven by risk preferences rather than ambiguity preferences? I recommend the following two points: (a) Clearly justify the reasoning behind the group approach; (b) Add an additional continuous analysis approach indicating whether the key results hold independent of the group definition based on risky decision-making.

      (2) k-parameter.

      The authors use the k-parameter that infers the expected high-payoff probability (e.g., page 11). On page 22, this is explained as: "the subjective value term K was assigned according to each participant's internal belief of the high-payoff rate under ambiguity, yielding a participant-specific estimate of expected value under uncertainty." I hope I have not missed anything, but I neither understood the role of this parameter nor how it was computed.

      (3) How were individual beliefs and models computed?

      A related but more general point is that it remained unclear how the authors computed internal beliefs and internal models in the study. The study contains many statements suggesting that the authors measured internal beliefs. For example:

      a) Abstract: "We show that individuals adopt distinct decision strategies that reflect different internal beliefs about unknown outcomes."<br /> b) Page 3: "We then inferred subjective belief parameters that captured how individuals internally interpreted the ambiguous probability mass and examined how these beliefs related to choice behavior, arousal dynamics, and neural activity."<br /> c) Page 16: "Together, these findings show that ambiguity does not evoke a uniform behavioral or physiological response across participants with different decision-making styles; instead, individuals rely on distinct internal models and computational strategies when forming decisions under ambiguity."<br /> d) Page 16: "Taken together, these results show that ambiguity aversion is not a uniform psychological bias, but a set of heterogeneous belief-driven strategies that shape how ambiguity is represented and acted upon."<br /> e) Page 18: "Ambiguity processing, therefore, reflects distinct belief-driven pathways rather than a single canonical mechanism."

      Based on the present data, analyses, and results, I don't think that the authors can draw these conclusions. Which analyses in the manuscript identify these internal beliefs, models, or strategies? How can we dissociate a unified strategy from a heterogeneous set of strategies based on the present results? My feeling is that the k-parameter might be related to this, but as explained above, I did not understand how it was computed and what it is supposed to reflect. The DDM analyses might also be targeted at this. However, it remains elusive how the DDM captures internal beliefs about ambiguity itself. My recommendation is that the authors more clearly explain (a) why the DDM is a useful model to study ambiguity, (b) what the different parameters exactly reflect about ambiguity processing, and (c) how the DDM captures internal beliefs and distinct belief-driven strategies in this context.

      (4) Statistical tests.

      4.1. Figure 2B: The authors summarize the number of participants with significant effects of ambiguity on choice behavior for each group. I recommend a statistical test at the second level that properly assesses the effects of ambiguity and group within a common statistical model. In my opinion, it is not enough to simply count the number of significant tests (from the first level) for each group.

      4.2. Figure 2C: For the analysis of response times, the authors might want to consider reporting the main effects of group and ambiguity.

      4.3. Figure 2D: The text on page 8 states that Figure 2D indicates that "aggressive investors showed no significant pupil modulation by ambiguity...". However, the figure and its caption indicate significant differences between ambiguous and non-ambiguous trials across all groups. Moreover, if the authors want to compare the groups, it is necessary to compare the groups to each other; a test against zero within each group would not be enough to demonstrate any group differences. In my mind, this would also be important for analyses in Figure 3C and D.

      4.4. Strictly speaking, for the statistical tests, it would be necessary to take into account that participants completed multiple sessions (within-subject variance is different from between-subject variance). Currently, each session is treated independently (page 19: "Each individual completed one to three experimental sessions. For data analysis, each session was treated as an independent participant, yielding a total of 108 sessions.")

      (5) Necessary quality control for pupillometry and EEG data.

      The task was performed in a virtual reality environment with a head-mounted display. The task was not isoluminant, and, to the best of my knowledge, participants were not instructed to avoid eye movements. The authors applied a GLM to control for luminance effects in the pupil data. For EEG, they used ICA to remove ocular and muscular artifacts. While these methods are established, they are usually applied to more controlled paradigms optimized for EEG and pupillometry. To demonstrate high data quality despite these issues, it is necessary to present quality-control analyses. Can the authors please indicate how many blinks had to be removed from the data? Could the authors please indicate how many blinks were removed from the data? Can the authors please show trial-level data (after preprocessing) for a few subjects?

      (6) Quality control for the DDM.

      The manuscript lacks systematic posterior predictive checks and parameter recovery for the DDM results. It is important to validate that the model accurately captures the data. Currently, we only see the model parameters, but it remains unclear whether the model performs well on the current data set. Moreover, if the authors aimed to test different strategies using the DDM, it might be useful to perform systematic model comparison.

      (7) Implications of the second experiment with collaborative task remain unclear.

      To me, the link between the main study and the second experiment on leadership and team performance is not obvious. In my opinion, this topic is beyond the scope of the present paper. Linking the two studies more comprehensively based on deeper theoretical grounds would likely be better suited for an independent manuscript.

    4. Author response:

      We read the Assessment as identifying two decisive gaps: (i) claims of group differences are supported by within-group tests rather than by a test of the group X condition interaction, and (ii) the drift-diffusion modeling is not yet validated or fit in a framework that permits group comparison. We agree with both, and we do not defend the current versions of these analyses. The revision will rebuild them rather than supplement them. We also agree with the reviewers that several of our conclusions about “internal beliefs” are currently stated more strongly than the analyses support, and these will be scaled back to what the modeling can carry.

      Below we first note factual errors and internal inconsistencies in the manuscript that the editors asked us to flag promptly, then summarize the planned revisions, then respond to each public review comment in turn.

      (1) Corrections and clarifications for the record

      Reviewer #2 identified one outright error in our text, and in re-checking the manuscript we found several further inconsistencies. We list them here so that they are on the record alongside the first version of the Reviewed Preprint. In each case the reviewers' reading of the manuscript is correct and the manuscript is at fault.

      (a) Aggressive investors and pupil modulation (Reviewer #2, 4.3). Our Results text states that “aggressive investors showed no significant pupil modulation by ambiguity,”. This is incorrect and contradicts our own Fig. 2d, which reports a significant, ambiguous vs. non-ambiguous difference in all three groups, including aggressive investors (T(31) = 2.74, P = 0.0304, Bonferroni-corrected). The figure and its statistics are correct; the text is wrong. The interpretive claim built on it - that aggressive investors show “blunted” arousal - is therefore unsupported and will be removed. The revised manuscript will state the correct within-group result and will test group differences directly rather than by contrasting significant against non-significant within-group tests.

      (b) Within-group versus between-group claims. Relatedly, our Results state that ambiguity-related pupil differences under the subjective model are “no longer significant relative to zero,” whereas the Discussion describes them as no longer differing across groups. These are different claims, and only the first was tested. This conflation runs through several of our physiological conclusions and is the same problem the Assessment identifies. It will be resolved by replacing these statements with explicit between-group and interaction tests.

      (c) Sample description (Reviewer #1, 5). The numbers are not contradictory but are certainly underspecified, and we accept that as written they cannot be reconciled by a reader. To state them plainly: 57 unique individuals each completed one to three sessions, yielding 108 sessions; 7 sessions were incomplete and 5 showed no response variability, leaving 96 sessions with usable behavior; of these, 95 had usable pupil data and 79 usable EEG. The three strategy groups are tertiles of these 96 sessions (32 each), and the smaller Ns in Figs. 2-4 (N = 31, 29, 25) reflect modality-specific exclusions within each tertile. The revision will include a participant/session flow diagram and a table reporting how many individuals contributed one, two, or three sessions, together with per-figure Ns.

      (d) DDM fitting procedure (Reviewer #1, 11). One clarification: the DDMs were fit separately for each participant, not separately per group (Methods, Section 4.6); group comparisons were then performed on the participant-level parameter estimates. We note this only for accuracy of the record. The reviewer's substantive objection is unaffected and we accept it: comparing parameters estimated in independent per-participant fits does not constitute a test of group differences within a common statistical framework, and it cannot test interactions.

      (e) Inconsistent specification of the EEG regressor. The trial-wise EEG covariate is described in one paragraph of Methods 4.6 as 8-13 Hz power averaged over 0-0.5 s, and in the text following Eq. 3 as a 1315 Hz difference over 0.1-0.4 s. The former corresponds to the analysis actually performed. We will correct Eq. 3's description and report the frequency band and window once, unambiguously.

      (f) Time-locking and figure/caption errors. As Reviewer #1 notes (comment 14), Methods describe stimulus-locked epochs (-0.25 to 1 s) while Figs. 2-3 are labeled relative to decision onset over -1 to 2 s. We will state for each analysis whether it is stimulus- or response-locked and harmonize axis labels accordingly. In addition: the Fig. 3 caption reads "N = 25 for ideal and aggressive; N = 29 for aggressive," where the second instance should read conservative; and Section 2.6 cites Fig. 1b and 1c for the leadership and team-performance results, which are Fig. 4e and 4f.

      (g) Delta/theta claims in the Discussion. Our Discussion attributes early delta- and theta-band enhancements to ideal investors. No such effects appear in our own time-frequency results, which report frontal beta-range and parietal alpha/low-beta clusters. These Discussion statements are not supported by the data presented and will be deleted. This also bears on Reviewer #1's comment 13, since it removes the only interpretation that depended on the low-frequency edge of our filter passband.

      (2) Summary of planned revisions

      (2.1) A single statistical framework with explicit interaction tests. All group comparisons will be replaced by unified models that include group, condition, and their interaction, with sessions nested within participants. Choice will be modeled with a generalized linear mixed model of the form invest ~ ambiguity ⨉ group + trial + (1 + ambiguity | participant/session); decision time with the corresponding linear mixed model including random slopes for ambiguity (Reviewer #1, 15). Time-resolved pupil and time-frequency EEG effects will be evaluated using cluster-based permutation tests on the group-by-condition interaction statistic rather than by aggregating within-group tests. Where our claim is that an effect is absent, most importantly, that ambiguity-related pupil differences vanish under subjective valuation, we will support it with equivalence testing and Bayes factors rather than with a non-significant P value, since a null result is not evidence for the null.

      (2.2) Continuous analyses as primary; grouping justified or abandoned. We accept that trichotomizing a continuous decision tendency requires justification that we did not provide. The revision will (i) define a continuous EV-consistency index and re-run every key analysis with it as a continuous predictor, establishing that the principal conclusions do not depend on the split. The “ideal” label implies a normative optimality we have not demonstrated and will be replaced with descriptive labels (EV-consistent, EV-exceeding, EV-shortfall).

      (2.3) Full specification, validation, and reliability of the belief parameter k. We agree the current description of k is inadequate. The revision will give the generative choice model in full, and explicitly state the estimation procedure and parameter bounds.

      (2.4) Rebuilt and validated drift-diffusion modeling. The boundary will be freed and estimated per participant/session, so that variance is no longer forced into the drift term. We will report whether the group effect on baseline drift survives.

      (2.5) Report why repeated sessions are treated as independent samples. We will conduct additional behavioral analyses to show repeated sessions from the same individual are sufficiently independent to be analyzed as unique samples. In particular, we will examine whether within-participant similarity across sessions is greater than cross-participant similarity. These analyses will provide direct evidence for whether sessions can reasonably be treated as separate samples rather than requiring sessions to be nested within participants.

      (2.6) Sharper framing and appropriately scaled claims. We will remove the framing that positions the field as treating ambiguity aversion as uniform and instead situate the work within literature that already documents heterogeneity in ambiguity attitudes. The distinction between ambiguity aversion and the internal representation of ambiguity will be made explicit, with the latter identified as our actual question. Claims about “internal belief models” will be restated in terms of what is estimated. Physiological hypotheses will be stated as specific, directional, falsifiable predictions rather than by appeal to the broad range of processes EEG has been linked to.

      (2.7) Quality control for pupillometry, EEG, and the VR context. We will report blink rates and counts, proportion of interpolated samples, trials and channels rejected, and ICA components removed. We will also report validate the luminance GLM.

      (2.8) The collaborative task. The Apollo Distributed Control Task preceded the Lottery Choice Task by design, so it must be reported as part of the protocol regardless of the leadership analysis. However, we agree that the leadership and team-performance analyses are not sufficiently motivated to carry the interpretive weight currently given them. They will be moved to a clearly labeled exploratory section, removed from the Abstract and framing, and presented without causal or trait-level interpretation.

      (3) Timeline and next steps

      The revisions above require refitting the behavioral, physiological, and computational analyses rather than adding to them, including hierarchical model fitting, parameter recovery, and new quality-control analyses. We therefore anticipate submitting the revised manuscript within approximately three months, and would welcome guidance if the editors would prefer a different schedule. We are content for the first version of the Reviewed Preprint to be published with this provisional response attached.

      We are grateful to both reviewers for the time invested in this manuscript. Several of the problems they identify are ones we should have caught ourselves, and the paper will be considerably stronger for their having been raised now rather than after publication.

    1. eLife Assessment

      This study presents valuable evidence that activating a lysosomal signaling pathway lowers blood glucose in diabetic mice, pointing to a potential new strategy for treating metabolic disease. However, the evidence for the proposed mechanism is inadequate: the glucose transporter's movement and glucose uptake were not measured in the physiologically relevant tissue, and the measurement that was performed used an unconventional technique. As a result, whether this transporter-dependent pathway explains the glucose-lowering effect has not yet been established.

    2. Reviewer #1 (Public review):

      Summary:

      This article purports to show that ML-SA8, a synthetic activator of the lysosomal TRPML1 channel, results in AMPK activation and glucose uptake in hepatocytes, and that this action has therapeutic potential for metabolic disease. The final figure shows that glucose levels are improved in db/db mice, although it is not entirely clear whether this is due to an effect on the liver, on other tissues, or on glucose production or uptake. The earlier figures try to make the case that SA8 causes activation and GLUT4 translocation and glucose uptake in liver cells; however, these data are not convincing. GLUT4 is expressed at such low levels in liver that it is likely not physiologically important. The authors use a fluorescent glucose analog to measure glucose uptake, and this molecule has been shown to enter cells largely by fluid phase endocytosis. Overall, this reviewer finds the premise misguided and the data unconvincing.

      Strengths and Weaknesses:

      The initial figures show phosphorylation of AMPK on Thr172, but no downstream effects are shown. Usually, to convincingly show that AMPK activity is increased, it would be appropriate to immunoblot phospho-ACC or some other substrate. This is minor.

      Lines 135-148: GLUT4 is not expressed at levels that are significant for physiology in liver cells, and its function in liver is not particularly relevant. The authors cite references 38-40 to support that it may be expressed at low levels in liver, but no knockout studies have been done to show that this expression is physiologically important.

      Figure 1e is not convincing. No controls are included to show the specificity of the antibody for immunofluorescent staining. No intracellular GLUT4 is visible in the unstimulated samples.

      In Figure 1f, again, the data are not convincing. The bands seem too sharp for GLUT4, which has 12 membrane-spanning domains as well as an N-linked glycosylation, so that it usually runs as a smear.

      Figure 1h. Data are not convincing. 2-NBDG is not a valid approach to measure glucose uptake. 2-NBDG enters cells largely via fluid phase endocytosis, and its accumulation is independent of known GLUT inhibitors such as cytochalasin B (Yazdani et al., MBoC 2022; PMID: 35921166; see also PMID: 42287154). The idea that such a bulky derivative of glucose could enter the transporter channel is not compatible with known structural data.

      Supplementary Figure 5 uses 2-NBDG glucose uptake again. This reviewer is not convinced that the data reflect transporter-mediated glucose uptake, as suggested by the authors. As well, although palmitate treatment of cells can cause an insulin-resistant-like phenotype in some cell types, this is not characterized in the present work. Finally, as noted, one would not expect hepatocytes to exhibit insulin-responsive glucose transport. Glycogen synthesis is the main insulin-regulated step that might be affected.

      The data in Figures 2b,c,f,g,k,l are not convincing. Again, 2-NBDG is used.

      For the glucose consumption measurements in other panels of Figure 2, the methods section states that cells were cultured in 10 mM glucose. What volume was used? It is difficult to believe that a monolayer of cells would consume very much of the glucose that is present in the culture medium. Data are shown as a percent of controls, and look reasonable, but it would be helpful to include absolute as well as relative units.

      In Figure 2, in experiments using the TRPML1 KO cells, no panel is shown to demonstrate knockout. The authors cite a previous paper for the construction of these cells, but the control immunoblot should still be shown here.

      In Figure 3, controls are missing in the BAPTA experiment in Figure 3a (only SA8-treated cells were treated with BAPTA and with EGTA). Again, it would be helpful to have p-ACC or some other readout of AMPK activity, and not just AMPK phosphorylation. 2NBDG is again used in this figure.

      Line 212-213 the text states "considering our finding that TRPML1-mediated Ca2+ release is essential for AMPK activation." This has not been shown. The work uses chelators and does not necessarily indicate a role for TRPML1. The drug may be specific, as suggested by the authors, but the way this phrase is worded is too strong. As well, AMPK was shown to be phosphorylated, but full activation towards its various substrates has not been shown.

      Figure 4cd suggests that GLUT4 expression is increased by 2 or 3-fold in the liver of DB+SA8-treated mice, compared to controls. This may be the case, but its abundance is still likely ~1000-fold less in liver compared to skeletal muscle or adipose tissue. This reviewer is still not convinced that this is physiologically relevant. The images in Supplementary Figure 8 suggest a larger increase, but it remains uncertain whether the staining really represents GLUT4.

      Data showing that blood glucose and HbA1c are reduced in SA8-treated mice are reasonable, and GTTs and ITTs are shown. Unfortunately, there are no insulin concentrations, and it remains uncertain whether glucose production is reduced or uptake is increased (or if both effects are present).

      In the discussion, the authors again state that GLUT4 is present in the liver and that it regulates hepatic glucose homeostasis, and they cite reference 63. This review article does not argue that GLUT4 acts in the liver to regulate hepatic glucose homeostasis, but that its actions in muscle and fat have secondary effects on the liver.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript contains interesting studies suggesting that pharmacological activation of TRPML1 could be useful to treat T2D by increasing glucose uptake via activation of AMPK. Preclinical studies suggest the inhibitor improved blood glucose in Db/Db mice. Ex vivo studies in cell lines examine both pharmacologic and genetic manipulations, both to activate and to inactivate TRPML1, and the results consistently suggest that TRPML1 activates AMPK and increases glucose uptake.

      Strengths:

      The manuscript is well written, and the studies are carefully performed.

      Weaknesses:

      All mechanistic studies were performed in transformed cell lines; conclusions would be stronger if performed in primary cells. The in vivo studies were only performed in male mice. Performing metabolic studies in both sexes is standard practice now. Whether the findings would extend to females was not tested and remains uncertain. Some controls are missing, such as plasma membrane loading controls for fractionation studies. The GLUT4 staining was performed after fixation and permeabilization, yet control cells appear to be devoid of intracellular (and all) staining, a confusing result that doesn't reflect the expected biology.

    4. Reviewer #3 (Public review):

      Summary:

      Zhu et al. present a proof-of-concept for targeting the lysosomal calcium channel MCOLN1/TRPML1endolysosomal ion channels to restore type 2 diabetes mellitus (T2DM). Using synthetic TRPML1 agonists (ML-SA8) and genetic manipulation, the authors demonstrate that TRPML1 stimulation triggers localized lysosomal calcium release. This calcium efflux sequentially activates CaMKKβ and phosphorylates AMPK at Thr172 in various cell models, including palmitic acid-induced insulin-resistant HepG2 cells. This signaling pathway promotes GLUT4 translocation to the plasma membrane and increases intracellular glucose uptake. When administered daily to diabetic db/db mice over six weeks, ML-SA8 lowers fasting and random blood glucose, improves oral glucose and insulin tolerance tests, reduces hepatic steatosis, and lowers serum ALT and AST levels.

      Strengths:

      Based on the TFEB-independent pathway activated by TRPML1 and the experimental approaches described by Medina's group (PMID: 31822666), the authors use a combination of pharmacological and genetic tools to dissect such an intracellular signaling pathway. Additionally, the animal experiments show consistent phenotypic improvements across independent metabolic parameters. The ability of ML-SA8 to restore glycogen deposition and clear hepatic lipid accumulation in db/db mice without causing weight loss or overt toxicity provides a strong rationale for exploring lysosomal targets in metabolic disease.

      Weaknesses:

      (1) The authors focus almost exclusively on hepatic GLUT4 to explain the observed glucose disposal. However, other glucose transporter isoforms such as GLUT2 dominate basal glucose transport. While the authors show increased AMPK phosphorylation in skeletal muscle and adipose tissue, they do not measure GLUT4 translocation or glucose uptake in these primary disposal organs. As a result, attributing systemic glycemic recovery primarily to hepatic GLUT4 translocation overlooks the major physiological roles of peripheral tissues.

      (2) In both HepG2 cells and mouse liver tissues, ML-SA8 treatment increases total GLUT4 protein expression in addition to plasma membrane localization. Because total protein pools expand, the enrichment of GLUT4 in plasma membrane fractions cannot be cleanly attributed to acute vesicular translocation alone. The manuscript does not explain the timescale or mechanism behind this rapid total protein upregulation, leaving a mechanistic gap between acute ion channel gating and protein expression.

      (3) While the in vitro specificity of ML-SA8 is well-controlled, the systemic animal experiments lack a specific rescue or knockout control. Small-molecule agonists administered intraperitoneally over six weeks can exert off-target effects. Without demonstrating that co-administering the TRPML1 inhibitor ML-SI5 blunts the therapeutic effect in vivo, or showing that ML-SA8 lacks efficacy in TRPML1-null mice, the definitive link between in vivo glycemic recovery and TRPML1 activation remains incomplete.

    1. eLife Assessment

      This valuable study advances our understanding of the physiological conditions that promote the formation of mitochondrial-derived compartments (MDCs). The genetically amenable yeast system allowed the authors to generate convincing data showing the rapid formation of MDCs under physiologically relevant conditions where the mitochondrial proteome needs to be expanded during metabolic remodeling. The identification of a key role for a yeast AMPK-related protein called Snf1 will be of interest for cell biologists and biochemists.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigated the formation of mitochondrial-derived compartments (MDCs) under metabolic adaptations. They hypothesized that MDCs may play a role in regulating the mitochondrial proteome under these conditions by removing excess and superfluous membrane proteins that may challenge mitochondrial proteostasis. They found that glucose restriction, carbon-source switching, and osmotic stress can stimulate MDC formation. Underlying these stressors is a common signaling pathway that involves Snf1-dependent derepression of mitochondrial biogenesis and rapid synthesis and trafficking of nuclear-encoded proteins into mitochondria. They then showed that MDC formation is attenuated in tom70/tom71 mutants, suggesting that the delivery of these proteins to mitochondria is critical. Data also suggested that HAP4-stimulated mitochondrial biogenesis promotes MDC formation, which is further enhanced by glucose restriction and is suppressed after prolonged adaptation.

      Strengths:

      The genetically amenable yeast system allowed the authors to generate convincing data showing the rapid formation of MDCs under physiologically relevant conditions where the mitochondrial proteome needs to be expanded to accommodate increasing metabolic function. MDCs therefore function to buffer spillovers of outer membrane proteins upon an abrupt protein influx. Overall, the data presented are of high quality. The conclusion is strongly supported by the data.

      I think this is a significant study as (1) it supported MDCs as a physiologically relevant mechanism of mitochondrial proteostasis; and (2) it offers a common mechanistic framework explaining the MDC phenomenon under many other conditions such as TOR inhibition and hydrophobic protein overloading previously published by this group. Although the precise mechanism of MDC formation and how MDC formation contributes to the overall proteostasis of mitochondria remain unknown, as the authors stated in the manuscript, the current work is a clearly identifiable milestone in this specific area of investigation.

      Weaknesses:

      Although the data are overall strong, weaknesses are mainly related to potential misinterpretation of the data.

      (1) I have reservations regarding the interpretation of some results. First, the authors concluded that MDC biogenesis is activated when glycolytic metabolism is altered. I disagree with this. The authors should distinguish between "loss of glycolysis" and "loss of glucose repression". The yeast S288C strains are GAL2 and can ferment galactose. Likewise, glycolysis is also supported by raffinose and sucrose. In a broad sense, these carbon sources do support glycolysis as long as sugar influx is maintained at a high level. However, these alternative carbons do not repress mitochondrial respiration like glucose. It is likely the derepression of mitochondrial respiration (which is stated in some sections of the manuscript) instead of loss of glycolytic metabolism that stimulates MDC formation. This needs to be made clear throughout the manuscript. As such, the statement that "Carbon-source switching" stimulates MDCs is not accurate and needs to be re-interpreted.

      (2) The explanation for the requirement of low glucose levels could be misleading. A complete lack of carbon sources and high concentrations of 2-DG may shut down global protein synthesis, cell cycle progression, and many other processes, including mitochondrial biogenesis. Glucose at 0.02% is not sufficient to cause glucose repression, as only the high-affinity but low-influx transporters are functioning. Under the low glucose conditions, mitochondrial respiration is also derepressed. In this scenario, low glucose simply plays a role in supporting cell growth without causing the repression of mitochondrial biogenesis.

      (3) The idea that MDCs are formed when protein load exceeds the capacity the organelle can accommodate is attractive. Early studies have shown that the mitochondrial compartment is expanded by several folds in volume when yeast cells are switched from fermentative to oxidative metabolism. Perhaps, space expansion takes longer than protein influx increase. It would be interesting to see whether there is a correlation between MDC frequencies and the delay in volume expansion. Long-term adaptation would solve this challenge, as it allows the cell to complete volume expansion.

      (4) HAP4 may primarily activate OXPHOS genes but not some MDC cargo proteins. The requirement for "metabolic remodeling" for full induction of MDC formation may be an overstatement. The authors should either reexamine the proteomic data to see whether known MDC cargos are not subject to HAP4 activation or have this discussed in the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      Price et al. present new work providing insight into the function and mechanisms of mitochondrial-derived compartments (MDCs) in yeast. The Hughes lab previously established that these large ~micron-sized structures are formed under a variety of conditions including amino acid stress (rapamycin, conA, cycloheximide), alterations in mitochondrial metabolites and lipids, or acute expression of specific outer membrane proteins. These stressors lead to the sequestration of outer membrane proteins (that can include mistargeted inner membrane proteins) that extend or tubulate into large multilamellar structures that ultimately target the vacuole in an ATG5/Dnm1 dependent autophagy related pathway for degradation. Initially reported in aging yeast a decade ago, it has now been accepted as a mechanism to remove excess mitochondrial proteins as a pathway distinct from mitophagy or the extraction of stalled precursors from the import translocon.

      In this study, the authors examined additional metabolic transitions they suspected would drive increased mitochondrial protein expression and promote MDC formation. Indeed, they show that glucose-restricted conditions (or a switch to galactose or incubation with 2DG) induced MDCs within 2 hours. This correlated with increased transcription/translation of mitochondrial precursors porin and OM45. Similar results were seen with osmotic shock, a process previously shown to induce mitochondrial gene expression. The metabolic or osmotic shift was shown to activate a yeast AMPK-type kinase called Snf1, which phosphorylates a key substrate Mig1 - an established repressor of mitochondrial gene expression. Loss of these pathways abolished the generation of MDCs under these conditions. As the key novel finding in the study, the authors explored the relationship/requirement for Snf1 and Mig1 using multiple approaches in different backgrounds and employing auxin-inducible degron tools for acute depletion. These data further support the hypothesis that excess mitochondrial outer membrane proteins result in MDC formation to facilitate their removal, at least transiently until the import machinery can adapt to the increased import demand. To test this more directly, they generated an inducible yeast strain to express a canonical transcription factor Hap4 that induces mitochondrial gene expression. In this system, induction of Hap4 expression also resulted in MDC formation. While not all previously reported MDC inducers act through Snf1/Mig1, the common feature is the transcriptional induction of mitochondrial protein expression.

      Strengths:

      The important aspect of this work is that the authors dissected the transcriptional signaling pathway that induces MDCs in a much more physiological metabolic transition, which complements the more common use of chemical compounds. They had previously shown that overexpression of individual outer membrane proteins could lead to MDCs, but here the Hap4 expression offers a new condition to show that the canonical induction of mitochondrial biogenesis leads to MDC shedding. Overall, the data are of high quality, the findings are clear, and the work provides important new insights into the regulation of MDC formation.

      Weaknesses:

      There are a few points that should be addressed.

      (1) MDCs are almost exclusively monitored through GFP-tagged TOM70, and the authors do not show the inclusion of any endogenous cargo. The evidence for their fate in the vacuole is through the appearance of cleaved, free GFP after 6 hours that is dependent on ATG5, Dnm1, Pep4, etc. Can the authors demonstrate the appearance of MDCs without expressing any GFP tags and instead monitor known outer membrane cargoes? In the case of Hap4 expression, the proteomics identifies some very highly induced mitochondrial proteins, and there surely must be some with antibodies that can detect the protein by IF and Western blot.

      (2) There is a very unexpected ~10X increased in a sporulation factor SPO21 upon induction of Hap4. I see no evidence of sporulation, and it's not long enough for stationary phase. Is the increased mitochondrial biogenesis driving a specific metabolic state of these cells that is signaling to other biology?

      (3) It is important to understand the kinetics and stoichiometry of outer membrane loading that drives MDCs, and their transit to the vacuole. This is why it would be highly informative to monitor some endogenous cargoes (previous point). In the review the authors cite (NRMBC, Pfanner lab 2019), it was stated that the import machinery is not generally increased upon metabolic induction of mitochondrial gene expression. Therefore, (pre-MDCs) the field concluded that the import machinery has a very high capacity for the rapid biogenesis of newly synthesized proteins, along with regulation through the phosphorylation of import receptors (ie; the work of Meisenger). Consistent with this, the Hap1 proteomics did not show any increases in the core import machinery, while ETC subunits and a large swath of mitochondrial proteins were elevated over 2-fold (I looked carefully through the Excel sheet). Since MDCs are induced transiently about 2 hours after glucose deprivation, and fully dependent on de-repression of Mig1, the authors are right to imply that this is coupled to the import of newly synthesized proteins.

      However, it seems to me that MDCs are being formed at very early stages of mitochondrial protein expression, not after they have necessarily "overloaded" the outer membrane. The Hap4 proteomics after 3.5hr of induction would suggest that the bulk of the mitochondrial proteins have been successfully inserted (no import failure) and are likely already functional (metabolizing). I'm trying to understand the percentage of the proteins that would be incorporated within MDCs, as the mitochondria appear to handle the bulk of their newly inserted proteins without issue. How can the authors adapt their "free GFP" assay to understand the stoichiometry of the transport of endogenous, newly imported outer membrane proteins to the vacuole?

      (4) As a last theoretical point for discussion: Can the authors exclude that MDCs are not functional or play a signaling role? Given the emerging work on SPOTs (Lena Pernas), and from the new evidence from Craig Thompson's lab that there can be very specific functional mitochondria (oxidizing vs reducing), it is possible that MDCs are not simply there to be degraded. They last at least 3 hours, which is a long time for yeast (budding cycle 90 min, 3 hours in glucose deprivation). Taking the data presented here very objectively, there is no direct evidence that the cargoes within MDVs reflect any failure to import, or that they are damaged in any way. The deletions of Tom70/71 have way too many pleotropic effects and essentially demonstrate only that the MDC cargoes came from the mitochondria. It could be helpful if the discussion also positioned these MDC mechanisms within the context of other aspects of selective mitochondrial-related compartments that have been emerging in the literature.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Price et al. report the physiological conditions and proteins involved in the formation of mitochondria-derived compartments (MDCs), specialized domains exclusively containing outer mitochondrial membrane (OMM) proteins, in budding yeast. Hughes and his colleagues have previously established MDCs as unique multilamellar membrane structures derived from the OMM that arise with both mitochondrial metabolic perturbation and hydrophobic OMM protein load. Whether cells undergo MDC formation in response to physiological changes in mitochondrial biogenesis remains to be explored. In this study, the authors sought to test if glucose restriction, carbon-source switching, and salt stress can induce MDC formation, and found that these situations, which naturally promote acute mitochondrial biogenesis concomitantly with metabolic transitions, trigger MDC induction. Under these conditions, loss of Snf1, an AMP-activated protein kinase that facilitates mitochondrial biogenesis under metabolic stress, almost completely abolished MDC formation. Snf1 induces MDC induction under metabolic stress via phosphorylating (suppressing) Mig1, a transcriptional repressor of mitochondrial biogenesis. Consistent with this idea, loss of Mig1 mostly rescues MDC formation under glucose restriction or salt stress in cells lacking Snf1. The authors further found that acute induction of Hap4, a core activator of mitochondrial biogenesis, is sufficient to trigger MDC formation even without metabolic stress. Finally, cells lacking Tom70 and Tom71, protein receptors of the TOM (translocase of the outer membrane) complex that mediate targeting of hydrophobic mitochondrial proteins, almost failed to form MDCs under glucose restriction. Correctively, the authors propose that MDCs act in the reduction of OMM protein load upon metabolic stress-induced acute mitochondrial biogenesis.

      Strengths:

      The experiments for this study are well-designed, and the resulting data are mostly convincing, with proper controls and significant statistics to support their conclusions. The paper potentially provides new insights into the physiology of MDC formation.

      Weaknesses:

      There are only a few new mechanistic advancements in this paper.

    1. eLife Assessment

      This important study uses dual-color super-resolution microscopy, quantitative clustering analyses and simulations, as well as biochemical approaches to investigate the nanoscale organization of mucins and trans-sialidases on the surface of Trypanosoma cruzi, the causative agent of Chagas disease. The evidence supporting a non-uniform distribution of these molecules into segregated nanoclusters and a more ordered non-clustered population is solid. Overall, this work provides a generalizable analytical framework that can help understand how parasite surface proteins are spatially organized. The main remaining limitations concern the mechanistic basis of the proposed organization and the need to exclude possible effects of sample preparation before fixation.

    2. Reviewer #1 (Public review):

      Summary:

      Escalante et al. employ super-resolution microscopy to achieve a clearer, nanoscale view of how trans-sialidases and mucins are organized on the Trypanosoma cruzi parasite membrane. Comparing the experimental data using clustering analysis with model-based simulations, they report two kinds of organizational states describing the non-uniform distribution of these two proteins: a segregated state where mucins and trans-sialidases form spatially distinct nano-clusters, and a non-clustered state where they share a proposed fibrillar network with more ordered, shorter-than-random separation distances. They also look at the oligomerization states of the two proteins to try and propose a mechanistic basis for the observed distributions.

      Strengths:

      The in-depth analysis of the distributions of both proteins coupled with model-based simulations brings out new insights into organizational principles underlying protein distribution on the membrane surface. The ability to resolve shorter-than-random separation distances even in the non-clustered state is to be highlighted and is a key take-away from this manuscript.

      Weaknesses:

      The authors propose the oligomeric state of mucins compared to the non-oligomeric trans-sialidases as a basis for explaining the distinct organization of these proteins. Although this hints at how segregation may occur, it does not inform us of how the more ordered non-clustered state could co-exist with the clustered segregated state and warrants further investigation.

      Overall, the analytical framework applied in this study to elucidate organizational principles for the non-uniform distribution of proteins can potentially be used in a wide-range of contexts across different organisms and systems. This study also lays the groundwork to understand mechanisms that spatially regulate how trans-sialidases act on their substrates. Going forward, it could be very interesting to look at how different kinds of mucins and trans-sialidases are organized with respect to one-another and amongst themselves. Also, the development of tools to observe the dynamics of these proteins live will likely provide further insights into the mechanism.

    3. Reviewer #2 (Public review):

      The manuscript describes a numerical analysis of the domains of the T. cruzi cell surface containing different proteins. It has the potential to be of great interest.

      I do not have the expertise necessary to comment on the image collection or analysis.

      I have one concern: the amount of manipulation of the cells prior to fixation; these were clearly stated in the methods, which is good.

      My concern is whether these manipulations prior to fixation alter the observations. The 'Labelling sialic acid acceptors' involves >6 centrifugations and >90 minutes incubation in PBS prior to fixation, and the 'immunostaining' protocol involves cells 'extensively washed with PBS' prior to fixation. I would like to suggest that the authors do controls in which they compare the pattern of anti-SAPA staining under four conditions.

      (1) Cells fixed in culture by the addition of paraformaldehyde to 4%, followed by blocking and PBS washes.

      (2) Cells fixed in culture by the addition of paraformaldehyde to 4% and glutaraldehyde to 0.2% followed by blocking and PBS washes.

      (3) Cells fixed by the 'labelling sialic acid acceptors' protocol.

      (4) Cells fixed by the 'immunostaining protocol'.

    4. Reviewer #3 (Public review):

      Summary:

      The authors present an innovative approach to tackle the lateral organization of mucins and trans-sialidases (TS) on the cell membrane of the organism Trypanosoma cruzi. By applying dual-color super-resolution microscopy (STORM), the authors report on a differential nanoscale distribution between mucins and TS on the cell membrane. Moreover, they find that 60% of mucins and TS are organized in nanoclusters with an inter-nanocluster distance following a random distribution. The remaining 40% of both proteins are organized in a non-random manner, and, using simulations, the authors claim that they are organized in rectilinear fibers.

      Strengths:

      The authors use dual-color STORM microscopy to unravel the protein nanoscale organization of mucins and TS on the cell membrane of Trypanosoma cruzi for the first time. They perform a dedicated analysis of the localizations and clustering of both proteins. Moreover, they perform, for every type of analysis on real data, simulations to compare their results for random organization. They also use an analysis approach together with simulations to propose that the lateral organization of both non-clustered proteins are within rectilinear fibers. They also complement their microscopy findings with BN-PAGE. Overall, the use of advanced microscopy techniques, corresponding data analysis and simulations is very solid and remarkable.

      Weaknesses:

      As the authors point out, they do not provide a molecular/biophysical mechanism explaining the non-random lateral organization of mucins and TS (both clustered and individual proteins).

    1. eLife Assessment

      A regulated cell death pathway that intentionally causes ferroptosis has not been previously described. The authors describe a useful finding that the ferroptosis inducers ML162 and erastin induce caspase-5-mediated cleavage of GSDME in mesenchymal-like ovarian cancer cells, resulting in pore formation. The data supporting the main claims are incomplete. This work will be of interest to researchers who study regulated cell death.

    2. Reviewer #1 (Public review):

      Summary:

      This report seeks to understand the mechanisms whereby the ferroptosis inducers ML162 and erastin cause cell death in several tumor cell lines. They present evidence that caspase-5 is activated and required for ferroptosis, but other caspases, including caspase-1 and -4, are not important. Surprisingly, caspase-5 cleaved and activated GSDME, instead of the expected gasdermin target GSDMD

      Strengths:

      The magnitude of effect for triggering ferroptosis by ML162 and erastin is strong, and the strength of inhibition by YVAD is also very strong, making these effects convincing. The lack of effect of DEVD, which inhibits apoptotic caspases, is also convincing. Also, the lack of effect of necrostatin is convincing. These negative results strengthen the positive results seen with YVAD.

      Inhibition by disulfiram is convincing.

      Caspase-5 knockout single-cell clones and the ability to complement these with caspase-5, but not catalytically inactive caspase-5 in Figure 5, is strong data.

      Weaknesses:

      (1) Prior publications have asserted that ferroptosis is caspase-independent. Can the authors repeat some of these experiments directly to reveal whether there was an error in the published work that resulted in missing the phenotype for a caspase in ferroptosis? In my experience, caspase inhibitors sometimes only delay cell death because they are not 100% effective, especially over hours of time. Can the authors repeat the prior experiments to reveal whether this caveat affected previously published data? At the least, the authors should use the z-VAD-fmk and Boc-D-FMK inhibitors to determine whether they give the same effects as YVAD to rule out a very unlikely possibility that these "pan-caspase" inhibitors do not inhibit caspase-5.

      a) The original report describing ferroptosis by Dixon and Stockwell, doi: 10.1016/j.cell.2012.03.042, shows that erastin treatment-induced ferroptosis is not affected by z-VAD-fmk in 3 cell lines.<br /> b) A later report from Dr. Stockwell states in data not shown that a different pan-caspase inhibitor (Boc-D-Fmk) does not block erastin-driven ferroptosis. Doi 10.1016/S1535-6108(03)00050-3<br /> c) An earlier 2008 report from Dr. Stockwell shows that z-VAD-fmk and Boc-D-fmk do not rescue cells treated with RSL-3 or RSL-5 treated cell lines derived from BJ cells. doi 10.1016/j.chembiol.2008.02.010<br /> d) A recent paper shows a delay of ferroptosis after RSL3 treatment by pan-caspase inhibitor Q-VD-OPh. Doi 10.1038/s41418-025-01514-7. The delay was about 8 hours in time, so cells were still dying.<br /> e) Gpx4 knockout cells or erastin or RSL3 treatment are unaffected by z-VAD-FMK. Doi 10.1038/ncb3064<br /> f) I encourage the authors to do more thorough searching of the literature to find more publications that have used caspase inhibitors.

      (2) The authors should discuss how mouse cells can undergo ferroptosis while they do not encode caspase-5, and the evolutionary conservation of caspase-5 in general. If caspase-5 is not encoded by an animal (as is the case with mice), can their cells undergo ferroptosis?

      (3) Disulfiram is not a specific inhibitor. It is a nonspecific inhibitor that modifies cysteine residues of many proteins. This should be described in more detail so the reader can appreciate the strengths and weaknesses of the inhibitor.

      (4) I encourage the authors to assess IL-1β processing by Western blot and show that this is inhibited by YVAD. Because ELISA can detect release of the pro form after lytic cell death by other mechanisms.

      (5) ASC knockdown in Supplementary Figure 3a for two cell lines is not sufficient to draw any conclusions in Figure 3a.

      (6) Caspase-5 can be more specifically inhibited by LEVD inhibitors. Can the authors show that these work as well?

      (7) I would like to see a positive control in Figure 5a to show what a strong caspase signal activity looks like.

      (8) Since caspase-3 is known to cleave GSDME, the authors need to assess whether caspase-3 is also activated, and whether other caspase-3 target proteins are also cleaved. There are many to choose from. Caspase-3 western blots, including with the cleaved caspase-3-specific antibody, are critical. This is in addition to the blot shown in Supplementary Figure 10. Positive controls should be included. It is important to continue to add controls to rule out caspase-3, with more than just negative data with DEVD inhibitors and the western blot in Figure S10.

      (9) The data in Figure 6c are not strong.

      (10) One would expect that any mode of activation of caspase-5 should lead to its proteolytic activity upon its preferred substrates, so LPS should cause caspase-5 to cleave GSDME and not GSDMD. Additional data to strongly activate caspase-5 with LPS should be investigated to see if this leads to GSDME cleavage and pyroptosis via GSDME and not GSDMD.

    3. Reviewer #2 (Public review):

      Summary:

      In the submitted manuscript, Akter et al use a series of ferroptosis inhibitors in mesenchymal-like ovarian cancer cells and discover that the ferroptosis inducers induce cell death that is inhibited by pyroptosis inhibitors, namely YVAD-fmk and disulfiram, which inhibit pore formation by gasdermin D (GSDMD). Remarkably, the authors also saw the release of IL-1β in response to ferroptosis inducers. Unexpectedly, they did not observe the involvement of caspase-1 but rather observed that caspase-5 was activated in response to the ferroptosis inducers. Moreover, they found that caspase-5 directly cleaves GSDME in response to the ferroptosis inducers, establishing CASP5/GSDME as downstream executors of ferroptosis.

      Strengths:

      These findings are interesting because only CASP1 is known to induce IL-1β maturation, and their data suggest that CASP5 rather than CASP1, is responsible for IL-1β activation in the context of ferroptosis inducers. Notably, CASP3 is the only caspase reported to be able to cleave GSDME, so the identification of CASP5 as a driver of ferroptosis in this context is a significant finding. They genetically show that loss of CASP5 and GSDME knockdown inhibits cell death in response to the ferroptosis inducers ML162 and Erastin, which is evidence that they play a role in this context.

      Weaknesses:

      The major findings in this paper are interesting, but the data presented do not robustly support the claims made in this paper. For example, they claim that CASP5 is responsible for the activation of GSDME by cleaving it directly to induce cell death. They try to rule out the involvement of CASP1, ASC, and CASP4 using siRNA targeting these genes, but the knockdowns are incomplete, and the loading controls are inconsistent. They also claim they do not see GSDMD or CASP3 cleavage and activation but use negative data to make that claim. It is unclear if the antibodies used can detect cleaved GSDMD or CASP3 as they do not include a positive control to show that they can indeed detect these activation events if they were occurring. This needs to happen in the same experiment - they need to show in the same experiment with the same lysates that they can detect CASP5, GSDME and IL-1β activation but not CASP1, GSDMD, CASP4, or CASP3 activation. Of course, they should include agonists for positive controls of CASP1, CASP4 and GSDMD activation, which are lacking in the current manuscript.

      Notably, the major evidence supporting a direct role for CASP5 cleavage of GSDME is one Coomassie gel using recombinant CASP5 and GSDME, but there were too many non-specific bands, and the full-length uncleaved protein could not be detected even in the untreated lanes. The authors need to show a gel where the protein can easily be identified and should also include a positive control protein like GSDMD to show the relative cleavage efficiency of GSDME compared to a known substrate. It would also be great to compare this to CASP3-mediated cleavage of GSDME. With recombinant proteins, calculating the catalytic efficiencies would be the best way to ascertain if this is biologically similar to other known substrates.

      The way that ferroptosis is defined, it is caspase-independent, and pyroptosis is defined as gasdermin-mediated cell death. Given that these agents lead to activation of CASP5/GSDME, it would be more accurate to say that these ferroptosis inducers also induce CASP5/GSDME-dependent pyroptosis, as opposed to them being the executors of ferroptosis. This can be a distinct mechanism/pathway from the ferroptosis pathway, as multiple cell death pathways can be initiated in cells. Consistent with this, ferrostatin-1 also inhibited cell death, likely due to inhibition of the ferroptosis signaling cascade. It is unclear if this pathway is upstream of the caspases. How these ferroptosis triggers selectively activate CASP5 and not CASP4 to induce GSDME cleavage is a major unresolved question. Notably, it is also unclear if this biology is specific to the mesenchymal-like cells used in this study or if it expands to other cells.

    4. Reviewer #3 (Public review):

      Summary:

      Akter et al. identify caspase 5 activation and Gasdermin E cleavage as a novel downstream executioner of ferroptotic cell lysis induced by erastin and ML162. These data are novel and very interesting to the wider cell death community.

      Strengths:

      Strengths of the study include the use and validation of findings in several mesenchymal ovarian cancer cell lines, the rigorous validation using small molecule approaches, siRNA-mediated silencing and CRISPR/Cas9-mediated knockouts with re-expression.

      Weaknesses:

      A weakness of the study is the fact that ferroptosis was not induced genetically (GPX4 ko) and, hence, off-targets of the mode of induction cannot be ruled out at this point (e.g. ML162 also targets TrxR1). Moreover, it would be vital to understand at which point in ferroptosis execution caspase 5 is activated in a time-resolved kinetic together with lipid ROS tracing to also obtain hints as to its possible activation.

      Conclusion:

      Despite the weaknesses described, this is a very interesting, timely, and well-executed study with the described limitations. The work provides important mechanistic insights into the interplay between ferroptosis and pyroptosis with possible consequences for inflammatory responses.

    1. eLife Assessment

      This study presents an important empirical analysis of the acoustic space of vocal repertoires across primates and humans, providing new data that challenge the widely held hypothesis that the expansion of the human vocal space, driven by modifications of the vocal tract, was a key prerequisite for the evolution of spoken language. The evidence is convincing in its technical implementation and acoustic measurements, offering helpful comparative insights into voiced vocalizations. However, the theoretical framing and interpretation would benefit from further discussion.

    2. Reviewer #1 (Public review):

      Summary:

      The authors conducted a comparative acoustic analysis of primate vocal repertoires, focusing on the assumption that speech and language evolution required and involved an expansion in the acoustic space of voiced vocalizations from non-human primates to humans. Results challenge this idea. The study compiles and analyzes a large dataset of calls to quantify differences in vocal production space.

      Strengths:

      The study is technically sound, with a solid implementation of acoustic measurements and a valuable new dataset that brings empirical rigor to test a dominant, yet hitherto strictly theoretical, notion about what speech and language evolution entailed. It provides concrete comparative acoustic data across species to disprove that speech and language required an increase in the range of voiced calls, and thus, by extension, of vowels. The approach is methodologically rigorous and directly engages with the relevant data, rather than relying on untested presumptions of what great apes "ought" to be able to do or not.

      Weaknesses:

      The theoretical contextualization should be strengthened and updated, as several aspects contain inaccuracies, most notably by equating voiced calls or vocalizations with speech (overlooking the critical role of consonants, as human languages typically show vowel:consonant ratios of 1:4 or greater) and misrepresenting the premises and current status of the neural (Kuypers-Jürgens) hypothesis.

      The discussion drifts into speculative territory on features like syntax and co-articulation that fall outside the paper's scope and data, and it does not sufficiently engage recent evidence on vocal learning and consonant-like capacities in great apes.

      Minor issues include incomplete sampling justifications, imprecise terminology, and reliance on references that have been critiqued in more recent work.

    3. Reviewer #2 (Public review):

      This study examines the evolutionary context of the emergence of human speech. The authors address the widely held hypothesis that the expansion of the human vocal space, resulting from modifications of the vocal tract, was a key prerequisite for the evolution of spoken language.

      To test this hypothesis, the authors quantified the acoustic space of human speech, non-linguistic vocalizations, and musical vocalizations and compared it with that of nonhuman primates, chimpanzees, bonobos, and chacma baboons.

      The authors found that speech and song occupied significantly less volume in the acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volume of speech and song was not statistically distinct from that of non-human primates. Accordingly, the authors conclude that the evolution of human speech did not depend on an expansion of the human vocal acoustic space.

      I find the analysis presented in this manuscript highly convincing. It is conducted at a contemporary scientific standard, and the results provide strong support for the authors' conclusions. I particularly appreciate that the authors explicitly discuss the limitations of their approach. For example, they acknowledge that MFCCs cannot capture all aspects of acoustic structure.

      I have only three minor comments:

      First, the authors may wish to briefly summarize the main findings of the study by Anikin et al., as it represents the central reference for the present work. A concise summary in two or three sentences would help readers who are not familiar with that study.

      Second, I would appreciate a brief explanation of why the authors chose this particular statistical approach.

      Third, the authors could briefly mention that the Chacma baboon dataset provides a very comprehensive representation of the vocal repertoire of this species, although a small number of rare vocalizations are not included. I am not sure whether a similar limitation also applies to the chimpanzee and bonobo datasets, but if so, it would be useful to mention this as well.

    1. eLife Assessment

      The ability to measure autophagic flux in vivo in response to physiological stresses remains a challenge for investigators in the field; the novel mouse model present in this work is significant and should prove valuable to investigators in addressing shortcomings of existing models. In addition, the development of the microplate reader approach to permit semi-high-throughput analysis of samples is convincing and a significant advance. The application of these approaches to measuring autophagy in multiple tissues is appreciated but raises some questions that need to be answered and highlighted, including what new biological insight has been generated for tissues under study, how overall autophagy versus rates of flux are determined, and how the sex of the animal affects outcomes. The neuronal populations under study should be reassessed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors develop a GFP-LC3-RFP autophagy reporter under the control of the Rosa26 locus to measure autophagic flux in mouse embryos as well as adult tissues. While image quantification is consistently used, the authors also develop a semi-high-throughput assay for measuring autophagic flux using a microplate reader. Additionally, the authors cross these mice with a Cre-inducible Atg5-deletion mouse model, allowing the investigation of how autophagy flux is affected upon loss of Atg5. With this model, they demonstrate that loss of Atg5 leads to an increased ratio of GFP/RFP intensity in multiple tissues, including the brain, revealing that the brain undergoes basal autophagy. They further go on to show that the increase in GFP/RFP intensity upon Atg5 loss is greater in adult tissues compared to their embryonic counterparts. The development of an animal model, along with quantitative tools to measure the model, will have a high impact on the field. However, the analyses from the data presented do not fully justify the conclusions.

      Strengths:

      (1) A mouse model to better measure autophagy.

      (2) The plate-reader-based method to quantify autophagy across tissues.

      (3) Assessment of autophagy in many different tissues.

      (4) Crossing the reporter mouse with the Atg5f/f mouse to assess basal autophagy.

      Weaknesses:

      (1) While the tool is of high impact, there is little new biological or mechanistic insight provided in these studies.

      (2) The quantification and normalization method is unclear, making it difficult to compare across tissues accurately.

      (3) Differential expression across cell types is not well documented or taken into account for comparisons.

      (4) There is no consideration for sex as a biological variable.

    3. Reviewer #2 (Public review):

      Summary:

      The aim of the authors was to measure starvation-induced and basal autophagy in vivo across several tissues and developmental stages. For this, they developed a novel mouse model expressing the GFP-LC3-RFP reporter. They also aimed to provide a more high-throughput method for autophagy flux measurements than assessment by imaging and developed an assay based on a microplate reader.

      Strengths:

      (1) Good validation of the mouse model. The knock-in strategy is well explained and illustrated.

      (2) The model has potential to be applied to a wide range of research questions. The Cre-dependent expression allows for customization of KO timing, which will be beneficial in developmental studies.

      (3) The authors presented consistent findings using two different methods to quantify autophagy, strengthening the robustness of their results.

      (4) The authors demonstrated the validity of the high-throughput method (microplate reader).

      Weaknesses:

      (1) The comparison of neuronal populations in different areas of the brain is not ideal. In the cerebellum, Purkinje cells were chosen, which are rare and not representative of this tissue, as well as functionally very different from the neurons in the hippocampus and cortex that they were compared to.

      (2) The explanation of the GFP-LC3-RFP construct and specifically if/how autophagosome formation can be measured and distinguished from flux could be clearer.

      Conclusion:

      The work presented is thorough, and the authors achieved their goals for this study. The effort used to further investigate unexpectedly high basal levels of autophagy in the brain is well appreciated and adds value to this paper. The conclusions of the authors are mostly very well supported by the data provided. The well-structured description of the results, along with clear figures, allows the reader to comprehend the authors' reasoning in reaching their conclusions.

      The presented mouse model has great potential for a lasting positive impact on the research field of in vivo study of autophagy. The method of utilizing a microplate reader will also benefit future research where semi-high throughput is an advantage. Together, the information provided in this study not only presents new methodology that will allow the investigation of new research questions, but also provides novel information about in vivo autophagy flux at the selected developmental stages that opens up new follow-up research questions.

    1. eLife Assessment

      This manuscript presents valuable experimental results describing the localisation and regulation of casein kinase CK1δ during the cell cycle. During the G2 phase of the cell cycle, the autophosphorylation of the C-terminal tail stabilises and inhibits the kinase which might protect CK1δ in the subsequent G1 phase. The results may be of interest to researchers working on casein kinase function and regulation but due to lack of some controls remain incomplete.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      We appreciate the authors have provided answers to many of the points we raised, and the changes made to their manuscript, which we think strengthen the overall evidence presented. However, we find that some important controls are still missing across experiments.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Antibody cross-reactivity was only tested against CK1ɛ, but should also be tested against CK1α, which is abundant in U2OS cells, and is also known to be involved in cell-cycle regulation.

      b. Fig. 1: Statistical analyses are missing from the analysis.

      c. Fig. 2: No colocalisation analysis shown for figure 2, only some arrowheads pointing to puncta. Appropriate colocalisation statistics are important since for practical reasons, only a few representative images can be shown on the figure.

      d. Fig. 6: Even if the figure is illustrative, it is important to show centrosome staining to visualise CK1ẟ's recruitment to the centrosome in G2/prophase, especially since this information is used to propose the model in figure 7.

      e. For all figures: Please mention the number of independent biological replicates in the figure legends (1, 2, 6). For figure 1, if there are 3 independent biological replicates, the quantification should take all of them into account (as opposed to the data points corresponding to 10 cells), and statistics must be done appropriately, taking those independent replicates into account. Same for the colocalisation analysis in figure 2 once you include it.

      (2) Shortcomings in biochemistry experiments:

      On CalA control, this is not a matter of confirming that CalA treatment works in principle, but rather to confirm that CalA treatment worked in this specific replicate. Aliquots may lose potency (e.g. with freeze-thaw cycles / exposure to light), hence checking for enrichment of phospho-proteins is essential to confirm the treatment was successful in this particular instance. In the worst-case scenario, the company may have sent the wrong compound altogether! A positive and a negative control is the basis for every experiment to make meaningful interpretation. On a separate note, many experiments have control and siRNA or compound treatments on two different gels - this should be rectified as they are meaningless if different exposures have been selected for different immunoblots.

      (4) As the authors mention, the kinase is not fully inactive when tail phosphorylated. Recent research has also suggested that tail-phosphorylated CK1ẟ may show increased catalytic activity for a few select, specific substrates, in the co-occurrence of pT220 (Cullati et al., 2022; Cullati et al., 2024). It is thus tricky to directly infer that phosphorylated CK1ẟ is inhibited, when no positive control for CK1ẟ inhibition was shown in the evidence presented. It would be necessary to either nuance your claim or include a positive control for CK1ẟ inhibition. Please revise statements in the manuscript accordingly.

      (5) It would be important to include statistical analyses for the immunofluorescence data in Fig. 1 and 2.

      (9) The authors mentioned "In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable." Unfortunately, the data from biochemical analyses presented in the eLife publication is uninterpretable due to a lack of loading controls.

      (10) While the data presented in Penas et al. strongly suggests a link between CK1ẟ stabilisation and the APC/C-Cdh1 complex, it is the only study to have shown it. Given that (1) science relies on data reproducibility and (2) your proposed model relies heavily on the relationship between CK1ẟ stabilisation and the APC/CCdh1 complex, it would be appropriate to include the investigations mentioned in our original comment.

    3. Reviewer #2 (Public review):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U20S cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 WT but not a phospho-null mutant strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin and Okadaic Acid) and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      The manuscript has improved overall, but some sections are still inconclusive and require clarification.

      Major comments:

      Figure 1 is inconclusive. CK1 nuclear staining is highly similar in untreated cells and in cells treated with CHX + PF670462. The reduction in centrosomal staining in these cells is barely significant. However, the authors draw very strong conclusions from these data sets. In panel B, the cells appear to have been fixed incorrectly, and the anti-PCNT shows a strong background signal. Not convinced that immunofluorescence is the best approach to look at protein dynamics in vivo.

      In Figure 2, panel B, the authors should co-stain the centrosomes of cells that co-express CRY1 and CK1, as some of these dots may represent the centrosomes.

      Figures 3B, please provide information on the non-phosphorylable CK1a mutant (it is mentioned as a variant in which all serine and threonine residues in the C-terminal tail were replaced by alanine). Specify the number of sites mutated and their exact position. Is this non-phosphorylable CK1a version catalytically active? Treatment of samples with inactivated PPase should be used as a control.

      Strengths:

      The authors reveal that the activity and abundance of dephosphorylated and phosphorylated CK1δ are regulated in a cell cycle-dependent manner. This suggests that these different pools are associated with distinct physiological functions.

      Weaknesses:

      Unfortunately, some of the data are inconclusive, and there is no data/information linking the cell cycle regulation of CK1δ to its function during the cell cycle.

    4. Author response:

      General Statements.

      We thank all reviewers for their careful evaluation and constructive comments.

      Fidel Serrano recently completed a related study on the role of CK1δ in the circadian clock which is published in eLife (https://doi.org/10.7554/eLife.110786.1). During these studies, we became interested in the physiological role of CK1δ autophosphorylation, whose functional significance has remained unclear.

      In the present manuscript, Fidel discovered that the autophosphorylated, auto-inhibited form of CK1δ accumulates specifically during mitosis. These findings provide a physiological context for CK1δ autophosphorylation that has remained elusive for many years. They constitute the central foundation of our model, which further integrates previous findings on APC/C function and activity with the data presented in our eLife study.

      Several suggestions and questions raised by the reviewers concern the regulation of CK1δ in cultured cells, which are predominantly in G1. Many of the corresponding experiments and analyses are already included in our eLife paper. We apologize that this overlap may complicate the review process. However, we are unable to publish identical datasets in both manuscripts.

      We therefore briefly summarize here the findings from the eLife study that are most relevant to the present work.

      - Overexpressed CK1δ is unstable, whereas CK1δ-K38R (catalytically inactive) and CK1δ-R178Q (altered specificity/activity) are comparatively stable (eLife Figs. 3A-E and 5B).

      - Assembly with PER2 stabilizes overexpressed CK1δ (eLife Figs. 4A-D).

      - Treatment with PF670462 similarly stabilizes the kinase (eLife Figs. 3F and 5D).

      - CK1δ binding sites in the centrosomal/Golgi area are present in excess even relative to overexpressed kinase (eLife Fig. 5C).

      The reviewers also raised questions about the PER2-CRY1 nuclear foci.

      Thermodynamically stable/persisting nuclear foci form upon coexpression of PER2 and CRY1. They were preliminarily characterized in the eLife study (eLife Fig. 2). For the purposes of the present work, they constitute a serendipitous and convenient tool for studying interactions of PER with CK1δ. To induce foci formation, we used a stable cell line expressing DOX-inducible mK2-CRY1. mK2-CRY1 accumulates only at relatively low levels because the protein without a binding partner is intrinsically unstable, as confirmed by Western blot analysis (eLife Fig. 4C). Coexpression of PER2 stabilizes mK2-CRY1 and promotes the formation of nuclear PER2-mK2-CRY1 foci.

      These foci contain elevated levels of endogenous CK1δ (eLife Fig. 2E and EV2), which accumulate gradually over a 24-hour period. This observation indicates that, at steady state, a fraction of endogenous CK1δ is continuously degraded in the absence of overexpressed PER2. Because endogenous CK1δ is synthesized at a relatively low rate, the unstable pool is small at any given time and therefore not readily detectable in conventional cycloheximide chase assays against the kinetically stabilized steady-state background of CK1δ.

      In our point-by-point response, we therefore refer the reviewers to the corresponding datasets and analyses presented in the eLife manuscript.

      Point-by-point description of the revisions.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      Involved in various cellular pathways and processes, Ser/Thr kinase CK1δ is thought to be constitutively active and its tail autophosphorylation is suggested as a putative inhibitory mechanism of CK1δ kinase activity. Here, the authors investigate CK1δ's phosphorylation status, in relation to its location and dynamics throughout the cell cycle. Immunofluorescence and biochemistry studies showed that the subcellular distribution of CK1δ is dynamic and in equilibrium between centrosomal and nuclear pools. The authors argue that CK1δ phosphorylation protects it from degradation. Finally, the authors show differences in CK1δ phosphorylation status and location throughout the cell cycle. Combining their findings with the existing literature, they propose a model of CK1δ phosphorylation status, location and dynamics throughout the cell cycle.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Lack of a positive control for centrosomal staining, eg. PLK1 (Fig.1A,C, 2B, 6A-B). This would also allow a colocalisation analyses between CK1δ and a centrosomal marker, which would build a more convincing body of evidence in favour of centrosomal recruitment of CK1δ in baseline conditions. A lack of a negative control for CK1δ IF? Most CK1 antibodies also cross-react with other isoforms.

      As suggested, we show centrosomal staining with a well-studied centrosomal marker pericentrin (PCNT; revised Fig. 1). Consistent with what has been well-described in the literature that CK1δ localizes to the centrosome (Sillibourne et al., 2002; Greer & Rubin, 2011; Greer et al., 2014) and with our observations, we confirm that CK1δ is indeed localized at the pericentral region.

      The CK1δ antibody does not crossreact, this is shown in Fig EV3A of the eLife paper.

      b. Lack of proper image quantification (Fig.1A,C, 2B, 6A-B). To establish a centrosomal recruitment of the CK1δ pools, colocalisation of CK1δ and a centrosomal marker must be quantified using a standard colocalisation quantification method and shown appropriately. Same comments regarding the colocalisation of CK1δ with PER2/mK2-CRY1-positive foci.

      Fig. 1: Colocalization analysis with pericentrin (PCNT) as centrosomal marker is provided.

      Fig. 2: The colocalization analysis of CK1δ (green) and mK-CRY1 (magenta) shows that in all PER2-untransfected cells (50 cells evaluated), CK1δ is concentrated at a single, mK2-CRY-negative spot, corresponding to the pericentrosomal region, which we have established in Figure 1. In contrast, in all PER2-transfected cells (15 cells evaluated) CK1δ localizes to mK2-CRY1 positive nuclear foci. The fields-of-view we have provided are representative of the typical phenotype of CK1δ localization in the presence or absence of PER2/mK2-CRY1-positive foci. Here it should be noted that CRY1 foci formation is strictly dependent of co-expression with PER2, (eLife paper).

      Fig. 6: This figure presents an analysis of fixed cells. Because only a small fraction of asynchronously growing cells is in mitosis at any given time, the number of cells assigned to individual mitotic stages is necessarily low (typically fewer than 10 cells per stage). The purpose of this figure is therefore primarily illustrative: to document and confirm the known subcellular localization of CK1δ during the different stages of mitosis rather than to provide a comprehensive quantitative analysis.

      c. Lack of proper statistical analyses (Fig.1, 2, 6). As immunofluorescence constrains one to only show few images per condition at best, statistical analyses on broader image analysis data are essential to measure the significance of the changes observed on the representative images provided.

      See answer to previous question.

      d. Lack of biological replicate numbers (Fig.1, 2, 6). Was it from 3 independent biological replicates?

      Fig.1 and 2: three independent biological replicates.

      Fig.6 is just illustrative to show the already known localization of CK1δ across mitosis. The localization shown in the four panels is seen in all cells at the respective cell cycle stages.

      e. NOTE: if available, using a confocal microscope would be best to provide optimal Z-axis resolution. This would provide you with more accurate colocalisation data, like CK1δ recruitment to the centrosome or PER2/mK2-CRY1-positive foci.

      Characterization of PER-CRY foci is published in the eLife paper in Fig. 2 and EV2.

      (2) Shortcomings in biochemistry experiments

      Lack of loading controls in Fig.3A-C and 4A-B. Until they are added, no conclusions can be safely interpreted from the experiments. Myc is a good turnover control for CHX across all experiments. Enrichment of phospho-proteins upon CalA treatment?

      Loading controls are provided in this revision. Enrichment of phospho-proteins following CalA treatment has been demonstrated and published in numerous previous studies (Cegielska et al., 1998; Ishihara et al., 1989; Rivers et al., 1998).

      (3) The authors infer from the literature that overexpressed CK1δ is unassembled, without checking it in any experiment. It could be that CK1δ overexpression drives overexpression of its binding partners, in which case most of overexpressed kinase could actually be assembled. Since CK1δ assembly is of great importance to the study's conclusions, it should be confirmed experimentally.

      In our eLife study, we show that overexpressed CK1δ is not fully assembled with stabilizing binding partners such as PER2. In fact, centrosomal/Golgi binding partners remain always available in excess even relative to overexpressed CK1δ, as demonstrated by the MG132-induced increase in centrosomal accumulation (eLife Fig. 5C). However, binding is dynamic and the affinities and concentrations are such that a substantial fraction of overexpressed CK1δ remains unbound and is therefore subject to degradation.

      (4) The authors infer from the literature that phosphorylated CK1δ is inactive - which is not a given. All Western blotting experiments should include confirmed CK1δ substrates like PER2 or DVL3 to confirm that phosphorylated CK1δ is indeed inhibited. As added benefit, blotting for known CK1δ substrates will act as confirmation that the FLAG-tag on overexpressed CK1δ/ε does not impact its kinase function, and that CK1δ/ε inhibitor treatments like PHF670 have indeed worked. Furthermore, because CalA and OA inhibit many phosphatases, inferences on CK1δ/ε activity upon these treatments should be taken catiously.

      The activity of phosphorylated CK1δ is severely attenuated by autoinhibition, although the kinase is not completely inactive. This has been demonstrated repeatedly in the literature and does not require re-establishment in the present study. We previously showed this directly in a PNAS study by Marzoll et al. (https://doi.org/10.1073/pnas.2118286119). At some point, one must rely on established published findings unless there is compelling evidence supporting an alternative interpretation.

      (5) Figure 1 requires further biochemical analyses to support immunofluorescence data. Western blotting analyses showing fluctuating CK1δ levels in nuclear vs. Centrosomal (https://pmc.ncbi.nlm.nih.gov/articles/PMC7618310/) fractions in the different treatment conditions would help illustrate the redistribution of CK1δ pools between nuclear and centrosomal areas. Additionally, adding a cytoplasmic fraction in each treatment condition would help visualise the amounts of unrecruited CK1δ in overexpressed conditions.

      Microscopy is a well-established and widely accepted approach for assessing subcellular localization. In many cases, it is superior to biochemical fractionation assays, which rarely yield completely pure nuclear or centrosomal fractions. The fractionation experiments suggested by the reviewer could provide additional supportive evidence, but they are not required to substantiate the conclusions presented here.

      (6) Figure 2 requires further biochemical analyses to support immunofluorescence data showing colocalisation of CK1δ with PER2, using for example immunoprecipitation to show that pulling down PER2 also pulls down CK1δ, and vice versa. Optimally, an additional Western blotting experiment showing total / phospho-PER2 and CK1δ levels for each condition in cytoplasmic vs. nuclear fractions would consolidate the evidence from immunofluorescence and immunoprecipitation.

      Association of CK1δ with PER2 has been demonstrated extensively in numerous previous studies (Aryal et al., 2017; Narasimamurthy et al., 2018; Philpott et al., 2020; Cao et al., 2021; Marzoll et al., 2022; An et al., 2022) including our own eLife publication.

      (7) The authors consistently interpret data showing CK1δ phosphorylation as CK1δ tail phosphorylation (p.8-11). A band shift in CK1δ signal on a Western blot does not in any way show the tail specifically is phosphorylated. Indeed, although CK1δ tail phosphorylation is widely recognised, several identified sites on other CK1δ domains, e.g.: the kinase domain, can be phosphorylated by CK1δ and other kinases. The authors' conclusions must therefore not mention tail-specific phosphorylation unless domain-specific phosphorylation is established. This could be done by comparing signal from antibodies specifically recognising known phospho-sites on the CK1δ tail against total CK1δ signal in experiments of Fig.3, 4, and 5.

      The electrophoretic mobility shift is caused by phosphorylation of the CK1δ tail. This has been demonstrated in numerous previous studies. For example, we generated a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines. In this mutant, CK1δA, inhibition of phosphatases by CalA no longer induces an electrophoretic shift (Figure 3B).

      (8) The use of anti-FLAG antibody in all figures looking at overexpressed CK1δ is not the most appropriate choice, as the anti-CK1δ antibody works perfectly for Western blotting and immunofluorescence studies. This adds more variables and hinders comparisons with endogenous CK1δ. If CK1δ is successfully overexpressed, wild-type levels should be negligible compared to overexpressed CK1δ levels, so the need for a FLAG antibody is not justified. Using an anti-CK1δ antibody in all experiments instead of anti-FLAG would confer them greater solidity.

      The amount of endogenous CK1δ is limited, and we therefore use endogenous detection only when it is scientifically necessary. FLAG-tagged constructs, in contrast, provide a robust and convenient experimental system. In the absence of evidence that the FLAG tag introduces artifacts or alters the observed behavior of the kinase, we do not consider it necessary to repeat these experiments using endogenous detection alone.

      (9) The authors base themselves off a correlation between CK1δ phosphorylation status and its observed stability to establish a causal relationship between both ('phosphorylation protects the overexpressed kinase from degradation' p. 7; 'tail phosphorylation protects the overexpressed kinase from rapid degradation' p.9). However, no experiments performed suggest there is a causal relationship occurring. This could be done by looking at specific known CK1δ tail phospho-sites using phosphosite-specific antibodies. The authors could assess whether CK1δ stability is impacted when overexpressing wild-type vs. phospho-dead mutant CK1δ.

      In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable.

      (10) The authors consistently link APC/CCDH1 with CK1δ's putative degradation, when none of their study touched on APC/CCDH1 activity or its involvement in CK1δ dynamics. To support these conclusions, fig.5 requires further biochemistry studies. To show APC/CCDH1's interaction with CK1δ, I suggest the authors (1) pull down CK1δ by immunoprecipitation and check for APC/CCDH1, phosphorylated CK1δ, and ubiquitin levels in enriched samples for each cell cycle stage in untreated vs. MG132-treated conditions. To consolidate this, the authors could (2) pull down CDH1 and check for phosphorylated and total CK1δ levels in pulled-down samples. Finally, they should investigate whether reducing CDH1 (3) expression (siRNA knock-down) and (4) CDH1 activity (specific E3 ligase inhibitor) have any impact on CK1δ levels at each cell cycle stage.

      These data were previously published in Penas et al. (doi: 10.1016/j.celrep.2015.03.016). For example, Fig. 5C shows that siRNA-mediated depletion of Cdh1, the G1-specific cofactor of APC/C, stabilizes overexpressed CK1δ, as well as the established APC/C-Cdh1 substrate Cyclin B1.

      (11) Methods section lacks a subsection detailing image acquisition, including the type of microscope used (confocal vs. widefield), magnification used, whether acquisition settings were kept consistent throughout conditions / technical/biological replicates.

      Is provided in the revised manuscript.

      (12) Methods section must include a subsection detailing immunofluorescence image analysis parameters and statistical analyses, e.g.: criteria for centrosomal localisation categories in Fig.1, criteria for categorising cells in different stages of mitosis (Fig.6), overall threshold / criteria stringency.

      Is provided in the revised manuscript.

      Minor comments:

      The Reviewer has raised about 50 minor points, counting the individual remarks and sub-point and sub-sub points. We have addressed a number of these comments where they were scientifically relevant or helpful for improving clarity. Most of the questions raised concern published and generally accepted data. We therefore do not believe that a point-by-point response to every individual minor remark is constructive or necessary.

      (1) PF670 was shown to selectively inhibit CK1δ/ε over 42 common kinases (TOCRIS), however it is not well-characterised regarding the remaining 474 kinases encoded by the human genome. There is thus a possibility for PF670 inhibition overlap between CK1δ/ε and other kinases, which should be mentioned, and the authors' conclusions should he more nuanced as a result.

      (2) CHX is a protein synthesis inhibitor, which means it does not selectively target CK1δ expression, but that of every protein within the cell. This should be addressed and controlled for, if possible. A potential way to go about it would be to knock-down CK1δ using siRNA and compare untransfected controls with select post-transfection timepoints to assess the impact of inhibiting CK1δ expression on total CK1δ levels.

      (3) All unshown Western blotting replicates should be included in the supplementary materials.

      (4) Figure 1:

      a. A-B: experiment is missing PF670 alone and CHX alone conditions to control for direct effects of either compound, versus combined.

      b. A-B: in the text, please address the fact that >50% cells show unclear or no centrosomal pattern in untreated conditions.

      c. C: the authors claim that CK1δ-FLAG levels are decreased in PF670-treated conditions because PF670 inhibits excess CK1δ autophosphorylation, inducing its subsequent degradation. Before this claim is made, a proteasome inhibitor like MG132 should be included within the experiment, to show that this decrease in CK1δ expression under PF670 treatment can be rescued with MG132 treatment. This would plead in favour of excess CK1δ degradation and exclude the possibility that 4h of PF670 treatment simply reduces the rate of CK1δ expression. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ degradation occurs via APC/CCDH1-mediated ubiquitination of CK1δ.

      d. C-D: experiment is missing pre-DOX induction control to confirm that CK1δ-FLAG is indeed being overexpressed.

      (5). Figure 2:

      a. A: should include non-transfected control panels to confirm the success of PER2 overexpression.

      b. B: Although observed in a preprint from the same team, there is no peer-reviewed evidence that establishes mK2-CRY1 as a reliable indicator of PER2 overexpression. 2A shows a correlation between both but does not exclude the fact that PER2 must be included in the imaging of the 2B panels, instead of using mK2-CRY1 as a proxy readout of PER2 expression and location. Appropriate analyses would then be required to show colocalisation of CK1δ with PER2 in the highlighted puncta.

      c. 'PER2-dependent foci': cannot be said of the data unless PER2 dependency has been validated in those images. Please see above point to resolve this.

      (6) Figure 3:

      a. B: the use of kinase-dead CK1δ mutant does not allow to fully separate direct autophosphorylation from phosphorylation by other kinases, unlike what the authors mention: CK1δ kinase activity may be required for the phosphorylation of certain sites by other kinases. In this case, the decrease of phosphorylation in the kinase-dead CK1δ mutant would not only result from inhibited autophosphorylation, but also from reduced phosphorylation by other kinases. Please adjust your conclusions accordingly (p.8).

      b. C: poor visualisation of CK1δ overexpressed condition, especially showing critically reduced signal at the 6min CHX timepoint, compared to its kinase-dead homolog. If this change is present in every replicate performed, please address it in the text. If not, perhaps you may have to display another representative replicate in the figure.

      c. The authors claim that the increased levels of overexpressed with CalA treatment suggest 'that full or partial phosphorylation of the CK1δ tail stabilizes both active and inactive forms of the kinase' (p.8). Importantly, tail phosphorylation was never shown, so this should be corrected. Additionally, it could be that CalA treatment considerably increases the rate of CK1δ expression - which would also match data in fig.4A, since CalA and CalA+PF670 treatments alone drastically increased CK1δ levels. This should be checked by pre-treating cells with CHX before applying CalN, or by treating cells with CHX and CalN simultaneously, and interpreted accordingly.

      (7) Figure 4:

      a. A: CK1δ levels in CHX-free controls from both 1h pre-treatments look much higher compared to untreated controls, which the authors interpret as an indication that 'most of the newly synthetised overexpressed kinase was degraded in untreated cells' (p.9). However other explanations are not explored: since loading controls are not provided, it may be that sample loading in the gel is simply off. Importantly, it is also possible that pre-treatments increased CK1δ expression before CHX application. Please make sure you touch on each

      b. A-B: the authors claim that CK1δ-FLAG and CK1ε-FLAG levels are decreasing with CHX treatment because they are being degraded. Adding a panel with a proteasome inhibitor like MG132 would solidify this argument. Rescue of CK1δ/ε degradation under CHX treatment would show that the loss is indeed mediated by the UPS. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ/ε degradation occurs via its ubiquitination by APC/CCDH1.

      c. B, D: blot in B does not match the quantification trends in D. E.g.: 60min CHX + CalA + PF670 condition shows clearly lower CK1ε signal compared to its 0min CHX control. Please ensure the biological replicate you display on the figure is indeed representative of your results.

      d. A, C: 'Hyperphosphorylated CK1δ remained stable throughout the CHX chase [...], indicating that tail phosphorylation protects the overexpressed kinase from rapid degradation' (p.9). Meanwhile this is true, unphosphorylated CK1δ in the CalA+PF670 treatment condition was also stabilised, showing that CK1δ phosphorylation may not be required for kinase stabilisation. This is an important point and should be addressed in the data interpretation. On another note - and as mentioned above -, tail phosphorylation specifically is not shown and cannot be inferred unless domain-specific phosphorylation is investigated.

      e. B, D: CK1ε-FLAG levels decrease with CHX treatment compared to its baseline in the CalA+PF670 condition, which is not the case for CK1δ-FLAG (A). Thus, the data shown does not support the authors' conclusions 'CK1ε turnover is regulated in a similar manner to CK1δ'. Please address this issue and adjust your conclusions accordingly.

      f. C-D: performing statistical analyses on protein band intensity in different conditions would be interesting to establish the significance of those changes.

      (8) Figure 5

      a. B: non-arrested controls mentioned in the text are missing from the figure.

      b. B: 'the accumulated CK1δ remained dephosphorylated' (p.10). The experiment is missing important positive / negative controls of CK1δ phosphorylation status to conclude whether CK1δ remained phosphorylated or unphosphorylated across conditions. One or the other cannot be concluded from the blot as it is. Using phospho-specific antibodies may also help to visualise phosphorylation status.

      c. B, D-E: blots should include CDH1 phosphorylation levels (hyperphosphorylated CDH1 is inactive), CK1δ substrates (indicators of CK1δ activity), and relevant phosphatase substrates to create a cohesive picture of changes in mitosis and support fig.7.

      d. D-E: please address the differences in CK1δ profile between overexpressed and endogenous CK1δ in G2/M phase.

      (9) Figure 6

      The authors say: 'similar results were observed when using U2OStx cells and staining for endogenous CK1δ'. However, CK1δ's subcellular distribution pattern is different in endogenous vs. overexpressed conditions in the telophase-cytokinesis and post-mitosis stages. In the telophase-cytokinesis stage, endogenous CK1δ seem to form nuclear hotspots, while overexpressed CK1δ is more diffuse. In the post-mitosis stage, overexpressed CK1δ shows a clear polar pattern in the nuclear periphery, while endogenous CK1δ shows a diffuse pattern similar to that of overexpressed CK1δ described in fig.1 as 'unassembled' by the authors. Please address this and adjust your conclusions appropriately.

      (10) Figure 7

      a. The authors infer APC/CCDH1's activity levels or relationship to CK1δ from existing literature only. Since it is a central mechanism of their study, literature-only components are insufficient for a summary figure. To include these elements in the figure, the authors must include investigations of APC/CCDH1's activity levels and involvement in CK1δ degradation at different stages of the cell cycle in their study. Please refer to point 10.

      b. Similar comment for CK1δ assembly status and activity levels. Please refer to point 4.

      c. Similar comment for phosphatase activity levels in different stages of the cel cycle. By observing phosphorylation status in known substrates of established CK1δ phosphatases, one can easily confirm phosphatase activity levels.

      d. Does not consider the fact that some results were different in endogenous vs. overexpressed CK1δ models. Please nuance your claims.

      (11) Reference missing p.12 paragraph 1: 'Yet, overexpression of CK1δ consistently accelerates the circadian clock, implying that kinase availability can influence clock speed. This finding suggests that CK1δ activity may be regulated not only by catalytic mechanisms but also by its spatial availability within the cell'. The facts stated there are not covered in the study's findings and is not referenced with a published study.

      (12) Grammar mistakes / typos to report:

      p.3 paragraph2: 'the kinases undergo futile cycles of phosphorylation and dephosphorylation'

      p.26 Fig.1A legend: 'Endogenous CK1δ was detected by IF to'? Unfinished formulation

      p.8: 'PP1 was previously suggested as a major PPase of CK1δ/ε'

      p.8: 'fewer phosphorylation sites are targeted by other kinases'

      Figure 4C legend: 'Overexpressed unphosphorylated CK1δ [instead of CK1ε] is degraded with a half-life of about 15 min'

      Figure 4E: Ponceau staining

      (13) Nomenclature inconsistencies to report:

      p.4 paragraph2: 'protein PER2', then p.4 paragraph3 'PERIOD2'.

      'CK1δ' used throughout the article's body text, but 'CSNK1D' is used in IF panel legends. The gene name was never introduced in the main text, nor has it been explained in the figure legends, so perhaps go for CK1δ for all mentions, including in figures.

      'FOV' nomenclature is unclear, please define in the figure legend.

      Figure 2B: what are the arrowheads pointing to? - please clarify in the figure legend.

      Mislabelled figures: figure 3 in the text refers to figure 4 in the figure section, and figure 4 in the text to figure 3 in the figure section.

      Figure 6A-B: abbreviated mitosis stages 'Pro' and 'Meta-Ana' should either be defined in the figure legend or put in full writing within the figure.

      Reviewer #1 (Significance):

      The paper will appeal to those working on CK1 biology, including cell cycle and circadian rhythms.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      In this study, authors aimed to address functional links between CK1 activation, subcellular localization and protein stability, which is an important biological question. This manuscript is a follow up of a recent study by the same team, which is currently deposited at Biorxiv and a fraction of data seems to overlap, which is somewhat confusing. Overall, the concept that the stability of CK1d is dynamically controlled across the cell cycle is interesting. On the other hand, how this is functionally connected to the circadian cycle described in the previous study remains unclear.

      The data presented in this study are highly preliminary and lack a number of essential controls, which weakens an otherwise interesting concept.

      Major issues:

      One of the main conclusions of the study is that CK1d is degraded predominantly in the nucleus while the centrosomal pool is protected from the degradation. This concept is interesting but unfortunately, the experimental evidence supporting this model is very limited. Authors previously showed that neither inhibition of proteasome or treatment of cells with PF670462 significantly influenced levels of endogenous CK1d but both treatments promoted accumulation of the tagged and overexpressed FLAG-CK1d. Absence of the phenotype at the level of endogenous protein clearly raises question whether this may be just an artifact of overexpression, tagging or both combined.

      As shown in our previously published eLife study, CK1δ is synthesized at a relatively low rate from its endogenous locus. Free CK1δ continuously shuttles between the cytosol and the nucleus. Although association of CK1δ with the centrosomal/Golgi area is dynamic, nuclear export followed by binding to centrosomal/Golgi structures constitutes the major sink for the kinase at steady state. Consequently, the fraction of unbound CK1δ that is targeted for degradation in the nucleus at any given time is very small and cannot be detected in cycloheximide chase assays against the much larger background of kinase stabilized by association with centrosomal/Golgi binding sites. To reveal that CK1δ is subject to degradation at all, we had to overexpress the kinase. Under these conditions, the fraction of unbound CK1δ increases substantially, allowing nuclear degradation to be detected experimentally.

      The statement that degradation occurs in the nucleus is based on weak data with CK1d-NES construct that was claimed to accumulate at higher levels compared to the wild type CK1d. However, that experiment used quantification of microscopic data in which two distinct regions (nucleus, cytosol) were compared, which is technically challenging. Including a reporter for normalizing for the transfection efficiency would strengthen conclusions of that experiment. Overall, quantification by immunoblotting may be more accurate.

      As suggested, we performed CHX time-course experiments with NES- and NLS-tagged CK1δ constructs followed by immunoblot analysis (shown in revised Fig. 4E).

      Specificity of the microscopy staining was not validated Fig. 1. Confirmation of the staining specificity by RNAi or KO approaches is essential. This antibody from Abcam has been discontinued which makes it impossible to reproduce this experiment.

      The authors should therefore attempt to demonstrate localization using other available antibodies against CK1d.

      The antibody is available through Thermo Fisher in the United Kingdom: https://www.fishersci.co.uk/shop/products/100-ul-mouse-monoclonal-af12g4-casein-kinase/13070202#

      It is specific for CK1δ and does not recognize CK1ε, as demonstrated in our eLife study.

      Furthermore, FLAG-tagged CK1δ detected with anti-FLAG antibodies, as well as endogenous CK1δ detected with the commercial antibody, localize to the pericentrosomal region and relocalize upon PER2 overexpression to mCRY1-containing nuclear foci. Together, we consider these data compelling evidence for the specificity of the antibody staining.

      The authors also failed to demonstrate that the dots represent centrosomes. To do so, they must perform co-staining with a robust centrosome marker.

      As requested, Pericentrin (PCNT) was used as a centrosomal marker and is now provided in the revised Fig. 1.

      This would help classify approximately 50% of the cases that are currently assessed as "potential" but are in fact inconclusive.

      Centrosomal colocalization with PCNT improved the robustness of the quantification.

      The quantification in the experiment is incorrect. Since the percentages are reported, both columns should add up to 100%. In Panel 1B, approximately 10% of the cells are missing, while in Panel 1D, there appear to be about 20% more cells.

      As indicated in our previous figure, the y-axis represents the number of cells, not percentages. We analyzed approximately 100 cells, which may have led to the misunderstanding that the values refer to percentages. As well, we have replaced this figure with a revised Figure 1 which shows that CK1δ co-localizes with PCNT as a centrosomal marker.

      Fig. 2B shows four different fields based on which authors come to conclusion that co-expression of CRY and PER2 promotes re-localisation of CK1d to nuclear foci. This is an interesting possibility but it is hard to conclude without any quantification. What was the fraction of cells that expressed CRY and PER2 that showed this phenotype? What was the fraction of cells that did not show this phenotype although both of these proteins were expressed?

      It would help to label cells expressing PER2 either by expressing it as a fusion protein or by co-expressing a marker protein ideally from the same plasmid.

      All cells in this stable cell line express mK2-CRY1. The cells were transiently transfected with PER2, and therefore only a fraction of the cells received the PER2 expression construct. In our eLife paper, we showed that all PER2-expressing cells stabilized mK2-CRY1 and formed nuclear foci. We demonstrate here that every cell containing such nuclear foci also showed accumulation of endogenous CK1δ within these structures. We did not observe a single cell with PER2-induced nuclear foci that lacked endogenous CK1δ accumulation.

      Cells that did not express PER2 were identified by the absence of nuclear foci and by low levels of mK2-CRY1, which was homogeneously distributed throughout the nucleus, as also shown in the eLife manuscript. In these cells, endogenous CK1δ was concentrated at a single discrete structure that, based on data shown in Fig. 1, corresponds to the pericentrosomal region. Under these conditions, neither CK1δ nor mK2-CRY1 was detected in nuclear foci.

      Fig. 5 suggests that massive phosphorylation of CK1d, which is responsible for its mobility shift on SDS-PAGE, is most likely linked with mitosis. Authors should use established markers to estimate a fraction of mitotic cells in their G2/M fraction. In principle, there are two possible explanations for the doublet observed with CK1d staining. Ether CK1d exists in two pools with different phosphorylation states in mitosis, or perhaps more likely, this fraction contains G2 cells where CK1d is not yet modified and mitotic cells where CK1d is fully phosphorylated. Performing a shake-off experiment yielding a pure fraction of mitotic cells could help to distinguish between these two options.

      As described in the main text and the methods section, the fractions were prepared by mitotic shake-off to further enrich our sample for rounded cells that are loosely attached during mitosis.

      Degradation of CK1d in telophase/cytokinesis when APC/Cdh1 becomes active is not apparent in Fig 6. The signal at mitotic spindle is missing, but there is still plenty of signal remaining in the cells. It is possible that the signal is just redistributed in the cell and the data shown do not support degradation of the protein. Authors could film cells expressing fluorescently labeled CK1d and quantify the signal during progression through mitosis and mitotic exit. The statement that "Following nuclear envelope reformation and mitotic exit, CK1d localized primarily to the single centrosome in each daughter cell" is incorrect. First, authors cannot deduce from the fixed cells whether they have just formed the nuclear envelope and exited mitosis.

      The reviewer is, of course, absolutely correct in the points raised.

      First, we cannot deduce from fixed cells whether they have only recently exited mitosis. This was not our intention. We merely selected cells in G1 and referred to them as “post-mitotic,” without intending to imply that these cells had just exited mitosis. To clarify this point, we changed the previous statement:

      “Following nuclear envelope reformation and mitotic exit, CK1δ localized primarily to the single centrosome in each daughter cell (Fig. 6A, 4th column)”

      to:

      “In G1, CK1δ localized primarily to the single centrosome (Fig. 6A, 4th column),” and replaced in column 4 of Fig. 6A and B the label “post-mitosis” with “G1.”

      Furthermore, we cannot deduce from fixed cells whether, or to what extent, CK1δ is degraded upon mitotic exit. This was neither the intention nor the conclusion drawn from Fig. 6. Rather, we show in Fig. 3 that phosphorylated CK1δ is not degraded, and in Fig. 5 that CK1δ is predominantly hyperphosphorylated during mitosis, leading us to conclude that this phosphorylated pool of CK1δ is stable. In G1, CK1δ is dephosphorylated and unassembled kinase is degraded. We currently have no data regarding the kinetics of CK1δ dephosphorylation, assembly with centrosomal/Golgi structures, versus degradation of unassembled dephosphorylated CK1δ.

      The data shown in Fig. 6 serve merely to illustrate the subcellular distribution of CK1δ, which is consistent with previous reports. In Fig. 7, we present a model that attempts to integrate the new findings reported here together with the data from our recent eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation. Of course, we do not claim that this model does by no means represents a final verdict, and many important questions remain open. However, we believe that the model provides plausible novel concepts and mechanistic ideas that have not been proposed in previous publications and therefore merit publication, as they provide a basis for further investigation and discussion.

      Live-cell imaging:

      Live-cell imaging of fluorescently tagged CK1δ throughout mitotic progression and mitotic exit could, in principle, provide additional insight, but such experiments are technically extremely challenging. Moreover, the central idea of our model is that as much CK1δ as possible is preserved throughout the cell cycle, whereas degradation selectively targets unassembled and potentially harmful kinase. When CK1δ is expressed at physiological levels, which would require tagging the endogenous locus, the fraction of kinase degraded upon mitotic exit is expected to be very small and therefore likely below the threshold for reliable quantification by fluorescence microscopy. Similarly, although overexpressed CK1δ undergoes substantial degradation in G1, the kinase is simultaneously synthesized at a high rate. Hence, quantitative interpretation of overexpressed CK1δ levels during mitotic exit by microscopy (without CHX) would still be difficult.

      Second, the images of interphase cells constantly show multiple dots (probably surrounding the centrosome), which is a pattern that likely corresponds to Golgi rather than a single centrosome.

      The reviewer is correct. Indeed, CK1δ localization to the Golgi apparatus is well established in the literature. In our original wording, we did not explicitly distinguish between Golgi and centrosomal localization, which may have been somewhat misleading. In the revised version, we therefore refer more cautiously to the “pericentrosomal region” rather than strictly to the centrosome.

      The data shown in Fig. 6 serve to illustrate the known subcellular distribution of CK1δ. In the model presented in Fig. 7, we attempt to integrate the new findings reported here together with the data from our eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation.

      It is unclear to which figure points the paragraph "Tail phosphorylation protects CK1δ/ε from degradation". I assume that one figure is missing.

      Figs. 3 and 4 were accidentally swapped, and we apologize for this error. The paragraph in question refers to Fig. 4, which shows the cycloheximide-induced degradation kinetics of CK1δ/ε.

      Minor points:

      CK1 kinase inhibitor PF670462 should not be named as PF670 as this causes confusion. Authors should either use the full name of the compound or just call it as CK1 inhibitor with providing details in the methods.

      PF670 has been changed to PF670462.

      Fig. 3B is discrepant with the figure legend. Figure shows CK1e but legend says kinase dead CK1D-K38R

      The captions to Figs. 3 and 4, as well as the references to these figures in the text, are correct. However, the actual Figs. 3 and 4 were inadvertently swapped during figure assembly. We apologize for this mix-up.

      The authors` interpretation of CK1 involvement in checkpoint is incorrect. The authors state that CK1 activity decreases p53 function promoting recovery, but Inuzuka et al (ref. 51) showed that inhibition of CK1 leads to this outcome.

      We thank the Reviewer for noting this mistake. We are no experts in p53 regulation, which is rather complex. CK1 decreases MDM2 stability and hence enhances p53 function.

      We corrected the statement and placed it in the right context: “CK1 phosphorylation triggers β-TrCP-mediated degradation of MDM2 and activates p53, thereby enhancing p53-dependent responses involved in checkpoint signaling and DNA repair (Inuzuka et al., 2010; Winter et al., 2004). After DNA repair, CK1δ has…”

      Reviewer #2 (Significance):

      It is generally assumed that CK1 is constitutively active, which is likely an oversimplified view; in a physiological context, some degree of regulation can be expected. Demonstrating that there are several pools of CK1 that are differently regulated at the level of protein stability during the cell cycle would be a significant advance in our understanding of CK1 functions.

      Reviewer #3 (Evidence, reproducibility and clarity):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U2OS cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin), and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      Overall, they propose that the activity and abundance of CK1 are regulated during the cell cycle. However, this claim would require several experiments to support it.

      Major comments:

      - As presented, some of the data are inconclusive. Co-staining with a centrosomal marker is required to determine whether or not Ck1 localises to the centrosomes. A large proportion of cells exhibit "potential" (their term) centrosomal staining, so a centrosomal marker is essential before any conclusions can really be drawn.

      In our revised Figure 1, we confirm this centrosomal staining using an antibody against pericentrin (PCNT).

      - Figures 3 and 4 have no loading controls, and these two figures have been mixed up in the text.

      The reviewer is right, we have corrected the Fig. 3 and Fig. 4 mix-up.

      Loading controls are now also provided.

      - A mobility shift on SDS-PAGE does not prove that a protein is phosphorylated. The authors should provide experimental evidence that the mobility shift is really due to phosphorylation. As they are inactivating phosphatases using CalA, it is likely the case, but they should prove it. Furthermore, the authors did not map any phosphorylation sites in this study, so they do not know whether CK1 phosphorylation occurs in the tail (as they assert) or elsewhere.

      We provide data in Fig. 3B showing that a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines does not undergo an electrophoretic mobility shift upon CalA treatment. These results demonstrate that the CalA-induced mobility shift is caused by phosphorylation of the CK1δ C-terminal tail.

      - The figure legends in general are limited and lack crucial information. For instance, in Figures 3C and 3D, how was the half-life of CK1 determined?

      We have adapted the figure caption:

      (C) Densitometric quantification of n=3 Western blots (see A) shown as mean ± SD. Overexpressed unphosphorylated CK1δ is degraded with a half-life of about 15 min. Both CalA and CalA + PF670462 treatments stabilize the kinase. (D) Densitometric quantification of n=3 Western blots (see B) shown as mean ± SD.

      - The CK1 regulatory model presented in Figure 7 is not supported by the data. What experimental evidence, for instance, shows that CK1 is inactive during mitosis? To make this claim the authors should directly assay its activity.

      We show that the majority of CK1δ is phosphorylated and therefore auto-inhibited during mitosis. In the original version, we referred to this fraction as inactive. In the revised manuscript, we refer to the phosphorylated kinase as auto-inhibited.

      Reviewer #3 (Significance):

      This study may be of interest to researchers working on cell cycle regulation.

    1. eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are compelling, and supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

    3. Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are largely supported by strong genetic and biochemical evidence, although some claims regarding the dispensability of fatty acid synthesis pathways remain incomplete. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We kindly ask that the editorial assessment be revised to focus on the major conclusions. Alternatively, we would respectfully suggest revising the second sentence to something akin to “The main conclusions are largely supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve.”

      Public Reviews:

      Reviewer #1 (Public review):

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

      We thank the reviewer for these positive comments.

      Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      To clarify, we conclude that the essential function of ACP in blood-stage P. falciparum parasites includes a critical stabilizing interaction with pyruvate kinase II. Apicoplast ACP presumably still plays a central biochemical role in FASII pathway function, but that role in FASII is dispensable for blood-stage parasites.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages. However, there are currently some important experimental issues/flaws, missing experiments that induced wrong interpretations and thus do not support some important claims of the study, notably for the role of FASII and the interaction between ACP and PKII.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We elaborate on these points and address the reviewer’s critiques below.

      Therefore, at this point, the study is only partial and would require major additions and/or important text edits/revisions before being considered for acceptance.

      We note that the manuscript has already been accepted for publication in accordance with the current eLife publishing model.

      Major points:

      From the graph of P. falciparum growth, we can see that in the lipid-rich condition, where both FabH KO and ACP KO can survive, the addition of mevalonate was essential for the growth of ACP KO. Along with the other evidence (PKII association, DNA levels...), we therefore agree that PfACP is involved in the mevalonate pathway.

      To clarify, our model is that ACP supports IPP synthesis by the apicoplast nonmevalonate/MEP pathway indirectly by stabilizing and thus supporting function by pyruvate kinase II that supplies the pyruvate and NTPs required for IPP synthesis by the MEP pathway.

      The authors claim that the FASII pathway is inactive/not essential in the P. falciparum blood stage. However, the authors have not shown any evidence on whether ACP is or not involved in the FASII pathway during the asexual blood stage.

      To clarify, there is overwhelming data in the prior published literature that we cite (including refs. 13, 14, and 32) to establish that FASII is dispensable for blood-stage Plasmodium growth in vivo in rodent parasites and in vitro culture in human parasites. Prior studies also strongly support a role for apicoplast ACP as the central scaffold for FASII-mediated acyl chain synthesis. However, our and prior studies support the conclusion that essential ACP function in blood-stage parasites is independent of its role in FASII.

      As currently designed, the experiments presented cannot conclude on that point for several reasons. Indeed, it was previously shown that (i) the expression of the protein from the FASII pathway are all present in blood stages and are significantly upregulated in patients that are under under "nutrient starvation" (Daily et al. Nature 2007), (ii) that, growing parasites under similar low lipid conditions in vitro induces an activation/upregulation of FASII, which can be measured by stable isotope precursor labelling and lipidomics (Botté et al. 2013).

      We are aware of these prior studies and cite and discuss the Botté et al. 2013 reference in our manuscript, which provided isotope-labeling evidence to support FASII activity in low-lipid growth conditions for P. falciparum. We note that neither study addresses whether FASII activity is required for growth in low-lipid conditions.

      (iii) that growing the PfFabI KO line under deprived lipid conditions leads to parasite death (Amiar et al. 2020), indicating that the FASII pathway can become critical, if not essential, depending on the host nutritionnal content together correlating patients' data and metabolic adaptation for the same reasons in the related parastie Toxoplasma gondii (Amiar et al. 2020, Krishnan et al. 2020, Liang et al. 2020, Primo et al. 2021, Charital et al. 2024, Dass et al. 2024, Bitew et al. 2025).

      All of the studies cited by the reviewer focus primarily or exclusively on Toxoplasma gondii parasites. We agree with the reviewer that these and other studies provide strong evidence that FASII activity contributes to growth of T. gondii parasites, including roles for apicoplast ACP that appear to differ from what we have unveiled for P. falciparum malaria parasites. We acknowledge and discuss these differences from T. gondii in the final section of the Discussion section and think that exploring these differences will be a fascinating area for future study.

      The Amiar et al. 2020 paper cited by the reviewer is the only study we are aware of that has directly tested the ability of a ∆FASII parasite (in this case, ∆FabI) to grow in low-lipid conditions. We acknowledge that they observed little to no growth of ∆FabI parasites in these conditions. Our growth assays with ∆ACP and ∆FabD parasites indicated a different outcome in which both WT and ∆FASII parasites grew similarly in low-lipid conditions. Our results thus contrast with the prior study. As noted below, the minimal lipid growth conditions explicitly reported in the methods section of the Amiar et al. paper are identical to those used in our study: fatty acid-free BSA, 30 µM palmitic acid, and 45 µM oleic acid (all sourced from Sigma) with daily media changes. Thus, the basis for these differences is unclear and additional follow-up work will be needed to explore and resolve these differences.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

      Here, the authors are expecting to show that FabH (and thus the FASII pathway) is not essential in an experiment that is not designed to be in low lipid conditions but rather in lipid rich conditions: Such high lipid conditions of culture in this study is granted by daily feedings with high fatty acid supplement (30-90 uM palmitic acid and 30-60 uM oleic acid). These fatty acid concentrations were used previously by Mitamura et al. (2005) and Miichi et al.(2007) to replace non-determined supplements such as Serum or Albumax supplement to grant similar growth by a completely controlled culture medium.

      This means the concentrations above do not represent limited fatty acid concentrations, especially not with daily feeding (representing an excess supplied amount of lipids, unlike regular 48h feedings) that allowed the authors to easily reach very high non-physiological parasitaemia of more than 20%!! Amiar et al. previously showed essentiality of FabI in P. falciparum in the limited fatty acid culture at a lower concentration (<30uM 16:0, <45um 18:1), than the Mi-Ichi et al. controlled medium with regular 48 h culture feeding. Therefore, with the current experimental settings, the FAH KO is placed in high lipid conditions, thus preventing any conclusion on its essentiality under low lipid conditions.

      The basis for the reviewer’s statements here is unclear, as this critique and the conditions it describes do not conform to the published conditions reported in the Amiar et al. 2020 paper. The methods section of that study for “Plasmodium falciparum growth assays” explicitly states (page e7):

      “Media was replaced daily, sub-culturing were performed every 48 h when required, and parasitemia monitored by Giemsa-stained blood smears. Growth assays in lipid-depleted media were performed by synchronizing parasites before transferring trophozoites to lipid-depleted media as previously reported (Botte ´ et al., 2013; Shears et al., 2017). Briefly, lipid-rich AlbuMAX II was replaced by complementing culture media with an equivalent amount of fatty acid-free bovine serum albumin (Sigma), 30 µM palmitic acid (C16:0; Sigma) and 45 µM oleic acid (C18:1; Sigma).”

      We used identical culture conditions to those described above: fatty acid-free BSA in place of lipid-rich AlbuMAX, 30 µM palmitic acid, and 45 µM oleic acid. We thus obtained growth results that contrast with the prior study and suggest that additional, future studies will be required to understand and resolve these differences.

      Furthermore, it is too uncertain to conclude that ACP is only essential for the mevalonate pathway.

      Please see our response above that clarifies our model for ACP function in supporting pyruvate kinase II and the many apicoplast pathways that appear to depend on PKII.

      This would be a similar discussion to the Yeh et al. 2011 and the Swift et al., where induced Apicoplast knockout caused parasites to require IPP to survive, but there were always remnant apicoplast vesicles and thus the putative presence of an active FASII in the parasite, where de novo fatty acid synthesis could be maintained.

      It is extremely unlikely that FASII remains active upon apicoplast disruption and loss of the apicoplast genome. The apicoplast-encoded SufB is lost upon apicoplast disruption and can no longer participate in making Fe-S clusters. Without Fe-S synthesis, the apicoplast lipoate synthase (LipA) cannot make lipoate to activate pyruvate dehydrogenase (E2 subunit) and produce the acetyl-CoA needed for FASII activity. There is no experimental evidence that FASII remains active upon apicoplast disruption and loss of the apicoplast genome.

      Amiar et al. (2020) and Krishnan et al. (2020) showed that disruption of FASII and absence of de novo FA synthesis in T. gondii could be compensated by the exogenous supplementation of myristic acid, C14:0.

      As explained above, we acknowledge that FASII contributes to Toxoplasma gondii growth and includes functions that appear to differ from P. falciparum.

      Here, high fatty acid supplementation using commercially available fatty acids may include unexpected fatty acid species such as myristic acid in palmitic acid or oleic acid, since all commercially available fatty acids guarantee only >99% but not 100%. If P. falciparum requires a very, very low amount of myristic acid to survive, the amount of possible contamination, like 1 nM, may be sufficient to maintain their survival. Thus, ACP and FabH might be very important to generate de novo fatty acids within parasites, but this was not shown by the authors.

      As noted above, we used identical culture conditions and commercial sources of defined fatty acids to those reported in the Amiar et al. study. We do not see a basis for the reviewer’s critique that the two studies utilized differing culture conditions. Nevertheless, we agree that future studies are needed to understand and resolve these differences.

      Therefore, the manuscript currently contains incorrect conclusions on the potential essentiality/use of FASII, against current experimental evidence.

      As explained above, we do not see a basis for the reviewer’s critique here or for viewing one study as more or less definitive than the other, as identical culture conditions were used yet contrasting results were obtained for reasons that remain uncertain. Future studies beyond the scope of the present manuscript will be required to fully understand and resolve these differences.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      We either request more solid experimental evidence showing the absence of fatty acid synthesis at low fatty acid conditions by re-doing the growth assay in the lower fatty acid feeding conditions without daily feeding to clarify if the ACP and FabH are essential in the blood stage growth, or not; as well as showing the absence of fatty acid synthesis at low fatty acid conditions using isotope labelled precursor. Without these, the authors cannot conclude on this important point. Alternatively, toning down the text to acknowledge the possibility of FASII being active and critical under certain conditions would be acceptable.

      We note that the Amiar et. al 2020 study cited by the reviewer reported growth assays for WT and ∆FabI NF54 P. falciparum in low-lipid conditions that similarly lacked direct tests of FASII activity by isotope labeling.

      We agree that our results, which utilized distinct ∆ACP and ∆FabD NF54 (PfMev) lines, contrast with the results and conclusions of the Amiar et al. study. We fully agree with the reviewer that future studies, utilizing tandem growth assays and isotope-labeling metabolic flux assays (e.g., mass spectrometry), will be required to fully test and understand the dependence of FASII activity on the lipid content of the growth medium and the functional dependence of P. falciparum growth on FASII activity in low-lipid conditions.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

    1. eLife Assessment

      This important work identifies phlda2 as a specific marker for primordial cardiomyocytes in the adult zebrafish heart and demonstrates their essential role in myocardial morphogenesis and coronary vascularization, but not in heart regeneration. The conclusions are well supported by single-cell transcriptomics, new genetic tools, and cell-specific ablation experiments. The revised version strengthens these findings with analyses of the epicardium, long-term time points showing that the developmental defects persist, and a control confirming that loss of primordial cardiomyocytes alone does not trigger a regenerative response. Overall, the evidence is solid and provides insight into the difference between developmental and regenerative cardiac programs, and this work will be of interest for those studying cardiac development and regeneration.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the well-established role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4⁺ cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive? Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days post-treatment when they were detected?

      Strengths:

      The major strengths are the identification of a marker that enables manipulation of primordial cardiomyocytes and the tools generated by the team.

      Weaknesses:

      The major weakness is not considering the longer-term consequences of primordial layer ablation during development, as it is unclear whether the animals succumb to the acute cardiac defects observed or fully recover.

    4. Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the wellestablished role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      We thank the reviewer for this important suggestion. To address this possibility, we examined epicardial cells in primordial CM-ablated hearts using the epicardial marker tcf21. Surprisingly, we found that epicardial cells rapidly expanded following primordial CM ablation, with an increase already detected at 5 days post-treatment. This increase persisted through at least 30 dpt, when epicardial cell abundance remained higher than that in control hearts. These findings suggest that the coronary vessel defects are unlikely to result from a reduction in epicardial cell number. In addition, we investigated whether loss of the primordial CM layer permits abnormal inward migration of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation. While we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype, our results indicate that the vascular defects are not attributable to reduced epicardial cell abundance or abnormal epicardial invasion. We have included these experiments in Page 10 and 11, and Fig. 3I-3J of revised manuscript.

      “Because epicardial cells play critical roles in coronary vessel development, we next asked whether the vascular defects observed following primordial CM ablation were secondary to alterations in the epicardium. We found that tcf21+ epicardial cells revealed an increase rather than a decrease in epicardial cell abundance in primordial CM-ablated hearts (Fig. 3I3J). These findings suggest that the impaired coronary vessel organization is unlikely to result from a reduction in epicardial cell number, although we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype.”

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      We appreciate the reviewer for this insightful suggestion. Because primordial cardiomyocytes form a continuous single-cell-thick layer at the ventricular surface, we examined whether their ablation affects the spatial distribution or promotes inward migration of epicardial cells or coronary endothelial cells. Using tcf21 and deltaC reporters, we carefully analyzed the localization of these cell populations within the myocardium. We did not observe any evidence of abnormal inward migration or ectopic localization of either epicardial cells or coronary endothelial cells following primordial CM ablation (Fig. 3J–3K). These results indicate that loss of the primordial CM layer does not disrupt tissue compartmentalization or lead to inappropriate cellular invasion into the myocardial interior. We have included these experiments in Page 11, and Fig. 3J-3K of revised manuscript.

      “In addition, because primordial cardiomyocytes form a continuous layer at the ventricular surface, we investigated whether their ablation permits abnormal invasion of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation (Fig. 3J-3K). Thus, loss of the primordial CM layer does not appear to disrupt tissue compartmentalization or permit abnormal cellular invasion into the myocardium.”

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4<sup>+</sup> cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      We thank the reviewer for this important suggestion. To further address the relationship between primordial cardiomyocytes and gata4<sup>+</sup> cardiomyocytes during heart development, we examined their spatial and cellular relationship in juvenile zebrafish hearts. Consistent with our observations during regeneration, we did not detect any overlap between phlda2<sup>+</sup> cardiomyocytes and gata4<sup>+</sup> cardiomyocytes in the juvenile heart (7-8 wpf). These results indicate that primordial CMs and gata4<sup>+</sup> proliferative CMs represent distinct cardiomyocyte populations during both heart development and regeneration. We have included these experiments in Page 13, and Fig.5E of revised manuscript.

      “Consistently, no overlap between phlda2<sup>+</sup> and gata4<sup>+</sup> cardiomyocytes was observed in the wpf juvenile zebrafish heart (Fig. 5E).”

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

      We appreciate the reviewer for this important suggestion. To determine whether ablation of primordial cardiomyocytes alone is sufficient to activate regenerative signaling, we treated adult phlda2:mCherry-NTR;gata4:EGFP fish and gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. Under these conditions, we did not observe any induction of gata4 expression following primordial CM ablation (Fig. S5). These results indicate that loss of phlda2<sup>+</sup> cardiomyocytes alone is not sufficient to trigger regenerative gata4 activation in the absence of injury, further supporting that primordial CMs are dispensable for activation of the regenerative response. We have included these experiments in Page 11 and Fig. S5 of revised manuscript.

      “To determine whether ablation of primordial cardiomyocytes is sufficient to activate regenerative signaling, we first treated adult phlda2:mCherry-NTR;gata4:EGFP fish and control gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. We did not observe any induction of gata4:EGFP expression following primordial CM ablation (Fig. S5), indicating that loss of phlda2<sup>+</sup> cardiomyocytes is not sufficient to trigger regenerative gata4 activation in the absence of injury.”

      Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive?

      We thank the reviewer for this important question. We found that efficient ablation of phlda2<sup>+</sup> primordial cardiomyocytes does not affect overall survival of zebrafish under standard laboratory conditions. Although these animals exhibit clear defects in cardiac structure and coronary vascular organization, they remain viable during the experimental period, indicating that the primordial CM layer is not essential for survival. While we did not assess detailed physiological parameters such as cardiac function, swimming behavior, or long-term fitness in this study, the observed structural abnormalities suggest that subtle functional consequences may exist. These aspects will be important directions for future investigation. We have included the description on page 14, paragraph 2.

      “Although primordial CM ablation does not affect survival under laboratory conditions, the observed structural defects may have functional consequences on cardiac performance and overall physiological fitness, which warrant further investigation.”

      Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days posttreatment when they were detected?

      We appreciate the reviewer for this important question. To determine whether primordial cardiomyocytes regenerate following ablation during development, we performed Mtz-mediated ablation in juvenile zebrafish (7–8 wpf) and examined the hearts at extended time points after treatment. We found that phlda2<sup>+</sup> cardiomyocytes did not recover even at 90 days post-treatment, indicating a persistent loss of this population following developmental-stage ablation (Fig. S6). Importantly, we further assessed whether the cardiac defects observed at earlier time points resolve over time. We found that the abnormalities in trabecular and compact myocardium, as well as coronary vessel organization, persisted at 90 days post-treatment and did not show evidence of recovery (Fig. S4C). These findings demonstrate that the observed defects are not transient developmental delays but represent long-lasting structural alterations of the heart following primordial CM ablation.

      Major Comments:

      (1) Figure 1: A more detailed characterization of the three CM populations would be helpful in the text as well as a new Supplemental Excel Data Sheet with the top unique genes expressed in each.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have expanded the description of the three cardiomyocyte populations in the Results section to provide a more detailed characterization. We have revised the description on page 8. In addition, we have added a new Supplemental Excel Data Sheet (Table S1) listing the top differentially expressed genes for each CM cluster.

      “Notably, Cluster 2 showed additional enrichment for pathways involved in mitochondrial respiratory chain assembly, ATP synthesis, TCA cycle, and ribosome biogenesis, suggesting a relatively higher metabolic and biosynthetic activity state. In contrast, Cluster 1 was enriched for GO terms associated with mitochondrial stress responses, protein degradation, and cytoprotective pathways, indicating a stress-adapted cardiomyocyte state. Cluster 3 showed reduced enrichment of metabolic pathways, consistent with an immature metabolic profile, and was further enriched for genes involved in muscle development and epithelial morphogenesis, suggesting a role in cardiac morphogenesis and tissue organization.”

      (2) Figure 3: Lower magnification views of control and MTZ-treated hearts are needed for all reporters shown (cmlc2, gata4, deltaC) at 10 days post-treatment. These "whole heart" views will enable the reader to get a gross sense of how disrupted heart development is following ablation of the primordial layer.

      We appreciate the reviewer for this helpful suggestion. We have added low-magnification whole-heart images for all reported markers (cmlc2, gata4, and deltaC) to better illustrate the overall cardiac morphology following ablation of the primordial layer. We have included these data in Fig. S3 of revised manuscript.

      (3) It should also be stated clearly in the Results section text what day post-fertilization Mtx treatment began. A section on Mtx treatment, including timing (what day was it applied and what day was it washed out) and dose, should be added to the methods section.

      We thank the reviewer for this helpful suggestion. In response, we have now clearly stated the timing of Mtz treatment in the Results section in page 9, 10 and 11 of revised manuscript.

      “To address the role of phlda2<sup>+</sup> cells during heart development, we performed the following experiments using a standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to 7-8 wpf juvenile zebrafish.”

      “To address the role of primordial cells during heart regeneration, we perform the below experiments using standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to adult zebrafish (4–6 months post-fertilization).”

      In addition, we have added a Mtz treatment section in the Methods, which now includes detailed information on dosage, duration, and washout schedule to ensure full reproducibility of the experiments.

      “Mtz treatment

      For conditional ablation of phlda2<sup>+</sup> cardiomyocytes, zebrafish expressing phlda2:mCherryNTR were treated with 10 mM metronidazole (Mtz) for 12 hours per day for three consecutive days. Fish were washed out and maintained in fresh system water after each daily treatment. For developmental analyses, juvenile zebrafish (7–8 weeks post-fertilization) were subjected to Mtz treatment as described above. Following completion of Mtz exposure, fish were maintained under standard conditions. Hearts were collected at 10 days post-treatment for assessment of gata4 activation, and at 30 days post-treatment for analysis of myocardial structure and coronary vessel development. For regeneration experiments, adult zebrafish (4–6 months old) were similarly treated with Mtz for three consecutive days with daily washout. Three days after the final Mtz treatment, ventricular apex resection was performed. Hearts were harvested at 7 days post-amputation (dpa) for analysis of gata4 activation and early regenerative responses, and at 30 dpa for evaluation of myocardial regeneration and coronary vessel revascularization.”

      (4) What happens to the heart 2 and 6 months post-treatment? Are there long-term consequences to primordial layer ablation or do the defects seen at 10 days post-treatment eventually resolve?

      We thank the reviewer for this important question. To determine whether the developmental defects observed following primordial CM ablation are transient or persist long term, we performed additional analyses at later time points after Mtz washout. First, we found that phlda2<sup>+</sup> cardiomyocytes failed to recover following Mtz-mediated ablation. In adult phlda2:mCherry-NTR fish, phlda2<sup>+</sup> cells remained absent 30 days after Mtz washout (Fig. 5B). Similarly, when juvenile fish (7–8 wpf) were treated with Mtz and subsequently allowed to recover, phlda2<sup>+</sup> cells were still not restored at 90 days post-treatment (Fig. S6). Second, juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated primordial CM ablation and analyzed 90 days after Mtz washout. We found that vascular abnormalities persisted long after ablation of primordial CMs (Fig. S4C). Coronary vessels remained disorganized and fragmented, indicating that the vascular phenotype does not resolve over time. Together, these findings demonstrate that primordial CM ablation causes long-lasting defects and that the abnormalities observed are not transient developmental delays. Instead, loss of primordial CMs results in persistent cellular and vascular defects that remain evident months after Mtz treatment. We have included these experiments in Fig. S4C and Fig. S6 of revised manuscript.

      “Notably, these vascular abnormalities persisted at 90 days post-treatment, indicating that the defects do not resolve during subsequent cardiac growth and maturation (Fig. S4C).”

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (5) Also, does the primordial layer come back in these animals where the lineage is ablated during development? Or is it permanently lost as shown in Figure 5 when it is ablated during adulthood?

      We appreciate the reviewer for raising this important question. To determine whether the primordial layer can be reestablished following ablation during development, we treated juvenile phlda2:mCherry-NTR fish (7–8 wpf) with Mtz and examined hearts 90 days after treatment. We found that phlda2<sup>+</sup> cardiomyocytes remained absent at this late time point (Fig. S6), indicating that the primordial layer does not recover following developmental-stage ablation. We have included these experiments in Page 12, and Fig. S6 of revised manuscript.

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (6) Figure 4: Need to show that phlda2 reporter fluorescence in lost/reduced following Mtz treatment during adulthood before apex amputation.

      We thank the reviewer for this important suggestion. We agree that confirming efficient ablation of phlda2<sup>+</sup> cardiomyocytes in adult fish prior to regeneration analysis is essential.

      In our study, we have already demonstrated in Fig.5B that Mtz treatment in adult phlda2:mCherry-NTR fish results in efficient and sustained loss of phlda2<sup>+</sup> cells, with no detectable recovery at 7 and 30 days post-treatment. These data confirm robust ablation of the primordial CM population following Mtz treatment in adults. Therefore, additional redundant imaging prior to apex resection was not performed in Fig 4.

      (7) Figure 5: It is interesting that primordial CMs do not regenerate following apex amputation or genetic ablation. This result suggests that primordial CMs are only important during development and dispensable during adulthood? This result also makes me question whether primordial CMs are actually required for heart development, which is why it is important to address whether the fish recovers.

      We appreciate the reviewer for this insightful comment. Our data indicate that primordial cardiomyocytes are essential for proper heart development, as their ablation during juvenile stages leads to significant structural and vascular abnormalities. Importantly, we further examined long-term outcomes and found that these defects do not resolve over time. Juvenile zebrafish subjected to primordial CM ablation failed to recover phlda2<sup>+</sup> cardiomyocytes even at 90 days post-treatment, and coronary vascular abnormalities also persisted at this late stage (Fig. S4C). These findings indicate that the observed developmental defects are not transient delays but instead reflect long-lasting structural alterations of the heart. In addition, we found that adult zebrafish similarly fail to regenerate primordial cardiomyocytes following genetic ablation (Fig. 5B), further supporting the limited regenerative capacity of this population. Together, these data demonstrate that primordial cardiomyocytes are required for proper cardiac development, and their loss leads to persistent defects that are not reversed during subsequent growth or regeneration.

      Minor:

      Line 234: Did the authors mean to write Cluster 3 (instead of Cluster 2)?

      We thank the reviewer for pointing out this error. We confirm that this was a labeling mistake, and “Cluster 3” is correct. The text has been corrected in the revised manuscript.

      Line 265: There is a typo of some sort in the phrase, "54.7% reduction closed to the ventricular wall".

      We thank the reviewer for pointing out this error. We have changed the description on page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      Is there a corollary lineage in mammals? This should be addressed in the Introduction or Discussion.

      We thank the reviewer for this insightful suggestion. At present, a direct corollary lineage to zebrafish phlda2<sup>+</sup> primordial cardiomyocytes have not been clearly defined in mammals. However, mammalian hearts also contain heterogeneous cardiomyocyte populations with distinct developmental states and metabolic profiles, including immature or embryonic-like cardiomyocytes that persist in specific regions during development and early postnatal stages. These populations may share functional similarities with the zebrafish primordial CMs in terms of developmental organization and maturation roles. We have now discussed this point in the Discussion and emphasized that whether a comparable lineage exists in mammals remains an important open question for future studies. We have included the discussion on page 15 of the revised manuscript.

      “Although a direct corollary of phlda2<sup>+</sup> primordial cardiomyocytes has not yet been identified in mammals, mammalian hearts contain heterogeneous cardiomyocyte populations with immature states. Whether these populations represent a functional equivalent of zebrafish primordial CMs remains an open question and needs further investigation.”

      Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

      We appreciate the reviewer for this important suggestion. In the revised manuscript, we have added supplemental data to support the scRNA-seq analysis, including gene expression tables for each cardiomyocyte cluster (Table S1), and full GO-term enrichment results (Table S2). In addition, we have revised the Results section to improve the clarity and description of the scRNA-seq findings. Please see detailed responses below for point-by-point clarification.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors did not provide enough data to demonstrate the quality of their scRNAseq. They only mentioned that they obtained "high-quality" transcriptomics. Specific parameters such as how many total reads and reads per cell should be provided.

      We thank the reviewer for this important suggestion. In the revised manuscript, we have added detailed sequencing quality metrics to the Methods section in Page 6 of the revised manuscript. The dataset contains 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.

      “The newly generated scRNA-seq data yielded 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.”

      (2) The authors utilized cmlc2:EGFP fish to perform scRNASeq. It will be helpful to include feature plots of cmlc2 and EGFP transcripts.

      We thank the reviewer for this suggestion. We have now included the description in Page 8, and feature plots of cmlc2 transcripts and EGFP reporter expression in the scRNA-seq dataset as a supplementary figure (Fig. S1A and S1B).

      “The expression of cmlc2 transcripts and EGFP reporter signal in the single-cell dataset further confirmed the enrichment of cardiomyocytes (Fig. S1A and S1B)”

      (3) The authors show that notch 3 is in cluster 3 of cardiomyocytes and suggest that this reflects elevated NOTCH signaling activity. The authors might consider using RNAScope to further validate that Notch 3 is expressed in cardiomyocytes. It will be also helpful to confirm phlda2 expression patterns during zebrafish heart development and regeneration by RNAScope.

      We appreciate the reviewer for this helpful suggestion. To further examine the spatial expression patterns of these genes, we analyzed publicly available spatial transcriptomic data from zebrafish hearts. We found that phlda2 is enriched in the outer region of the heart during both uninjured and regenerating conditions. Similarly, notch3 and actn1 were also predominantly localized to the outer heart region in the uninjured heart. These findings are consistent with our scRNA-seq results and support the spatially restricted signature of Cluster 3 cardiomyocytes. We have included these data in Page 8, and Fig. S2A-S2C of revised manuscript.

      “To further validate their spatial distribution, analysis of previously published spatial transcriptomic data revealed that phlda2, notch3, and actn1 were predominantly expressed in the outer region of the heart (Fig.S2A-S2C).”

      (4) The authors did not provide any data as supplemental tables to support their analyses of scRNAseq and GO-term analysis.

      We thank the reviewer for this suggestion. In the revised manuscript, we have added new supplemental tables providing full support for the scRNA-seq and GO-term analyses (Table S1 and S2), including lists of differentially expressed genes for each cardiomyocyte cluster and the corresponding GO enrichment results.

      (5) The description of the phenotype in Fig. 3A and B is not clear, especially for the sentence "trabecular muscle formation was severely impaired showing an approximately 54.7% reduction close to the ventricular wall.". The authors might consider using a bracket to show the compact muscle and trabecular muscle and the distance to the ventricular wall.

      We appreciate the reviewer for this suggestion. We have added brackets in Fig. 3A to label the compact and trabecular myocardium for improved clarity. Regarding “distance to the ventricular wall,” we found that this measurement varies substantially across different regions within the same heart, making a single distance-based metric unreliable. Therefore, we quantified trabecular muscle using the trabecular area fraction (trabecular area/total ventricular area) within a defined region of interest as a robust and unbiased indicator. The reported 54.7% reduction refers to this area fraction, and we have revised the text accordingly for clarity in Page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      (6) It is not clear how the authors quantify the vessel "length"/ventricular area and found that there is no difference (Fig. 3G). The vessel length is significantly shorter in the images (Fig. 3F) as the author also indicated that the vessels are fragmented.

      We thank the reviewer for this important comment. Coronary vessel “length” was quantified by selecting a fixed region of interest (ROI) within the ventricular area, followed by skeletonization of deltaC:EGFP<sup>+</sup> vessels using ImageJ. The total vessel length within the ROI was measured and normalized to the ROI area to obtain vessel length density. Although the representative images (Fig. 3F) show a more fragmented vascular pattern, the total summed vessel length within the defined region was not reduced. This indicates that primordial CM ablation primarily affects vascular organization rather than overall vessel length within the ventricular area. We have updated the figure legend for clarity of the revised manuscript

      “Vessel length was measured within a fixed region of interest (ROI) after skeletonization of deltaC:EGFP<sup>+</sup> vessels in ImageJ.”

    1. eLife Assessment

      This important study identifies PRRT2 as an auxiliary regulator of Nav channel slow inactivation in vitro and in vivo, demonstrating that PRRT2 facilitates entry into, and delays recovery from, the slow-inactivated state. The revised manuscript has been substantially strengthened, providing compelling evidence that PRRT2 is relevant to normal brain physiology and disease pathophysiology, providing a mechanistic link between PRRT2 mutations and episodic neurological phenotypes. Overall, this study will be of interest to ion channel biophysicists and neurophysiologists, particularly those studying channelopathies.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing and Senior Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The manuscript by Lu and colleagues demonstrate convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      Comments on revised version.

      The manuscript by Lu and colleagues has been revised sufficiently to address all my prior concerns.

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

    3. Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".<br /> PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last (20th) compound APs in panels B and C.

    4. Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      (4) The mechanistic separation between trafficking of PRRT2 and its gating effects is not clearly resolved.

      (5) Additional studies with Nav1.6 should be carried out.

      Comments on revised version.

      These comments have been addressed in the revised version.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      We sincerely thank the reviewer for the meticulous evaluation of our revised manuscript and for the constructive and insightful comments.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      We appreciate the reviewer’s suggestion. In the revised manuscript, we have clarified that PRRT2-mediated regulation of Nav1.2, together with its regulation of Nav1.6, may contribute to neuronal and network excitability in the cortex. Please refer to Page 15, Line 439-440.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      We thank the reviewer for this comment. We have corrected the timescale description in the Discussion accordingly. Please refer to Page 13, Line 382.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".

      PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      We appreciate the reviewer’s constructive suggestion. We agree that PRRT2 is unlikely to be the sole regulator of neuronal excitability or Nav channel slow inactivation. In the revised Discussion, we have clarified that PRRT2-dependent regulation of Nav channel slow inactivation represents one mechanism among several that fine-tune neuronal excitability. Other mechanisms, including regulation of potassium channels such as Kv7.2/Kv7.3, and other ion channel- or signaling- dependent pathways, may also contribute to excitability control in both PRRT2-positive and PRRT2-negative neurons. We have revised relevant paragraph to provide a more balanced discussion of alternative mechanisms. Please refer to Page 15, Line 430-438.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last compound APs in panels B and C.

      We appreciate the reviewer’s valuable suggestion. We have now added representative traces of the first and last compound action potentials, corresponding to the 1st and the 100th responses, respectively, to panels B and C of Figure 7-figure supplement 2. Please refer to Page 51, Figure 7-figure supplement 2.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Page 19, line 546: a sampling rate of 10kHz was used for recording Nav currents. Because the Nav channel isoforms examined in this manuscript (Nav1.1, Nav1.2, Nav1.4, Nav1.5, and Nav1.6) exhibit extremely rapid activation and inactivation kinetics, the temporal resolution provided by a 10 kHz sampling rate may not be optimal for detailed kinetic analysis. Although I am not requesting additional experiments, the authors may wish to choose a higher sampling rate (e.g., 50 kHz) in future Nav channel studies.

      We thank the reviewer for this helpful suggestion regarding the sampling rate for Nav current recordings. We will adopt higher sampling rates for rapid kinetic analyses of sodium currents in our future work.

      (2) Typo: Page 24, line 712-713: should "a Digidata (Molecular Devices, 1332A)" be "... 1322A"?

      We thank the reviewer for pointing out this typo. We have corrected “Digidata 1332A” to “Digidata 1322A” in the relevant Method section of revised manuscript. Please refer to Page 25, Line 722.

    1. eLife Assessment

      The authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. This work represents an important contribution to our understanding of the global burden of pertussis. The evidence presented is compelling, with strengths including routine sampling irrespective of symptoms and rigorous qPCR methodology. The study is a significant and much-needed contribution, which sheds light on an often-overlooked dimension of pertussis transmission and opens avenues for future research and policy consideration.

    2. Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants in particular. The authors use a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka Zambia and they utilized qPCR based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high burden low resource settings and they advocated for an enhanced molecular surveillance strategies to capture full pertussis infection including those that might have gone undetected.

      Strengths:

      The strength are the use of innovative study design especially the longitudinal approach and routine sampling rather than symptom driven testing that minimizes bias in the study. The methodology were also rigorous and transparent by evaluating IS481 signal strength to classify pertussis detection and conducts retesting to assess qPCR reliability. There was also important epidemiological insights and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low resource setting.

      Weaknesses:

      These includes reliability on qPCR based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanation for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there are limited generalizability as the study was done in a single urban site in Zambia. There is also lack of functional immune data.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Comments on revised version:

      I appreciate the authors' attention to my comments during the revision process and still believe that their work represents an important contribution to our understanding of pertussis epidemiology. In most cases, the authors have done a thorough job of either addressing or responding to my comments. However, I do not believe the authors engaged sufficiently with two of the queries raised in my previous round of comments. The two queries were about the vaccination status of the mothers and engagement with literature on asymptomatic transmission. I still think they matter and ask the authors to consider them again.

      I do not think the authors can rule out two alternative explanations: (1) recent introduction of pertussis and low vaccination coverage amongst study mothers, or (2) recent introduction of a breakthrough strain (either w.r.t. the vaccine or prior infection) and higher vaccination coverage/infection-derived immunity amongst study mothers. Depending on which mechanism was mostly driving the observed patterns in Zambia, i.e.,

      a. long-running, widespread, unreported transmission;<br /> b. transmission started recently, and vaccination was low amongst study mothers;<br /> c. a breakthrough strain is causing the current rise (here we'd still want to know about vaccination status); and<br /> d. something else that I have not considered

      would have implications for how the results are interpreted, and potentially far-reaching implications for the broader pertussis community. All of that is to say, I think the authors were too quick to dismiss these concerns (even if they disagree with my assertions).

      In their reply, the authors largely dismissed concerns about not knowing the mother's vaccination status, stating in their reply that, "our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity."

      However, they also stated that, "Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009" and "As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa."

      I don't disagree with the authors' conclusion that pertussis is clearly spreading in Zambia. I also don't disagree that there's clearly evidence for minimally symptomatic, infectious mothers spreading infections to children. Both of these findings matter for Zambia and for our broader understanding of pertussis. However, I don't see how the authors can so confidently conclude that low vaccination rates, coupled with a recent introduction, high vaccination rates, coupled with a breakthrough strain, or high infection-derived immunity, coupled with a breakthrough strain, couldn't be what's driving the increase. The authors do hedge in places and also state in the discussion that their findings don't line up with expectations related to WP/infection-derived immunity, "This corresponds to a mean return frequency of one infection per 14.8 years, which is much shorter than the presumed duration of immunity from natural infection or the whole-cell pertussis vaccination used in Zambia (70, 71)." But, my read of the paper is that the authors are pushing way to ward for a preferred hypothesis that is not more favored than other alternatives.

      Secondly, I asked about placing this work in the context of other studies on asymptomatic transmission, but realize that I did not list any specific papers. Two worth considering are Warfel et al. 2014 and Althouse and Scarpino 2015. Restating for the editor, the Warfel study found that WP facilitated rapid clearance in a non-human primate experimental infection study (admittedly with small sample sizes and many other caveats). Many took that as evidence that WP would also block transmission (admittedly experiments Warfel did not run). If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research. While not incompatible with the Warfel et al. results, it would negate most of the importance of their finding that WP blocked transmission. From what I can see, the authors do not even cite Warfel et al. 2014, which is a serious gap regardless of whether the authors agree or disagree with the findings. A quick sidebar, the authors seem to duplicate Craig et al. 2020 10.1093/cid/ciz531, listing it as both citation 9 and 38.

      In Althouse and Scarpino, they found evidence of a rise in asymptomatic/underreported/subclinical transmission following the switch from WP to AP. While not as directly relevant to the current study as the Warfel paper (so I leave it to the authors to decide whether citing this paper is important), Althouse and Scarpino discuss asymptomatic transmission at length and also assume that WP conferred strong protection against transmission, so their results (along with dozens and dozens of other studies assuming similar WP/infection-induced immunity protection and durability) would also need to be reinterpreted in the context of this study. The authors should engage with the implication of their results in the context of past modeling studies and what we think we know about vaccine-/infection-derived immunity.

      Going back to my earlier points, unvaccinated mothers and the recent introduction of pertussis, or vaccinated/infection-induced immune mothers with a breakthrough strain, would both explain the current results and be compatible with Warfel et al., Althouse and Scarpino, and a sizable number of other studies. Instead, if transmission from WP- or naturally infected mothers is common (in the absence of a breakthrough strain), that would really change the landscape of pertussis epidemiology. The authors have not convinced me that they can make this conclusion. Hence, why I think it's important that the authors engage more actively with those hypotheses and with relevant debates in the pertussis literature on asymptomatic transmission. I think it's appropriate for the authors to present their preferred hypothesis, but, absent other data, they should also present plausible alternatives that are consistent with past publications.

      References:

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC medicine, 13, 1-12.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787-792.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We again thank the editor and reviewers for their detailed attention to our work. In our previous revisions we endeavored to address the principal concerns raised by reviewers that we were capable of addressing. We recognize that asymptomatic pertussis transmission represents a particularly thorny area of epidemiology and public health, where multiple (and sometimes overlapping) mechanisms have been proffered even as empirical evidence remains thin, particularly in low-resource settings such as sub-Saharan Africa. A key finding of our work is that prospective surveillance in such a low-resource setting revealed abundant evidence of otherwise unobserved asymptomatic incidence. Moreover, as we note in our prior revisions, this finding is supported by recent work in South Africa and elsewhere (Kayina et al., 2015; Moosa et al., 2019, 2025). As such, we believe that further prospective surveillance in similar settings would be highly informative, a point that we have sought to emphasize in our present manuscript (and associated commentary).

      While we broadly agree with many of the concerns raised by the reviewers, we believe that we are unable to significantly strengthen the present work through further revisions. We do, however, wish to respond to several points raised in these reviews. Of note, a reviewer raises the possibility of a "breakthrough strain" without reference to existing literature. We agree that we cannot test this hypothesis, and though it is not incompatible with our own findings, it is, however, not consistent with recent molecular surveillance in South Africa (Moosa et al., 2023). The reviewer also raises the potential of low adult vaccination coupled with recent reintroduction. This hypothesis relies on our investigation looking "at the right place and the right time", and further does not explain how immunologically naive adults would have escaped morbidity. We have adopted what we believe is a more parsimonious interpretation of our results (i.e., that asymptomatic infection represents evidence of previous immune exposure), though we agree that a more thorough exploration of this particular issue is warranted, particularly in light of our persistently colonized mothers. In addition, we have noted similar studies in sub-Saharan Africa that also found widespread evidence of asymptomatic pertussis, which we believe is inconsistent with a “right time, right place” interpretation.

      A reviewer also pointed to the work of Warfel et al. (2014) and Althouse and Scarpino (2015). We are familiar with both of these studies and agree with their broad relevance to the field (our apologies for omitting Warfel et al.). The reviewer states that, "If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research." Critically, we believe that our prospective field study of human patients in a real-world public health system complements previous research, including animal trials and simulation studies. Simply put, given that our study was unable to establish the prior vaccination or exposure status of participants, we do not claim to have shown evidence for transmission despite wP vaccination or prior infection.

      Regarding Althouse & Scarpino (2015), we believe that, for the majority of readers, the most compelling analysis in their paper was the examination of genome sequences that pointed to substantial asymptomatic transmission in the US. This conclusion emerged from their population model, which required that “births” (representing transmission events) exceeded “deaths” (representing recovery of infectious individuals) in order to be consistent with the sequence data. Unfortunately, this paper does not provide a detailed explanation of their methods and data sources, nor is this work directly reproducible through, for example, an open-access code/data repository. We have explored the availability of US genome sequences over the time period of their study and were able to find only 36 sequences: 2 from the pre-vaccine era, 8 from the wP vaccine era, and 26 from the aP vaccine era. Given this notable imbalance in the number of sequences (and thus sequence diversity) that was biased in favour of the most recent time period, is it then surprising that the “birth rate” in their model had to exceed the “death rate” in order to match the genetic diversity in the data? Based on a careful inspection of this work, we do not consider its conclusions to represent a gold standard against which all subsequent studies should be judged. We also note that genomic surveillance and analysis of pertussis remains sparse relative to other fields, though recent works have added dramatically to the corpus of available sequences (Bridel et al., 2022).

      Finally, we note that our previous revisions addressed several concerns raised in the present reviews. For example, we previously sought to address reviewers' about our presentation of the strength of our evidence. In this regard, we broadly agree with the reviewers, and we now state that our results "suggest that pertussis transmission occurs between minimally symptomatic mothers and their newborn infants." We believe this largely addresses a present reviewer's concern that 'the mother-to-infant transmission pathway should be framed as "highly suggestive" rather than "confirmed"'. We also note that our results examine three different threshold Ct values (survival analysis, Fig 4), a point that we believe partially addresses a reviewer's suggestion to "including a sensitivity analysis using a stricter cut-off" and concerns about "the decision to use a Ct<45 threshold, as this is higher than standard clinical cut-offs". Indeed, we discuss the issue of clinical cut-offs (and their appropriateness) at some length in the section, "Test reliability, disease surveillance, and public health where we state, "We recognize that such weak and potentially ambiguous signals may not be appropriate for clinical diagnosis. However, our results demonstrate that they nonetheless contain valuable information about pathogen presence and infection intensity that can (and should) be leveraged for disease surveillance." We have also included in the present work a detailed discussion of qPCR sensitivity and efficiency that we believe should interest others working in pertussis surveillance.

      We do not view our own research as the "last word" in this rather controversial subject. In this spirit, we have attempted to present our work transparently, state our claims carefully, and underscore future activities that we believe would benefit the pertussis research community going forward. For example, we agree that further attention to shared exposure and functional immune data among low-resource communities could provide valuable insights into epidemiology and ecology of pertussis. However, we also believe the trade-offs of including one set of activities over another should be clearly acknowledged by researchers, clinicians, and public health officials. To simply state that we must measure more fails to account for the very real resource constraints that we all face.

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC Medicine, 1–12. https://doi.org/10.1186/s12916-015-0382-8

      Bridel, S., Bouchez, V., Brancotte, B., Hauck, S., Armatys, N., Landier, A., Mühle, E., Guillot, S., Toubiana, J., Maiden, M. C. J., Jolley, K. A., & Brisse, S. (2022). A comprehensive resource for Bordetella genomic epidemiology and biodiversity studies. Nature Communications, 13(1), 3807. https://doi.org/10.1038/s41467-022-31517-8

      Kayina, V., Kyobe, S., Katabazi, F. A., Kigozi, E., Okee, M., Odongkara, B., Babikako, H. M., Whalen, C. C., Joloba, M. L., Musoke, P. M., & others. (2015). Pertussis prevalence and its determinants among children with persistent cough in urban Uganda. PLoS One, 10(4), e0123240.

      Moosa, F., du Plessis, M., Weigand, M. R., Peng, Y., Mogale, D., de Gouveia, L., Nunes, M. C., Madhi, S. A., Zar, H. J., Reubenson, G., & others. (2023). Genomic characterization of Bordetella pertussis in South Africa, 2015–2019. Microbial Genomics, 9(12), 001162.

      Moosa, F., du Plessis, M., Wolter, N., Carrim, M., Cohen, C., von Mollendorf, C., Walaza, S., Tempia, S., Dawood, H., Variava, E., & others. (2019). Challenges and clinical relevance of molecular detection of Bordetella pertussis in South Africa. BMC Infectious Diseases, 19, 1–11.

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018: Findings from the PHIRST study. Journal of Infection, 106550.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787–792. https://doi.org/10.1073/pnas.1314688110


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants, in particular. The authors used a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka, Zambia, and they utilized qPCR-based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants, thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high-burden low-resource settings, and they advocated for enhanced molecular surveillance strategies to capture full pertussis infection, including those that might have gone undetected.

      Strengths:

      The strengths are the use of innovative study design, especially the longitudinal approach and routine sampling, rather than symptom-driven testing that minimizes bias in the study. The methodology was also rigorous and transparent by evaluating the IS481 signal strength to classify pertussis detection and conducting retesting to assess qPCR reliability. There were also important epidemiological insights, and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low-resource settings.

      Weaknesses:

      These include reliability on qPCR-based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanations for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there is limited generalizability as the study was done in a single urban site in Zambia. There is also a lack of functional immune data.

      Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Weaknesses:

      While I am quite enthusiastic about the work, I am concerned that a number of likely relevant confounders were not discussed and that the broader implications of their findings were not well grounded in the existing literature. For example, I could not find information on the vaccination status of the mothers in the study. Given the conclusions about asymptomatic transmission and the durability of immunity, it is important to know the vaccination status of the mothers. Moreover, did the authors have other metadata on the mother/infant dyads, e.g., household size, vaccination status of household members, etc.? Given the potential implications of more widespread asymptomatic transmission associated with pertussis infection, I believe the authors should better couch their results in the context of the broader debate around asymptomatic transmission.

      We appreciate the reviewers' detailed feedback. We provide an overview of our responses here and we address specific recommendations below. In light of reviewers’ comments, we have revised our manuscript in order to improve the clarity of our presentation and to better situate our results within the context of the existing literature. Unfortunately, as the field study has been concluded, many of the reviewers’ recommendations are not possible. These include additional testing (i.e., culture or serology) or sequencing. We have updated the manuscript to more clearly indicate our knowledge regarding maternal vaccine status and immunological immunity of study participants. We have also provided a more comprehensive overview of existing pertussis studies, including genomic surveillance and details regarding sub-Saharan Africa and Zambia in particular. Finally, we have revised the formatting of Figure 4 (survival analysis) to more clearly highlight differences between mothers and infants and to better align with the text, and note that the underlying results are unchanged.

      A particular concern raised in the reviews that we wish to address is the recommendation of culture- or serology-based tests as "confirmatory". We have revised the manuscript in light of this feedback to better reflect our own position on this matter. We believe these recommendations do not adequately account for important trade-offs between testing sensitivity and specificity that are widely recognized in both clinical practice and epidemiology (Enøe et al., 2000; Florkowski, 2008; Swift et al., 2020). When the results of different testing methodology disagree, rarely is one method, a priori, correct. Rather, the disagreement may point to specific test limitations or important biological questions about the study system.

      In the case of pertussis detection, cell culture is recognized for its very low sensitivity, while serological detection is complicated by debate around appropriate threshold levels and time horizons for seroconversion and subsequent decay (Lee et al., 2018; van der Zee et al., 2015). Furthermore, while anti-PT antibodies are a common target of serological detection, these are not reliable correlates of protection (Mills, 2001; Wilk et al., 2019), nor are they reliably generated in response to colonization (de Cellès & Rohani, 2024; Graaf et al., 2020). Overall, the detailed relationship between exposure, carriage, transmissible infection, and the dynamics of anti-PT serology remains poorly characterized (Craig et al., 2020; de Cellès et al., 2025).

      While we agree that these are important questions in epidemiology and public health, we nonetheless wish to highlight that there is no “free lunch": each additional test and protocol comes with additional cost and complexity that should be evaluated based on the specific goals of the intended surveillance. In our case, the repeated sampling of longitudinal surveillance serves as a low-cost "confirmatory" testing regime. We disagree that cell culture would have added value to the present study and would not recommend its addition to future studies (primarily due to low sensitivity). While we agree that before-and-after serology of mothers would have added important context to the present study, we nonetheless expect that significant ambiguity would have surrounded any such results (e.g., Moosa et al. (2025)).

      One area that we strongly agree warrants further attention is the household dynamics in pertussis transmission, particularly in low-resource settings where crowding is common. In our study we were not able to rule out environmental and/or shared transmission events, though our survival analysis did demonstrate a greater impact of mothers on infants than vice versa, results which suggest a causal mechanistic role. In previous studies we detailed the demographics of household size, number of children, and mothers' age (Gill et al., 2021; Gunning et al., 2020), though we have not conducted formal analyses of these important covariates here. We also note that the code and data are freely available, allowing for others to build on our work.

      We believe that future prospective studies are an invaluable tool for directly tracking pertussis disease transmission, including both community and household studies. We have argued here for the value of qPCR-based community surveillance, which could integrate into existing public health activities. Regarding household studies, we note that a key challenge in implementing these studies is selecting an appropriate sampling interval and duration to best capture epidemiological linkages. Our results suggest that qPCR-based real-time population-level surveillance could be used to initiate such a prospective household study during a pertussis outbreak so that a higher sampling frequency (e.g., weekly swabs) could be gainfully employed over a shorter time period.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Enhance Validation of qPCR Findings:

      We address these comments above at greater length. We also note that the term “false positive” is rather ambiguous here, as we lack a clear distinction between carriage and transmissible infection. We note that our manuscript includes a considerable discussion of qPCR validation, including negative controls and sample retesting. We agree that test sensitivity and specificity remains an important question, and have endeavored to clearly indicate these concerns throughout the manuscript.

      (2) Clarify Transmission Dynamics:

      While we agree that this is an important question, we lack the relevant sequence data to test it. Notably, we suspect that any such phylodynamic linkage would require genomic sequencing due to the relatively low genetic diversity observed in pertussis (see, e.g., population-level estimates of time to most common ancestor in Lefranc et al., (2022). However, sequencing pertussis genomes remains resource-intensive, and we expect that the deployment of such sequencing at scale would be cost-prohibitive in low-resource settings.

      We have revised our manuscript to underscore uncertainties around shared/household exposure. We also direct the reviewer’s attention to our survival analysis, where a notable asymmetry exists between mothers and infants. Here, mothers’ prior qPCR signals exhibit a larger impact on their infants than infants on their mothers (Fig 4). If shared exposure was the principal cause of the observed increase in hazard in ego from alter, then we would expect no such asymmetry between mothers and infants.

      (3) Expand Discussion on Public Health Implications:

      As noted above, we have revised our introduction and expanded our discussion to better account for the existing literature. And, while we are hesitant to put forward specific recommendations based on our study, we feel confident in stating that active pertussis surveillance in low-resource settings is A) almost entirely absent B) possible to achieve, and C) necessary to resolve long-standing questions about pertussis epidemiology at regional and national levels. We have endeavored to clarify these points, particularly within the discussion.

      (4) Address the Role of Immunity More Directly:

      While we lack such immune data, we have pointed to recent work from the notable PHIRST study in South Africa, as well as highlighted ambiguities surrounding these data.

      Reviewer #1 (Recommendations for the authors):

      (1) Do we know the vaccination status of the mothers in the study? … If these data are not available, I think that the paper must be re-framed to acknowledge that all the conclusions are statistical in nature, based on publicly available vaccine coverage data from Zambia.

      We do not have information on the immunization status of mothers, though we cite national rates for Zambia across the relevant time period. We have clarified this point in the revised manuscript. We have endeavored to clearly acknowledge that many of our conclusions are statistical in nature and to clearly quantify the strength of evidence.

      We strongly disagree that anomalously low vaccination rates amongst mothers (i.e., relative to national averages) would materially alter the interpretation of our findings.

      Overall, our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity. Indeed, some have argued that the preponderance of mild/asymptomatic infections in mothers is, of itself, evidence of prior immunological exposure (Fine & Clarkson, 1982).

      (2) Do we know anything about rates of pertussis in Zambia, especially in the study site?

      We address this important question in the discussion. In particular, we state that: “As a populous, middle-income and primarily urban country, Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009”.

      (3) I couldn't find information in the paper related to the severity of infection in the infants. It's mentioned in the section describing results in Figure 5, but I only saw analyses with symptoms (as opposed to severe symptoms). Do you have outcome data from infants testing positive?

      This question was addressed in more detail in our previous work (Gill et al., 2021), which we briefly summarize in the Introduction. We also show the frequency of severe symptoms in Fig 5 (bottom panel), and detail mild versus serious symptoms in our subgroup analysis (Fig 7D).

      (4) Do we know anything about vaccine-resistant strains of pertussis in Zambia?

      We are not aware of any such work. As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa.

      (5) While I believe sequencing is beyond the scope of the current study, the authors should comment on the potential utility of sequencing elements of the pertussis genome and use that to demonstrate causality and direction of transmission more strongly.

      We believe that existing literature has addressed the potential of sequence data and phylodynamics to infer transmission, particularly for pathogens with high mutation rates such as RNA viruses. To date, research into the phylodynamics of pertussis has focused exclusively on population-level dynamics (Lefrancq et al., 2022), where estimates of time to most recent common ancestor (TMRCA) are long, indicating low genomic variability at the scale of countries and years. To our knowledge, no work on pertussis has directly inferred transmission chains from sequence data. Given the existing evidence, we expect that any such work would require genome-level sequencing, which would likely be cost-prohibitive in low-resource settings.

      (6) … However, it would be helpful to understand more about how your results fit into the broader story around pertussis resurgence. … if the infant cases were all mild, they might never have been captured in surveillance data sets.

      We believe that a key result of our study is the remarkable mismatch between country-level symptoms-based surveillance and prospective surveillance, which demonstrates that such mild cases have almost certainly not been captured. These findings are mirrored by recent work in South Africa (now addressed in our Discussion, see Moosa et al. (2025)). We believe that prospective surveillance, particularly in under-surveilled regions, is critical to understanding pertussis transmission writ large, which we have attempted to communicate throughout our discussion.

      (7) Relatedly, if there are still high rates of asymptomatic mother-to-infant transmission with whole cell vaccination, then why is there an observed drop in infant pertussis following whole vaccination in most countries?

      In previous work, we demonstrated that some infants in this cohort exhibited asymptomatic infection (Gill et al., 2021). We note that a drop in pertussis incidence amongst infants after the roll-out of the whole-cell vaccine is not contradictory with our findings. We want to clarify that our results, and evidence that mother-to-infant transmission can occur, does not imply that the whole-cell vaccine fails to protect against transmission.

      We have previously used epidemiological evidence to infer the population-level impacts following the roll-out of whole-cell pertussis infant immunization. For example, we observed an increase in the inter-epidemic period that, together with the drop in infant cases, are consistent with a reduction in transmissible infections (Broutin et al., 2010; Rohani et al., 2000).

      References

      Broutin, H., Viboud, C., Grenfell, B. T., Miller, M. A., & Rohani, P. (2010). Impact of vaccination and birth rate on the epidemiology of pertussis: A comparative study in 64 countries. Proceedings of the Royal Society B: Biological Sciences, 277(1698), 3239–3245.  https://doi.org/10.1098/rspb.2010.0994  

      Craig, R., Kunkel, E., Crowcroft, N. S., Fitzpatrick, M. C., Melker, H. de, Althouse, B. M., Merkel, T., Scarpino, S. V., Koelle, K., Friedman, L., Arnold, C., & Bolotin, S. (2020). Asymptomatic Infection and Transmission of Pertussis in Households: A Systematic Review. Clinical Infectious Diseases, 70(1), 152–161. https://doi.org/10.1093/cid/ciz531

      de Cellès, M. D., & Rohani, P. (2024). Pertussis vaccines, epidemiology and evolution. Nature Reviews Microbiology, 1–14. https://doi.org/10.1038/s41579-024-01064-8

      de Cellès, M. D., Wong, A., Dalby, T., & Rohani, P. (2025). Natural immune boosting biases pertussis infection estimates in seroprevalence studies. Nature Communications, 16(1), 8883. 

      Enøe, C., Georgiadis, M. P., & Johnson, W. O. (2000). Estimation of sensitivity and specificity of diagnostic tests and disease prevalence when the true disease state is unknown.  Preventive Veterinary Medicine, 45(1–2), 61–81.

      Fine, P. E. M., & Clarkson, JacquelineA. (1982). The recurrence of whooping cough: Possible implications for assessment of vaccine efficacy. The Lancet, 319(8273), 666–669.  https://doi.org/10.1016/S0140-6736(82)92214-0 

      Florkowski, C. M. (2008). Sensitivity, specificity, receiver-operating characteristic (ROC) curves and likelihood ratios: Communicating the performance of diagnostic tests. The Clinical Biochemist Reviews, 29(Suppl 1), S83.

      Gill, C. J., Gunning, C. E., MacLeod, W. B., Mwananyanda, L., Thea, D. M., Pieciak, R. C., Kwenda, G., Mupila, Z., & Rohani, P. (2021). Asymptomatic Bordetella pertussis infections in a longitudinal cohort of young African infants and their mothers. eLife, 10, e65663. https://doi.org/10.7554/elife.65663

      Graaf, H. de, Ibrahim, M., Hill, A. R., Gbesemete, D., Vaughan, A. T., Gorringe, A., Preston, A.,  Buisman, A. M., Faust, S. N., Kester, K. E., Berbers, G. A. M., Diavatopoulos, D. A., & Read, R. C. (2020). Controlled Human Infection With Bordetella pertussis Induces Asymptomatic, Immunizing Colonization. Clinical Infectious Diseases: An Official Publication of the Infectious Diseases Society of America, 71(2), 403–411.  https://doi.org/10.1093/cid/ciz840

      Gunning, C. E., Mwananyanda, L., MacLeod, W. B., Mwale, M., Thea, D. M., Pieciak, R. C., Rohani, P., & Gill, C. J. (2020). Implementation and adherence of routine pertussis vaccination (DTP) in a low-resource urban birth cohort. BMJ Open, 10(12), e041198.

      Lee, A. D., Cassiday, P. K., Pawloski, L. C., Tatti, K. M., Martin, M. D., Briere, E. C., Tondella, M. L., Martin, S. W., & Group, C. V. S. (2018). Clinical evaluation and validation of laboratory methods for the diagnosis of Bordetella pertussis infection: Culture, polymerase chain reaction (PCR) and anti-pertussis toxin IgG serology (IgG-PT). PLoS One, 13(4), e0195979.

      Lefrancq, N., Bouchez, V., Fernandes, N., Barkoff, A.-M., Bosch, T., Dalby, T., Åkerlund, T.,  Darenberg, J., Fabianova, K., Vestrheim, D. F., Fry, N. K., González-López, J. J.,  Gullsby, K., Habington, A., He, Q., Litt, D., Martini, H., Piérard, D., Stefanelli, P., … Brisse, S. (2022). Global spatial dynamics and vaccine-induced fitness changes of Bordetella pertussis. Science Translational Medicine, 14(642), eabn3253.  https://doi.org/10.1126/scitranslmed.abn3253 

      Mills, K. H. G. (2001). Immunity to Bordetella pertussis. Microbes and Infection, 3(8), 655–677. https://doi.org/10.1016/s1286-4579(01)01421-6 

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018:  Findings from the PHIRST study. Journal of Infection, 106550. 

      Rohani, P., Earn, D. J., & Grenfell, B. T. (2000). Impact of immunisation on pertussis transmission in England and Wales. The Lancet, 355(9200), 285–286. https://doi.org/10.1016/S0140-6736(99)04482-7 

      Swift, A., Heale, R., & Twycross, A. (2020). What are sensitivity and specificity?  Evidence-Based Nursing, 23(1), 2–4.

      van der Zee, A., Schellekens, J. F., & Mooi, F. R. (2015). Laboratory diagnosis of pertussis.  Clinical Microbiology Reviews, 28(4), 1005–1026.

      Wilk, M. M., Allen, A. C., Misiak, A., Borkner, L., & Mills, K. H. G. (2019). The immunology of Bordetella pertussis infection and vaccination. In Pertussis: Epidemiology, Immunology & Evolution. Oxford University Press.

    1. eLife Assessment

      This study provides a useful single-cell atlas of the Clytia hemisphaerica planula, complemented by an updated medusa dataset, ultrastructural analyses and in situ validation of expression patterns. The cross-stage comparison and cluster-similarity framework offer a promising basis for investigating cellular diversification across the life cycle. The evidence has the potential to be convincing, but is currently incomplete due to documentation deficits and insufficient cross-referencing with prior work. It should be relatively easy to address these issues, which should make this a valued resource for cnidarian researchers and for colleagues studying cell-type evolution.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript further explores the single-cell atlas of Clytia hemisphaerica by incorporating the planula larva. It compares the cell clusters with the previously established atlas of the medusa. It identifies similarities and differences between the two life stages.

      Strengths:

      The manuscript provides an important set of single-cell data that have not been assessed previously: the Clytia planula. The data is further supplemented with high-quality EM-based histology and an extensive in situ hybridisation of selected genes.

      Weaknesses:

      The detailed analysis does not go deep into the comparison between stages, nor does it provide an analysis of genes within the clusters; it could be described as remaining overall rather superficial.

    3. Reviewer #2 (Public review):

      Summary:

      The generation of alternate stages in the life cycle of a single species requires vast remodeling of the cellular complement of the individual during metamorphosis from one stage to another. In this paper, the authors provide a detailed description of single-cell RNA-seq data derived from the planula stage of the hydrozoan model Clytia hemispherica and compare this to an expanded dataset from the medusa stage to assess changes in transcriptomic identity of cell types between these two phases of the life cycle. The paper further includes valuable TEM data illustrating fine anatomy of cell types present at both stages investigated, and documentation of the retention of epithelial polarity from the planula through to the polyp stage, using a reporter line.

      Strengths:

      The study provides a solid and convincing transcriptomic characterization of planula cell types (including in situ validations and a planula-to-polyp mapping of epithelial polarity), and introduces a potentially valuable method for evaluating cluster similarity.

      Weaknesses:

      The work suffers from insufficient documentation of methodological approaches and missing code, lack of clarity regarding clustering resolution and nomenclature (thereby hindering cross-referencing with prior papers), and unclear plans for public, fully annotated data release.

      Full Review:

      The single-cell transcriptomic data analyzed include both previously published and newly generated data: two additional medusa libraries and two additional planula libraries were generated and integrated with the data from https://doi.org/10.1126/sciadv.abh1683, and https://doi.org/10.1126/sciadv.adv1159. The original release of the planula dataset in their 2025 Science Advances paper did not include analyses of all cell types. Here the authors provide this analysis for the planula stage. However, as both the number of clusters and the nomenclature of the clusters changed, this leads to some confusion and inability to cross-reference the two papers. There is no explanation given for the re-processing of the planula dataset in the current paper, and the fact that only some of the data is new is buried in the supplement, which is not referenced in the main document, while the text within the main article suggests that the entire dataset is new. The fact that the dataset in the current analyses contains fewer cells than presented in their Science Advances paper further adds to this confusion. The current paper would benefit from greater transparency in the origin of the data analyzed.

      The authors do try to apply the same nomenclature for the updated medusa dataset that is present in their 2021 Science paper. For example, the previously identified 'bioluminescent cells' are identified as 'gas-m8'. A look-up table that has all of the cluster id's cross-referenced would be useful (i.e. new: 8 = gas-m8 = previous: 28 = BC = "Tentacle GFP cells"). The inability to easily cross-compare with the published data is a major weakness of the current work and would benefit greatly from consistency between the three papers. Indeed, the clustering resolution is quite different across all three papers, and the current work does not adequately address how the clustering resolution was selected here. As an updated atlas, one would expect the entire transcriptomic diversity to be included here, so that the previous work can be transferred to the updated genomic mapping resource used in the current work. Nonetheless, presenting a unified nomenclature for moving forward would benefit the community as a whole and would increase the impact of the current work substantially.

      The paper also includes new TEM data of the planula cell types. The authors attempt to correlate transcriptomic profiles with these anatomical data through in situ hybridizations that provide spatial distribution of the profiles. While the TEM data are valuable to catalog the presence of cells with different morphologies within the planula, the association with the transcriptomic profiles is somewhat speculative. These valuable anatomical data should be provided at a high enough resolution to zoom in and see the details, and further description could be provided. For example, the paper states that vacuolated cells are characteristic of the basal gastrodermal cells adjacent to the mesoglea; please identify the vacuoles in Figure 4f/h for the reader.

      A novel method for reconstructing cluster similarity relationships is applied to grouping clusters into cell categories within the same life cycle stage, and also for matching cell types between stages. This is a valuable contribution to the field that is worthy of further evaluation. This is, however, difficult, as the methods for which DESeq2 was applied ("see code for details") are not present in the provided code, nor is it adequately described how the "binary matrix of marker gene presence/absence" was constructed. Similarly, there are additional details of other parts of the data analysis that are missing from the provided code, and the provided supplementary material is not referenced in the main document. More rigorous documentation of the methods is warranted.

      The description of the transcriptomic profiles present in the planula is solid, and the attempt to associate these profiles with anatomic locations and putative morphology provides a foundation onto which further studies can be developed. Mapping of the planula ectoderm through to the polyp stage is also an important step forward in characterizing the life cycle, and the evidence for the retention of the oral/aboral ectodermal axis is convincing. The paper falls short in describing the updated medusa dataset and could benefit from a minor restructuring of the paper. Introducing the new medusa data only after the planula dataset is fully described would mediate the shallower treatment of the updated medusa dataset, where only 22 of the original 36 transcriptomic states are recovered. In this way, the focus will shift onto the cross-life cycle stage comparisons, and it could be argued that the lower resolution of the medusa dataset is justified in order to simplify the comparisons.

      It will be essential that the datasets that are presented in this work be made available for public exploration in a fully annotated format. It is currently unclear how the authors intend to do this; however, there are many repositories available for this. The UCSC Cell Browser hosted at cells.ucsc.edu is one very good option if the authors do not wish to develop an interactive tool themselves. It is imperative that the gene annotations which correspond to the dataset, and the cluster annotations that are presented in this paper, are available and easily connected to the released dataset.

    1. eLife Assessment

      This study offers valuable insights into the genetic and evolutionary basis of the starvation response by confirming the hypothesis about the mito-nuclear etiology of this trait. The level of evidence is currently incomplete but could be improved if controls were designed better, particularly by including analysis of variation in the initial, common population prior to all treatments. Overall, this work would be interesting to a broad audience, beyond the Drosophila community, if the analysis of the candidate genes against Human ortholog loci were to be conducted with more careful controls.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript presents a genome-wide investigation of the genetic architecture underlying adaptation to prolonged starvation in Drosophila melanogaster, using an E&R experimental design maintained across 60 generations. Four starvation-selected (SS) and four matched control (C) populations were whole-genome resequenced, and two complementary analytical frameworks, selective sweep inference combined with low-heterozygosity mapping, and a diffusion-based drift-filtering approach, were applied to identify genomic regions under selection. As a result, the authors report (1) 62 high-confidence sweep-low-heterozygosity regions encompassing 255 genes, and (2) 3,578 SNPs with allele-frequency shifts exceeding neutral drift expectations shared across all four SS replicates, mapping to 578 genes. Mitochondrial pathways are identified as prominent targets, with a 13.9-fold enrichment of nuclear-encoded mitochondrial genes among candidates and differentiation at the mitochondrial origin of replication. Finally, the authors demonstrate that human orthologs of starvation-responsive fly genes are enriched for highly differentiated variants in four human populations from the 1000 Genomes Project.

      Strengths:

      (1) The experimental design with four evolution replicates provides proper control for false discovery.

      (2) The phenotypic characterisation is thorough. The approximately 3-fold increase in starvation survival and 1.5-fold increase in TAG content provide a clear physiological basis for interpreting the genomic findings, and the observation of increased adult longevity adds a meaningful life-history dimension to the results.

      (3) The mito-nuclear analysis is one of the more novel contributions of this paper. The implicated picture of coordinated mito-nuclear remodelling under sustained nutrient deprivation is compelling.

      (4) The comparative analysis connecting fly selection candidates to human population differentiation is ambitious and adds evolutionary breadth to the study.

      Weaknesses:

      (1) Ne estimation is derived from controls only, not from selected populations

      The entire drift-filtering framework rests on estimates of effective population size obtained from allele-frequency variance among the four control replicates (Ne = 530 for autosomes, Ne = 461 for the X chromosome). This is justified by assuming that divergence among control populations reflects neutral drift alone, a reasonable assumption for C populations maintained on standard food.

      However, the starvation-selected populations experienced 75-80% mortality per generation as an explicit design feature of the selection regime. This severe, recurrent demographic bottleneck would substantially reduce the effective population size within SS lines relative to controls. The authors do not acknowledge this discrepancy, nor do they attempt to estimate Ne within SS replicates or assess the sensitivity of their drift thresholds to plausible reductions in Ne. If Ne in SS populations is appreciably lower than in controls, the drift thresholds derived from the control-based Ne will underestimate the amount of neutral drift occurring in SS lines. Consequently, some allele-frequency shifts that are driven by the repeated bottleneck could be misclassified as candidate loci, inflating the apparent number of selection targets. This is the most consequential methodological concern in the paper. The authors should either estimate Ne separately for SS populations, implement a sensitivity analysis varying Ne over a biologically plausible range, or, at a minimum, provide a thorough discussion of how downward bias in SS Ne would affect their results and conclusions.

      (2) Lack of consideration about binomial sampling noise due to the pool size in the modeling

      With only 100 individuals pooled per population, binomial sampling from the pool contributes a non-trivial additional source of variance to allele-frequency estimates, on top of genetic drift and sequencing error. This is a well-documented issue in Pool-seq data. Critically, the Kimura diffusion framework used for drift modeling does not appear to explicitly incorporate this binomial sampling noise component, an omission that could affect the calibration of drift thresholds, particularly for low-frequency alleles. The authors should discuss whether and how pool-size-induced sampling variance is accounted for in their drift model.

      (3) No benchmarking against established Pool-seq analysis tools

      The authors use Pool-HMM for sweep detection and a custom diffusion-based drift framework for allele-frequency analysis, with PoPoolation (v1) used only for Tajima's D calculations. However, the study does not benchmark its candidate SNP sets or sweep regions against well-established Pool-seq analysis frameworks such as PoPoolation2, which provides CMH tests and FST estimation specifically designed for replicated Pool-seq E&R data, or R/poolSeq, which implements drift-aware testing purpose-built for this experimental design. The authors should either benchmark their approach against at least one established alternative or provide explicit justification for why their custom framework is preferable and how it compares in sensitivity and specificity.

      (4) Absence of negative controls in the human PBS comparative analysis

      A critical missing element in this comparative analysis is a negative control: the authors do not test whether equivalent enrichment is observed in populations with no particular history of famine or nutritional stress, such as European or East Asian populations from the 1000 Genomes Project. The inclusion of at least one negative-control population triplet is necessary to support the cross-species interpretation as stated.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use an Evolve-and-Resequence approach in Drosophila to study the genomic basis of adaptation to long-term starvation. Replicated selection lines and control populations are sequenced and analyzed to identify signals of selection, which are then related to starvation-related phenotypes. The general experimental design is appropriate, and the combination of genomic and phenotypic data is a clear strength of the study.

      Strengths:

      The strongest aspect of the work is the experimental evolution framework combined with population genomic inference across replicate populations. The observed parallelism across replicates supports the robustness of at least a subset of the detected selection signals. However, several key methodological details are either unclear or insufficiently justified. In particular, both the maintenance of control populations and demographic assumptions are not fully described, and the treatment of structural variation (e.g., segregating inversions) is not sufficiently addressed. The phenotypic analyses are broadly appropriate and replicated but would benefit from access to raw data.

      Weaknesses:

      The human ortholog enrichment analysis is an interesting component of the study, but it should be interpreted more cautiously. As currently presented, it is based on correlational signals of differentiation and is therefore sensitive to potential confounding factors. While the analysis may point to intriguing patterns consistent with conserved genetic architecture, the evidence is not sufficient to support strong claims of conserved starvation/malnutrition-related polygenic adaptation in humans. Framing this component more explicitly as exploratory would strengthen the manuscript. In its current form, this analysis is somewhat less conclusive than the experimental evolution results in flies.

      Overall, the study provides a useful dataset and a reasonably solid analysis of starvation adaptation in experimental Drosophila populations, but several methodological clarifications and a more balanced framing of the cross-species comparisons would strengthen the manuscript.

    4. Reviewer #3 (Public review):

      Summary:

      This study tries to identify the genetic signatures of adaptation to starvation conditions. For this, outbred populations of Drosophila melanogaster were selected for starvation resistance by using the 20% surviving adults after starvation to start the next generation. This was done for 60 generations while parallel populations were kept under control conditions. At the end of the experiment, starvation-selected flies showed increased survival, longevity, and TGA storage. DNA poolseq data from control and starvation populations were compared to identify genomic regions with low heterozygosity and signatures of selective sweeps, and SNPs with differences in allele frequency. The candidate regions point to mitochondrial and metabolic pathways as the targets of selection for starvation resistance.

      The authors replicate the experimental design, selection approach, data collection, and analyses from Hardy et al 2018 (https://doi.org/10.1093/molbev/msx254), which also investigated adaptation to starvation conditions but used a different Drosophila melanogaster population. In this sense, the current study recapitulates most of the findings from Hardy et al. (2018). The analyses of the mitochondrial results, including the overlap with human data, are the novelty of this paper. However, those analyses are not very well justified. The fact that this study is almost identical to Hardy et al is not clearly stated nor discussed in the manuscript.

      Strengths:

      The authors made use of an experimental evolution approach to identify the genetic basis underlying adaptation. This is a powerful approach that has proven very successful in the past. They used a good number of replicates (four per condition), an appropriate depth of sequencing, and quantified higher-order phenotypes to validate the claim that the populations had evolved increased starvation resistance.

      Weaknesses :

      Although the findings of this study seem credible based on the known biology of starvation resistance, there are several aspects of the experimental design that weaken my confidence in the results. The points below should be clarified, and the limitations of the experimental design and analyses need to be included in the discussion.

      (1) Pooled genomic data were collected for the four replicates at the end of 60 generations of selection, and four replicates were kept under control conditions. No data were collected at the beginning of the experiment, which is the current standard in Evolve and Resequence experiments. To infer the genomic regions underlying adaptation to starvation, evolved control and starved cages are compared. Although this will identify regions that are possibly truly caused by adaptation to starvation stress, the available data doesn't allow to determine, for example: a) whether the differences between control and starvation regimes are due to changes in control cages relative to the starting population, combined with no changes in starvation cages relative to the starting population; b) whether the differences across replicates are due to different genomic composition at the start of the experiment that could have been amplified by drift.

      (2) Selection was applied by starving flies until ~80% of the population died. The 20% surviving flies were used to seed the next generation. The control populations, on the other hand, were propagated using the whole population. Given that only the starvation populations were subject to such a strong bottleneck, it is not possible to disentangle whether the genomic signatures at the end of the experiment are due to this, and not necessarily to starvation resistance. For example, the low heterozygosity blocks and the very great changes in allele frequency could be a natural result of such a bottleneck. A proper comparison would have been to select a random 20% of the control individuals to seed every generation.

      (3) The analyses that involve human populations are poorly justified, and the enrichment tests are not clearly explained. There is no evidence of signatures of selection for starvation resistance in human datasets (as mentioned in the text, line 112), and yet the authors claim that their analyses that identify branch-specific alleles for a set of four human populations serve as a dataset for it. I don't think the results of this analysis and further overlap with candidate genes identified in the Drosophila experiment support the conclusion that polygenic adaptation of metabolic pathways is conserved across species (line 303).

      (4) The conclusion that adaptation to starvation conditions is repeatable is not justified by the data. The overlap across replicates is very low in every metric.

      (5) The methods are poorly described. In most of the sections, there is not enough information to be able to replicate the experiments or the analyses. Several of the analyses presented in the results are not described in the methods. Without this information, it is very difficult to assess whether the analyses were correctly done or whether the results are robust.

    1. eLife Assessment

      This valuable study combines a chromosome-level genome assembly with population resequencing and demographic modelling to reassess the status of the Formosan landlocked salmon, a critically endangered salmonid at the southern edge of the genus. The assembly and the evidence for extensive lineage-specific chromosomal fusions are genuinely well done, and the discovery of an overlooked, genetically distinct population in Hehuan Creek is sound and of clear management relevance. Evidence for the headline claims is incomplete: species rank is asserted but nowhere argued, and gene flow is excluded using a coalescent model that contains no migration parameter. Support for the population-viability conclusions is also not complete, as the demographic mechanisms invoked in the discussion are not borne out by the supplementary tables.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting paper on an important topic, the taxonomic and conservation status of some unusual salmonid populations in Taiwan.

      Strengths:

      The first part of the manuscript is quite strong: the authors sequence and build a reference genome and conduct a phylogenomic analysis. They examine chromosome structure and rearrangements, test for loss-of-function mutations, and do a proteomic analysis. As a stand-alone, this could serve as its own manuscript, perhaps for a more specialized journal.

      Weaknesses:

      I find this manuscript rather disjointed. The first part of the manuscript is related to phylogenomics of the taxon in question, compared to other nearby species from Japan. The authors go on to describe chromosomal rearrangements, sex-chromosome location, and proteomics. All of these fit within a paper about taxon-level issues. I do find the proteomic analysis perhaps unnecessary. I'm not sure we learn much of substance through this analysis, which is highly speculative.

      PSMC analysis seems highly questionable for taxa with such strong genetic structure. If historical Ne and past changes in structure are confounded, what does this analysis provide? I recommend deletion of the analysis included in Figure 1e.

      The second portion of the manuscript deals with population structure of three O. formosanus populations, based on RADSeq data. This part reads as a separate manuscript, in my opinion. I think the authors are trying to squeeze too much into one manuscript.

      For the second part on population genomics, not enough detail is provided to evaluate the methods, results, and interpretations. For example, not enough detail is provided about each of the three Taiwan populations, the stocking history, and the demographic data collection. The only information available is a brief paragraph in the introduction. Was the Luoyewei (L) population stocked from a brook derived from this population or from Qijiawan (Q)? Why do three L fish have such different levels of MLH? Are these stocked from somewhere else? Are the rest of the fish from one pool, and maybe one family (this would also explain the extremely low contemporary Ne)? Only 17 fish were examined from L, and apparently from one site in the stream; more detail is needed. Are L, Q, and H currently isolated? What is the stocking history? The authors conclude that the Hehuan (H) population has more genetic variation and is likely the result of an unknown native population that bred with stocked fish (which arise from Q). This story does align with the genetic results, but again, more detail is needed. Are there alternative explanations? A more careful treatment would be helpful.

      The demographic modeling is not convincing. Not enough detail is provided, and the lack of individual identification of fish makes it so the modeling is very general. It is hard to place too much stock in these vital rate estimates. The methods were fishing, snorkeling, and some electrofishing. Scales were used for ageing, and catch curve analysis was employed. Overall, this is an underdeveloped portion of the paper that is important, but not convincing as written.

    3. Reviewer #2 (Public review):

      Summary:

      Lee et al. is a comprehensive conservation genomics study that combines a chromosome-level genome assembly (sex-specific, too), population resequencing, coalescent species delimitation, and simulations to reassess the evolutionary status and conservation outlook of the Formosan landlocked salmon, Oncorhynchus formosanus. The authors showed a distinctive genome structure, replete with chromosome fusions and an unusual placement of the sex-determining gene sdY. Across sampling sites, they observed variable levels of genetic diversity, but in a way that was surprising given previous census numbers and conservation history. In particular, the authors report a previously unrecognised native population in Hehuan Creek, and conclude that Hehuan is more resilient to typhoon disturbance than the long-protected Qijiawan population - motivating stream-specific rather than range-wide conservation.

      Overall, this is a well-written paper that combines a number of elements that are timely and relevant. It uses state-of-the-art techniques to reach its conclusions and is generally performed to a high standard. It describes a critically endangered species that poses its unique conservation challenges. There are a number of things to like, as well as some substantial shortcomings in this paper.

      Strengths:

      The genomic resource is excellent. The assembly is well validated (97.3% anchored to 25 scaffolds, 95.7% BUSCO), and the authors generated a separate male assembly specifically to resolve the sex-determining region, allowing XY-shared and Y-specific contigs to be distinguished on coverage rather than inference. This is truly well done, and at a high standard. The synteny evidence for telomere-to-telomere fusions involving at least 14 ancestral chromosomes, against two in O. m. masou, is convincing.

      The Hehuan Creek result is the paper's most valuable contribution. Elevated heterozygosity, short and infrequent runs of homozygosity, and private alleles absent from the Qijiawan broodstock are difficult to reconcile with a purely reintroduced origin. The contrast with Luoyewei is a clean and useful cautionary case for hatchery supplementation.

      Weaknesses:

      (1) The species-rank claim is featured in the abstract, but it is made with any level of rigour in the paper. "New species" appears once, in the abstract (l. 32). The Results conclude only that O. formosanus is a distinct evolutionarily significant unit (ll. 188-191), which itself can be well-justified, but it's far from a taxonomic rank (see author's own ref 10). No species concept is explicitly named anywhere, and the taxon is referred to across the manuscript as a subspecies (l. 68), a "new species" (l. 32), and an ESU (l. 189) in turn.

      (2) Gene flow is asserted, not tested, and two divergence estimates disagree twentyfold. The abstract reports "no detectable gene flow for ~50,000 years." That figure is a divergence time from BPP under the A00 model, which contains no migration parameter; a model that cannot fit gene flow cannot report its absence. Separately, Figure 1b shows a split at 1.15-5.09 Mya (Figure 1b), while the ddRAD coalescent places the same split at ~50 kya (Figure 1d). The explanation offered (ll. 417-421, "differing temporal sensitivity of genomic markers") is not a mechanism.

      (3) The placement of sdY is unresolved, and the paper's own figures conflict with its text. Figure S10 and Table S6 both make O. formosanus chr13 homologous to O. m. masou chr32, whereas reference 28 - on which the authors rely - places the sdY contig on O. m. masou chr7, whose O. formosanus homologue is chr5 (Table S6). These cannot both be correct, and Figure S10's caption compounds the confusion by attributing chr13 to masou and omitting the chr32 track entirely.

      (4) The population-viability model's stated mechanisms are contradicted by the authors' own supplementary tables. The Discussion attributes Qijiawan's vulnerability to "lower juvenile survival, decreased fecundity, and narrower terminal age class representation" (ll. 522-525). Table S10 gives Qijiawan higher age-0 survival (0.202/0.616 vs 0.184/0.615); Table S11 gives it higher fecundity at every reproductive age (7.68/19.27/7.86 vs 6.10/8.95/4.14); Table S9 gives it a broader terminal age class (5.0% vs 0.8% age-3 in November). The only parameter favouring Hehuan is age-1 survival - 0.087 (95% CI 0.000-0.180) versus 0.131 (0.093-0.187) under typhoon, and 0.054 (0.000-0.167) versus 0.087 (0.047-0.149) at baseline. Both Qijiawan intervals include zero and overlap Hehuan's, yet a reported extinction odds ratio of 4.48 rests on this difference.

      (5) The two streams were not measured equivalently, and every asymmetry favours the conclusion.

    1. eLife Assessment

      This important study combines cryo-EM, biochemical, and cell-based assays to examine how Gβγ interacts with and potentiates PLCβ3. The authors present evidence for multiple Gβγ interaction surfaces and argue that Gβγ primarily enhances PLCβ3 activity after membrane recruitment rather than serving mainly as a membrane-recruitment factor. Following additional experimental support, the evidence in support of their conclusions is convincing.

    2. Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper showing the membrane bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Major concerns

      My most major concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is of course the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help figure 1, is an evolutionarily conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This also will help orient the reader to the identity of the mutated residues assayed in figure 3.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent tesmer structure with PI3K.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated as of course that can effect any of the three surfaces. My main question is that from the way fig 2A is oriented that the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      (4) To help reader interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Comment on revised version.

      After revision the authors have addressed all of my concerns.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification.

      Comments on revised version.

      The authors have reasonably addressed my comments.

    4. Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solution and cryoEM studies of the PLCβ3 / Gβγ complex to delineate the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate. The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation in accord with other PLCβ enzymes.

      Weaknesses:

      (1) The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits but will be investigating this in future studies.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself, and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper, showing the membrane-bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Thank you for the positive feedback.

      Major concerns:

      My biggest concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors' main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is, of course, the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this, any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help Figure 1 is an evolutionary conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This will also help orient the reader to the identity of the mutated residues assayed in Figure 3.

      We agree that crosslinking can capture non-physiologically relevant interfaces. However, because we do not observe any crosslinking between Gβγ and a PLCβ3 variant that retains a cysteine in the X–Y linker or between PLCβ3 and any other cysteine in the Gβγ heterodimer, we believe it is site-specific.

      The question about sequence conservation in the Gβγ–PLCb3 interfaces is interesting and we have included this information in Figures S8 and S9.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent Tesmer structure with PI3K.

      We agree that the orientation of Gβ in the crosslinked structure is different. We include a comparison of this reconstruction to other published Gβγ–effector complexes as Figure S6.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated, as of course that can affect any of the three surfaces. My main question is that, from the way Figure 2A is oriented, the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      Thank you for pointing this out, and we expanded our analysis to include this residue. The R199A and R199E mutations had no defects in basal or Ga<sub>q</sub>-stimulated activities. However, R199A had 3-fold lower activation by Gβγ, while R199E was not activated in this assay. The R199E mutation also had significantly decreased agonist-dependent BRET with Gβγ and decreased PI(4,5)P2 hydrolysis, while retaining robust recruitment to the plasma membrane by Gα<sub>q</sub>. These data is included in the main text and Figures 2-4.

      (4) To help the reader's interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Thank for the suggestion. In revised Figure S3, we show the Gβγ–PLCb3 D892-PH<sub>cys</sub> complexes determined in this study at different contour levels.

      Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify and map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on the membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification. The manuscript is generally strong, but some issues need to be addressed.

      Thank you for the positive comments.

      Major comments:

      (1) BMOE/BM(PEG)2 crosslinking may enforce a non-native docking geometry, potentially compromising the physiological relevance and precision of the Gβγ-PLCβ3 interface as described. Although a >50% 1:1 crosslinked complex is formed and remains active, the solution maps show lower local resolution for Gβγ, consistent with a dynamic, potentially heterogeneous, interface. One interface is captured via a single engineered cysteine pair (PLCβ3 E60C-Gβ C271), which could potentially bias the pose. It would be helpful if the authors could provide additional orthogonal support (e.g., alternative crosslinked sites) and bolster the clarification of its uniqueness and relevance.

      We did attempt to isolate other crosslinked complexes. PLCβ3-D892 self-crosslinked under all reaction conditions, while PLCβ3-D892 XY<sub>Cys</sub>, which retains an endogenous cysteine within the X–Y linker (C516), did not result in any crosslinked product when incubated with Gβγ. Only the PLCβ3-D892 E60C crosslinked to Gβγ. With the exception the C68S mutation at the C-terminus of Gg to eliminate its prenylation site, all endogenous cysteines were retained in both Gβ and Gγ. Indeed, Gβ contains two solvent-exposed cysteines in its canonical effector binding surface (C204 and C271), but we did not observe any crosslinker density involving C204. While we cannot exclude the possibility that crosslinking occurred between PLCβ3-D892 E60C and other residues in Gβγ, we were unable to identify any 2D classes corresponding to these alternative conformations. These observations, together with the high efficiency of crosslinking, are consistent with a stable and persistent interaction.

      (2) In the crosslinked structure, the authors report that GβD228 interacts with PLCβ3 R199 and K183. In Figure 2A, R199 appears closer to Gβ D228 than K183, yet only K183 is functionally tested. Testing R199 (e.g., R199E/R199A) would strengthen the structure-guided validation of this interface.

      We agree, and functional analysis of PLCb3 R199E is included in the revised manuscript (see Figures 2-4).

      (3) The mutagenesis strategy appears inconsistent across figures/assays, which makes it difficult to interpret phenotypes and directly link the functional data to the proposed interfaces. For example, in Figure 2E, we see R185L but R215E, while residue L40 is mutated to Gly in the IP accumulation assays but to Glu/Lys (L40E/K) in the BRET assays (Figures 3B/3D/3F). The authors should (i) clearly justify the rationale for each substitution (conservative vs charge-reversal, interface disruption, etc.) and (ii), where possible, test the same mutants across assays (or provide evidence that alternative substitutions yield consistent conclusions).

      Mutagenesis experiments were initially carried out independently in the Lambert and Lyon Labs. As the study progressed, additional mutants were identified and/or designed based on results from both groups. The residues subject to mutagenesis are overall consistent across the different assays, with differences in the identity of the mutation varying in some cases. The L40G mutation is one such example, where given its modest impact on Gβγ-mediated activation in the IP accumulation assay, more impactful changes were made (L40E and L40K) for the BRET and signaling assays. In the revision, we now state that mutations were designed to maximally disrupt the three observed interfaces, such as by changing the size of the side chain and/or introducing charge reversal mutants.

      Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solutions and cryoEM studies of PLCβ3 / Gβγ, trying to understand the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate, although this is not completely convincing. Although these sites may not naturally be operative, the authors might want to develop the potential role of these sites.

      The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation, in accord with other PLCβ enzymes, but not for PLCβ3, and again, the authors might want to develop this point further.

      Thank you for the suggestions. We are investigating whether the other PLCb isoforms contain multiple Gβγ binding sites and the relative importance of preactivation by Ga<sub>q</sub> for a manuscript in preparation.

      Weaknesses:

      (1) I'm confused as to why the authors feel that their mechanism is distinct from the two-state enzyme, the synergistic activation proposed by Ross in 2011, using a primarily thermodynamic argument. As written, the authors appear to be very reliant on structural and BRET studies that do not give the details that would disprove this interpretation. The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits.

      The reconstitution experiments are under extremely artificial conditions, using nM-µM of purified proteins and liposomes that contain up to 30% PI(4,5)P2. Under these conditions, we think the increased activity is due to interfacial activation promoted by Gβγ binding to the lipase once it is associated with the liposome surface. This would be sufficient to account for the dose-dependent increase in both PLCb2 and PLCb3 activity as a function of Gβγ concentration. Given the higher basal activity of PLCβ2 and its decreased sensitivity to activation by Ga<sub>q</sub>, one possible explanation is that this isoform differs in its autoinhibition and/or structure of its proximal CTD that Ha2’ displacement is not a prerequisite for activation. In addition, Gβγ may also be a direct activator of PLCβ2. Further studies, ideally in cell-based systems, are needed to answer these questions.

      (2) In a recent study, McKinnon presents a model showing that Gαq and Gβγ activate PLCβ3 by two distinct pathways and that activation by Gβγ occurs through membrane recruitment. It is not surprising that the authors find that this is not true since the pelleting method used by McKinnon is subject to error. The authors should directly address the limitations of this previous work and the changes in proteoliposomes with sedimentation that alter partition coefficients. Although the inability of Gβγ to drive membrane binding is in accord with the quantitative studies of Scarlata, showing that the affinity of PLCβ3 to Gβγ is fairly weak as compared to the intrinsic membrane partition coefficient.

      We have added some of the limitations of proteoliposome sedimentation experiments to the discussion.

      (3) It was proposed many years ago that in signaling complexes Gαq - Gβγ may not have to fully dissociate when binding PLCβ, but rather shift their relative orientation when binding to PLCβ to allow activation. Is their model consistent with this? Is it possible that PLCβ3 keeps Gβγ from diffusing to enhance the rate of Gq / Gβγ re-association?

      Our crosslinked complex is compatible with simultaneous binding of a Gα<sub>q</sub>-Gβγ heterotrimer to the PLCb3, without disrupting the observed interface. If Gαq were to interact with the Gβγ molecules bound to the PH or EF hands, the interaction would be mediated by the N-terminal helix of Gα<sub>q</sub>. It is possible Gβγ–PLCβ3 interactions may slow heterotrimer reassociation, but this may be complicated by the intrinsic GAP activity of the lipase.

      (4) The authors find that Gβγ binds multiple sites, and it is clear that the PH domain site is the primary one in accord with previous work. Could these weaker sites be an artifact of the elevated concentrations used in cryoEM and BRET assays?

      While more studies have focused on the PH domain as a Gβγ binding site, our data does confirm the EF hands are also functionally relevant. To our knowledge, the role of the EF hands has not been investigated in this capacity until very recently, and so we hesitate to label them primary or secondary. It is possible the EF hands may be a lower-affinity site for Gβγ and the protein concentrations needed in cryo-EM drive complex formation. However, it is also possible the concentration of free Gβγ adjacent to an activated receptor may be high enough to saturate the PH and EF hand binding sites.

      (5) Although their assays infer differences in binding affinities, it would strengthen the paper if the authors could estimate the association energies of these different binding sites. This estimation would also address the concern stated above.

      We appreciate this suggestion and quantifying the affinities of the Gβγ–PLCβ3 interactions is the subject of future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please correct PIP2 to the correct PIP2 (many examples throughout).

      These have been corrected.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Figure S1B: The lane-condition labels above the third gel appear to be incorrect, as both lanes are marked identically (Gβγ +/+, PLCβ variant +/+, BMOE +/+) despite clearly different banding patterns. Please confirm.

      We have confirmed the markings above the gels are correct.

      (2) In Figure S2, the authors show three fitted models, but the helical density for Gβγ cannot be seen in two of them at the displayed contour level. The authors should provide views of the map at different contour levels (thresholds) to better support the model fitting. Otherwise, it is difficult to assess whether the Gβγ subunit could adopt alternative orientations (i.e., whether it may be rotated) within the density.

      We have included a new figure (Figure S3) that provides images of the maps at different contour levels.

      (3) Page 5: "where PLCβ3 is increased by the overexpression of either Gβγ or Gαq" should be revised to "where PLCβ3 activity is ...".

      This sentence has been corrected.

      (4) Figure 2: Please label residue R215 in Figure 2A/2B (or the relevant structural panel), since R215E is tested in 2E but the position is not shown.

      R215 is now included in Figure 2C.

      (5) Page 19, Figure 2 legend: "Changes ... Figure S3" should be "Changes ... Figure S5".

      We have corrected this figure call.

      Reviewer #3 (Recommendations for the authors):

      The studies seem well carried out, although more details regarding the BRET controls and the significance of the values should be included.

      We have revised the captions to provide more details about the experimental controls and a brief description of significance. Individual p-values are included in the supplemental tables.

    1. eLife Assessment

      This fundamental work uncovers an unexpected lysosomal function for NINJ2 and links it to ferroptosis and cancer biology. The evidence supporting the conclusions appears to be convincing. This work will be of general interest to the community of ferroptosis and cancer biology.

    2. Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of western blot data and further discussion of mechanistic questions would strengthen the study.

      The findings are likely to have broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      Comments on revised version.

      The authors have addressed all of my comments and questions. I have no further concerns.

    3. Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing co overexpression and positive correlation of NINJ2 with ferritin genes in iron addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:<br /> NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc⁻ pathways is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      (1) The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of Western blot data and further discussion of mechanistic questions would strengthen the study.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      (2) The findings are likely to have a broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      We thank the reviewer’s comment. We have discussed the potential implications of these findings for cancer treatment in the “Discussion” section.

      Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing cooverexpression and positive correlation of NINJ2 with ferritin genes in iron-addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:

      NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture, is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc</sup>-</sup> pathways, is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

      Areas for improvement:

      While the study is strong, several issues should be addressed for mechanistic depth and general relevance.

      (1) Although NINJ2 is shown to interact with LAMP1 and LAMP1 knockdown rescues ferritin levels, it remains unclear whether the NINJ2-LAMP1 interaction is required for lysosomal protection. The authors could: a) Map the NINJ2 domain required for LAMP1 interaction and test whether an interaction-deficient mutant fails to protect against LMP and ferroptosis. b) Rescue NINJ2 KO cells with wild-type versus mutant NINJ2 to establish causality.

      We thank the reviewer’s comments. Ongoing work in our laboratory is focused on elucidating the molecular mechanism by which the NINJ2-LAMP1 interaction regulates lysosomal membrane integrity and ferroptosis, and we anticipate reporting these findings in a future publication.

      (2) The conclusion that NINJ2 suppresses ferroptosis relies primarily on RSL3 and Erastin sensitivity. A direct assessment of ferroptosis would hence the study, such as:

      (a) Include ferroptosis rescue experiments using ferrostatin 1 or liproxstatin 1.

      (b) Assess lipid peroxidation directly (e.g., C11 BODIPY staining) to strengthen the ferroptosis claim.

      We thank the reviewer for this thoughtful comment. We agree that ferrostatin-1 or liproxstatin-1 rescue experiments, together with direct analysis of lipid peroxidation, would provide complementary evidence for ferroptosis. We will incorporate these additional experiments in future studies to further strengthen the mechanistic basis of our findings.

      (3) The manuscript discusses lysosomal ferritin degradation but does not directly examine NCOA4, a central mediator of ferritinophagy. It would be good to: a) Test whether NCOA4 knockdown rescues ferritin loss and ferroptosis sensitivity in NINJ2 KO cells. b) This would clarify whether NINJ2 acts upstream of canonical ferritinophagy pathways or via an alternative mechanism.

      We appreciate the reviewer's thoughtful suggestion. Defining the contribution of NCOA4 to NINJ2-mediated ferritin degradation is an important question that could further clarify the underlying mechanism. Addressing this issue will require a comprehensive set of additional experiments, which will be addressed in the future studies.

      (4) The study is entirely cell-based, despite references to inflammatory and tumor phenotypes in Ninj2-deficient mice. While not strictly required, even limited in vivo validation (e.g., ferroptosis markers or iron accumulation in existing Ninj2 KO tissues) would substantially strengthen the manuscript.

      We thank the reviewer for this insightful suggestion. We agree that in vivo validation of ferroptosis markers and iron accumulation in Ninj2-deficient tissues would further strengthen our conclusions. However, these experiments will require substantial additional investigation, which will be pursued in the future studies.

      (5) Finally, most imaging data (e.g., Galectin 3/LAMP1 colocalization, PLA signals) and immunoblot data are presented qualitatively. The authors should provide the qualifications of Western blots and other measurements.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) What mechanisms might underlie the regulation of LAMP1 transcript levels by NINJ2?

      A clear mechanism by which NINJ2 regulates LAMP1 transcripts has not been elucidated and warrants further investigation. Nevertheless, several possibilities can be considered. First, LAMP1 transcription is known to be regulated by TFEB (transcription factor EB), a master regulator of the lysosomal–autophagy pathway. Upon lysosomal membrane permeabilization (LMP), TFEB translocates to the nucleus and activates a broad set of lysosome-related genes, including LAMP1. Notably, phosphorylation of TFEB by mTORC1 at Ser211 inhibits its activity by preventing nuclear translocation. Thus, it would be of interest to determine whether NINJ2 modulates TFEB phosphorylation status and subcellular localization. In addition, the tumour suppressor p53 has been reported to engage in complex crosstalk with TFEB in regulating basal autophagy. Interestingly, p53 expression is increased in NINJ2-KO cells. It is therefore plausible that NINJ2 regulates TFEB activity through p53, or alternatively modulates the p53–TFEB signaling axis more broadly to maintain lysosomal integrity.

      (2) Does Ninjurin1 play a similar role in regulating lysosomal membrane permeabilization (LMP)?

      At this moment, it remains unclear whether NINJ1 plays a role similar to that of NINJ2 in regulating LMP. In fact, our previous studies demonstrated that NINJ2 physically interacts with NINJ1 and may antagonize NINJ1-mediated pyroptosis. Furthermore, NINJ1 has recently been reported to promote ferroptosis by interacting with the xCT cystine/glutamate antiporter (PMID: 38464226), a function that contrasts with the protective role of NINJ2 against ferroptosis identified in the present study. These findings suggest that NINJ1 and NINJ2 may have distinct, or even opposing, functions in regulating cell death pathways. Nevertheless, further studies are required to determine whether NINJ1 also participates in the regulation of LMP and to define its relationship with NINJ2 in maintaining lysosomal membrane integrity.

      (3) What are the potential clinical implications of these findings, particularly in the context of cancer progression or therapeutic targeting?

      Targeting NINJ2 may have important clinical implications in cancer therapy. Given its role in maintaining lysosomal membrane integrity, inhibition or loss of NINJ2 could promote lysosomal membrane permeabilization (LMP), thereby sensitizing cancer cells to ferroptosis through increased intracellular labile iron accumulation and disruption of redox homeostasis, ultimately enhancing tumor cell killing. As such, NINJ2 inhibition may represent a strategy to selectively destabilize lysosomal function in cancer cells and improve responsiveness to ferroptosis-inducing agents or other combination therapies that exploit oxidative stress vulnerabilities. Indeed, we previously developed a peptide that targets NINJ2. Whether this NINJ2-targeting peptide can sensitize cancer cells to ferroptosis therefore warrants further investigation.

      (4) The authors should provide quantification of all Western blot data throughout the manuscript to enhance data robustness and reproducibility.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Reviewer #2 (Recommendations for the authors):

      (1) Controls for knockdown efficiency of NINJ2 (Figure 2D) should be shown.

      NINJ2-KO MCF7 cells were generated previously and published in the article (PMID: 38325550) along with sequencing confirmation.

      (2) In Figure 2A legends, the concentration of LLOMe is 1mM or 1µM - need to be clarified?

      The concentration for LLOME is 1µM. This typo has been corrected.

      (3) In the figure legends section, "Figure 4" is missing.

      We thank the reviewer’s comment. Figure 4 has been added to the Figure legends.

      (4) Some description of NINJ2 ko cells generation should be included in the materials section.

      In the Materials and Methods section, we briefly described how these cell lines were generated and cited the original publication (PMID: 38325550)

      (5) The manuscript would benefit from a schematic model figure summarizing the proposed NINJ2-LAMP1-iron-ferroptosis axis.

      In the revised manuscript, we provided a model to elucidate the role of NINJ2 in modulating lysosomal membrane integrity and iron homeostasis.

      (6) Some sections of the Introduction are lengthy and could be streamlined to focus more directly on lysosomes and ferroptosis.

      We have streamlined the introduction.

      (7) Statistical methods should clarify whether data meet assumptions for Student's t-test and whether multiple comparisons were corrected where applicable.

      Statistical analysis has been added to the figure legends when applicable.

    1. eLife Assessment

      This valuable study uses genetic, behavioural, neuronal activity and cell-based assays to show that the Drosophila ionotropic receptor IR20a contributes to the detection of arginine and low-sodium salt, and suggests that peripheral receptor combinations may help integrate these nutrient cues. The evidence for a role for IR20a and associated co-receptors in these responses is solid, but the evidence that distinct receptor assemblies mediate receptor-level multimodal integration is incomplete, as this central model is inferred from functional data and leaves some discrepancies between behaviour, physiology and prior work unresolved. The study will be of interest to sensory neuroscientists and chemosensory biologists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigated the function of a Drosophila chemosensory receptor, IR20a, using genetics, neuronal histology, calcium imaging (in vivo and in cultured cells), and behavioral approaches. They provide evidence that this receptor functions in the detection of the amino acid arginine and of low salt (NaCl) concentrations, functioning in different combinations with "co-receptor" IRs, IR25a and IR76b.

      Strengths:

      The experiments are generally very well-performed and clearly presented, using established methodology. While, unsurprisingly, some puzzles remain (mentioned below), the work provides one of the clearest lines of evidence for the combinatorial coding of sensory information at the periphery through the combined action of distinct sets of chemosensory IRs.

      As taste neurons have long been recognized to express many different combinations of IRs and Gustatory Receptors (GRs), this study will be of interest to chemosensory biologists in general, particularly those studying invertebrate model systems (though co-expression of different families of taste receptors is a feature of mammalian taste cells).

      The precise molecular mechanisms remain unclear: there is no direct evidence here for protein complex formation (though this is likely), the stoichiometry of such complexes, or how subunits interact to confer or suppress sensory sensitivity. Nevertheless, these receptors, and the authors' success in reconstituting functionality in cultured cells, might make these a powerful model to explore such questions in the future.

      Weaknesses:

      Given the particular interest of the data from the heterologous reconstitution in cultured cells, the authors should be quite explicit about the nature of the quantification of the S2 cell responses. It is unclear whether the cited "n" refers to numbers of cells or something else, and whether all or only a fraction of (transfected) cells gave responses.

      There has been some prior work on the context-specific role of IR76b in amino acid-sensing and salt sensing by Ganguly and colleagues (Cell Reports 2017), who also implicated (weakly) a contribution of IR20a in contributing to the amino acid-sensing role. In that work, the authors focussed principally on the labellum and used electrophysiology rather than calcium imaging. The present manuscript appears rather dismissive of the earlier results (only mentioning them in the Discussion), and the authors could be a bit more generous about what was previously determined, where they have confirmed previous findings, where their results diverge, and why this might be. Similarly, the original functional analysis of IR76b (Zhang Science 2013) argued this was a low-salt sensor by itself, which is at least partially corroborated here; it remains unclear how this role relates to the low-salt detecting function of a potential complex of IR20a/IR25a/IR76b. It would be useful to have a summary model of the possible variety of complexes of IRs in different types of sensory neurons, as supported by the results in this and previous studies.

      The discord between the lack of requirement for IR20a for physiological responses to arginine in tarsi versus the necessity for behavioral responses is puzzling (though might reflect a labellar role for IR20a). There appears to be a trend of a decrease in calcium signal in tarsi to 100 mM arginine, which is the highest concentration tested (Figure 2A, C). Would a statistically significant decrease be observed with lower arginine concentrations? (A more substantial experiment would be to perform calcium imaging in the labellar IR20a neurons, or their axonal projections in the SEZ; this is not necessary, but the authors should at least acknowledge that their imaging of tarsal responses, while convenient, only examines a tiny fraction of the entire IR20a neuron population.

      The authors argue for synergistic responses to arginine and NaCl mediated by IR20a/IR25a. It's not clear to me to what extent there is synergism. In Figure 5A, 10 mM arginine or 10 mM NaCl individually lead to c.30-40% PER, and then when both are presented together in the "Mix" (presumably both compounds at 10 mM?), PER rises to c.60%. Is this really synergism, or rather simple additivity of behavioral responses to two attractive compounds? The authors could discuss this more thoroughly. Similarly, in Figure 5G the authors show that 50 mM arginine does not evoke a significant response in S2 cells expressing IR25a/IR20a, but in Figure 4 it would seem likely that a 50 mM dose would produce a significant response (the response to 25 mM arginine in Figure 4F is already elevated above the control, albeit not statistically significant). Is this just a batch effect of the experiments performed at different times (so they are not directly comparable)?

      The legend title to Figure 6 implies cooperation between tonic and state-modulated pathways, but I don't see specific evidence for "cooperation". Rather, as in the results text, they seem to work in parallel, so this analysis seems slightly peripheral to the main focus of the manuscript. It's ultimately unclear how the IR20a/IR76b/IR25a low salt sensor and the sensor containing IR56b functionally interact at the behavioral level. Here, a graphical summary, as mentioned above, of the different salt sensing neurons, the receptors they use, and the behaviors they control could be useful to establish the current knowledge and highlight open questions for the future.

    3. Reviewer #2 (Public review):

      Summary:

      This study identifies IR20a-expressing gustatory neurons in Drosophila as a multimodal sensory population integrating amino acid (arginine) and low-salt signals through combinatorial IR20a/IR25a/IR76b receptor assemblies. The proposed model of peripheral-level signal integration and synergistic enhancement of feeding preference is potentially significant, as it expands current understanding of gustatory coding beyond single-modality labeled lines.

      Strengths:

      Overall, the findings are conceptually interesting and suggest a novel framework for multimodal taste integration, but some mechanistic interpretations remain incompletely supported by direct evidence.

      Weaknesses:

      (1) Although the authors demonstrate co-expression of IR20a, IR25a, and IR76b in the same GRN population, this evidence is insufficient to support the proposed model of distinct receptors coexisting within individual neurons. Additional molecular or structural data would be required to distinguish whether these subunits assemble into complexes.

      (2) Given that IR76b has already been established as a sodium/salt sensing channel, the novelty of this study relies on the proposed role of IR20a in conferring multimodal integration and synergy. However, it remains unclear whether this represents a fundamentally new sensory mechanism or a re-interpretation of known IR76b-dependent salt responses in a different neuronal context.

      (3) Line 127:<br /> -The statement that there is no overlap between IR20a-GAL4 and GR64f-LexA or GR66a-LexA is not sufficiently supported by the presented imaging data. In particular, the resolution and clarity of the confocal images in Figure 1 appear suboptimal, making it difficult to confidently assess co-localization. The authors are encouraged to provide higher-resolution images or additional quantitative co-localization analysis to substantiate this conclusion.<br /> -In addition, the images shown in Figure 1 F1-F2 suggest possible partial overlap between IR20a and GR66a signals, which appears inconsistent with the authors' statement of no co-expression. This discrepancy should be clarified.

      (4) Lines 138-141:<br /> There appears to be a discrepancy between imaging and behavioral data: IR20a is reported as dispensable for arginine-evoked neural responses, yet IR20a mutants show significantly reduced attraction to arginine in behavioral assays. The authors should clarify how behavioral deficits arise in the absence of detectable changes in calcium imaging,

      (5) The manuscript proposes that IR20a functions in combination with IR25a to mediate multimodal detection of arginine and low NaCl. However, the specific role of IR25a in this context remains unclear.

      (6) The authors report that co-expression of IR20a and IR25a confers synergistic responses to combined arginine and NaCl stimulation, whereas the inclusion of IR76b abolishes this response (Figures 5E-K). This is an intriguing and potentially important finding; however, the mechanistic basis for this suppression is not clearly explained.

      (7) The authors propose that IR56b mediates state-dependent modulation of low-salt preference. However, the current data do not clearly distinguish whether IR56b acts as a real nutrient state sensor or just functions as a downstream modulatory component within a broader feeding circuit. Additional evidence linking IR56b activity changes to upstream metabolic state signals would be necessary to support the interpretation that IR56b functions as a primary state sensor.

      (8) The manuscript suggests that IR20a and IR56b define two parallel and functionally independent pathways mediating nutrient detection and state-dependent preference, respectively. However, this conclusion is not fully supported by the current dataset. While the two receptors are shown to be expressed in distinct neuronal populations, the possibility of indirect interactions or convergence at downstream circuit nodes has not been excluded. Given that both pathways ultimately influence feeding behavior, it remains possible that they converge at higher-order interneurons or shared neuromodulatory circuits.

      (9) In the state-dependent feeding assays (Figure 6), using H2O as a control introduces a severe masking effect. Salt-deprived flies actively suppress pure water intake to avoid osmotic shock, which artificially inflates the Preference Index (P.I.) for salt due to the denominator effect. To cleanly isolate salt preference from the thirst/osmotic drive, the authors will need to utilize an "isosmotic sucrose vs. isosmotic sucrose + salt" paradigm (Jaeger et al., 2018, eLife; Puri et al., 2026, PNAS).

    4. Reviewer #3 (Public review):

      Summary:

      Drosophila, like other animals, use sophisticated taste systems with specialized chemoreceptors to identify gustatory cues in their environment. Multiple gustatory cues associated with a food source are often encountered simultaneously, but our understanding of how this sensory information is detected and integrated remains incompletely understood. This valuable study investigates how salt, amino acids, or their combination are detected by specific combinations of peripheral Ionotropic Receptors, leading to behavioral attraction. The authors show that distinct combinations of IR76b, IR25a, and IR20a confer sensitivity to salt, arginine, or both. They also show striking evidence that cells co-expressing IR25a/IR20a display a synergistic response to a mixture of sub-activating concentrations of these tastants. Together, these experiments lead to the conclusion that combinatorial expression of different subunits and synergistic responses to taste mixtures facilitates integration of taste cues beginning in the periphery. However, in its current form, key methodological details are missing or inadequately described, which complicates interpretation. Additionally, characterization is heavily focused on the population of IR20a+ neurons in the tarsi, while the response properties of the newly-identified, functionally distinct population in the labellum are investigated only through behavioral analysis, limiting the description of potentially additional IR20a complexes. Ultimately, more in-depth biochemical characterization of the IR complexes described will be required to fully support the conclusion that combinatorial assembly of distinct IR20a receptors enables peripheral integration of taste mixtures.

      Strengths:

      The authors characterize the expression pattern of IR20a in the tarsi as well as in the labellum, a tissue for which IR20a expression has been a point of debate. Multiple levels of analysis, including behavioral assays, physiological recordings, as well as ectopic and heterologous expression systems, are used to characterize the response properties of different combinations of IR subunits, demonstrating remarkably consistent behavior of the IR-complexes across cell types. Well-controlled genetic analysis and the use of multiple behavioral assays provide additional support for their results, including the surprising demonstration of synergistic responses to mixtures of tastants that supports the idea of peripheral integration of gustatory inputs. This report also identifies a distinct IR, IR56b, required for starvation-enhanced responses to salt.

      Weaknesses:

      (1) The title states that IR20a integrates L-arginine and salt signals via distinct subunit assemblies, though the paper lacks direct evidence that IR20a serves as a multimodal tuning receptor in distinct functional assemblies. Heterologous expression shows that co-expression of IR20a/IR25a confers sensitivity to Arg, IR76b confers sensitivity to NaCl, and IR20a/IR25a/IR76b co-expression confers sensitivity to both Arg and NaCl. This seems to be interpreted to mean that all three subunits are assembling into a single complex. However, current results do not show any difference in salt response when IR76b is expressed alone compared to alongside IR25a+IR20a. Without more direct evidence for co-assembly of all three subunits, it is equally plausible that the responses observed represent activity of distinct IR25a/IR20a and IR76b receptors for Arg and salt, respectively. In this model, genetic disruption resulting in expression of either IR25a or IR20a alone with IR76b could disrupt its activity or membrane trafficking (as seen here and in previous studies) while co-expression of both IR20a and IR25a relieves this inhibition by sequestering IR20a/IR25a into a distinct complex from IR76b. Direct biochemical characterization, for instance in the form of co-immunoprecipitation or FRET, will be required to differentiate between these possibilities.

      (2) Key methodological details are missing throughout the manuscript. For instance, incomplete genotype and staining information is provided for images in Figure 1, making it difficult to interpret what is being shown. Additionally, for the calcium imaging methods, what is the imaging speed? How are max values calculated (is this the average of several images or just a single maximum)? How are ligands diluted and delivered to cells, and were they applied in a manner that allowed for subsequent washout?

      (3) The composition of the S2 imaging bath buffer requires clarification. As described, the bath buffer appears to lack any Ca2+ or other IR-permeable cations. If this is indeed the case, more detail should be provided about why this bath buffer was selected and what this means for the source and mechanism of calcium responses observed, since it would not reflect direct IR-mediated transduction. It is also notable that addition of water gives such a detectable change in the tarsal preps.

      (4) Visualization of IR20a driver activity in the labellum is interesting. Previous descriptions of labellar expression of IR20a range from no expression to expression in bitter neurons, so the current data linking IR20a to a different population of IR76b+ neurons warrants careful analysis in light of this discrepancy. However, some of the strongest presented evidence for expression is found in Figure 1, where the images are quite small, making it difficult to distinguish the morphology and sensillar innervation pattern of the cells labeled by the IR20a driver. In Figure 1A, several of the arrows do not appear to be associated with any visible fluorescence. It is similarly difficult to assess overlap. Including higher-resolution images and/or validating labellar expression, using antibodies, in situ hybridization, RT-PCR, or transcriptomics would strengthen these claims.

      (5) Similarly, Figure 2 shows that IR20a is not required for Ca2+ responses to AAs or KCl in the legs, but is required for behavioral preferences and PER responses in the labellum. This suggests that IR20a receptors may function differently in different tissues, though direct evidence is lacking. Calcium imaging from a weakly expressed driver may be difficult, but electrophysiological recordings from relevant labellar sensilla or ectopic/heterologous reconstitution of the molecular receptors found there would give important insights into the response properties of these other IR20a receptor type(s) and could provide evidence for additional IR20a-containing complexes. The current paper focuses exclusively on IR25a/IR20a/IR76b, which do seem to reliably reproduce the Arg/NaCl responses observed in the tarsi, but even for the tarsal neurons it is unclear that this represents an exhaustive list of all the relevant IR20a-interacting subunits coexpressed in these cells. For instance, Koh et al., 2014 (PMID: 25123314) found several additional IR driver lines, including IR56b, were active in the 5v/s tarsal sensilla.

    5. Author response:

      We thank the editors and three reviewers for their careful evaluation and constructive feedback. We are pleased that our identification of IR20a as a multimodal tuning receptor required for both low-salt and arginine sensing in a distinct gustatory neuron population was recognized as a valuable contribution to sensory coding. We agree with the major points raised and outline our planned revisions below, organized thematically.

      (1) Receptor assembly and integration model

      Our genetic, heterologous, and calcium imaging data show functional cooperation among IR20a, IR25a, and IR76b but do not demonstrate physical association. In the revision we will replace terms like “distinct subunit assemblies” and “peripheral integration” with more cautious language such as “functional receptor combinations.” We will state explicitly that our data cannot resolve whether the three IRs form a single heteromeric complex, and we will discuss the alternative possibility that IR76b and the IR20a/IR25a pair function as separate receptors within the same neuron. This interpretation better accounts for the response patterns observed upon co-expression. Throughout the text and in a revised summary figure, we will clearly differentiate elements directly supported by data from those that remain inferential.

      (2) Tarsal calcium imaging versus behavior

      The mismatch between tarsal calcium imaging and behavioral arginine responses will be addressed directly. We will explain that tarsal recordings sample only a small subset of IR20a neurons, whereas proboscis extension and feeding assays predominantly engage the more numerous labellar sensilla, whose neurons may carry different receptor compositions. We will generate labellar imaging where feasible; if additional functional data cannot be obtained, we will acknowledge this limitation rather than overinterpreting the tarsal results.

      (3) Synergy versus additivity

      We will discuss behavioral and cellular data separately. Recognizing that true synergy is difficult to demonstrate in feeding and proboscis extension assays, we will adopt conservative terminology when interpreting those experiments. For cellular data, where mechanistic insight is stronger, we will present the evidence for functional synergy and discuss why the outcomes may differ between the cellular and organismal levels.

      (4) Contextualizing prior IR76b and IR20a literatures

      We will expand the Introduction and Discussion to cite more fully the works establishing IR76b as a low-salt sensor and IR20a’s role in amino-acid sensing, including the earlier report that IR20a overexpression can inhibit IR76b-dependent salt responses. We will clarify how our single-cell imaging and loss-of-function data obtained in the native context refine models derived from ectopic expression. In its endogenous setting, IR20a marks neurons narrowly tuned to amino acids such as arginine, and IR20a is strictly required for low-salt detection. These findings contrast with the earlier view that IR20a functions broadly as an amino-acid sensor or as a salt-response blocker.

      (5) State-dependent modulation and IR56b

      Our data do not identify IR56b as the direct molecular sensor of internal state. We will reframe IR56b as a necessary component for state-dependent modulation of low-salt preference and retract any claim that it is the sensor itself. We will also clarify that IR20a and IR56b define genetically separable peripheral pathways that may converge on downstream circuits.

      (6) Methodological and presentation issues

      We will address the following points raised across reviews:

      (1) Co-localization: Higher-resolution confocal images and co-localization analysis will be provided.

      (2) Summary model figure: A new figure will illustrate the distinct functions of IRs in low-salt and amino-acid taste, clearly indicating which aspects are directly supported and which are inferential.

      (3) Feeding-assay control: We will either include an isosmotic sucrose control to avoid the water confound or explicitly discuss this limitation and temper the interpretation of feeding-preference results.

      (4) S2 cell quantification: Complete details on response criteria, responder fractions, and statistical reporting will be added.

      (5) Figure and supplementary corrections: All noted errors in figure legends, scale bars, citations, and supplementary-file mismatches will be fixed.

      We are confident these revisions will bring our mechanistic claims into close alignment with the evidence and substantially improve the manuscript. We again thank the editors and reviewers for their detailed and helpful comments.

    1. eLife Assessment

      This Review delineates postmenopausal age-related lobular involution (ARLI) in the breast that remains only partially resolved. In contrast to the conventional view of persistent lobules as passive residual structures, this work defines them as an actively maintained senescence-immune reserve niche. Inclusion of operational definitions for the reserve state in human breast tissue would strengthen the article. The work will be of interest to scientists working in the fields of breast disorders.

    2. Reviewer #1 (Public review):

      Summary:

      This Perspective proposes a conceptual model in which incomplete age-related lobular involution (ARLI) in the breast reflects an actively maintained senescent-immune "reserve niche," rather than simply passive failure of lobular regression after menopause. The authors aim to integrate breast cancer epidemiology, mammary gland biology, cellular senescence, immune surveillance, and comparative reserve-tissue systems to explain why persistent postmenopausal lobules are associated with increased breast cancer risk. The manuscript is ambitious, creative, and potentially useful in shifting attention from residual epithelial quantity alone toward the microenvironmental state of persistent lobules.

      Strengths:

      A major strength of the manuscript is its forward-looking synthesis. The authors bring together several areas that are often considered separately: ARLI as a tissue-level risk marker, inflammatory features of incompletely involuted breast tissue, senescence biology, macrophage-mediated remodeling, and the menopausal transition as a potential window of biological plasticity. The model is conceptually interesting and, if supported by future evidence, could stimulate new approaches to risk stratification and prevention focused on the perimenopausal period.

      Weaknesses:

      However, the current manuscript often presents the proposed model with more certainty than the available evidence supports. The evidence clearly supports associations among incomplete ARLI, inflammatory or immune features, and breast cancer risk, but it does not yet demonstrate that senescent cells maintain persistent lobules, that immune clearance failure causes incomplete involution, or that a self-sustaining senescent-immune "niche lock" exists in human breast tissue. Much of the mechanistic framework is extrapolated from other tissues, postpartum involution, or general senescence biology. These are reasonable sources for hypothesis generation, but the manuscript would be stronger if it more clearly distinguished established observations from inference and speculation.

      The senescence component of the model requires stronger and more direct support. Several claims about senescent burden in the aging breast appear to rely on general senescence literature or mammary aging studies that do not directly demonstrate senescence in persistent human TDLUs. This distinction is important because the manuscript's central model depends on senescent cells being spatially and functionally linked to incomplete ARLI.

      The epidemiologic evidence also requires a more balanced treatment. Although several studies support incomplete ARLI as a breast cancer risk-associated phenotype, other cohorts and quantitative approaches have reported attenuated or null associations. This mixed evidence is acknowledged, but it is treated largely as a caveat rather than incorporated into the central argument. For readers, this uncertainty is important for interpreting the strength and generalizability of the proposed model.

    3. Reviewer #2 (Public review):

      Summary:

      This review constructs a novel theoretical framework to elucidate incomplete postmenopausal age-related lobular involution (ARLI) in the breast. Differing from the conventional view of persistent lobules as passive residual structures, the work innovatively defines them as an actively maintained senescence-immune reserve niche. It comprehensively integrates multidisciplinary evidence from breast epidemiology, stromal biology, cellular senescence and immune surveillance, as well as cross-tissue research findings, and identifies menopause as a core biological turning point regulating ARLI and relevant breast cancer risk, providing a new theoretical perspective for subsequent breast cancer risk assessment and preventive intervention research.

      Strengths:

      This study presents an original, logically rigorous, and well-organized research hypothesis. It innovatively breaks through the traditional cognitive perspective of ARLI and adopts a multidisciplinary and cross-tissue analytical approach to sort out relevant biological mechanisms systematically. The proposed theoretical framework is insightful, with good theoretical innovation and potential translational value for guiding breast cancer risk evaluation and targeted prevention strategies.

      Weaknesses:

      The manuscript currently serves primarily as a conceptual framework rather than a rigorously evidenced synthesis. Its central argument relies heavily on cross-sectional correlations and theoretical analogies to other organ systems, lacking operational definitions for the reserve state in human breast tissue.

    4. Author response:

      We thank the editors and reviewers for recognizing the originality and potential value of the proposed framework. The reviews rightly ask us to distinguish three things more sharply: what is established directly in human breast tissue, what is inferred from mammary and aging studies, and what remains hypothesis. We agree, and the revision will make that distinction explicit throughout.

      We will define the proposed reserve state operationally and specify the findings that would distinguish active niche maintenance from passive persistence. Throughout, we will treat passive persistence as a legitimate competing hypothesis rather than a settled question. Heterogeneity and immune or inflammatory associations will be presented as consistent with active maintenance and causal directionality, not as establishing them. We will also integrate the mixed epidemiologic evidence more centrally into the argument, rather than treating it as a caveat.

      We will reassess the evidence for senescence in the aging breast and describe it more precisely, correcting or narrowing statements that outrun the data. The figures will be revised so that observed inflammatory and immune-regulatory features are clearly separated from proposed senescence- and SASP-mediated mechanisms. We will also clarify the limits of our analogies: postpartum involution and cross-tissue reserve systems will be presented as sources of candidate mechanisms and testable predictions, not as direct evidence for the proposed mechanism in human ARLI. Finally, we will frame the translational implications as contingent. They depend on first demonstrating that senescent cells are enriched near persistent lobules, identifying the relevant cell types and immune states, and establishing causal relevance.

      We appreciate the reviewers' constructive suggestions. We believe these revisions will preserve the conceptual contribution of the model while making its evidentiary status, the alternative explanations, and its falsifiable predictions substantially clearer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

      Thank you for your critical review and insightful comments.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      (b) Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.

      - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function. Thank you.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

      Thank you.

      Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

      Thank you for your critical review of our manuscript and for the very constructive comments and suggestions to improve the manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (There is some redundancy below with my comments in the public review, but I've included here for clarity and further explanation.)

      The authors have addressed my primary concern, as treatment with PLK inhibitor reduces phosphorylation of T301, while also providing some comment on relative impact of PLK-mediated KIN-G phosphorylation.

      It is notable that phosphorylation of T301 was reduced by only ~27%, while phosphorylation of S569 was reduced by ~100% in the presence of PLK inhibitor. The authors note that this might be explained by slower dephosphorylation of T301. In the absence of a phenotype, and with cell doubling continuing unabated in presence of the inhibitor, it is intriguing that more loss is not observed. An alternative explanation is that an alternate kinase might also be able to phosphorylate T301 in the absence of PLK activity, and the authors should consider that possibility.

      We added a sentence in the main text to suggest an alternative explanation.

      The model shown in figure 8 should include a dephosphorylation step, per the authors comments in the text regarding the small fraction of T301 that is phosphorylated and proposal of a phosphorylation/dephosphorylation cycle. The Fig 8 legend needs to have some commentary on the role of phosphorylation, as phosphorylation is the center point of this paper.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function.

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Fig 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Fig 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces motility of KIN-G. These statements are contradictory.

      (3) p. 8, and Fig 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      (4) Fig 7B. Please explain labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      (5) Fig 4, 5, and 7: "% Cells" is reported. Please indicate what number of cells total were examined.

      These minor comments have already been addressed in the previous revision.

    2. Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

    3. Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations.

    4. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

    5. eLife Assessment

      This important study provides new insights into the regulation of cell organization and division in Trypanosoma brucei through the phosphorylation-dependent control of a kinesin motor protein by a polo-like kinase. The authors present convincing evidence, combining rigorous biochemical, cell biological, and imaging analyses, demonstrating that phosphorylation modulates kinesin localization and function, thereby influencing cellular organization and cytokinesis. The findings advance our understanding of the molecular mechanisms governing trypanosome cell division and will be of broad interest to researchers studying trypanosomes, cytoskeletal regulation, and eukaryotic cell division.

    1. eLife Assessment

      This study provides an important contribution to retinal regeneration research by using overexpression of pro-neural factors to reprogram fetal human RPE cells into retinal neurons. The authors provide solid evidence of fetal RPE reprogramming into neural and photoreceptor-like states using scRNA-seq and imaging validation; however, there are concerns regarding comparisons between the effectiveness of different combinatorial transcription factor codes.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors identified transcription factor combinations capable of inducing retinal neuronal programs in cultured fetal human retinal pigment epithelial (RPE) cells. Using a pooled screening strategy, single-cell RNA sequencing, lineage barcoding, and immunohistochemical analyses, they identified ASCL1 and NEUROD1 as an effective combination for inducing retinal neuron-associated transcriptional states. This work aims to advance the development of therapeutic approaches for retinal regeneration by exploring the plasticity of RPE cells.

      Strengths:

      A major strength of the study is the comprehensive experimental design. The combination of transcription factor screening, lineage tracing, single-cell transcriptomics, and molecular validation provides a detailed characterization of the cellular responses to reprogramming factor expression.

      Weaknesses:

      All experiments were performed using fetal human RPE cells. Because fetal RPE remains relatively immature and retains proliferative capacity, it remains unclear to what extent the observed responses reflect true reprogramming of differentiated RPE cells versus activation of developmental plasticity already present in fetal tissue. The absence of adult human RPE controls limits assessment of the generality and translational relevance of the findings.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting study that explores how human RPE could be used as a source for new retinal neurons. This is a welcome addition to the field of retinal regeneration, which is currently focused almost exclusively on the regenerative capacity of Müller glia cells. The line of inquiry is firmly rooted in findings from amphibian and embryonic chick model systems and advances a fetal human retina RPE-based screening system as a rich resource for insights into human RPE biology, including as a potential stem cell source.

      The authors investigate the potential of fetal human RPE cells to be reprogrammed into retinal neurons using overexpression of pro-neural factors. While this is a critical knowledge gap in the field of retinal regeneration with significant promise for developing regenerative therapies, several methodological concerns impact the interpretation of results. Firstly, while the authors sought to evaluate factors that enhance RPE reprogramming when co-expressed with ASCL1, nearly all co-expression constructs tested failed to achieve appreciable expression of ASCL1, leaving a central hypothesis of this study largely untested (Major concern 1). Second, although the authors were able to detect a cluster of photoreceptor-like cells in their screen, they were unable to identify which reprogramming construct generated this cluster (Major concern 2). Finally, an essential control that definitively demonstrates the value of combinatorial transcription factor reprogramming is missing (Major concern 3).

      In summary, the authors establish a valuable new paradigm for culturing and reprogramming fetal human RPE, and even more importantly, demonstrate successful reprogramming to neural fates. However, the discussion and interpretation of results needs to be modified significantly to make it clear that (i) the outcome of many co-expression paradigms remains effectively unknown/untested due to failed over-expression of ASCL1, and that (ii) the reprogramming construct giving rise to photoreceptor-like cells could not be conclusively identified from their initial screen.

      Strengths:

      (1) Powerful new screening system advanced for exploring the regenerative potential of human fetal RPE cells.

      (2) Co-expression vector system for testing additive effects of proneural transcription factors.

      Weaknesses:

      Major concerns:

      (1) The authors executed a screen for combinations of factors that can enhance ASCL1-mediated reprogramming of RPE into retinal neurons. However, the expression level of ASCL1 was remarkably low in virtually all co-expression paradigms (see Figure 3C). Notably, the reprogramming combination with the highest potency (ASCL1 + NEUROD1) was also the one exhibiting the highest level of ASCL1 expression. The "failed" reprogramming of most of the co-expression constructs (ASCL1+LMO1, ASCL1+EZH2, and ASCL1+RAX2) is potentially a false negative resulting from low transgenic expression of ASCL1.

      (2) The authors' interpretation is that the overexpression of NeuroD1 and Ascl1 generated a new cluster that expressed markers of photoreceptors such as RXRG and RCVRN (see Figure 3D). However, there does not actually appear to be any overlap between the ASCL1+NEUROD1 cluster (orange dots, left panel) and the cells expressing markers of photoreceptors (yellow/green/purple?/black? dots, right panel; yellow being ASCL1-EZH2, green being ASCL1, purple being ASCL1-and black being control - though color coding here is admittedly somewhat confusing). Thus, the photoreceptor-like cluster of interest actually seems to correspond to gray cells that were unmapped/exposed to an unknown programming cocktail. So, it remains completely unknown which reprogramming construct generated this cluster.

      (3) To conclusively establish the additive role of NEUROD1 in reprogramming, it would be prudent to compare ASCL1 + FA directly to ASCL1+NeuroD1+FA. This control was not included but is needed for a more complete interpretation of results.

    1. eLife Assessment

      This important study presents an ultrastructural atlas of extracellular vesicles and non-vesicular particles in Drosophila olfactory sensilla, using cryofixation-based serial block-face electron microscopy to catalogue roughly 7,800 particles across four antennal regions. The evidence is compelling: the cryofixation minimizes fixation artifacts, and the rigorous, well-validated methodology and morphometric analyses provide strong support for the structural classifications and spatial distributions, though some conclusions require moderation or additional evidence. The work provides a potentially foundational resource for researchers studying extracellular-particle signaling in sensory organs.

    2. Reviewer #1 (Public review):

      Summary:

      Using cryofixation and serial block-face electron microscopy (SBEM), P. Vijayakumar and K. Cauwenberghs characterize extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) within native Drosophila olfactory sensilla. The study provides a unique and valuable dataset comprising approximately 7,800 extracellular particles, systematically describing their morphology, size, density, and distribution. The ultrastructural analysis across different sensillum classes offers insights into the potential biogenesis and functions of these extracellular particles.

      Strengths:

      Cryofixation preserves EVs and NVEPs within native tissue conditions. The detailed quantification of a very large dataset provides a unique source of information on extracellular particle number, categories, and distribution. The expertise of the group in the method and the tissue explored, as well as their detailed quantification, provide confidence in the dataset and observations.

      Weaknesses:

      Major comments

      (1) As the authors state, the rare observation of EV budding or MVB release events suggests that these are transient processes, whereas EVs and NVEPs are retained for relatively long periods within the sensillum lumen. The current analyses may overinterpret steady-state vesicle abundance as differences in vesicle production.

      (2) Given the above conclusion, differences between sensillum classes may be somewhat overstated.<br /> a) The absolute number of EVs and NVEPs per sensillum is highly variable, even within the same sensillum class (Figure 3D). For example, a substantial proportion of coeloconic sensilla have an empty lumen (Figure 3C). Consequently, expressing the data as ratios or proportions (Figures 3B, 4C, and 4E) may exaggerate differences between sensillum classes and should therefore be interpreted with caution.<br /> b) The rate of EV/NVEP production is unknown. For a similar rate of production across sensillum classes, Figure 3E suggests that the differences in lumen morphology and size may largely explain variation in EV density and distribution.

      That said, I agree that ab1 sensilla display a striking enrichment of large cargo-filled EVs compared with the other sensillum classes, while coeloconic sensilla display enrichment in small dense filled EVs (Figure 4C). Together, large and cargo-filled EV observation provides strong support for differences in EV biogenesis between ab1 sensilla and the other sensillum classes. In that context, I also think the EV size distribution shown in Figure 4 - Figure Supplement 1 should be moved into the main figure, as it demonstrates that the majority of EVs in the ab1 lumen are relatively large and are therefore likely to represent microvesicles. Could you clarify why ab1 sensilla are only included in Figure 4 and not analysed in Figure 3?

      (3) Approximately 10% of ORNs appear to be degenerating in 6-8-day-old flies, which seems unexpectedly high. This contrasts with the relatively infrequent occurrence of auxiliary cell apoptosis or complete sensillum degeneration. In these "degenerating ORNs", the authors state that the hallmarks of ORN apoptosis are restricted to the dendrites. As hallmarks, they state dendrite truncation, fragmentation and blebbing. Rather than apoptosis, I wonder whether these observations might instead represent ciliary truncation and ectosome shedding, followed by degradation of the shed ciliary membrane into EVs. Ciliary truncation and ectosome shedding, followed by ciliary regrowth, are dynamic processes that have been described across multiple species. This interpretation could explain large EVs that remain in the lumen long after the cilium has regenerated. It would reconcile this article with the general agreement that cilia are a prime site for the budding of EVs across species. Additional evidence supporting apoptosis of the ORNs would help distinguish between these possibilities. Otherwise, I believe the author should reconsider their interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      This paper presents a large structural survey of extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) in the olfactory sensilla of Drosophila melanogaster. Using high-pressure freezing and serial block-face SEM, the authors avoid many of the artifacts associated with conventional fixation and analyze more than 7,800 particles across 352 sensilla. The manuscript maps the distribution of these particles, describes their morphological heterogeneity, and examines their likely origins across different sensillum classes in both normal and degenerating tissue.

      Strengths:

      The strongest aspect of the paper is the imaging. Preservation is painstakingly controlled. The cryofixation appears to preserve the sensillum lymph in a more convincing native state than standard preparation methods, giving this work gravitas. Further, the authors characterized thousands of particles, further making this data strong.

      The figures are strong. They are clear, easy to read, and generally well designed; I think people will use this paper as a model for how to present complex data in a concise and straightforward manner. The manuscript is careful in how it presents the dataset and does not overinterpret the descriptive observations. As an ultrastructural resource, this paper will be useful to the field. The identification of auxiliary support cells as major secretory sites, together with the striking accumulation of EVs in degenerating tissue, will provide a useful starting point for future work.

      Weaknesses:

      The main point that could use more clarification is the vesicle categorization. In particular, the distinction between "dense," "cargo-filled," and "double EVs" is not always easy to follow from a biological perspective. Some additional discussion of how the authors think these categories relate to one another, and whether they are intended as purely morphological groupings or as distinct biological classes, would strengthen the manuscript.

    4. Reviewer #3 (Public review):

      Using cryofixation-based serial block face electron microscopy of several subregions of the Drosophila antenna, the authors segment and assemble a high-resolution atlas of the anatomical structure, density, and spatial distribution of extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) in different Drosophila olfactory sensilla types. This systematic and thorough description is an important prerequisite to understanding the function of extracellular particles in intercellular signaling in the nervous system. Additionally, they describe examples of putative biogenesis events (budding/fusion), as well as neuronal and axonal cell degeneration events and measure the changes in particle accumulation in these altered microenvironments.

      Overall, this is a significant and comprehensive analysis and represents an invaluable resource to this burgeoning field. The authors assemble an important dataset and their claims match the level of evidence provided.

      Strengths:

      (1) The authors use segmentations from four different patches of the antenna to provide a systematic ultrastructural survey of extracellular particles in native insect sensilla. The dataset captures the diversity of sensilla types and reconstructs ~7800 particles.

      (2) We commend the authors for making the EM volumes available in the public Cell Image Library with accession numbers. It would be helpful to the community to also make the segmentations for this great resource easily accessible.

      (3) The sample preparation technique appears to minimize typical artifacts associated with chemical fixation, as evidenced by the high reported sphericity of EVs.

      Specific points:

      (1) In Figure 3C, the authors should include a continuous measure of particle distribution in the sensilla. Currently, the authors define three categories of particle localization. In the five examples shown in Figure 3A, the spatial distribution of these particles appears quite distinct across classes. For example, EVs/NVEPs in large and small basoconic sensilla are largely restricted to the area proximal to the base, with a limited number located more distally. In contrast, intermediate sensilla show a marked concentration of particles more distally.

      (2) The conclusion of different EV ratios across sensillum classes stems from a Kruskal-Wallis of p = 0.0476, with none surviving pairwise comparisons. This is not a strongly supported conclusion and is probably better characterized as a trend.

      (3) The statement "selective enrichment of large, cargo-filled vesicles within the ab1 lumen suggests specialized EV-mediated communication adapted to the coordination demands of this neuronal population" seems speculative for a Results section without supporting functional evidence. It would seem better suited for the Discussion.

      (4) Figure 4: Criteria for defining the classes of EVs.<br /> a) The authors should explain the rationale for classifying EVs using relative density rather than absolute density? We would expect EVs with similar contents to have similar electron density (similar darkness in the images). Would classifying them relative to the background, which itself might vary across sensilla or regions, create a possible confound, especially when comparing across sensillum classes?<br /> b) The two example images (in Figure 4A) of the cargo-filled EVs appear to have different densities themselves. Do the cargo-filled ones also display systematic differences in density and, if so, why is this another class instead of being a subcategory within the dense and lucent classes (i.e. dense with/without cargo, lucent with/without cargo)? The dense and lucent classes are defined by their density, whereas this is a more structural property.<br /> c) Regarding "Double" and "Ball-and-Socket" EVs, does the density vary between the two particles involved (e.g., does the inner structure consistently differ in density from the outer)?

      (5) What was the rationale for the 200μm and 1000μm size cutoffs? A continuous distribution of maximum particle sizes would provide a clearer understanding of the data.

    1. eLife Assessment

      This study addresses a valuable question with implications for the development of EEG-based neurofeedback interventions for pain. However, the strength of evidence is incomplete because the principal conclusions rely on post hoc responder analyses and methodological ambiguities that weaken the support for the claimed causal relationship between gamma modulation and pain reduction.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate whether EEG neurofeedback (NFB) can be used to increase spontaneous parieto-occipital gamma oscillations and thereby reduce experimentally induced pain. Healthy participants were randomly assigned to active or sham neurofeedback and completed three consecutive neurofeedback blocks with concurrent EEG measurements and phasic painful stimulation. The study addresses a relevant question regarding the causal role of spontaneous gamma oscillations in pain perception and the potential of neurofeedback as a non-pharmacological pain intervention. While the reported findings appear consistent with an association between increased gamma power and reduced pain in a subset of participants, the current analyses do not provide sufficient support for the strong causal conclusions drawn by the authors.

      Strengths:

      (1) The study addresses an important and timely research question with potential implications for EEG-based neurofeedback approaches to pain modulation.

      (2) The sample size is relatively large for an experimental EEG neurofeedback study and includes a sham-control condition.

      (3) The manuscript is generally well written and clearly organized.

      (3) The authors address an important methodological concern regarding EMG contamination of gamma-band activity by including additional EMG recordings in a subset of participants.

      Weaknesses:

      (1) The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      (2) The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as "responders" based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      (3) The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      (4) Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      (5) The neurofeedback implementation also raises questions. Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      (6) The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigated whether neurofeedback (NFB) training targeting spontaneous gamma oscillations (30-60 Hz) at the parieto-occipital region (Pz electrode) could reduce experimental pain perception. They randomized 88 healthy participants to active or sham NFB groups across two cohorts (44 each). Active NFB consisted of real-time feedback based on participants' own gamma power; sham NFB consisted of the preceding participant's gamma power. Participants completed three ~16-min sessions, and approximately 52% of active NFB participants showed increased gamma power in session 3 and were considered responders. Analyses restricted to these 23 responders (matched with 23 sham controls) showed reduced pain intensity, unpleasantness, and laser-evoked potential (LEP) amplitudes, with a significant negative correlation between gamma power and pain intensity after session 3.

      Strengths:

      (1) The distinction between spontaneous and stimulus-evoked gamma oscillations in pain processing is theoretically important.

      (2) The rationale for targeting spontaneous gamma via NFB is clearly articulated.

      (3) The study was sham-controlled, and the blinding was adequate.

      (4) The authors commendably ran a second cohort (n=44) with simultaneous posterior neck EMG recording to address the critical concern of muscle artifact contamination of gamma, in response to a previous review

      Weaknesses:

      (1) The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increased gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      (2) Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group -> gamma change -> pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      Minor points:

      (1) For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      (2) Was baseline gamma power comparable between groups?

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to test whether spontaneous gamma-band oscillations over the parieto-occipital region can be volitionally upregulated using EEG neurofeedback, and whether this upregulation reduces subsequent pain perception and nociceptive-evoked brain responses. Gamma-band activity has been repeatedly associated with pain processing, but most available evidence remains correlational, and previous attempts to modulate pain-related gamma activity using non-invasive stimulation have not produced robust analgesic effects. The present study therefore addresses an important question: whether real-time neurofeedback may provide a more effective way to train endogenous gamma activity and thereby influence pain.

      Strengths:

      A major strength of the study is the use of an active/sham neurofeedback design. The authors also combine subjective pain ratings with laser-evoked potentials, which provides converging behavioural and neurophysiological outcome measures. The manuscript is clearly written overall, and the study addresses a question of broad interest for pain neuroscience and neurofeedback research.

      Weaknesses:

      A number of aspects limit the strength of the conclusions. The first and most important issue concerns the interpretation of scalp gamma-band activity. Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      Second, the evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      A third limitation concerns the control condition and blinding. Participants were reportedly blinded to group allocation, and the credibility ratings appear similar between groups, which is reassuring. However, it is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

    5. Author response:

      Reviewer #1 (Public review):

      R1-Q1: The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      We thank the reviewer for raising this important and thoughtful point. We agree that the relationship between spontaneous gamma oscillations and pain perception remains a matter of active debate, and we will moderate these statements throughout the manuscript. We will acknowledge the ongoing debate and present the gamma-pain relationship as an active area of investigation rather than settled fact.

      R1-Q2: The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as 'responders' based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      We appreciate this careful critique. We wish to clarify the rationale behind our analytical approach and address the concern.

      A well-established finding in the neurofeedback literature is that a substantial proportion of participants are "non-learners" — individuals who, despite receiving real feedback, fail to achieve effective control over the targeted neural activity. This is not a failure of the intervention, but reflects individual differences in neurofeedback learning capacity. Our core research question is therefore: "Among individuals who can successfully learn to upregulate gamma oscillations, does this upregulation reduce pain perception and nociceptive brain responses?"

      To address the circularity concern and improve transparency, we will make the following revisions:

      - We will reframe the wording from "NFB increases gamma and reduces pain" to "Successful gamma upregulation via NFB is associated with reduced pain in those who achieve it." All causal language will be replaced with appropriately cautious, correlation-based terminology.

      - We will report full-sample results for completeness.

      R1-Q3: The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      We thank the reviewer for requesting this clarification. We will specify the exact criterion in the revised manuscript, i.e., a participant was classified as a responder if their post-intervention gamma power minus pre-intervention gamma power was positive (i.e., an increase in gamma power following the neurofeedback intervention).

      R1-Q4: Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      We thank the reviewer for these constructive suggestions. We will supplement and refine the methodological details in the revised manuscript, and we will make the analysis code and the experimental program (including the custom neurofeedback software) publicly available via an open repository.

      R1-Q5: Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      We will discuss the limitations of the discontinuous visual feedback and the viewing distance in the revised manuscript. We acknowledge these as valid methodological concerns and will address them as limitations in the Discussion.

      R1-Q6: The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

      We thank the reviewer for pointing out that the description of the muscle-confound analysis was insufficiently detailed. In the revised manuscript, we will clarify the EMG analysis procedures and explicitly report: (a) the exact number of participants contributing to the EMG analysis; (b) the cohort from which they were drawn; and (c) whether responder selection was performed before or after restricting the sample for EMG analysis.

      Reviewer #2 (Public review):

      R2-Q1: The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increase gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      As detailed in our response to R1-Q2, the responder analysis reflects a conceptually motivated subgroup defined by successful neurofeedback learning — a well-documented challenge in NFB research where many participants are non-learners. Our central question is whether successful gamma upregulation (among those capable of achieving it) is associated with pain reduction. We will make this rationale explicit in the revised manuscript. We will also: (a) transparently report full-sample results; (b) discuss potential biases introduced by subgroup selection (e.g., differences in attention, self-regulation); and (c) acknowledge the lack of preregistration for the responder analysis.

      R2-Q2: Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group → gamma change → pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      We will substantially revise the language throughout the manuscript to avoid causal claims. We also plan to conduct a formal mediation analysis (NFB group → gamma change → pain change) to test whether changes in gamma activity statistically mediate the observed pain reduction.

      R2-Q3: For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      We thank the reviewer for raising this point, and we will clarify both points in the revised manuscript. (a) Because group assignment was randomized, the first participant could in principle have been assigned to the sham group. To prepare for this possibility, we collected EEG data from one participant in advance (equivalent to pilot data) to serve as the sham feedback signal, had the first participant been assigned to the sham group. In the actual experiment, however, the first participant was randomly assigned to the active group, so this pre-collected dataset was never used. (b) We will also compare the discrepancy between actual gamma power and the sham feedback signal in the sham group, and report this result in the revised manuscript.

      R2-Q4: Was baseline gamma power comparable between groups?

      We thank the reviewer for this suggestion. We will report and compare baseline gamma power between the active and sham groups in the revised manuscript.

      Reviewer #3 (Public review):

      R3-Q1: Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      We thank the reviewer for this important suggestion. We will revise the manuscript to discuss more explicitly the inherent difficulty of recording pure gamma-band oscillations with scalp EEG. We will acknowledge that scalp gamma is not reliably observable in all participants, is vulnerable to contamination from multiple muscle sources, and that a single posterior neck EMG channel provides only limited control. We will also discuss the possibility that changes in posture, facial tension, breathing, and arousal may contribute to high-frequency scalp activity.

      R3-Q2: The evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      We appreciate the reviewer's careful consideration of this point. We will temper our conclusions, presenting the current evidence as demonstrating the feasibility of gamma-band neurofeedback training in a subset of participants and an association between successful training and pain reduction, rather than a demonstrated causal relationship. We will discuss individual differences (attention, suggestibility, relaxation ability, neurofeedback learning capacity) as potential confounds that cannot be fully disentangled from gamma-specific effects.

      R3-Q3: It is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      We will clarify that the study employed a single-blind design: participants were unaware of their group assignment. We will state this clearly in the revised manuscript.

      R3-Q4: The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      We will provide a stronger rationale for the selection of the two neurofeedback video scenarios. We will also place greater emphasis on the pain reduction observed in both groups, acknowledging the substantial nonspecific analgesic effects associated with the procedure.

      R3-Q5: The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      We thank the reviewer for this insightful comment. We will expand the discussion of why neurofeedback may produce effects beyond those achieved by gamma-frequency tACS. In particular, we will elaborate on the potential contributions of volitional control, attentional engagement, immersion, expectation, and self-regulatory processes, and discuss how these factors may be central to the observed pain reduction rather than merely incidental to the gamma modulation.

      R3-Q6: Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

      We thank the reviewer for this balanced assessment. We fully agree that the conclusions should be tempered, and we will revise the manuscript accordingly to reflect that the current evidence supports feasibility and association rather than established causal efficacy.

      Summary

      In summary, the planned revisions include:

      (i) full-sample results reported;

      (ii) moderating causal language throughout and adding a formal mediation analysis;

      (iii) clearly specifying the responder criterion and adding this to the preregistration;

      (iv) providing complete methodological documentation and publicly releasing all analysis code;

      (v) expanding the Discussion to address limitations regarding scalp gamma measurement, EMG control, visual feedback, viewing distance, single-blind design, NFB scenario rationale, and nonspecific effects;

      (vi) adding analyses on baseline gamma comparability and sham-feedback discrepancy.

    1. eLife Assessment

      This study provides important evidence that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its traditional acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. Together, the multidisciplinary results provide a coherent and convincing chain of evidence for an olfactory component of predator detection in this bat-insect system. The work will be of broad interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

    3. Reviewer #2 (Public review):

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank you and the reviewers for the thoughtful evaluation of our manuscript and for the constructive comments and important suggestions. We are encouraged by the recognition that the study is valuable and of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution. We are also grateful that the reviewers acknowledged the integrative approach of our work, including fecal metabarcoding, behavioral assays, electrophysiological recordings, chemical analyses, and field observations.

      We have carefully considered all comments and have revised the manuscript accordingly. In particular, we have made the following major revisions:

      (1) We clarified the biological origin of limonene and expanded the discussion of possible bat-associated sources, including snout secretions and microbial contributions, while avoiding overinterpretation of limonene as an exclusively endogenous mammalian compound.

      (2) We added more detailed descriptions of contamination controls, including instrument cleaning procedures, materials used for housing and handling, and blank-control results, to address the concern that limonene could have originated from human-associated or environmental contamination.

      (3) We added individual-level odor data and revised the presentation of terpenoid profiles to show more clearly where limonene was detected and how its relative contribution varied among samples.

      (4) We clarified the rationale for the concentration choices used in the electrophysiological and field experiments, emphasizing that these assays were designed to test physiological detectability and functional sufficiency rather than to reproduce exact natural emission concentrations or determine response thresholds.

      (5) We revised our interpretation of limonene more cautiously. We now state that limonene is sufficient to trigger avoidance responses, while we acknowledge that its natural ecological specificity, concentration dynamics, and interactions with other bat odor components require further investigation.

      (6) We corrected the terminology related to limonene enantiomers. Because our GC-MS and GC-EAD analyses did not use an enantioselective column, we replaced “(-)-limonene” with “limonene” throughout the manuscript, figures, legends, and supplementary materials.

      (7) We improved the figures and supporting materials by moving the figure illustrating predator-prey relationship into the main text, adding electrophysiological traces from all tested crickets, revising figure legends, clarifying the rationale for comparing bat body odor with air controls, and providing additional chemical-identification details in the supplementary materials.

      (8) We checked and clarified statistical annotations, including the exact adjusted P-values for relevant comparisons.

      We believe these revisions have substantially improved the clarity, rigor, and balance of the manuscript. We hope that the revised manuscript and the detailed point-by-point responses satisfactorily address all concerns raised.

      We are confident that our study represents an important contribution, as it fundamentally expands the traditional acoustic-centred view of bat–insect interactions by demonstrating that crickets can use olfaction as a complementary sensory modality to detect bat odors and initiate avoidance behavior. The multidisciplinary evidence, integrating behavioral assays, electrophysiology, chemical profiling, and field validation, provides convincing support for this novel olfactory pathway.

      eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

      We sincerely thank the editors for the positive and constructive assessment of our work. We greatly appreciate the recognition of our study’s value and its potential interest to multiple research communities.

      We have carefully considered all comments and have revised the manuscript accordingly. The revisions include clarifications of limonene’s biological origin and ecological specificity, strengthened contamination controls and individual-level odor data, clearer rationales for experimental concentration choices, and more cautious interpretation throughout the manuscript. All changes are addressed in detail in our point-by-point responses below.

      Importantly, the central conclusion of our study—that crickets can detect and avoid bat odor through olfaction, and that limonene is sufficient to trigger avoidance responses in both laboratory and field settings—remains robustly supported by the multidisciplinary evidence we present. We believe the revised manuscript now provides a clearer, more balanced, and scientifically rigorous account of our findings, and we hope it meets the standards of eLife.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are delighted that you found the study interesting and enjoyable. We also appreciate your kind remarks about the clarity of the manuscript, the logical flow of the experiments, and the quality of the figures.

      We are especially grateful that you recognized the value of our integrative approach. Combining behavioral assays, electrophysiology, chemical analysis, and field observations was central to our study design, and we are pleased that this approach resonated with you.

      You provided an accurate and comprehensive summary of our work. You confirmed that our main narrative is clear and logically coherent. Specifically, you followed our progression from establishing the predator–prey relationship, to demonstrating olfactory avoidance, to identifying limonene as an active compound, and finally to validating its behavioral effects in both laboratory and field settings.

      We have carefully considered all your constructive comments and suggestions. We address them in detail in our point-by-point responses below. Your feedback has been extremely helpful, and we believe the revised manuscript is substantially stronger as a result.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

      We sincerely thank you for your critical and constructive comments. Your questions regarding the identity, biological origin, and ecological specificity of limonene are insightful and have helped us substantially strengthen the manuscript. We address each of your points below.

      On the biological origin of limonene and whether mammals synthesize it de novo

      You raised an important point that limonene seems unusual for a mammal to emit endogenously. We fully agree. We agree that direct evidence for de novo synthesis of limonene in mammals is currently limited.

      However, limonene in bat body odor could still have biological origins. First, it may originate from skin- or gland-associated microbiota. Recent work has shown that skin-associated microorganisms can substantially shape bat volatile odor profiles (Sun et al., 2026, BMC Biology), and some microbes possess enzymes capable of terpene biosynthesis. Second, previous studies have independently reported limonene in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.). This suggests that its presence in bats is not unique to our study. Third, our own analyses detected limonene in hair and snout secretions, but not in faeces or blank controls. This pattern is consistent with a biological source associated with the body surface, rather than diet or environmental deposition.

      We have now expanded our Discussion to cover these possibilities more explicitly. We also emphasize that the exact source, i.e., endogenous, microbial, or otherwise, remains an open question that warrants future investigation. Please see lines 258–273 of the clean version of the revised manuscript, or see the excerpt below:

      “Although limonene reliably induced avoidance behaviour in crickets, two related questions still merit careful consideration. One question is whether the limonene we identified genuinely originates from bats or reflects contamination during sampling. Limonene is common in plants and numerous consumer products (Boncan et al., 2020; Schuman, 2023), making its endogenous production by a mammal seem unusual. Nevertheless, multiple lines of evidence militate against contamination. First, we adhered to rigorous protocols. For example, all instruments were cleaned with ethanol and oven-dried before each use; bats were housed in stainless-steel cages, and cloth bags had been rinsed with purified water. Second, limonene was absent from all blank controls, including empty-chamber air samples and clean swabs, and was not detected in bat faecal samples. In contrast, it was consistently identified in hair and snout-secretion samples from bats. Third, independent studies have similarly identified limonene in the secretions of other bat species (Faulkes et al., 2019; Zhang et al., 2022). Furthermore, emerging evidence indicates skin-associated microbes may contribute to bat volatile profiles, with some taxa possessing enzymes involved in terpene biosynthesis (Sun et al., 2026). Taken together, these observations point towards an endogenous or microbe-mediated source, although the exact biosynthetic pathway remains to be determined.”

      On contamination from human-associated sources

      You asked how we excluded the possibility that limonene entered our samples through handling, equipment, cleaning products, or other human-associated sources. We appreciate this concern and have addressed it in detail.

      We believe contamination is highly unlikely for several reasons. First, we followed strict protocols throughout. All instruments were cleaned with ethanol and oven-dried before and after each use. We used stainless-steel cages and cloth bags made of degreased bleached cotton washed with purified water. These materials are not sources of terpenes. Second, we ran multiple blanks. Limonene was not detected in any empty-chamber air controls or in blank cotton swabs. In contrast, it was consistently found in multiple bat snout-secretion samples. This clear difference between samples and blanks strongly argues against contamination. Third, we now report individual-level odor data (see new Figure 3; Supplementary Table 5.xlsx). These data show that limonene was consistently present across individual bats. It did not appear sporadically, as one would expect from accidental contamination.

      In the revised manuscript, we have added detailed descriptions of our contamination controls in the Methods section (Please see lines 405–407, 443–444, 471–477 of the clean version of the revised manuscript, or see the excerpt below). We have also included the blank-control results (Supplementary Table 5.xlsx), as you suggested.

      Lines 405–407: “Prior to sampling, all glassware was thoroughly rinsed with ethanol and dried in an oven at 120°C, and volatile odor collection was conducted in a dedicated odor-free room to minimize environmental contamination.”

      Lines 443–444: “The empty-chamber controls were used to account for potential background signals from the experimental system and to provide a baseline for comparison with bat odor extracts.”

      Lines 471–477: “Upon capture, bats were placed in clean stainless-steel cages and kept in groups consistent with their natural social associations during the brief interval prior to immediate odor sampling. Hair samples (10 mg per individual) were clipped from dorsal and ventral regions. Snout secretions were collected using sterile cotton swabs (CS15-005, Shenzhen SihuaBo Technology Co., Ltd., China), with two blank swabs as controls. These blank swab controls were included to account for potential volatile contamination from ambient air or the swab material itself (Supplementary Table 5).”

      On how crickets distinguish bat-associated limonene from environmental sources

      You raised a thoughtful question about ecological specificity. Given that limonene is abundant in mint, citrus peel, pine, and other non-threatening plants, how would crickets use it as a reliable indicator of bat presence?

      We agree with you completely. We do not claim that limonene alone serves as an unambiguous bat-specific signal. Instead, our interpretation is more nuanced. We argue that elemental perception represents one effective strategy within a broader olfactory toolkit. It does not exclude the importance of other cues.

      In our revised manuscript, we have softened our interpretation accordingly. We now state explicitly that limonene is sufficient to trigger avoidance under our experimental conditions, but we do not interpret it as a uniquely bat-specific keystone kairomone. We also discuss mechanisms that could help crickets reduce false alarms under natural conditions. These include concentration differences, temporal patterning (bats are active at night), spatial context (specific foraging habitats), co-occurrence with other bat-specific volatiles, and possibly enantiomeric composition. Please see lines 274–292 of the clean version of the revised manuscript, or see the excerpt below.

      Lines 274–292: “The second question is how crickets might distinguish bat-derived limonene from environmental sources of this compound, given its prevalence in mint, citrus peel, pine and other non-threatening plants (Boncan et al., 2020; Schuman, 2023). It seems implausible that crickets could simply rely on limonene per se to differentiate a bat from a leaf. Two non-exclusive mechanisms could help resolve this issue. First, limonene need not be the only olfactory cue mediating risk perception. Our findings establish the sufficiency of limonene as an avoidance trigger, but do not preclude a role for other odor components. The crickets’ antennal responses to other bat volatiles in our GC–EAD analyses suggest more complex peripheral perception. Additional compounds, either alone or in synergistic blends, may modulate the full behavioral response in nature. Therefore, elemental perception via limonene likely represents one effective strategy within a broader olfactory toolkit available to insects. Second, crickets may discriminate bat-derived limonene through context-specific cues (e.g., temporal and spatial patterning, co-occurrence with other bat-specific compounds) to minimize false alarms. Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays. Critically, our field data confirm that limonene exposure in nature robustly triggers an adaptive anti-predator response, irrespective of the precise discrimination mechanism.”

      We acknowledge that fully testing these ideas would require substantial additional work. We have therefore framed this as an important direction for future research, rather than as a resolved issue in the present study.

      We thank you again for these insightful comments. Your feedback has helped us present a more rigorous, balanced, and transparent account of our work.

      Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are especially gratified that you recognized our study as having a "complete chain of evidence" and as a "highly praised study." This recognition means a great deal to us, given the multidisciplinary nature of our approach and the effort required to integrate behavioral, electrophysiological, chemical, and field data into a coherent narrative.

      We also appreciate your accurate summary of our work. You correctly highlighted the broader context that while acoustic co-evolution between bats and insects is well documented, whether crickets can use bat scent as an early warning cue has remained unknown. Your summary confirms that our main findings are clear: the body odor of S. kuhlii triggers robust avoidance and electrophysiological responses in L. equestris, and that limonene alone is sufficient to elicit avoidance in the laboratory and suppress calling in the field.

      We are particularly grateful that you acknowledged the multi-sensory nature of cricket predator recognition, with keen olfactory, tactile, and auditory senses. We agree that crickets are an excellent model for studying multimodal predator detection, and we hope our study encourages further exploration of olfaction in this classic predator–prey system.

      We have carefully considered all your specific comments and suggestions. These include questions about the novelty framing of olfactory eavesdropping, the rationale for our concentration choices, and the presentation of electrophysiological comparisons. We address each of these points in detail in our point-by-point responses below. Your thoughtful feedback has been extremely helpful in improving the clarity and rigor of the manuscript.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      Thank you for this comment. You are absolutely right that olfactory eavesdropping across the vertebrate–invertebrate divide is not a new concept in itself, and we acknowledge that this phenomenon has been well documented in previous studies.

      However, we would like to clarify our intended emphasis. In the Introduction and Discussion, we have already stated that empirical examples combining chemical identification, electrophysiological validation, behavioral assays, and field confirmation within a direct predator–prey context remain relatively limited. This is especially true for the bat–insect system, where research has historically focused on acoustic interactions rather than olfaction.

      We did not intend to overstate the novelty of olfactory eavesdropping per se. Instead, our emphasis was on providing a complete chain of evidence in a vertebrate–invertebrate predator–prey system that has traditionally been viewed through an acoustic lens. In that sense, we believe our study adds a complementary olfactory perspective to this classic system, rather than claiming to have discovered olfactory eavesdropping as a novel phenomenon. We hope this clarifies our position, as already stated in the original manuscript (lines 71–78, lines 88–90, lines 235–241 of the clean version of the revised manuscript, or see the excerpt below):

      Lines 71–78: “However, a fundamental gap exists in understanding whether such olfactory eavesdropping can operate across the vast phylogenetic divide separating vertebrate predators and invertebrate prey (Apfelbach et al., 2015; Dicke and Grostal, 2001; Schoeppner and Relyea, 2005). Although olfactory interactions across broad taxonomic boundaries are widespread in nature, such as mosquitoes using host odors to blood-feed, elephants and moths sharing pheromonal components, and aroids chemically mimicking carrion to attract pollinating flies (Kang et al., 2023; Zaremska et al., 2022; Zhao et al., 2022), these interactions are primarily shaped by selective pressures tied to foraging, reproduction, or mutualisms, rather than by predation-related selection.”

      Lines 88–90: “The bat–insect system presents an ideal model to address these questions. Despite the clear importance of olfaction to both taxa and its established role in predator–prey ecology, whether it plays any functional role in the iconic bat–insect arms race remains unexplored.”

      Lines 235–241: “Beyond the specific bat–insect model, our work addresses a central question in sensory ecology: how chemical eavesdropping operates within predator–prey systems between phylogenetically distant taxa with fundamentally divergent olfactory systems (Adams et al., 2020; Emerson and Johnson, 2024; Kaupp, 2010). While intraphyletic kairomone detection is well-established (e.g., rodents avoiding carnivore odors, aphids responding to ladybug chemicals), compelling experimental evidence for such olfaction-mediated recognition across broad phylogenetic divides has been limited (Apfelbach et al., 2005; Ferrari et al., 2007; Tanis et al., 2018).”

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      Thank you for this question. We fully agree that knowing the natural concentrations and relative content of limonene in bat odor would be valuable. However, our experimental aims were not to mimic natural emission levels precisely. Instead, they were designed to answer two distinct questions: First, whether cricket antennae are physiologically capable of detecting limonene at all; and second, whether limonene alone is sufficient to trigger behavioral responses under controlled and semi-natural conditions.

      For the EAG experiments, we selected a concentration gradient (0.001%, 0.01%, 0.1%, 1%, and 10% v/v) following standard practices in insect chemical ecology and referencing a previous study (Tang et al., 2024). The goal was to establish dose-dependent antennal sensitivity, not to match a specific natural concentration. Our data clearly show that cricket antennae respond across a range of concentrations, with stronger responses at higher doses.

      For the field experiment, we used a 10% v/v limonene spray over 25 m<sup>2</sup> plots. We acknowledge that this concentration does not quantitatively reflect natural bat emissions. Natural odor plumes are highly dynamic and depend on airflow, turbulence, temperature, humidity, vegetation structure, and distance from the source. Accurately reconstructing these natural dynamics would require detailed quantitative measurements of bat odor release rates and plume modeling, which were beyond the scope of the present study. Instead, our field experiment was designed for a functional purpose: to test whether limonene could alter cricket calling behavior under semi-natural conditions, using a concentration sufficient to produce a detectable odor stimulus in the field.

      We also note that the field-applied concentration is comparable to what has been used in other chemical ecology studies testing the behavioral effects of single volatile compounds under natural or semi-natural conditions. In that context, our positive result supports the ecological relevance of limonene as an avoidance cue, without requiring that the exact applied concentration matches natural bat emissions.

      We agree that quantitative characterization of natural bat odor composition, limonene release rates, and ambient exposure concentrations is an important direction for future research. According to your comments, we have added this point to the revised Discussion as a clear future direction (please see lines 286–290 of the clean version of the revised manuscript, or see the excerpt below). We have also clarified in the Methods that our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly (please see lines 515–522, 556–558 of the clean version of the revised manuscript, or see the excerpt below).

      Lines 286–290: “Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays.”

      Lines 515–522: “For each antenna, a hexane control was first presented to establish baseline antennal activity. For the initial screening, limonene, undecane, pentadecane, and hexadecane were diluted to 10% (v/v) in hexane and delivered individually in a randomized order. To assess dose-dependent responses, limonene was further tested at five concentrations (0.001%, 0.01%, 0.1%, 1%, and 10%, v/v in n-hexane), following the concentration gradient used in a previous study (Tang et al., 2024). Following the initial hexane control, the five limonene concentrations were tested in a randomized order across trials. Each stimulus lasted 0.5 s, with an inter-stimulus interval of 1 min to allow full recovery of antennal responses.”

      Lines 556–558: “Our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly.”

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

      Thank you for this suggestion. We understand your point that comparing GC-EAD responses between bat body odor and snout secretions would be a more direct way to identify the anatomical source of active compounds.

      However, we would like to explain why we did not include this comparison in the current study.

      First, the purpose of Figures 1C and 1D (i.e., Figure 2C and 2D in the revised manuscript) was to answer a more fundamental question: whether bat body odor, as a whole, contains volatile compounds that are detectable by cricket antennae. Comparing with an odor-free air control was therefore the appropriate first step. It established the basic phenomenon of olfactory detection before we moved on to source attribution.

      Second, we did conduct chemical profiling of snout secretions, hair, and faeces using HS-SPME-GC-MS (presented in Figure 3). These analyses showed that limonene was consistently present in hair and snout secretions, but absent from faeces and blanks. This allowed us to identify snout secretions as the most likely source of limonene, without requiring GC-EAD recordings from secretion samples themselves.

      Third, we did not perform GC-EAD on snout secretions for practical reasons. The secretion samples were collected in very small amounts. They were almost entirely consumed during the HS-SPME-GC-MS chemical analyses, leaving insufficient material for additional GC-EAD testing.

      We agree with you that directly comparing GC-EAD responses to snout secretions versus whole-body odor would be an excellent experiment. It would further strengthen the source attribution and provide more direct evidence. We have noted this as a valuable direction for future studies in the revised Discussion.

      We hope this clarifies our rationale. Thank you again for your thoughtful suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I would suggest moving Supplementary Figure 1 into the main figures. It contains important information about the predator-prey relationship and the ecological relevance of L. equestris, so it seems too central to be placed only in the supplement.

      We agree with your suggestion. The predator–prey relationship and the ecological relevance of L. equestris are indeed central to the biological context of this study. We have therefore moved the original Supplementary Figure 1 into the main text as Figure 1. We have renumbered the remaining figures accordingly, revised the figure legends, and updated the Results text to better highlight this ecological context (revised manuscript, Figure 1 legend, lines 856–864).

      (2) Lines 134 to 136: Only one representative EAD trace is shown. I suggest showing all five traces, either in the main figure or as a supplementary figure, to better illustrate the reproducibility of the antennal responses across individuals.

      We agree. To better illustrate reproducibility across individuals, we have added EAD traces from all five tested crickets as Supplementary Figure 1. The figure legend now describes the sample sizes for both the bat odor treatment and the odor-free control (revised manuscript, Supplementary Figure 1 legend, lines 900–906).

      (3) Lines 144 to 153: It would be helpful if the authors reported which VOC collections contained limonene and in what amounts. Showing individual-level data for the bat odor samples, rather than only pooled or summarized profiles, would strengthen the conclusion that limonene is consistently associated with S. kuhlii body odor.

      We agree. To better show the consistency of limonene detection across individuals, we have added Figure 3B, which displays an individual-level terpenoid profile. Each stacked bar represents one VOC collection, and the limonene segment indicates its presence and relative contribution. We have also revised the Results section to state explicitly that limonene was detected in hair and snout secretions but absent from feces (lines 144–149 in the revised manuscript).

      Lines 144–149: “To identify the biological sources of bat body odor, we analyzed VOCs from hair, faeces, and snout (pararhinal gland) secretions of nine bats using headspace solid–phase microextraction coupled with gas chromatography–mass spectrometry (HS–SPME–GC–MS). Snout secretions and hair shared similar hydrocarbon-rich VOC profiles, whereas faecal volatiles were distinct (Figure 3A). Individual-level terpenoid profiles further showed that limonene was detected in hair and snout secretion VOC collections but was absent from faeces (Figure 3B and Supplementary Table 2).”

      (4) Figure 2A: I recommend adding representative chromatograms or VOC traces for feces, hair, and snout secretions. This would make the source comparison more transparent and would help readers assess the underlying chemical profiles behind the heatmap.

      We agree that representative chromatograms would make the source comparison more transparent. However, due to the analytical workflow, individual chromatograms were not retained in the dataset we received. As an alternative, we added Figure 3B, which shows individual-level terpenoid profiles. This allows readers to assess which samples contained limonene and how its relative contribution varied among sample types and individuals. In addition, we have provided the NIST retention index, quantitative ion, qualitative ion, and molecular weight for each terpenoid compound in Supplementary Table 2.

      (5) The chemical identification of the key compounds would benefit from more detail. The authors state that compound identities were confirmed by matching retention times and mass spectra to authentic standards, including a mixed standard injection. It would be useful to provide the retention times, match and reverse-match scores, blank traces, and, if available, sample-plus-standard co-injection data showing peak augmentation without the appearance of new peaks. For (−)-limonene specifically, the enantiomeric assignment would require an enantioselective method, such as chiral gas chromatography, unless this was already performed and not described.

      We fully agree that comprehensive identification evidence is important for transparency and reproducibility.

      Regarding the identification data: we have already confirmed compound identities by matching retention times and mass spectra to authentic standards, including a mixed standard injection. In the revised version, we will deposit the raw chromatographic data and identification details (retention times, match scores, and blank traces) in Figshare as supporting information.

      Regarding co-injection: we acknowledge that sample-plus-standard co-injection would provide even stronger confirmation. However, our bat odor samples were difficult to obtain and were almost entirely consumed during the GC–EAD and GC–MS analyses. We therefore could not perform additional co-injection validation. We have noted this limitation in the revised Materials and Methods (revised manuscript, lines 463–464).

      Regarding the enantiomer assignment of limonene: you are correct that determining the specific enantiomer requires a chiral GC column, which was not available in this study. To avoid overinterpretation, we have replaced “(-)-limonene” with “limonene” throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) This is a typical study in the field of chemical ecology. As long as the source of the active substances is determined and the biological activity has been detected, this research is complete, so Figure 2 seems to be unnecessary.

      We agree that identifying the source of the active compound and confirming its biological activity are central to this study. However, we believe Figure 2 serves an important purpose. Bat body odor could originate from multiple sources, i.e., hair, faeces, or snout secretions, and comparing VOC profiles across these sources helps us determine which source most likely contributes to the odor cues detected by crickets. This is especially important for limonene, because it is a plant-associated terpenoid and not a typical animal-derived volatile, as the reviewer #2 pointed out. Our analysis showed that hair and snout secretions shared similar VOC profiles, while faecal VOCs were distinct and lacked limonene. This supports snout secretions as the likely primary source. To make this purpose clearer, we have revised the Methods section to state that this analysis was conducted to investigate potential biological sources of bat body odor (revised manuscript, lines 494–495, 567–569).”

      Line 494–495: “To identify the biological source of the characteristic body odor, we compared the VOC profiles from hair, faeces, and snout secretions.”

      Line 567–569: “Principal component analysis (PCA) based on a binary (presence/absence) matrix was performed using the vegan package in R to compare profiles from hair, feces, snout secretions, and bat body odor.”

      (2) Figure 3B seems to be incorrect. The significance of n-hexane and 1% limonene is ***P < 0.001, while for 10% limonene, why is it only two stars, **P < 0.01? Please check it.

      We have rechecked the statistical analysis and confirmed that the original annotation was correct. The significance levels in original Figure 3B (Figure 4B in the revised manuscript) were based on Bonferroni-corrected paired t-tests comparing each limonene concentration with the n-hexane control. The adjusted P value was 0.0006 for 1% limonene (P < 0.001) and 0.00485 for 10% limonene (P < 0.01). The higher adjusted P value for 10% limonene reflects greater among-individual variation at this concentration, which may be due to differential sensitivity of individual antennae at higher doses. We have now added the exact adjusted P values in the revised manuscript, lines 163–167.

      Lines 163–167: “EAG responses to limonene were concentration-dependent (repeated-measures ANOVA, F(5, 25) = 24.95, P < 0.001, η<sup>2</sup>p = 0.83; Figure 4B), with both 1% and 10% limonene solutions eliciting significantly stronger responses than the hexane control (Bonferroni-corrected paired t-tests, 1%: t(5) = 10.77, P < 0.001, Hedges' g = 3.82; 10%: t(5) = 6.92, P = 0.005, Hedges' g = 2.46).”

      Finally, we would like to express our sincere gratitude to the editors and the reviewers for the thoughtful feedback. Your comments have significantly improved the quality and clarity of our manuscript. We hope the revised version satisfactorily addresses all concerns raised.

    1. eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on previous smaller studies, to comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound, and the exploration of different k-mer lengths and use of a more complete reference assembly strengthen confidence in the analysis while clarifying its limitations and the method is broadly applicable to large-scale short-read datasets for assessing copy-number variation and genomic repeat content. The findings are convincing in their scope and novelty for A. thaliana. It will be of interest to learn in future how the approach generalizes to species with substantially greater or lower repeat content.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

    3. Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on smaller studies to more comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound and broadly applicable to large-scale short-read datasets for assessing copy number variation and genomic repeat content. While convincing in its scope and novelty, the findings would be further strengthened with exploratory analyses of datasets from other species with more or fewer repeats in their genomes.

      We agree with the assessment and appreciate the Editor’s handling of our manuscript and careful consideration of our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      While it would be interesting to further investigate the mechanisms underlying cis-QTLs, we do not believe this dataset can distinguish between the two proposed models of cis-variation. As argued in the manuscript, the observed cis-association signals are linked to repeat-associated SNPs and therefore are most likely driven by variation in repeat copy number, supporting the latter hypothesis.

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      As a compromise, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      Our decision to use 12-mers to profile genome content in A. thaliana was based on an empirical evaluation of the sensitivity and specificity of different K-mer lengths for this application. Because our goal was to characterize intraspecific variation in genome content, we did not extend this analysis beyond A. thaliana. Nevertheless, we agree that alternative K-mer lengths may capture additional and potentially useful information and that using the method in other species would be interesting.

      Although we found generating a 14-mer dataset to be computationally infeasible, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      To our knowledge there is not a published gapless A. thaliana assembly, although several of the recent assemblies are pretty close. Although we continue to use the 1001 Genomes SNP dataset that relies on TAIR10 for the GWAS analyses, we now present Figure 2A and Figure S9 using the Col-PEK assembly (Hou, Wang, Cheng, Wang & Jiao 2022 Molecular Plant).

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      We concede that A. thaliana does not possess a typical plant genome and thus have moderated our claims throughout the paper.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

      We agree that the conclusions of our work are correlational, as they are based on GWAS associations and have not been validated through functional experiments. We would love to see tests of the hypotheses we presented, but we believe this is beyond the scope of the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The Snakemake workflow Kmer-it had a few minor bugs that prevented use with modern Snakemake versions, which I fixed as part of the review (submitted as a pull request). Please be sure to validate that the whole workflow works with the latest Snakemake version.

      We appreciate the Reviewer’s testing of our software and have updated it accordingly.

      (2) L162: Since you already map samples to a reference genome to filter organellar DNA, consider adding some form of filter or correction for duplication rate in paired-end data, which I have found to be one major source of batch effects as observed here between sequencing centers.

      We appreciate this suggestion. In earlier versions of the analysis, we removed duplicated reads, but doing so substantially weakened the relationship shown in Figure 1C. Because it is difficult to distinguish duplicate reads arising from genuine repeat copy number variation from those resulting from technical artifacts, we ultimately chose not to include duplicate removal in the final pipeline. However, we did remove the effect of the sequencing center using a linear model.

      (3) Have you used this approach with raw long-read data? Do you see any difference between Illumina and low-error-rate long read data (e.g., HiFi, R10 Nanopore with SUP calling)? This could be compared in a directly paired manner using some of the recent papers with HiFi/ONT data on 1001G accessions, e.g., Lian et al, 2024, Wlodzimierz et al, 2023, Teasdale et al, 2025. (I do not believe such a comparison is required to prove the utility of this method, but it could be of interest as such sequencing technologies become increasingly cost-competitive).

      We have not evaluated this approach using raw long-read sequencing data, although we agree that it would be an interesting direction for future work. As noted in the manuscript, we detected significant batch effects attributable to sequencing center, even among datasets generated with the same sequencing technology. Given these observations, we are cautious about comparing K-mer frequencies across sequencing platforms that have distinct error profiles, as technical differences could confound biological signals.

      Reviewer #2 (Recommendations for the authors):

      (1) I would strongly suggest modulating some of the claims made in the paper, especially with regard to "genome evolution in plants", given that the paper focuses on A. thaliana. Alternatively, the authors could test the method in a different dataset from a different species. This would alleviate some concerns regarding the choice of k-mer length and demonstrate robustness.

      We agree that A. thaliana is not representative of most plant genomes and have revised the manuscript to better reflect this limitation. Since the primary goal of this study was to characterize intraspecific variation in genome content within A. thaliana, we believe that extending the analysis to an additional species falls beyond the scope of the present work. We do not believe that our choice of K-mer length undermines the robustness of the approach. To evaluate this concern, we repeated the GWAS analyses using 10-mers and found broadly consistent results, which are presented in Supplementary Figure S19-21. These findings suggest that the major conclusions are not strongly dependent on the specific K-mer length selected.

      (2) A k-mer length of 12 is quite small. There are 4^12 possible 12-mers (~17 million), so you would expect that all 12-mers would occur in the A. thaliana genome by chance at least once. While larger k-mers may be more computationally expensive to compute, it would be worth repeating the analysis with a larger k-mer length.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      (3) Line 689 on page 31 should include the figure number.

      Thanks! Fixed.

    1. eLife Assessment

      This important study presents an innovative combination of experimental evolution and CRISPR screening to study the genetic basis of a polygenic trait, resistance against octanoic acid in Drosophila. The evidence supporting the authors claims is solid. The work will be of interest to the Drosophila community and researchers interested in functional testing of polygenic traits.

    2. Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting and the idea of narrowing down a large list of candidate loci by CRISPR based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow up experiments to validate the results.

      Comments on revised version.

      Weaknesses:

      - The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This point has been confirmed by the authors.

      - The statistical analysis suffers from some problems and insufficient description of the analyses performed.

      Has not been addressed in their response.

      - Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      This has now been included in the discussion. I would recommend that they make the distinction between genetic and adaptive architecture, as this matches their verbal description.

      - The reviewer would have liked to see more connection between the experimental evolution and GWAS data. As some D. simulans genotypes have similar resistance as D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      The reviewer is happy with the response.

      - At several places the authors discuss the challenge of studying a polygenic trait, but at the same time they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      The reviewer is not satisfied with the arm waving explanation of the authors. The important question is how much of the phenotypic variation is explained by the two candidate genes? The reviewer is inclined that based on the weak selection response, the variation is too little to be detected experimentally. Nevertheless, the overexpression of alkbh7 alone was sufficient to generate resistance levels similar to the ones in d. Melanogaster. Hence, it is not adequate to speak of small effects. This discrepancy requires more discussion.

      - It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      This point remains valid, in particular in the light of the discrepancy between the very limited selection response of alkbh7 and its large phenotypic effect after overexpression.

      Impact:

      - Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid-a trait that is rather easy to evaluate.

      The reviewer did not find the reply satisfactory.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Marconcini et al. report results of an ambitious study on the genetic mechanisms that contribute to resistance of Drosophila flies to the toxin octanoic acid (OA). This study was motivated by two observations: first, Drosophila sechellia, a close relative of D. melanogaster, has evolved specialized feeding on fruits of Morinda citrifolia, which contain high concentrations of OA and second, that artificial selection on Drosophila simulans, a sister species of D. melanogaster, can generate higher resistance to OA. Previous studies had performed genetic mapping studies between D. simulans and D. sechellia that implicated certain genomic regions in resistance to OA and, in particular, implicated several Osiris gene paralogs as contributing to resistance, though the molecular mechanisms of resistance remain unclear. In this study, Marconcini et al. performed two major experiments. First, they performed evolution-and-resequence on Drosophila simulans populations exposed to OA for 50 generations and identified candidate regions with excessive shifts in allele frequencies as candidate regions containing OA resistance genes in D. simulans. Second, they performed a CRISPR knock-out screen in a D. melanogaster cell line to identify genes that contribute to OA resistance and susceptibility.

      Evolve-and-resequence yielded many candidate genomic regions with extreme allele frequency shifts, which may be regions containing OA resistance genes, or linked genes, or regions that happen to show a strong shift in all replicate populations by chance. As the authors note, detecting significant shifts in allele frequencies is a challenging problem, and the authors use two measures of allele frequency shifts (the Cochran-Mantel-Haenszel method and Bait-ER) and perform simulations under neutrality to estimate a reasonable significance threshold. I am not entirely convinced by this method of estimating significance levels, because the simulations involve assumptions that may not be met by the real populations. I would think that a permutation test would provide an assumption-free method of estimating significance levels. I have tried to think whether there is something about the design of these experiments that would preclude the use of permutation tests (which are used widely for genome-wide studies, such as QTL), but I can't think of one. Perhaps the authors are aware of a reason permutation tests would be invalid here, and if so, they should state this reason.

      Significance thresholds have been estimated using a variety of approaches in the literature, including false discovery rate (FDR) control, simulations under neutral drift with predefined cut-offs, arbitrary significance thresholds, haplotype-based analyses, permutation tests, amongst other methods. Our choice was guided by a review (doi:10.1186/s13059-019-1770-8), which compared the performance of several of these approaches and found that the relatively simple assumptions underlying the Cochran–Mantel–Haenszel (CMH) test often performed as well as, or better than, more complex alternatives. As with most significance thresholds used in genome-wide analyses, the CMH threshold applied here is ultimately based on a degree of arbitrariness, although we aimed to be as stringent as possible. As a complementary method, we used Bait-ER; here the threshold used followed the recommendations provided in the original publication describing this method (doi:10.1111/jeb.14134).

      One reason that permutation-based approaches may be less widely adopted than theoretical null models is the extensive linkage disequilibrium among SNPs, which results in large blocks of correlated variants that cannot be considered independent observations (and therefore not shuffled). In principle, permutations could be performed at the haplotype-block level; however, defining haplotype blocks itself requires selecting thresholds or criteria that are often user defined or based on significance cut-offs. Although several tools are available for this purpose, our experience has been that the resulting block definitions remain sensitive to these choices and therefore introduce a comparable degree of arbitrariness. In reality, there is not yet a standard method in the field that has emerged as the “best” one.

      There is overlap between regions detected by the two methods, but the methods disagree for many regions. The authors state that a "majority of prominent peaks were found by both methods," but I am unclear on what "prominent" means here. It would be more helpful to be more quantitative about the extent of overlap.

      To quantify the agreement between the two methods, we calculated the overlap in genomic coverage (base pairs) between CMH and Baiter candidate regions. For G25, the two methods shared 1.90 Mb of candidate regions. The overlap encompassed 65.3% of the genomic span identified by CMH and 64.1% of the span identified by Bait-ER. For G50, the methods shared 4.77 Mb of candidate regions. The overlap encompassed 63.6% of the CMH candidate span and 99.9% of the Bait-ER candidate span. We have now replaced the admittedly qualitative statement the reviewer highlighted ("majority of prominent peaks were found by both methods”) with this information.

      The authors hypothesized that the response would be at similar genomic loci in all populations (line 222). It seems at least possible that epistatic interactions would lead to different combinations of alleles evolving in each population. I wonder if it would be possible to test whether there is heterogeneity in the responses across the replicate populations.

      We agree that epistatic interactions could lead to different allelic combinations being favored in different replicate populations. However, both BaitER and the CMH test are designed to detect parallel evolutionary responses across replicates, which was the focus of our study. One approach for testing population-specific responses is the LRT-2 test (doi:10.1534/genetics.118.301824). We did not pursue this analysis because it would likely generate many additional candidate loci, making interpretation more challenging, while providing limited additional insight into the repeatable genomic responses that were the primary focus of this work.

      The evolve-and-resequence method yielded many possible regions contributing to OA resistance in D. simulans, but perhaps too many regions to test directly or even to build sensible hypotheses about the genes involved. Thus, the authors performed a second experiment to try to narrow down the list of possible candidate genes. They performed a CRISPR knockout screen in a D. melanogaster cell line for genes that contribute to resistance or susceptibility to OA. The authors identify several limitations of this experiment, but they nonetheless identified several genes where knockouts contribute to OA susceptibility or resistance. Intersecting top hits with regions that experienced selection identified two "resistance" genes: kraken and Alkbh7. The selection hit at kraken is quite compelling, whereas the evidence at Alkbh7 is less strong because only two SNPs were marginally significant. Further functional assays, including gene knockouts in D. melanogaster and D. sechellia, provide some support for the claim that both of these genes can contribute to resistance to OA in flies.

      Beyond the few issues raised above, I do not have significant questions about methodology or the results. I do think, however, that the authors should be more conservative about the implications and significance of their results. For example, on line 139, the authors claim that this intersection approach provides a "powerful paradigm to investigate ecotoxicology." I am not sure I agree that the identification of two genes that may contribute to OA resistance, after a seemingly heroic selection experiment and CRISPR screen, suggests that this method is all that powerful. It seems that most of the genes that contribute to the selection response remain unidentified.

      We agree with the reviewer and have modified the sentence accordingly. While we believe that integrating the approaches discussed in this paper can provide valuable insights into the genetic basis of ecotoxicological traits, these approaches are not a panacea for traits with highly complex genetic architectures, such as OA resistance. Nevertheless, the identification of two candidate genes with some evidence of contributing to the trait represents a meaningful advance toward understanding its underlying genetic basis.

      Finally, given that one motivation of this project was to identify genes that contribute to evolved resistance to OA, I am surprised that the authors did not generate CRISPR alleles of kraken and Alkbh7 in D. simulans and then use these together with the existing alleles in D. sechellia to perform reciprocal hemizygosity tests to determine if these two genes actually contribute to evolved resistance in D. sechellia. This test is simpler to perform and may be more sensitive than the allelic replacement that the authors propose (lines 446-449).

      While generating null alleles for the candidate genes in D. simulans is beyond the scope of this revision, we note that we did attempt reciprocal hemizygosity tests using the mutants available in D. melanogaster. However, the resulting hybrids were recovered in low numbers, and these animals were rather weak, making them unsuitable for the severe OA exposure conditions employed in our assays. We agree that, where feasible, future studies should incorporate reciprocal hemizygosity tests, as they represent a powerful approach for validating the contribution of candidate genes to the trait of interest. We have revised the final sentence of the Results section accordingly.

      Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit, in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting, and the idea of narrowing down a large list of candidate loci by CRISPR-based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow-up experiments to validate the results.

      Weaknesses:

      The reviewer is not convinced of the conceptual idea behind their approach: the intersection of the two approaches implicitly assumes that null alleles (or at least compromised alleles) should be selected during experimental evolution. The reviewer considers this unlikely, and the authors made no attempt to test this implicit hypothesis in their data.

      We respectfully disagree with the reviewer’s interpretation of the conceptual idea behind our approach. Our strategy did not assume that experimental evolution selects for null alleles, but rather for any type of variant that could contribute to the trait being selected for (i.e., increases in OA resistance), pointing to candidate genes contributing to the trait. Like many evolve-and-resequence experiments, this approach identified hundreds of candidate genes. This is why we took an orthogonal, genome-wide CRISPR screening approach, where loss-of-function mutations could lead to increases or decreases in OA tolerance of cultured cells. However, we stress that the naturally selected alleles – which could be gain or loss of function – may have much subtler phenotypic effects than the null alleles used for functional validation.

      Along the same lines, it is not clear how to reconcile an upregulation of candidate genes in resistant flies with the knockout experiments.

      We respectfully disagree that these findings are difficult to reconcile. The observed upregulation of the candidate genes in the selected, resistant D. simulans is consistent with a role in promoting OA resistance, while the knockout experiments (whether in cultured cells or in whole animals) demonstrate that loss of gene function reduces resistance. These observations are complementary: increased expression is associated with enhanced resistance, whereas complete loss of function impairs it. The knockout experiments of kraken and Alkbh7 were intended as functional validation of gene involvement and do not imply that the alleles selected during experimental evolution are loss-of-function alleles (we rather hypothesize that the selected alleles are gain-of-function through some, as yet undetermined, mechanism).

      The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This is correct. As our results suggest that the phenotype is shaped by multiple genes, such that the effect of any individual gene is likely modest compared to their combined contribution. Consequently, detecting and validating the effect of a single gene requires more stringent OA conditions than those needed to observe the overall phenotypic response over the course of several generations. In addition, the shorter-term plate assay was more practical for higher temporal resolution of the analysis of mortality in the presence of OA.

      The statistical analysis suffers from some problems and an insufficient description of the analyses performed.

      Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      We agree that comparing the experimental evolution and GWAS results (from our previous work, doi:10.1093/g3journal/jkag032) is of considerable interest. We have now expanded the Discussion to explicitly discuss the relationship between the two datasets. Overall, the overlap between the approaches was limited, although two GWAS candidate genes, bez and CG13003, fall within genomic regions exhibiting significant CMH signals at generation 25 (but not generation 50) of the evolve-and-resequence experiment. (We note that these genes could not have been identified in our CRISPR screen, as they are not expressed in S2R+ cells). More broadly, differences between the GWAS and evolve-and-resequence results likely reflect the distinct evolutionary processes captured by each approach: GWAS maps standing phenotypic variation among isofemale lines, whereas experimental evolution tracks allele frequency changes under sustained selection. Understanding why some signals are shared whereas others are not – whether due to effect size, genetic background, epistasis, pleiotropic costs, or the contribution of initially rare variants – remain important open questions.

      The reviewer would have liked to see more connection between the experimental evolution and the GWAS data. As some D. simulans genotypes have similar resistance to D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      While D. simulans genotypes displayed a range of OA resistance levels, none approached D. sechellia levels of resistance (see Figure 2e from our previous work, doi:10.1093/g3journal/jkag032) (The reviewer might have conflated the data from our GWAS of D. melanogaster strain, shown in Figure 2b of that paper, where some lines of that species exhibit comparable resistance to D. sechellia under the conditions of that assay). Regardless, our experimental evolution data do not provide sufficient resolution to identify the specific favorable alleles underlying the response to selection. Instead, we detect genomic regions containing many linked variants whose frequencies change under selection. Consequently, a direct comparison between evolved genotypes and GWAS-associated genotypes is currently difficult. We note, however, that expression of the D. sechellia kraken allele in D. melanogaster did not produce a significant effect on resistance, suggesting that even the most promising candidate alleles might not have strong effects in isolation.

      At several places, the authors discuss the challenge of studying a polygenic trait, but at the same time, they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      Our results do not imply that kraken and Alkbh7 are major-effect loci or that they explain a substantial proportion of the phenotypic variation. Rather, our data indicate that these genes make measurable contributions to OA resistance, consistent with the expectation that complex traits are influenced by many loci of individually modest effect. The absence of these genes among the top GWAS candidates does not preclude their involvement, as GWAS and experimental evolution interrogate different aspects of the genetic architecture and differ in their power to detect loci of varying effect sizes and allele frequencies.

      It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      We agree that the significant peaks identified in the evolve-andre sequence experiment represent promising targets for future investigation. However, these peaks typically span large genomic regions containing tens to hundreds of genes, making it difficult to prioritize individual candidates based on the experimental evolution data alone. In this work, we focused our functional validation on genes independently supported by the CRISPR screen, which provided gene-level resolution. We fully acknowledge that additional causal genes are likely to reside within the selected regions and remain to be functionally characterized.

      Impact:

      Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid, a trait that is rather easy to evaluate.

      The reviewer appears to imply that a trait being straightforward to phenotype necessarily implies that its genetic basis should also be straightforward to resolve. Many classic complex traits, such as human height, are simple to measure yet have an extraordinarily complex, highly polygenic genetic architecture. We believe that OA resistance represents a similar challenge: while the phenotype is readily assayed, differences in assay conditions, genetic backgrounds, and the contribution of many loci of individually modest effect make its genetic basis difficult to dissect. We have strived to be cautious in our conclusions, in particular the evolutionary interpretations; nevertheless, to our knowledge, this is the first study to provide functional evidence supporting the contribution of specific genes to OA resistance in D. sechellia, combining both loss-of-function phenotypes and expression data. As emphasized by the title of our manuscript, we view the principal contribution of this work as demonstrating how complementary experimental approaches (both of which are fairly novel for study of toxin susceptibility/resistance genetics) can be integrated to prioritize and functionally evaluate candidate genes underlying complex adaptive traits.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers propose to include additional statistical and quantitative measures to strengthen the results. Furthermore, certain parts of the manuscript could be rephrased to make sure that the readers can understand more clearly the implications and significance of the results. To test whether the genes identified in the present manuscript (as contributing to octanoic acid resistance) are also involved in the evolution of the resistance, reciprocal hemizygosity tests using CRISPR alleles of kraken and Alkbh7 in D. simulans would be a plus, but such experiments are not required because the main focus of the paper is on the genetic basis of octanoic acid resistance, and not on evolution.

      We thank the Reviewing Editor for these constructive comments. We have revised the manuscript to clarify the interpretation and significance of our findings and have addressed the reviewers’ comments throughout. Regarding reciprocal hemizygosity tests, we agree that they would provide a valuable means of assessing the evolutionary contribution of candidate genes. However, generating the necessary reagents in D. simulans represents a substantial undertaking beyond the scope of the present study, particularly given the expected modest effects of individual loci underlying this highly polygenic trait. We have nevertheless revised the end of the Results section to mention reciprocal hemizygosity tests as an important direction for future work.

      Reviewer #2 (Recommendations for the authors):

      (1) Provide more details about the selection tests: which sequences were used? Please report p-values. It would also be important to discuss the possibility of false positives caused by a bottleneck in D. sechellia. A genome-wide analysis could help to see if the bottleneck increased the signal of positive selection.

      We are not entirely sure what additional analyses are being suggested. The sequences and methods used for the selection analyses are described in the Methods, and the statistical support for these analyses is reported in the Supplementary Material (MK test p-values and FUBAR posterior probabilities). We are also unclear as to how a genome-wide analysis would address the interpretation of the gene-specific selection analyses presented here. If the reviewer intended a different analysis, we would appreciate further clarification.

      (2) The significance level of the CMH test needs to be determined with the effective population size, not with the census size, as done by the authors. This is important, as it is not clear if the candidate genes remain significant after significance adjustment based on the effective population size.

      We thank the reviewer for this comment. In our analyses, effective population sizes were explicitly incorporated by Bait-ER to model the effects of genetic drift. The CMH significance thresholds, however, were obtained from neutral forward simulations following the recommended workflow for the method, which requires census population sizes rather than effective population sizes as input. To make these simulations as realistic as possible, we therefore used the observed census population sizes at each generation of the experimental evolution. Estimating generation-specific effective population sizes for use in an alternative simulation framework would require substantially more temporal data than are available here (e.g., sequencing many additional time points) and is beyond the scope of the present study.

      (3) Include the allele frequency trajectory across time for the candidate genes.

      We thank the reviewer for this suggestion. However, plotting allele frequency trajectories for the candidate genes is not straightforward because the evolve-and-resequence analysis identified broad linked genomic regions rather than individual causal variants. Each candidate region contains numerous SNPs spanning several kilobases, with each SNP exhibiting its own allele frequency trajectory across the 10 replicate populations. Consequently, there is no single representative trajectory for a given candidate gene, and plotting all SNPs within each region would be difficult to interpret. For kraken, however, we identified seven candidate regulatory SNPs and have plotted their individual allele frequency trajectories, which we include in Author response image 1 to illustrate the diverse patterns observed:

      Author response image 1.

      (4) Estimate the effect of candidate loci in the GWAS data (independent of significance).

      We thank the reviewer for this suggestion. However, we are not entirely sure what analysis is being proposed. In particular, it is unclear which variants the reviewer is referring to, as our candidate genes are associated with multiple linked variants rather than a single causal SNP. Consequently, we are unsure how the effect of a candidate locus should be estimated in the GWAS dataset. Moreover, there is no guarantee that the same variants are represented in both datasets. For example, variants filtered out during the GWAS because of low allele frequency may subsequently have increased in frequency during experimental evolution and therefore contributed to the evolve-and-resequence signals. We would appreciate further clarification of the analysis the reviewer has in mind.

      (5) Discuss the challenge of false positives (see: 10.1016/j.cub.2020.12.023).

      We thank the reviewer for this suggestion. We agree that false positives (as well as false negatives) are an important consideration when studying complex polygenic traits, both through the initial “screening” efforts (e.g., GWAS, experimental evolution) and follow-up functional validation (e.g., RNAi, mutant, overexpression analyses). We believe this issue is already addressed in the manuscript through our discussion of the limitations of the individual approaches, the polygenic nature of OA resistance, and our cautious interpretation of the functional validation results. Our conclusions are limited to identifying candidate genes that contribute to OA resistance, rather than claiming to have identified all of the loci or variants underlying the evolution of this trait. We are therefore not sure what additional discussion the reviewer has in mind and would appreciate further clarification if a specific point from the cited study is intended.

      (6) The figures with the Manhattan plots should be improved to indicate the overlapping genes in the Manhattan plots, rather than in the circle figures below. By using different colors this should be quite easy and clean.

      We thank the reviewer for this suggestion. However, we believe Author response image 2 more appropriately illustrate the overlap between the CMH and BaitER analyses. Although the Manhattan plots display individual SNPs, our candidate loci are defined by broader genomic regions comprising blocks of linked significant SNPs rather than by single variants. Simply highlighting SNPs within overlapping regions would therefore add visual complexity without providing additional biological insight beyond that already captured by the circos plots. As an illustration, we provide here an example of the generation 50 Manhattan plot with overlapping SNPs highlighted in red:

      Author response image 2.

      (7) A more focused discussion of the assumption that functional data from D. melanogaster can explain resistance in D. simulans or D. sechellia. At some places, epistatic interactions are mentioned, but the reviewer feels that the entire screen is based on the idea that similar effects are found across species, hence, this needs to be adequately reflected in the discussion.

      We thank the reviewer for raising this important point. We would like to emphasize that our study was not based on the assumption that functional effects identified in D. melanogaster or D. simulans necessarily explain the evolution of OA resistance in D. sechellia. Rather, our motivation stemmed from the longstanding difficulty of identifying individual genes underlying this highly polygenic trait using mapping approaches alone. We therefore sought to combine orthogonal experimental approaches to identify genes contributing to OA susceptibility (of D. melanogaster and D. simulans) and resistance (D. sechellia). We hoped, but did not assume, that genes supported by multiple independent lines of evidence would be informative for understanding the natural evolution of OA resistance in D. sechellia. Indeed, we acknowledge in the manuscript that the cell-based CRISPR screen, experimental evolution, and the natural evolution of D. sechellia occurred under very different selective contexts and timescales. As reflected in the title of our manuscript, our conclusions are intentionally framed around the identification of novel toxin resistance loci, while remaining cautious about their evolutionary interpretation.

      (8) Discuss that the controls in the RNAi test were quite variable. Could this reflect some problems with the assay?

      We do not believe that the variability among the control lines reflects a problem with the assay. Rather, the different controls (Gal4, UAS-RNAi etc.) represent distinct genetic backgrounds, each of which may exhibit a different baseline level of OA resistance. Indeed, in our recent GWAS of OA resistance (doi:10.1093/g3journal/jkag032), we observed substantial natural variation in OA resistance among D. melanogaster and D. simulans lines. Importantly, each RNAi line was compared with its corresponding genetic background control, so differences among control lines do not affect the interpretation of the individual RNAi experiments.

    1. eLife Assessment

      This study presents valuable findings on the functional consequences of nuclear envelope rupture caused by the depletion of nuclear pore complex subunits on the behaviour and organisation of chromosomes during mitosis in C. elegans embryos. The experiments are generally well-designed and executed; however, the evidence for some of the main conclusions is incomplete. This work is of potential interest to cell biologists working on cell division.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

    3. Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP-3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically co-depleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC-4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Co-depletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP-3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Co-depletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. Because loss of the SAC independently causes premature anaphase and genomic instability through well-established mechanisms unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SAC-dependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP-4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf-1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF-1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      We understand the more severe chromosome missegregation and DNA damage phenotypes in the npp-3 mdf-1 or npp-3 mdf-2 double RNAi may suggest that npp-3 RNAi is more sensitized to the loss of spindle assembly checkpoint (SAC) component MDF-1 or MDF-2 for chromosome protection, but may not clearly show that the localization of chromosomes to the nuclear periphery serves a chromosome protective function. Based on our current phenotypic analyses, we will tone down the title to “Prophase Chromosomes Relocalization to Nuclear Periphery in NPP-3/NUP205 Depletion Depends on the Spindle Assembly Checkpoint and Inner Kinetochore Proteins”.

      We have thought about artificially tethering the chromosomes to the nuclear periphery. However, without the same stimulus/defect and activation of the response pathway, the effects could be different. In future, to further clarify the functions of chromosome relocalization, we could analyze the chromosomal missegregation and DNA damage phenotypes in npp-3 lem-2 RNAi, which has a partial chromosomal nuclear periphery phenotype, to see if there is any quantitative relationship between chromosome nuclear periphery location and DNA protection.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      We have not performed partial npp-3 RNAi yet. To explore whether the level of nuclear envelope rupture correlates with chromosome localization to the periphery or if a minimal level of nuclear rupture is required, we have performed the lacI::GFP reporter assay in different NPP RNAi to indicate the nuclear permeability defects (in Fig. S1A and B). npp-2, npp4 or npp-5 RNAi also causes increased permeability, at a comparable level as npp-3, npp-7 and npp-13 RNAi, yet only npp-3, npp-7 and npp-13 RNAi causes chromosome relocalization. This may explain the peripheral chromosome relocalization response may be more closely linked to the specific NPC subcomplex’s (inner ring and nuclear basket) or individual NPP’s function, rather than the general permeability defect or rupture size.

      However, we have also attempted to use other methods to induce targeted nuclear rupture with specific rupture size by laser ablation (Author response image 1, 355 nm UV pulsed laser marked by the white rectangle), and we could occasionally observe all chromosomes relocalizing to all over the nuclear periphery after 660 s upon laser ablation in different strains. However, the chromosome relocalization phenotype is not very consistent (~33%, n =12). Importantly, the laser also introduces DNA damage at the rupture sites, as marked by HUS-1 (Author response image 2), complicating the interpretation, so we did not include this attempt in the manuscript.

      Author response image 1.

      Time-lapse imaging of embryos expressing LEM-2::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce nuclear envelope rupture (white rectangle). 0s is the time applying laser ablation. Scale bar 5 μm.

      Author response image 2.

      Time-lapse imaging of embryos expressing HUS-1::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce DNA damage (white rectangle). 0 s is the time applying laser ablation. Scale bar 5 μm.

      Another attempt to achieve different nuclear rupture sizes is by observing nucleus in different stages of embryos. By npp-3 feeding RNAi approach, we observed an increase of the proportion of the nuclear circumference with nuclear rupture during embryonic development (Fig. S3D). Yet, the chromosome periphery phenotype is displayed in all embryo stages, suggesting that the localization of chromosomes to the nuclear periphery occurs across a range of rupture sizes.

      Given these observations, it appears that the extent of nuclear rupture or permeability increase may not be the sole determinant of chromosome relocalization. Instead, disruption of certain NUP subcomplexes or NPPs might stimulate this process.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      We observed that chromosomes localize to the nuclear periphery in all cells during embryonic development until the embryo dies, as demonstrated by snapshots and time-lapse imaging (Fig. S1C and Fig. S1E).

      We focused on the P1 blastomere for our analyses because its cell cycle timing and division orientation is well characterized, which facilitates precise examination of chromosome localization and cell cycle dynamics under different perturbation conditions. We will add a sentence at the beginning of that result section to clarify this choice: "The chromosome nuclear periphery localization in npp-3 RNAi is consistent in all cells at different embryonic stages (Fig. S1C and Fig. S1E). For analysis in distinct perturbation conditions, we specifically examined the P1 cells, whose cell cycle timing and division orientation is well characterized and suitable for such studies.”

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (see response to point 2) We observed that npp-2 or npp-4 RNAi can cause nuclear rupture with certain permeability defect compared to wildtype (Fig. S1A and S1B), but not chromosomal nuclear periphery localization, suggesting that such chromosome localization to the nuclear periphery does not just depend on the nuclear rupture but could be a more specific phenotype related to individual NUP subcomplex’s (inner ring and nuclear basket) or NPP’s function. 

      For Y complex, we also now tested NPP-5 depletion, which shows permeability defect but did not show nuclear rupture marked by NPP-1:GFP or chromosomal periphery phenotype (Fig. S1A and S1B), which is consistent to the paper cited [1]. 

      To confirm the efficiency of NPP-2 depletion, we have observed the presence of smaller nuclei (as described in phenobank) by NPP-1::GFP marker (Fig. S1A) and a marked reduction in NPP2::GFP signal in npp-2 RNAi-treated embryos (unpublished data). These evidences support that NPP-2 depletion was effective. 

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      We agree with this suggestion and have decided to remove the section on AIR-1 in the manuscript text to avoid confusion. To clarify, air-1(RNAi) causes multiple centrosomes in only a small proportion of embryos (approximately 13%) [2] , primarily by affecting centrosome positioning [3]. and does not routinely lead to significant centrosome amplification.

      Our interest in AIR-1 was inspired by previous findings by Hachet et al., showing that AIR1/Aurora A localizes at sites without NPP-3 during mitotic entry—these sites are believed to correspond to centrosome locations [4]. This relationship prompted us to explore how AIR-1 depletion might influence NPP-3 localization at the nuclear envelope and its potential effects on chromosome positioning. Interestingly, we discovered that following air-1(RNAi) treatment, nuclei exhibit discontinuous NPP-3 localization on the nuclear envelope. Thereby, we could investigate the interplay between NPP-3 and chromosomal dynamics. Chromosomes localize to nuclear periphery specifically without NPP-3 (Author response image 3). This suggests a potential negative correlation of NPP-3 localization with the chromosomes. We did not imply any regulation by AIR-1.

      Author response image 3.

      Condensed chromosomes tend to localize at nuclear envelope sites without NPP-3 or NPP-7. (A) (C) Selected representative confocal images of all chromosomes localizing at the nuclear periphery in the control and air-1(RNAi) embryos expressing H2B::GFP and mCherry::NPP-3 (A) and GFP::NPP-7 and mCherry::H2B (C). The upper right image is the zoom-in view of the nucleus. Scale bar 10 μm. (B) (D) The intensity of GFP and mCherry is normalized to the average intensity along the nuclear envelope and plotted in the control and air- 1(RNAi) nucleus from (A) and (C).

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      Based on our time-lapse data in Figure 1D, most chromosomes appear to condense at or near the nuclear periphery, but there are still some chromosomes inside the nuclear space at approximately -600 seconds before NEBD, and then subsequently moving to the periphery during condensation (with full condensation at 100 s past NEBD). This suggests that initial condensation may occur both within the nucleus and at the periphery. Over time, chromosomes condense and cluster at the nuclear envelope. However, we did not separately measure the condensation level of individual chromosomes at the nuclear periphery versus in the middle of the nucleus, which could be challenging in live cells. So far, we cannot separate the chromosome nuclear periphery phenotype and the condensation.

      To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      We recognize that npp-3 RNAi embryos exhibit broad developmental defects, and therefore global transcriptional changes. Our RNA-seq analysis revealed that differentially expressed genes did not display positional bias within the genome. NPP-3 depletion downregulates many pathways, including pathways related to RNA polymerase II activity and cell cycle regulation, consistent with a global transcriptional downregulation. Thus, while the transcriptomic data are broad, it is consistent with the chromosome condensation phenotype.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      In our co-depletion experiments, the images for npp-3(RNAi) and X(RNAi) (Fig. 4) are displayed side-by-side under the same exposure conditions and scaled the same way with the same intensity thresholds to facilitate comparison. Despite that, some differences in the background intensity can be observed across samples. Thus, the normalised mCherry::NPP-3 signal intensity (subtracting the background) the single and double RNAi samples will be quantified and added to supplementary figure S4.

      To ensure consistency of double RNAi, we used ligation-based RNAi constructs designed to simultaneously target both NPP-3 and X, aiming to achieve comparable knockdown efficiencies. Nonetheless, variability in RNAi efficiency is a recognized limitation, and we interpret our data within this context. To further validate the knockdown, we will also assess the efficiency of the other gene X using the corresponding fluorescent reporter (see response to Reviewer 3, point 2).

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      Inactivation of MDF-1 or MDF-2 suppresses both the chromosome nuclear periphery localization and extended duration of interphase to prophase and prometaphase in npp-3 inactivation. We have performed the analysis to assess the timing and extent of chromosome condensation in the single and double depletion conditions (Author response image 4). The chromosome condensation dynamics is similar in the mdf-1 RNAi and npp-3 mdf-1 double RNAi, as well as the control group, indicating that MDF-1 is also involved the premature chromosome condensation phenotype caused by NPP-3 depletion.

      Author response image 4.

      The dynamic changes of chromosome condensation parameter, in which 30% of pixels in the ROI analyzed is below the threshold scaled intensity (<77), in the different groups. The sample size is 5. Error bars show mean ± SEM.

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      We have aligned the nuclear envelope reassembly (NER) time and highlighted the time point in the images (Fig. 5A). Our data show that in control embryos, this duration is approximately 780 seconds, while in npp-3(RNAi) embryos, it extends to about 930 seconds. This difference is statistically significant, as determined by one-way ANOVA (or appropriate nonparametric/mixed tests). The representative image is consistent with the quantification presented in Fig. 5B. We could add the corresponding videos to the supplementary information.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

      It is correct that the average embryonic lethality observed in npp-3(RNAi) embryos reachs 95% (Fig. S6). Given the broad developmental defects and such high baseline lethality, the genetic interaction with mdf-1 is modest.

      Nevertheless, our findings highlight that in the npp-3 mdf-1 double RNAi condition, we observe significantly increased rates of lagging chromosomes (Fig. 5D), micronuclei formation (Fig. 7A), and elevated DNA damage (Fig. 7B). These effects support the idea that MDF-1-, MDF-2dependent chromosome anchoring to the nuclear periphery (and condensation) plays a positive role in NPP-3 depleted cells.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP- 3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically codepleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Codepletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

      The DNA damage evidence is based on lagging chromosomes in 2-cell stages, the HUS-1 reporter, and micronuclei in embryos at the 20-30 cell stage. It is noted that npp-3 RNAi is pleiotropic and also causes DNA damage. There is additional DNA damage caused by loss of MDF-1 and MDF-2 in npp-3 RNAi, but we agree that it is difficult to say whether the effect is additive or not, complicating the interpretation. Thus, we will tone down in our title to describe the dependency of the chromosomal nuclear periphery phenotype (see response to Reviewer 1 point 1).

      As for the premature chromosome condensation phenotype in NPP-3 depletion, we hypothesize it may result from accumulation of factors such as BAF-1 at the nuclear periphery, which could facilitate chromatin condensation. To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization (also see response to Reviewer 1 point 6).

      Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP- 3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Codepletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. , they unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SACdependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      We agree that the current data are largely correlative. In the early C. elegans embryos, single depletion of SAC components MDF-1 or MDF-2 does not affect the mitosis duration or chromosome segregation. When spindles are defective, the functional SAC delays progression through mitosis [5]. Depleting SAC components such as MDF-1 in npp-3 RNAi indeed impacts multiple processes, including prophase and prometaphase duration, chromosome repositioning and condensation, making it challenging to disentangle effects specifically to related chromosome positioning.

      To address this, future experiments involving NPP-3 LEM-2 co-depletion, which has been shown to partially impair peripheral chromosome positioning, will be utilized to assess whether disruption of partial peripheral localization results in increased DNA damage, micronuclei formation, or HUS-1 foci accumulation, independent of SAC function (see response to Reviewer 1 point 1). 

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      We acknowledge that in our double RNAi experiments, the knockdown efficiency of the npp-3 is checked by imaging (see response to Reviewer 1 point 8), whereas that of the second gene was not validated in each condition. To address this, we confirmed the effectiveness of certain partner gene depletions, e.g. HCP-3, by examining the levels of the respective proteins using available GFP-marked strains or immunofluorescence (Author response image 5). Additionally, for genes like HCP-3, KNL-1, and BUB-1, we also assessed the functional consequences on chromosome segregation, where severe defects observed (Fig. S4C) can support effective depletion. However, for some strains, we do not have GFP makers and will need to check the RNA levels. 

      Author response image 5.

      Representative confocal image of GFP::HCP-3 and mCherry::H2B at the NEBD time point of P1 cell in the control, single and double RNAi. Scale bar, 5 μm.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      We agree that these chromatin modifications and transcriptional alterations could be related to the nuclear transport failure and developmental delay in npp-3 disruption. We will discuss this possibility in results and discussion. 

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      We agree that the evidence for SAC activation during prophase is indirect. The prophase extension (100 s) in npp-3 RNAi is a functional assay to support SAC activation, and the suppression in npp-3 mdf-1 double RNAi suggests dependency. The changes in MDF1/MDF-2 intensities are modest. Biochemical analyses of MCC assembly in C. elegans mixed cell cycle stage embryos is challenging to demonstrate SAC activity in prophase. 

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf- 1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF- 1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

      We did not include the BAF-1/CENP-C bridge hypothesis in the abstract or the model, and we will discuss this as a speculative mechanism rather than an established dependency.

      References:

      (1) Rodenas, E., Gonzalez-Aguilera, C., Ayuso, C. & Askjaer, P. Dissection of the NUP107 nuclear pore subcomplex reveals a novel interaction with spindle assembly checkpoint protein MAD1 in Caenorhabditis elegans. Mol Biol Cell 23, 930-944 (2012).

      (2) Schumacher, J.M., Ashcroft, N., Donovan, P.J. & Golden, A. A highly conserved centrosomal kinase, AIR-1, is required for accurate cell cycle progression and segregation of developmental factors in Caenorhabditis elegans embryos. Development 125, 4391-4402 (1998).

      (3) Kotak, S., Afshar, K., Busso, C. & Gonczy, P. Aurora A kinase regulates proper spindle positioning in C. elegans and in human cells. J Cell Sci 129, 3015-3025 (2016).

      (4) Hachet, V. et al. The nucleoporin Nup205/NPP-3 is lost near centrosomes at mitotic onset and can modulate the timing of this process in Caenorhabditis elegans embryos. Molecular Biology of the Cell 23, 3111-3121 (2012).

      (5) Encalada, S.E., Willis, J., Lyczak, R. & Bowerman, B. A spindle checkpoint functions during mitosis in the early Caenorhabditis elegans embryo. Molecular Biology of the Cell 16, 1056-1070 (2005).

    1. eLife Assessment

      This useful study employs an innovative chemoproteomic and multi-modal experimental approach to support VDAC2 as a key functional mediator of STX effects on mitochondrial bioenergetics. However, concerns remain regarding the physiological and disease relevance, including the reliance on cell lines, lack of loss-of-function validation, unresolved binding specificity, and limited justification for focusing on POMC neurons in the context of Alzheimer's disease. As a result, the evidence is incomplete for a link between VDAC2 and Alzheimer's pathogenesis, and aspects of the binding data raise alternative interpretations, including a potential role for other VDAC isoforms.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Qiu et al. examine the effects of the estrogen mimic STX on mitochondrial function and its interaction with VDAC2 in PMOC neurons.

      Strengths:

      The authors employ a broad range of molecular, cellular, and chemoproteomic approaches with generally sound methodology.

      Weaknesses:

      The work suffers from major conceptual and experimental issues that substantially limit its scientific impact.

      Major Concerns

      (1) Lack of Rationale.<br /> The study provides no justification for investigating sex specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in PMOC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      (2) Weak Link to AD Pathogenesis.<br /> Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      (3) Unclear Relevance to AD Contexts.<br /> While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in PMOC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex specific AD phenotypes.

      (4) Interpretation of Competitive Binding Data.<br /> The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

    3. Reviewer #2 (Public review):

      Summary:

      STX is a non-steroidal, CNS-selective estrogenic compound with neuroprotective effects in stroke and Alzheimer's disease models, but its molecular target has remained unknown for nearly 20 years. In this study, the authors identify VDAC proteins as the direct mitochondrial targets of STX using chemoproteomics, single-cell qPCR, electrophysiology, and metabolic flux analyses. They further show that VDAC2 is the primary functional target in female POMC neurons, linking STX-mediated VDAC modulation to enhanced mitochondrial bioenergetics and neuroprotection.

      Strengths:

      This study is strengthened by its innovative chemoproteomic approach, in which the authors developed a novel bifunctional STX probe (BF-STX) containing a photo-crosslinkable diazirine group and an alkyne handle to capture transient STX-protein interactions in living cells. The experimental design is further reinforced by rigorous controls, including no-UV negative controls and competition assays with excess unlabeled STX, which provide convincing evidence that VDAC1, VDAC2, and VDAC3 are genuine STX-binding targets rather than nonspecific artifacts. Finally, the authors validate the STX-VDAC interaction using multiple complementary approaches, including chemoproteomics, single-cell qPCR, planar lipid membrane electrophysiology, and Seahorse metabolic flux analyses, providing strong mechanistic support for their conclusions.

      Weaknesses:

      While the study provides convincing evidence that STX directly modulates VDAC function, several limitations remain. Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance. In addition, the exact structural binding site of STX on VDAC remains unresolved, and no loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects. The non-linear dose-response at higher STX concentrations also requires further investigation.

    4. Author response:

      Reviewer #1:

      We thank Reviewer #1 for the thoughtful critique. We have revised the manuscript to clarify the physiological rationale for the experimental system, the potential relevance of mitochondrial STX–VDAC signaling to neurodegeneration, and the appropriate scope of our conclusions.

      (1) Lack of Rationale

      The study provides no justification for investigating sex-specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in POMC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      We thank the reviewer for raising this important point. We agree that the original manuscript did not sufficiently explain the rationale for using POMC neurons.

      Our rationale is primarily physiological rather than disease-specific. POMC neurons are a well-characterized, metabolically sensitive, and estrogen-responsive neuronal population in which membrane-initiated estrogen signaling and the actions of STX have been extensively characterized. They therefore provide a physiologically relevant neuronal system for identifying the molecular mechanisms through which STX regulates mitochondrial function.

      The potential relevance to neurodegeneration is supported by evidence that hypothalamic and POMC neuronal function can be disrupted in neurodegenerative disease models. Do and colleagues (2018) reported hypothalamic neurodegeneration, increased inflammatory and apoptotic markers, and reduced POMC neuronal populations in 3xTg-AD mice and further showed that exercise attenuated hypothalamic apoptosis and restored POMC neuronal populations (Do, Laing et al. 2018). In addition, Shen and colleagues (2016) demonstrated disruption of POMC/MC4R signaling in APP/PS1 mice and showed that restoration of this pathway improved synaptic function (Shen, Tian et al. 2016).

      These studies do not establish POMC neurons as a primary site of AD pathology, but they demonstrate that POMC-related neuronal systems can be vulnerable to neurodegenerative processes. This provides a broader biological context for investigating mitochondrial mechanisms in this neuronal population.

      There is also a strong physiological rationale for examining estrogen-sensitive mechanisms in POMC neurons. These neurons are established targets of 17β-estradiol and are highly responsive to metabolic and mitochondrial state. STX is a non-steroidal estrogenic compound that activates membrane-initiated estrogen signaling, and our previous studies demonstrated neuroprotective and mitochondrial effects of STX in a neurodegenerative disease model.

      Thus, POMC neurons were used because they provide a well-defined estrogen-responsive and metabolically sensitive neuronal population in which STX signaling can be mechanistically investigated—not because we consider them an initiating site of AD pathology.

      We have revised the Introduction and Discussion accordingly. The revised manuscript now emphasizes the physiological significance of STX–VDAC signaling for neuronal mitochondrial function and presents its potential relevance to neurodegeneration as an important area for future investigation.

      (2) Weak Link to AD Pathogenesis

      Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      We agree that the present experiments do not establish VDAC2 as an etiological or pathological driver of AD. This was not the objective of the present study.

      The primary goal was to identify the molecular target(s) through which STX influences mitochondrial function. Using BF-STX chemoproteomic capture and competition experiments together with single-cell gene-expression analysis, planar lipid membrane electrophysiology, and mitochondrial metabolic analyses, we identify VDAC proteins as mitochondrial targets of STX and demonstrate functional effects of STX on VDAC channel properties and mitochondrial bioenergetics.

      Our previous studies demonstrated neuroprotective and mitochondrial effects of STX in the 5xFAD model (Lee, Bostick et al. 2025), providing a neurodegenerative context that motivated the present mechanistic investigation. However, the current findings should not be interpreted as establishing VDAC2 as an AD pathogenic mechanism.

      We have revised the manuscript accordingly. The STX–VDAC interaction is now presented principally as a mitochondrial mechanism with potential relevance to neuronal physiology and neurodegeneration. Whether this pathway contributes to the neuroprotective actions of STX in disease models will require direct experimental testing.

      (3) Unclear Relevance to AD Contexts

      While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in POMC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex-specific AD phenotypes.

      We agree that the present experiments do not directly establish the relevance of the STX–VDAC interaction to mitochondrial dysfunction in AD.

      The current study was designed to identify and characterize the molecular mechanism underlying the mitochondrial actions of STX, rather than to test this pathway in a specific neurodegenerative disease model. Our findings demonstrate that STX interacts with VDAC proteins, modifies VDAC channel properties, and alters mitochondrial bioenergetics.

      The study builds on our previous findings in the 5xFAD model, in which STX reduced amyloid-β-associated pathology and affected mitochondrial function (Lee, Bostick et al. 2025). The identification of VDAC proteins as STX targets therefore provides a mechanistic foundation for future studies examining whether this pathway contributes to neuronal protection under neurodegenerative conditions.

      We have revised the Discussion to acknowledge the absence of direct disease-model validation. Future studies using conditional or neuron-specific manipulation of VDAC isoforms will be important for determining the physiological and neuroprotective significance of STX–VDAC signaling in vivo.

      Accordingly, our conclusions now emphasize what is directly supported by the present experiments: STX interacts with VDAC proteins and modulates VDAC channel function and mitochondrial bioenergetics. The relevance of this mechanism to neurodegenerative disease remains to be established.

      (4) Interpretation of Competitive Binding Data

      The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

      We thank the reviewer for highlighting this important point. We agree that the original manuscript placed too much emphasis on VDAC2 based on the chemoproteomic data.

      Our chemoproteomic experiments identified VDAC1, VDAC2, and VDAC3 as STX-interacting proteins. Competition with unlabeled STX produced a particularly clear reduction in VDAC3 labeling, including complete loss of the VDAC3 signal at three molar equivalents of unlabeled STX. We therefore agree that VDAC3 represents an important candidate STX target.

      We have revised the Results and Discussion so that the competition experiment is no longer interpreted as demonstrating preferential or exclusive binding to VDAC2.

      However, the persistence of VDAC1 and VDAC2 labeling may also be influenced by properties of the BF-STX photoprobe as alkyl diazirine probes can preferentially photolabel membrane proteins (Kleiner, Heydenreuter et al. 2017). We now present this only as a possible technical consideration rather than an explanation established by our data.

      Our subsequent emphasis on VDAC2 was based on the integration of several observations. Single-cell qPCR demonstrated that Vdac2 is the predominant transcript in the native POMC neurons examined (revised Figure 3B), with an approximate expression hierarchy of Vdac2 > Vdac3 >> Vdac1. In addition on a technical note, recombinant VDAC3 is more difficult to reconstitute reliably into artificial membranes because of its lower stability in detergent.

      The revised manuscript therefore recognizes all three VDAC isoforms as candidate STX targets but does not claim that VDAC2 is the exclusive or preferential target. Direct comparisons of STX binding and functional modulation among the three isoforms will be required to establish isoform selectivity.

      Reviewer #2:

      We thank Reviewer #2 for the positive assessment of our chemoproteomic strategy and multidisciplinary characterization of the STX–VDAC interaction. We also appreciate the reviewer’s identification of important limitations and directions for future investigation.

      (1) Physiological Relevance of Cell-Line Experiments

      Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance.

      We agree that the use of immortalized neuronal cell models represents an important limitation.

      However, these models provided the cellular material, reproducibility, and experimental control required for chemoproteomic target capture and Seahorse metabolic measurements. Importantly, we complemented these studies with single-cell analysis of native POMC neurons and electrophysiological characterization of recombinant VDAC channels.

      Nevertheless, these approaches do not substitute for direct demonstration of STX–VDAC signaling in primary POMC neurons or in vivo. We have therefore revised the Discussion to emphasize that establishing the physiological significance of this mitochondrial pathway will require validation in primary neuronal preparations and whole-animal models.

      (2) Structural Binding Site of STX on VDAC

      The exact structural binding site of STX on VDAC remains unresolved.

      We agree. BF-STX chemoproteomics identifies proteins interacting with STX in a cellular environment but does not resolve the amino acid residues or structural pocket responsible for binding.

      We have clarified this limitation in the Discussion. Determining the STX-binding site will require complementary approaches such as targeted mutagenesis, direct binding measurements with purified VDAC proteins, and structural studies in membrane-like environments, including lipid nanodiscs. Such studies should also determine whether STX recognizes a conserved feature among VDAC isoforms or exhibits isoform selectivity.

      (3) Lack of VDAC2 Loss-of-Function Experiments

      No loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects.

      We agree that VDAC loss-of-function experiments would provide an important additional test of causality.

      Our present evidence is convergent: chemoproteomics identifies VDAC proteins as STX-interacting targets; single-cell analyses demonstrate VDAC expression in POMC neurons; electrophysiological studies show that STX modifies VDAC channel properties; and metabolic analyses demonstrate STX-dependent changes in mitochondrial bioenergetics. Together, these findings support a STX–VDAC mechanism but do not establish that VDAC2 alone is necessary for the mitochondrial or neuroprotective actions of STX.

      We have revised the manuscript accordingly and explicitly identify the absence of loss-of-function experiments as a limitation.

      Because whole-body VDAC2 loss-of-function is associated with severe developmental consequences, future studies will require conditional or neuron-specific approaches. Parallel evaluation of VDAC1 and VDAC3 will also be important because all three isoforms were identified by chemoproteomics and potential functional redundancy may complicate single-isoform manipulations.

      (4) Non-linear Dose-Response at Higher STX Concentrations

      The non-linear dose-response at higher STX concentrations also requires further investigation.

      We agree that the non-linear concentration-response warrants further investigation and have revised the Discussion to avoid interpreting the STX response as a simple monotonic concentration-response relationship.

      One possibility is that STX engages multiple mitochondrial targets with different apparent affinities and opposing effects on respiration. At lower concentrations, STX may preferentially engage a higher-affinity target, potentially VDAC, whereas higher concentrations may recruit lower-affinity targets that constrain this response. Consistent with this possibility, our BF-STX dataset (Supplemental Tables) identified several mitochondrial proteins involved in oxidative phosphorylation and metabolite transport, including ATP5F1C, NNT, SLC25A4, SLC25A5 and SLC25A3. ATP5F1C is required for efficient mitochondrial ATP production (Fiorillo, Scatena et al. 2021), and estrogenic regulation of ATP synthase has been reported (Massart, Paolini et al. 2002, Moreno, Moreira et al. 2013). NNT, SLC25A4, SLC25A5 and SLC25A3 also regulate mitochondrial respiration, redox balance, and ATP production (Mayr, Merkel et al. 2007, Lopert and Patel 2014). However, BF-STX enrichment does not establish direct STX binding, relative affinity, or functional modulation of these proteins. We therefore present the multi-target explanation only as a hypothesis. Direct binding and concentration-dependent target-engagement studies will be required to determine whether the non-linear response reflects recruitment of a lower-affinity mitochondrial target, concentration-dependent effects on VDAC itself, or downstream mitochondrial feedback.

      We thank the editors and reviewers again for their constructive comments. The revised manuscript more clearly defines the physiological rationale for the experimental system, establishes the scope of the STX–VDAC mitochondrial mechanism supported by our data, and places its potential relevance to neurodegeneration in an appropriately forward-looking context.

      References cited in the response

      Do, K., B. T. Laing, T. Landry, W. Bunner, N. Mersaud, T. Matsubara, P. Li, Y. Yuan, Q. Lu and H. Huang (2018). "The effects of exercise on hypothalamic neurodegeneration of Alzheimer's disease mouse model." PLoS One 13(1): e0190205.

      Fiorillo, M., C. Scatena, A. G. Naccarato, F. Sotgia and M. P. Lisanti (2021). "Bedaquiline, an FDA-approved drug, inhibits mitochondrial ATP production and metastasis in vivo, by targeting the gamma subunit (ATP5F1C) of the ATP synthase." Cell Death Differ 28(9): 2797-2817.

      Kleiner, P., W. Heydenreuter, M. Stahl, V. S. Korotkov and S. A. Sieber (2017). "A Whole Proteome Inventory of Background Photocrosslinker Binding." Angew Chem Int Ed Engl 56(5): 1396-1401.

      Lee, H.-J., Z. Bostick, J. Doherty, T. L. Swanson, M. J. Kelly, J. F. Quinn, N. E. Gray and P. F. Copenhaver (2025). "Neuroprotection against beta-amyloid toxicity by the novel estrogen receptor modulator STX requires convergent signaling pathways." Frontiers in Molecular Neuroscience Volume 18 - 2025.

      Lopert, P. and M. Patel (2014). "Nicotinamide nucleotide transhydrogenase (Nnt) links the substrate requirement in brain mitochondria for hydrogen peroxide removal to the thioredoxin/peroxiredoxin (Trx/Prx) system." J Biol Chem 289(22): 15611-15620.

      Massart, F., S. Paolini, E. Piscitelli, M. L. Brandi and G. Solaini (2002). "Dose-dependent inhibition of mitochondrial ATP synthase by 17 beta-estradiol." Gynecol Endocrinol 16(5): 373-377.

      Mayr, J. A., O. Merkel, S. D. Kohlwein, B. R. Gebhardt, H. Böhles, U. Fötschl, J. Koch, M. Jaksch, H. Lochmüller, R. Horváth, P. Freisinger and W. Sperl (2007). "Mitochondrial phosphate-carrier deficiency: a novel disorder of oxidative phosphorylation." Am J Hum Genet 80(3): 478-484.

      Moreno, A. J., P. I. Moreira, J. B. Custódio and M. S. Santos (2013). "Mechanism of inhibition of mitochondrial ATP synthase by 17β-estradiol." J Bioenerg Biomembr 45(3): 261-270.

      Shen, Y., M. Tian, Y. Zheng, F. Gong, A. K. Y. Fu and N. Y. Ip (2016). "Stimulation of the Hippocampal POMC/MC4R Circuit Alleviates Synaptic Plasticity Impairment in an Alzheimer's Disease Model." Cell Rep 17(7): 1819-1831.

    1. eLife Assessment

      This revised study offers valuable insights into how nuclear export influences protein condensates and TDP-43 phase behaviour. The findings are solid and highlight several noteworthy observations for the field, such as RNA-dependent stabilisation of TDP-43 condensates and the inhibition of nuclear export in an ALS organoid model. The work suggests a possible mechanistic link between nuclear export and TDP-43 aggregation in ALS/FTD; however, the exact mechanistic connection remains unclear, as much of the evidence is based on indirect observations within sensitised model systems. Although the authors carefully acknowledge these limitations and moderate their conclusions, further validation in more physiologically relevant models will be necessary to demonstrate a direct causal role of nuclear export in regulating pathological TDP-43 aggregation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.<br /> The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.<br /> The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.<br /> The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.<br /> The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.<br /> The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      (6) No functional neuronal readout has been provided for the organoid model.<br /> The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      (7) The abstract and title continue to overstate the mechanistic conclusions.<br /> Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.<br /> The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.<br /> The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

    3. Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:<br /> a) What is the smear in Figure S1 after VLX treatment?<br /> b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.<br /> c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.<br /> d) Why did the authors choose cow lover cytosol for this study?<br /> e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.<br /> f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      (6) Several figure legends require clarification:<br /> a) In the section stating "Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.<br /> b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.<br /> c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.<br /> d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."<br /> e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.<br /> f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.<br /> g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."<br /> h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R², p-value).<br /> i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.<br /> j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

    4. Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.

      The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      We agree with the reviewer that the TDP-43 2KQ mutant is a non-physiological variant. To address this concern, we have removed all statements that extrapolate our findings to disease pathogenesis. As reflected in the revised title, we now present this study as an investigation of factors that modulate TDP-43 phase transition and aggregation using a sensitized model system, rather than as a direct disease model. The Abstract and Discussion have also been revised accordingly.

      The choice of the 2KQ mutant was dictated by the requirements of the screening strategy. To identify modulators of TDP-43 phase transition, it was necessary to use a TDP-43 variant that reliably undergoes phase separation in cells within an experimentally practical timeframe. In this context, we believe the 2KQ mutant provides a suitable and justified experimental tool. We have revised the manuscript to clearly distinguish observations made with this engineered construct from conclusions regarding physiological or disease-associated TDP-43.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.

      The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      We agree with the reviewer and have removed all speculative statements regarding the relative contributions of cytoplasmic aggregation and RNA splicing defects to disease pathogenesis. The organoid section has also been revised to focus solely on the experimental findings. Specifically, we only present evidence that inhibition of nuclear export in a sensitized organoid model promotes the accumulation of cytoplasmic phosphorylated TDP-43 without making broader claims regarding its role in disease pathogenesis.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.

      The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      We have now added some sentences on page 11 to acknowledge the limitation of our experiments. It reads as “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.”

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.

      The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      We thank the reviewer for this point. We have now explicitly mentioned in the discussion that the effect of XPO-1 on anisosome dynamics is likely mediated by an indirect mechanism. We also acknowledge that we do not fully understand why over-expression of XPO-1 causes TDP-43 to accumulate in gel-like structures in the cytoplasm. Although we did not check whether overexpression of XPO1 increases cargo export, we cited studies showing that over-expressed XPO1 disrupts the normal distribution of cargos between nucleus and cytoplasm on page 7. To avoid confusion, we also revised the result part on page 7, emphasizing on the difference rather than the similar increase in puncta size by opposing manipulations.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.

      The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      We removed the RNA-seq analysis because both the reviewers and editors agreed that the modest transcriptional rescue should not be overinterpreted. Upon reconsideration, we believe that reinstating these data would not address the reviewer's principal concern, namely whether nuclear export inhibition reduces TDP-43 aggregation in organoids. We have therefore chosen to limit our conclusions to the direct observation supported by the current data, namely a reduction in cytoplasmic phosphorylated TDP-43 immunoreactivity. We also add a sentence to acknowledge that “Whether the reduction in pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined” on page 10.

      (6) No functional neuronal readout has been provided for the organoid model.

      The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      We thank the reviewer for this helpful suggestion. We agree that assessments of neuronal survival or function would be important if the manuscript were claiming that nuclear export inhibition improves neuronal health or rescues disease phenotypes in the organoid model. However, in response to the reviewers' comments regarding the physiological relevance of the homozygous K181E organoids, we have substantially revised both the framing and interpretation of this section.

      Specifically, we have revised the statement to read, "These results imply that maintaining TDP-43 in the nuclear demixed liquid state might diminish p-TDP-43 accumulation but whether the reduction of pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined," thereby limiting our conclusion to the direct experimental observation. We no longer make claims regarding disease modification or functional rescue in the organoid model. Given this revised scope, we believe that additional measurements of neuronal survival or function, while certainly of interest, are not essential to support the conclusions presented in this study. We have also revised the conclusion to explicitly acknowledge that the functional consequences of reducing p-TDP-43-positive puncta (whether this can be translated to reduced aggregation) remain to be determined by future studies (page 10).

      (7) The abstract and title continue to overstate the mechanistic conclusions.

      Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      We have revised the title of the paper, reframing it as a screen that reveals modulators of TDP-43 phase separation. The last sentence of the abstract is also revised accordingly. It now reads as “These findings identify multiple modulators of TDP-43 phase transitions in a sensitized model system and establish a framework for further dissecting the link between nuclear transport and TDP-43 phase dynamics.” We also tone down our conclusions and discussions.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.

      The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      While we agree with the reviewer that studies relying on siRNA should provide sufficient information regarding knockdown efficiency and specificity, we respectfully disagree that this should be a major concern in the present study. As explained in the manuscript, we deliberately chose not to pursue mechanistic studies using chronic XPO1 knockdown because prolonged depletion of this essential nuclear export factor is likely to produce secondary effects that could complicate data interpretation. Instead, we employed multiple chemically distinct XPO1 inhibitors to achieve acute inhibition, thereby minimizing indirect consequences while providing a more appropriate approach for mechanistic analysis.

      We agree that assessing knockdown efficiency is technically straightforward. However, because our mechanistic conclusions are based primarily on acute pharmacological inhibition rather than siRNA-mediated depletion, we prioritized experiments that directly addressed the central mechanistic questions raised by the reviewers, particularly the semi-permeabilized cell assay. Moreover, the XPO1 inhibitors used in this study are well-characterized, highly specific compounds that have been extensively validated and widely used in the literature. We therefore believe that our experimental strategy provides a reliable basis for the conclusions presented.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.

      The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      We thank the reviewer for this helpful suggestion. We have now added a sentence on page 11, which state that “our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires validation.”

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

      We thank the reviewer for his/her appreciation of our previous revision. We hope that the new changes now satisfactorily address the remaining concerns.

      Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      We thank the reviewer for acknowledging the strength and the potential significance of our study.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:

      (a) What is the smear in Figure S1 after VLX treatment?

      We thank the reviewers for the positive assessment. We do not know why VLX treatment causes a fraction of TDP-43 to migrate slowly. We suspect that it may form detergent-insoluble aggregates. However, we cannot be sure whether this occurred during drug treatment or sample preparation. We now add a sentence in the figure legend to clarify this point.

      (b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.

      We reasoned that the reduction in TDP-43 was probably not caused by protein elimination because the puncta could be reformed when permeabilized cells were incubated with exogenously added cytosol and ATP/GTP. We have revised the text to avoid this confusion. The revision on page 8 reads as “The reduction in TDP-43 signal probably resulted from a shift of TDP-43 from a phase-separated high fluorescent state into a soluble state with reduced fluorescence intensity (Zhang et al., 2026). We attributed this phenotype to the depletion of cytosolic factors and ATP during cell permeabilization because it is known that anisosome formation and maintenance require HSP70, a cytosolic ATPase (Yu et al., 2021).”.

      (c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.

      Since TDP-43 protein was apparently still in the nucleus after cell permeabilization (see above) but became invisible, the best interpretation is that the protein is shifted into a soluble state, which reduces the fluorescence intensity substantially. We have revised the text to clarify this point. We also cited a recent study showing that EGFP-alpha-synuclein oligomerization/aggregation enhances its fluorescence intensity in cells.

      (d) Why did the authors choose cow lover cytosol for this study?

      The main reason is because we have access to a large amount of cow liver cytosol that is known to have activities in in vitro reconstitution assays. We now cite several papers from us that reported the use of the same cytosol in other in vitro assays in the method (page 14). 

      (e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.

      We now revise the schematic in Figure 6A and include more details in the method and figure legend.

      (f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      The total TDP-43 level was shown by immunostaining in green in Figure 7. This was used as a control to show that the increase in p-TDP-43 was not simply caused by an overall increase in its protein level. We have added the quantification to Figure 7C. We also discuss the reduced nuclear TDP-43 in organoids bearing the disease mutation in the main text.  

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      We agree that the structures reformed after incubating permeabilized cells with cytosol and ATP/GTP are not fully characterized. Due to their small size, we could not see the typical ring-shaped anisosome morphology. FRAP experiment is also tricky. Due to these issues, we have revised the text to acknowledge that we do not know the exact identity of these structures. We speculate that they are anisosome-related because like anisosome formation, it depends on cytosolic factor and energy (page 9). It is worth noting that whether these structures are anisosomes is not the main conclusion of this experiment. We conclude from this experiment that TDP-43 was still in the nucleus after cell permeabilization (not degraded). The fact that we could not see the protein likely because the protein was in a low-fluorescence soluble state.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      Since the formation of anisosome requires HSP70, a cytosolic chaperone that likely needs to be imported into the nucleus, we included cytosol and ARS/GTP in our in vitro reaction. We revise the description in the result part to improve clarity (page 8-9).

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      As discussed in Yu H et al., Science 2021, proteins in anisosomes cannot be stained by antibodies due to an antibody accessibility issue. It was mentioned in our paper as “since antibody staining could not conclusively demonstrate the sequestration of endogenous XPO-1 in anisosomes due to an antibody penetration barrier {Yu, 2021 #937}.” We now revise this section completely to better clarify this point. We could see overexpressed XPO1 in anisosome because it has a mCherry tag.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      The full characterization of the iPSC-derived organoids is presented in a second paper that is posted in BioRxiv (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v2), which is cited. This manuscript reports not only the accumulation of p-TDP43, but also other ALS-related phenotypes including cell death, gene transcriptional changes, cryptic exon inclusion etc. in mutant organoids.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      The longer treatment (24 h) was used in the chemical genetic screen in which different drugs may act with different efficiency. To maximize our chance of detecting more drug effect, we used a longer treatment scheme. For later follow-up experiments involving Spuatin-1, Bortezomib, TRP, because the phenotype appears quickly. To avoid secondary effects from long treatment, we shortened the treatment to 3-5 hours. For LMB treatment, we used long treatment to reveal the steady-state phenotype (anisosome enlargement in size and reduction in number has reached maximum). This time point was determined in Figure 5A-C. In contrast, shorter treatment (5 h) was to reveal early changes that might be causal to the end-point phenotypes (e.g. the anisosome fusion phenotype could be detected as early as 5 h post-treatment). We have added some explanations in the result section to make this point clear.

      (6) Several figure legends require clarification:

      We thank the reviewer for pointing out the inconsistencies. We have corrected the outstanding issues, as explained below.

      (a) In the section stating “Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.

      Thanks for pointing out this error. This is now corrected.

      (b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.

      Due to the large sample size, each concentration was analyzed once. We have added other information to the figure legend. We also remove the redundant inaccurate information from the figure legend. The treatment time shown in the figure is correct as it was also indicated in the method.

      (c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.

      We thank the reviewer for noticing the discrepancy and apologize for the error. We have corrected the figure legend to indicate that 4-6 anisosomes were analyzed for each condition. In Figure 2G, each data point represents the initial fluorescence loss rate averaged from the first 10 sec after reverse photobleaching.

      (d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."

      In Figure 3B, we used Z score >2 as the threshold. This is now mentioned in the legend and defined in the method. There is no Figure 3G.

      (e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.

      We have changed the figure legend throughout the paper accordingly.

      (f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.

      In Figure 5B, the graph reflects data collected from 3 independent replicates. To ensure reliable baseline measurement, for each experiment, two independent control samples were included, which is why it has 6 data points. In Figure 5H, each dot represents a randomly selected imaging field. We now mention the total number of fields analyzed.

      (g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."

      These are all fixed. Thank you for pointing this out.

      (h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R<sup>2</sup>, p-value).

      We now include the cell number in the legend and R<sup>2</sup> and p-value in the figure.

      (i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.

      We now revise the experimental scheme in Figure 6A to better explain the experiment and the sequence of different events. For Figure 6D, permeabilized cells were incubated with cytosol with or without ARS/GTP for 40 min. This information is added to the figure legend.

      (j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

      To improve clarity, we change mean spot volume to p-TDP-43 puncta mean volume. This refers to the average volume of segmented phosphorylated TDP-43-positive puncta.  

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

      We thank the reviewer for this helpful suggestion. We have added a few sentences in the discussion (page 10) to acknowledge the limitation of the semi-permeabilized cell assay. Specifically, we mentioned that “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.” We also revise our manuscript throughout to avoid over-interpretation.  

      Recommendations for the authors:

      Editor's notes:

      The value of the work is clear.

      We also recognise that it may not be possible/practical to get around the 'incomplete' appellation attached to this body of work, by further experiments. However, there may be scope here to retreat from the less well supported mechanistic claims-by editing the title, abstract and discussion and thus earn a 'solid' descriptor on a revised paper that remains a useful addition.

      The authors are best placed to consider creatively how to achieve this.

      We thank the editors for this helpful suggestion. We have revised the manuscript extensively to address every single concerns of the reviewers.

      Reviewer #1 (Recommendations for the authors):

      This revision addresses some minor concerns and adds one mechanistic experiment of genuine value. However, the major deficiencies of the original submission persist: the exclusive reliance on non-physiological TDP-43 model systems without validation in more disease-relevant contexts, the unresolved asymmetry between the two XPO1 perturbation phenotypes, the thin organoid section that now has fewer readouts than before the revision, and overstatement of the mechanistic conclusions in the abstract and Discussion. The manuscript in its current form still does not provide methods, data, and analyses that sufficiently support the primary claim that nuclear export is an established key regulator of TDP-43 phase transitions with mechanistic and disease relevance.

      As mentioned before, we have clarified the interpretation of the data, tempered conclusions where appropriate, and revised the text to explicitly acknowledge the limitations of the current study. We hope that these changes satisfactorily addressed the reviewer’s concern.

      Reviewer #2 (Recommendations for the authors):

      I would suggest adjusting the title to match the data, which shows so much more than just an effect of nuclear export.

      We thank the reviewer for this suggestion. We have changed the title to “Cellular modifiers of TDP-43 phase transition and cytoplasmic aggregation”

      When introducing the semi-permeabilized cell-based in vitro assay, it would be helpful to add a short statement describing what this model resembles, and what advantage it can bring to use this system in the context of the study.

      We have revised this section extensively and hope that improves the clarity. See marked text in page 8-9.

    1. eLife Assessment

      The study presents valuable findings of a new E. coli cell-free protein synthesis (eCFPS) system that has been simplified by reducing the number of core components from 35 to 7; furthermore, the findings communicate a simplified 'fast lysate' preparation that eliminates the need for traditional runoff and dialysis steps. The system's robustness is exhibited by its applicability to nanoluc, an ubiquitous protein and to more challenging proteins like vimentin and the active restriction endonuclease Bsal. The evidence is convincing and supports the main claims on efficiency backed up by investigations on the mechanisms. The paper will be of interest to scientists in cell and molecular biology, microbiology, biotechnology and protein synthesis.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The authors presented a simplified E. coli cell-free protein synthesis (eCFPS) system reduces core reaction components from 35 to 7, improving protein expression levels. They also presented a "fast lysate" protocol that simplifies extract preparation, enhancing accessibility and robustness for diverse applications.

      Strengths:

      The authors present a valuable new protocol for eCFPS, which simplifies its application.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have made a convincing argument that the current system of in vitro translation using E. coli extracts can be significantly optimized to work with much lesser components, while maintaining activity. They have showcased their improved activity using not only physical but also functional readouts.

      Strengths:

      The experiments are designed in a very logical and easy to understand manner, which makes it easier not only to follow the paper, but also reproduce the results. Functional assays with the synthesized proteins are a good way to demonstrate functionality and applicability of the system. They also benchmark their system against a commercial kit to show superior performance of their system.

      Weaknesses:

      The production of the lysate requires special instrumentation, limiting accessibility.

      Comments on previous version:

      Thank you to the authors for addressing the concerns both textually and experimentally. This work has significant value.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to overcome the challenges associated with complex, conventional prokaryotic cell-free protein synthesis (CFPS) systems, which require up to thirty-five components, by developing a streamlined and efficient E. coli CFPS platform to encourage broader adoption. The main objective was to reduce the number of reaction components from thirty-five to seven, while also developing an accessible 'fast lysate' preparation protocol that eliminates time-consuming runoff and dialysis steps. The authors also sought to demonstrate the robustness and translational quality of this streamlined system by efficiently synthesising challenging functional proteins, including the cytotoxic restriction endonuclease BsaI and the self-assembling intermediate filament protein vimentin.

      Strengths:

      This study presents several key strengths of the optimised E. coli cell-free protein synthesis system in terms of its design, performance and accessibility.<br /> - The reaction mixture has been dramatically simplified, with the number of essential core components successfully reduced from up to thirty-five in conventional systems to just seven.<br /> - The "fast lysate" protocol is a significant advance in terms of procedure.<br /> - The system's ability to synthesise challenging, functional proteins is evidence of its robustness.

      Comments on previous version.

      The authors have adequately addressed my previous concerns.

    5. Author response:

      The following is the authors’ response to the previous reviews

      The revisions in this version are minor and primarily include the addition of RT-qPCR validation experiments. In addition, the benchmarking data against commercial systems have now been incorporated into the main manuscript.

      We also sincerely appreciate Reviewer 2 and Reviewer 3 for their highly encouraging evaluations and recognition of our system's robustness. To fully address the remaining mechanistic queries from Reviewer 1 and the benchmarking concerns from the editors, we have performed quantitative RT-qPCR to directly measure transcript levels and have integrated our commercial benchmarking data into the revised manuscript.

      (1) The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

      To directly validate our claims regarding transcriptional efficiency, we performed quantitative RT-qPCR to determine the transcription levels of the reporter gene in both systems.

      First, we established no-reverse-transcriptase (no-RT) controls to verify complete DNA template removal. The Ct values for these controls remained above 34, confirming the absence of plasmid DNA contamination in our RNA samples.

      Second, transcript levels were calculated using the comparative 2<sup>-ΔΔCt</sup> method, normalized to the standard initial system (100 ng/μL T7) at 30 min. The optimized system achieved a 16.56-fold increase (P < 0.001) in transcript levels. In contrast, supplementing the initial system with high concentrations of T7 RNA polymerase (400 ng/μL) only yielded a 2.87-fold increase (P < 0.01)—which is nearly 6-fold lower than our optimized system.

      These findings perfectly mirror our protein-level titration assays (Figure S3C). Supplementing the initial system with excess T7 RNA polymerase fails to rescue either transcript accumulation or protein expression. This mutual validation confirms that transcription is severely bottlenecked in traditional systems due to rapid nucleotide degradation or inhibitory reaction environments. By streamlining the reaction buffer to seven core components and omitting runoff/dialysis, our system successfully relieves these systemic bottlenecks. We have incorporated these new qPCR findings into Figure 3B, the Methods, and the Results sections of the revised manuscript.

      (2) Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems...

      We appreciate the editor’s emphasis on establishing standard performance benchmarks. To address this important point, we would first like to highlight that our manuscript already contains extensive, rigorous benchmarking against typical cell-free platforms widely utilized in the literature. This includes detailed head-to-head comparisons with both our 35-component "initial" system and the classical, widely established "PEP-based" system across multiple expression kinetics and western blot analyses (as shown in Figure 4 and Figures S3–S4).

      To fully embrace the editor's valuable recommendations regarding standard commercial performance, we are very pleased to formally integrate our commercial benchmarking data into the revised manuscript as Figure S3C.

      To maintain technical neutrality, we have omitted specific brand names, presenting it generically as "a high-end commercial cell-free system." The data demonstrate that our optimized system significantly outperforms this commercial alternative in both expression speed and final absolute yield, reaching an absolute productivity of 0.46 mg/mL compared to approximately 0.21 mg/mL for the commercial kit.

      We are grateful for the guidance from the editors and reviewers, which has significantly strengthened the scientific rigor of our work.

    1. eLife Assessment

      This is an important study that comprehensively determines the consequences of DNMT3A mutations on human neuronal development in culture. The data derived from multiple different mutations in iPSC and hESC derived neurons and organoids are convincing and well controlled. This work will be of interest to researchers who study chromatin mechanisms of brain development and to those interested in DMNT3A mutations in Tatton-Brown-Rahman Syndrome.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include 1. Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      (3) Also, some errors are noted and detailed in the recommendation section.

      Comments on the latest version:

      I have reviewed the revised manuscript and the authors' responses to the reviewers' comments. They addressed the comments adequately.

    3. Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A which cause of Tatton-Brown-Rahman Syndrome (TBRS) disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with rigorous experimental design and novel findings with potential impact across better understanding OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors clarify that neuronal and synaptic gene de-repression is the more prominent direct consequence of mCG loss, while PIK3/AKT/mTOR pathway upregulation is not itself directly linked to differentially methylated regions, suggesting an indirect relationship between DNMT3A LOF and increased proliferative signaling. This framing is reasonable, but the mechanism linking DNMT3A mutation to mTOR activation remains unresolved, and the manuscript would benefit from being explicit about this gap. Relatedly, the rapamycin rescue, while demonstrated across multiple models including 904 and Del1 (Supplementary Fig. S3e-f), remains limited to proliferation readouts. Whether mTOR inhibition also rescues the downstream neurogenesis, maturation, or network phenotypes is an important open question that the authors appropriately frame as motivation for future work.

      (2) The claim that H3K27me3 compensates for mCG loss is supported by prior work (Lii et al. 2022), which demonstrated increased PRC2 component expression and H3K27me3 gain at sites of DNA methylation loss in Dnmt3a knockout mouse neurons, and by data showing that PRC2 subunits (SUZ12, EED, EZH2) are significantly more highly expressed in D-NPCs than V-NPCs. Together, these findings provide a plausible molecular basis for why dorsal progenitors may be better equipped to maintain repression when DNA methylation is lost, and they make the EZH2 overexpression rescue in V-NPCs more interpretable. Yet, a formal distinction related to two competing, potentially underlying mechanisms, between active compensation, in which EZH2 is recruited to specific loci in response to methylation loss, and functional redundancy, in which higher baseline Polycomb occupancy in dorsal cells simply becomes the dominant repressive mark once mCG is reduced, has not been resolved.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Comments on revised version.

      The authors have responded to the reviewer's comments satisfactorily, and the manuscript has been much improved.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include: Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states.

      We thank the reviewer for their constructive feedback. We previously validated our 882 models with whole genome sequencing and teratoma formation upon mouse fat pad injection, while the parental human embryonic stem cell line (WA01 hESCs) used for P904L variant knock-in was validated by our Genome Engineering Stem Cell (GESC) core upon derivation of that variant knock-in model. We have now added both karyotyping and pluripotency staining (SOX2/OCT4) for all other hPSC lines as (new) Supplementary Figure S17 and included further description in our Methods section under “hPSC Model Generation and Culture” (pg. 19, lines 6-7).

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      We thank the reviewer for their insightful suggestions and have detailed our responses in the “recommendations to the authors” section.

      (3) Also, some errors are noted and detailed in the recommendation section.

      We thank the reviewer for catching these errors and have since corrected them, with detailed responses below.

      Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A, which cause Tatton-Brown-Rahman Syndrome (TBRS), disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation, where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability, potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with a rigorous experimental design and novel findings with a potential impact on a better understanding of the OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types, and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed, and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors partially address this tension by noting that DNMT3A LOF alone is insufficient to initiate neuronal differentiation, i.e., V-NPCs upregulate neuronal and synaptic genes while retaining progenitor identity, implying that transcriptomic priming and commitment to differentiation are decoupled. However, the relationship between the proliferative phenotype and the epigenetic priming phenotype remains mechanistically unresolved. The manuscript documents mTOR pathway upregulation at the protein level and identifies shared DEGs that include proliferative regulators, but it does not establish whether mTOR-driven proliferation and mCG-loss-driven neuronal gene de-repression/enhanced differentiation are causally linked or represent two independent consequences of DNMT3A LOF.

      We thank the reviewer for their comment and agree that this phenotype, whereby progenitors exhibited both increased proliferation and hallmarks of gene expression associated with neuronal differentiation is striking and interesting, given that these are typically antagonistic paradigms during normal development.

      We documented that these phenotypes involve upregulated expression of both neuronal/synaptic and proliferative genes in V-NPCs (Figure 2d), with concomitant loss of repressive DNA methylation at regulatory elements associated with these genes (Figure 2f, Supplementary Data 5). In this work, DNMT3A mutation had a more prominent role in de-repressing neuronal and synaptic gene expression to promote hallmarks of neuron differentiation, while playing a relatively less central role in direct regulation of proliferation genes, as seen from the relative prominence of neuronal/synaptic- versus proliferation-related GO terms in our Supplementary Data 5 table (pg. 6, lines 16-19).

      To examine the mechanisms underlying increased V-NPC proliferation in our TBRS models, we assessed a potential relationship with the PIK3/AKT/mTOR pathway, as this is implicated in increased proliferation resulting from DNMT3A-associated mutation in myeloid leukemia (Dai et al., 2017, PMID: 28461508). In our work, DNMT3A mutation increased the expression and/or phosphorylation of mTOR signaling pathway targets specifically in V-NPCs (Figure 1q-r, Supplementary Figure S3a-d). However, while TBRS mutation directly affected repressive DNA methylation at a suite of cell proliferation-related genes, these did not include the PIK3/AKT/mTOR pathway genes themselves, suggesting an indirect relationship between altered DNA methylation and increased mTOR signaling.

      We have since incorporated discussion of how DNMT3A-mediated gene repression and levels of PIK3/AKT/mTOR pathway signaling may be interacting, providing a framework for future studies to identify how these related OGID gene mutations may converge mechanistically (pg. 5, lines 19-21; pg. 16, lines 8-10).

      (2) Relatedly, the rapamycin rescue experiment is a valuable proof-of-concept for the PIK3/AKT/mTOR convergence but is limited to a single dose in a single model (882) with a single readout (Ki67+ proliferation). Given the prominence of mTOR pathway convergence in the manuscript as a potential shared therapeutic avenue across OGIDs, the data supporting this claim are somewhat preliminary. It remains unknown whether mTOR inhibition rescues downstream phenotypes (neurogenesis, gene expression, neuronal maturation) or whether less severe TBRS models respond similarly. This might also help tackle the first comment above. e.g., if mTOR inhibition rescued proliferation but not the transcriptomic priming, that would support two independent mechanisms.

      We thank the reviewer for their comment. We explored both the overall levels and phosphorylation of proteins involved in PIK3/AKT/mTOR signaling in the 882, 904, Del1, Del2, and KO V-NPC models (Figure 1q-r, Supplementary Figure S3a-d), finding specific increases of all proteins. We showed that rapamycin addition reversed the increased proportion of KI67+ proliferating cell nuclei resulting from 882 mutation in V-NPCs in main Figure 1s, while demonstrating that rapamycin also reduced the proportion of KI67+ nuclei observed in both less severe 904 and Del1 V-NPC models (Supplementary Figure S3e-f).

      We agree that understanding whether rapamycin treatment can rescue TBRS neuronal phenotypes would be very interesting, as previous work on Tuberous Sclerosis Complex has utilized rapamycin and other mTOR inhibitors to effectively reverse TSC-related alterations of neuronal morphology and neuronal hyperexcitability (Buttermore et al., 2025, PMID: 40792287). Future studies examining convergent mechanisms and therapeutics for OGIDs should examine how similarly targeting this and related pathways rescues altered neuronal morphology, maturation, and function, as we have demonstrated that TBRS mutation has subsequent consequences for V-IN differentiation, maturation, and function. This point has been detailed in the discussion section on pages 15-16.

      (3) The claim that H3K27me3 compensates for mCG loss is an important mechanistic point, but the current data do not distinguish between active compensation, in which EZH2 is recruited in response to methylation loss, and functional redundancy, in which H3K27me3 is independently established and becomes the dominant repressive mark once DNA methylation is reduced. The EZH2 knockdown/inhibition experiments show that H3K27me3 is sufficient to maintain repression at hypo-DMR sites, but they do not establish that H3K27me3 gain is itself a response to methylation loss. Because H3K27me3 profiling was performed only in the severe 882 model, it is also unclear whether H3K27me3 gain scales with DNMT3A LOF severity, as a compensatory model would predict. Finally, the EZH2 overexpression rescue is performed in V-NPCs, whereas the compensation model is developed primarily in D-NPCs, making it difficult to assess whether the same mechanism operates in the lineage where it was originally inferred.

      We thank the reviewer for the opportunity to clarify our findings and experimental reasoning. A previous study using a conditional Dnmt3a knockout mouse model (Li et al., 2022, PMID: 35604009) demonstrated increased expression of multiple PRC2 components following the loss of Dnmt3a. This study demonstrated that sites which lost DNA methylation gained H3K27me3 in postnatal neurons upon Dnmt3a loss. Therefore, we hypothesize that the gain of H3K27me3 likely occurs in response to loss of DNMT3A methylation.

      While we did not perform CUT&Tag for H3K27me3 in our less severe models, we did validate gene expression changes following EZH2 knockdown and inhibition in both the R882H (Figure 4g-h) and P904L (Supplementary Figure S8b) models, finding that gene expression was unchanged in the model with the less severe DNMT3A mutation (P904L). Based upon these findings, we hypothesized that compensatory H3K27me3 may occur only upon severe DNMT3A loss, as seen in the dominant-negative R882H model. Furthermore, as H3K27me3 compensation was more prominent in D-NPCs, we hypothesized that this might be sufficient to prevent de-repression and aberrant neuronal gene repression upon loss of DNMT3A-mediated repression in D-NPCs. However, since TBRS mutation caused the most prominent de-repression of neuronal gene expression in V-NPCs, we also tested whether EZH2 overexpression could reverse this, finding that it partially suppressed this dysregulated neuronal gene expression. To better clarify this logic and the findings, we have made text edits to this results section and referenced Li et al., 2022 in both the results (pg. 9, lines 16-21; pg. 10, lines 4-7,10-12) and discussion (pg. 16, lines 12-15).

      (4) The narrative framing of dorsal neuron development as unaffected by DNMT3A LOF is somewhat at odds with the data presented. The 882 D-NPCs show substantial DNA methylation changes, and TBRS D-INs exhibit what the authors describe as "substantive transcriptomic differences" involving persistent expression of pluripotency and progenitor genes, which seems to be a distinct but potentially significant phenotype. The impact of DNMT3A loss between ventral and dorsal lineages might be more accurately framed as divergent in nature rather than specific to a certain population.

      We thank the reviewer for their comment. While TBRS mutations appear to have a significantly stronger effect on V-NPCs and subsequently V-INs, both transcriptomic and methylation alterations do also occur upon TBRS mutation in D-NPCs and D-INs, as noted in Supplemental Figure S4d, S11, and Supplemental Data 2. However, we observed substantially greater molecular alterations in V-NPCs/V-INs, a lack of overt cellular phenotypes in D-NPCs where assayed, and a lack of functional consequences in matured D-INs, suggesting a more significant requirement for DNMT3A in regulating the differentiation and subsequent maturation of cortical inhibitory interneurons during embryonic and early pre-natal development, the developmental periods that we can readily model in hPSC-derived neurons.

      It should also be noted that these hPSC differentiation models do not recapitulate post-natal deposition of non-CpG (mCA) DNA methylation, a mechanism disrupted postnatally by TBRS-associated mutations in our prior work in murine models (Harrison Gabel; e.g. Beard et al., 2023, PMID: 37952155), which we have now added in the results section (pg. 7, lines 8-11). Therefore, we hypothesize that if we could sufficiently mature D-INs to a state that modeled postnatal development and recapitulated this non-CpG methylation, we might be able to detect cellular and functional phenotypes in later stage D-INs. To avoid misinterpretation, we have altered the language in the results section to confirm that there are both transcriptomic and methylation changes in our D-NPCs/D-INs, but that these are not accompanied by cellular phenotypes or neuronal dysfunction (pg. 6, lines 1-2; pg. 7 line 3; pg. 7, lines 20-23; pg. 8, lines 1-2; pg. 8, lines 18-23; pg. 9, lines 1-4).

      (5) SST stainings are not entirely convincing. They appear mostly nuclear, and some instances localized to rosettes in organoids, whereas the protein is largely confined to processes and is expected to be found outside progenitor-rich zones like rosettes.

      We agree that the perinuclear SST staining detected in these young ventral telencephalic-patterned organoids at day 30 differs somewhat from the more process-localized and cytosolic signal seen in later stage organoids in other studies. This may be related to the use of different commercial SST antibodies across studies but also likely reflects SST immunoreactivity in newborn neurons near the onset of SST expression. For example, immature SST-immunoreactive neurons in the early postnatal rat cortex exhibit predominant SST staining in perinuclear cytoplasm and short processes (e.g. Fig. 3 in Lee et al, PMID: 9664223) while acquiring more cytosolic and process-localized staining as postnatal neuron maturation occurs. Evaluation of immunopositivity for other markers of neurogenesis (ASCL1) and immature neurons (TUJ1) is also congruent with these findings for SST, with TBRS-associated mutations increasing in the fraction of cells in V-NPCs/V-ORGs that express these three markers.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Weaknesses:

      (1) Li et al., 2022 (referred to in the manuscript) seems to already show the interplay between H3K27me3 and Dnmt3a discussed in this study i.e., that in the absence of DNA methylation, there is an expansion of polycomb-like repression. These data should be better acknowledged in the paragraph 'Repressive H3K27me3 compensates for severe loss of DNA methylation' (page 9), given it supports the data presented in this manuscript and suggests this as a common mechanism in the interplay between these two repressive marks, as it is well established in the literature.

      We thank the reviewer for this suggestion. We have now added Li et al., 2022 to both the results section (pg. 9, lines 16-20) and our discussion section (pg. 16, lines 12-13).

      (2) The authors should acknowledge that the omics data come from a mixed population of cells.

      We thank the reviewer for their comment. We have validated that the established 2-D differentiation methods we used in this study generate cell populations with >85-90% enrichment for the desired progenitor and neuronal cell type, based upon marker expression, but acknowledge that these are bulk -omics data obtained from cells that may represent a mixed population and have now detailed this in the methods section under “Sequencing” (pg. 21, lines 16-18).

      (3) The authors are encouraged to further discuss whether the overgrowth observed in ventral GABAergic cultures or organoids compares to the overgrowth observed in diseased patients. One expects MRIs to have been performed in patients and that these could be harnessed to discern if overgrowth occurs in the cortex or ventral regions of the brain.

      We thank the reviewer for their suggestion and do note that at least one published study documents increased cortical thickness in the MRIs of TBRS patients (Jiménez de la Peña et al., 2024, PMID: 37795572); however, to our knowledge studies have not examined regional or cell type-selective overgrowth of cortical tissue in TBRS patients. Future clinical studies examining the nature of the neuronal progenitor overgrowth and resulting consequences for patient brain imaging would be of interest to better understand TBRS-associated etiology of brain overgrowth and its manifestations.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A potential explanation for the ventral organoid's sensitivity compared to the dorsal can be provided. I understand that the PRC2-mediated methylation loss was more pronounced in the ventral than in the dorsal. However, this raises the question: why is this so? Are the PRC2 components more highly expressed in ventral organoids or in GABAergic neurons than in glutamatergic neurons?

      We thank the reviewer for this question. Upon examining RNA-sequencing data for control V-NPCs and D-NPCs, we found that expression of the PRC2 subunits SUZ12, EED, and EZH2 is significantly elevated in D-NPCs relative to V-NPCs (see Author response image 1). This could contribute to restoring epigenetic repression in D-NPCs in the context of DNMT3A mutation, to yield the more limited transcriptomic changes and greater gain of H3K27me3 we observed in TBRS D-NPCs, despite their comparable levels of global mCG loss). This could also account for our finding that overexpression of EZH2 in V-NPCs could normalize some gene expression changes observed in the TBRS models.

      Author response image 1.

      (2) DNMT3A is a major enzyme that mediates non-CpG methylation during synaptogenesis. There is no discussion of this phenomenon or any related analysis. I suppose the brain organoids may be too immature to detect non-CpG methylation. Even then, this can be described in the results sections.

      We thank the reviewer for this comment. While prior work has examined an important role for DNMT3A in postnatal deposition of non-CpG methylation (Christian et al. 2020; Beard et al. 2023), we found that our hPSC models, as the reviewer suggests, were too immature to detect appreciable levels of non-CpG methylation. Accordingly, we have focused our work here on neurodevelopmental requirements for DNMT3A; this work demonstrated that many of the functional alterations of neuronal network activity originate from altered neurogenesis and neuronal maturation of cortical GABAergic interneurons. We have now included a description of the important role of non CpG methylation by DNMT3A during postnatal development and have indicated that the immaturity of hPSC-derived neurons precludes our ability to study non-CpG methylation using these models in the results (pg. 7, lines 8-11) with present language in the discussion (pg. 15, lines 8-11).

      (3) Given the Figure 4 results showing H3K27 methylation and EZH2 knockdown rescue the DNMT3A mutation, the authors argue that EZH2 and DNMT3A regulate a similar set of genes and that the two diseases are related. To claim this, there should be data showing that gene misregulation upon loss of EZH2 and DNMT3A is similar.

      Given our data, it is premature to draw direct relationships with the dysregulated genes we detected in TBRS versus the molecular basis of Weaver Syndrome; therefore, we have attempted to better reflect the potential but currently unproven relationship between these OGIDs and altered the respective language in the results section (pg. 10, lines 10-12) to reflect this. Understanding convergent mechanisms of OGIDs involving both DNMT3A (TBRS) and EZH2 (Weaver Syndrome) gene mutations will be a compelling topic for future studies, as our work here suggests that they may regulate similar gene suites.

      (4) Stronger phenotypes were observed in the R882H mutant compared to the P904L mutant. However, this cannot be translated to the functionality of the mutant protein or disease severity because one is on iPSCs and another is made in hESCs. The limitation should be noted to avoid misleading.

      We agree that the severity of each pathogenic mutation can be influenced by its presence on a different iPSC vs hESC background; accordingly, our study design indicated the hPSC background for each model and employed paired isogenic controls to clearly define the consequences of each TBRS mutation relative to a model with an identical genetic background but lacking the mutation. Our finding that the R882H mutation resulted in more severe epigenomic and transcriptomic consequences than the P904L mutation is congruent with findings made in prior work (Beard et al., 2023 PMID: 37952155; Russler-Germain et al., 2014 PMID: 24656771).

      (5) It is unclear what measurements were done for DNMT3A level quantification shown in Figure 1e- f.

      Protein quantification for DNMT3A models was performed by western blot, shown in Supplemental Fig. S16. We have now included reference to whole blots in the methods section under “Cellular Phenotyping” (page 20 line 13).

      (6) The results section of the paper for Figures 6 and 7 cites the wrong figure numbers and panels.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      (7) Some typos for PI3K (meaning PIK3) in several places, including the Abstract.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      Reviewer #2 (Recommendations for the authors):

      (1) There is a figure numbering discrepancy in the manuscript - the text references six main figures, but the figure pages include seven, with the MEA network data apparently mislabeled.

      We thank the reviewer for catching these errors and have corrected them in the new manuscript draft.

      (2) The nomenclature "D-IN" for dorsal immature neurons is potentially misleading, as "IN" conventionally denotes interneurons, which these glutamatergic cells are not.

      While we agree that IN could be interpreted as interneurons, this nomenclature is defined at an early point in the manuscript and was used to allow readers to easily identify compare findings made in D-NPCs and D-INs (and, as a counterpart, findings made in V-NPCs versus V-INs).

    1. eLife Assessment

      This important work presents a novel computational framework for modeling macroscopic traveling waves in the mouse cortex by integrating open-source connectomic and transcriptomic data into a spiking network model. This approach allows the computational model to assign excitatory/inhibitory connections based on neurotransmitter profiles and extends simulations to the 3D domain. The authors present results that demonstrate how spatiotemporal dynamics such as slow oscillations (0.5-4 Hz) emerge and self-organize at the whole-brain scale. This study provides convincing initial insights into the structural basis of traveling waves at the whole-brain scale, and allows future links between connectome-driven spiking neural networks and whole-hemisphere imaging in the mouse.

    2. Reviewer #1 (Public review):

      I thank the authors for their thoughtful and thorough responses, which address my concerns. Their two methodological changes: (1) the switch to Poisson stimulation and (2) the new LFP estimation pipeline, together with the expanded parameter-grid sweep and Kuramoto synchrony analysis, substantially strengthen the manuscript. The Poisson spike train better approximates the stochastic subcortical drive cortex receives in vivo and removes the artificiality of the original protocol (Point 1.2). The LFP pipeline directly resolves my concern about the disconnect between simulated voltages and experimental signals; showing that the macroscopic wave structure persists in the LFP-like proxy clarifies the framework's practical relevance (Point 1.6). The expanded per-band sweep addresses my worry that the Allen-connectivity advantage was confined to a narrow regime, and acknowledging the small delta-band difference is a more convincing presentation (Point 1.5). The Kuramoto analysis connects dynamics across scales and gives a clear, quantitative account of the non-monotonic coupling dependence (Points 1.4, 1.7). Finally, I appreciate that the remaining connectivity-realism issues (Points 1.3, 1.8) are now stated explicitly as limitations with concrete future directions. I agree that incorporating them is beyond the scope of the present study, and their upfront acknowledgement is appropriate.

    3. Reviewer #2 (Public review):

      Summary:

      This work presents a spiking network model of traveling waves at the whole-brain scale in mouse neocortex. The authors use data from the Allen Institute to re-construct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in neocortex.

      Strengths:

      Overall, the results are interesting and shed new light on the dynamic organization of activity across neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      Weaknesses:

      (1) Description of Algorithm 1: While the Methods section clearly explains the density parameter \rho, the statement on line 358 concerning the "ideal" average number of connections is a little unclear. The authors should explicitly clarify that \rho is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity.

      (2) Lines 102-103: The \rho parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      (3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      (4) Line 217: Can the authors clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013), or differ from them?

      (5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      (6) Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detections. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      (7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      (8) Line 287-288: The authors suggest at this point that they do not have enough information to estimate time delays due to axonal conduction along white matter fibers. However, experimental data from white matter connections typically includes information about fiber length, which does enable estimating conduction delays. These estimations have been previously implemented for Allen Institute connectome data in the mouse (Choi and Mihalas, PLoS Comput Biology, 2019) and human connectome data (Budzinski et al., Physical Review Research, 2023).

      (9) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Comments on revised version.

      In this response and revised manuscript, the authors have addressed all points raised in the first round of review. In response to Point 2.7, however, is it not the case that the Allen dataset has the axonal lengths?

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1.1) The manuscript “Realistic coupling enables flexible macroscopic traveling waves in the mouse cortex” by Sun, Forger, and colleagues presents a novel computational framework for studying macroscopic traveling waves in the mouse cortex by integrating realistic brain connectivity data with large-scale neural simulations.

      The key contributions include: (1) developing an algorithm that combines spatial transcriptomic data (providing detailed neuron positions and molecular properties) with voxelized connectivity data from the Allen Brain Atlas to construct neuron-to-neuron connections across 300,000 cortical neurons; (2) building a GPU-accelerated simulation platform capable of modeling this large-scale network with both excitatory and inhibitory HodgkinHuxley neurons; (3) extending phase-based analysis methods from 2D to 3D to quantify traveling wave activity in the realistic brain geometry; and (4) demonstrating that realistic Allen connectivity generates significantly higher levels of macroscopic traveling waves compared to simplified local or uniform connectivity patterns.

      The study reveals that wave activity depends non-monotonically on coupling strength and that slow oscillations (0.5-4 Hz) are particularly conducive to large-scale wave propagation, providing new insights into how anatomical connectivity enables flexible spatiotemporal dynamics across the cortex.

      The authors leverage two existing dense datasets of spatial transcriptomic data and connection strength between pairwise voxels in the mouse cortex in a novel way, allowing for the computational model to capture molecular and functional properties of neurons as determined by their neurotransmitter profiles, rather than making arbitrary assignments of excitatory/inhibitory roles. Additionally, the author’s expansion of 2D phase dynamics to 3D phase gradient analysis methods is important and can be widely applied to calcium imaging, LFP recordings, and likely other electrophysiological recordings.

      Thank you for the accurate summary of our manuscript and list of strengths.

      (1.2) The model’s Allen connectivity approach overlooks critical aspects of real cortical dynamics. Most importantly, it excludes subcortical structures, especially the thalamus, which drives cortical traveling waves through thalamocortical interactions. The authors’ method of electrically stimulating all layer 4 neurons simultaneously to initiate waves is artificially crude and bears little resemblance to natural wave generation mechanisms.

      We agree that excluding subcortical structures, especially the thalamus, is an important limitation of the current model. Because adding these structures would substantially expand the scope and complexity of the model, we now state this limitation more explicitly in the Discussion and leave it as a future extension:

      “For simulations, we choose to randomly stimulate the total population of layer 4 neurons as a way to mimic subcortical input and generate traveling waves, which can be unrealistic. Subcortical structures, such as the thalamus, are vital to cortical dynamics like slow-wave activity [1] and are known to regulate traveling waves [2]. Therefore, a direct and important future improvement would be adding subcortical structures to the model.”

      We also agree that the original constant-current stimulation was too artificial. We therefore replaced it with a 10Hz Poisson spike train delivered to layer-4 excitatory neurons across the isocortex, which more closely mimics the stochastic input that cortex receives from subcortical regions such as the thalamus. The revised stimulation protocol is described in the Results:

      “Therefore, to mimic the stochastic drive that cortex receives from subcortical regions like thalamus, we deliver a 10Hz Poisson spike train to layer-4 excitatory neurons (Figure 2a), since layer 4 is the canonical thalamocortical input layer [3]; each Poisson event applies a fixed voltage bump V<sub>stim</sub>, and no other external input is applied.”

      as well as in the Methods 4.5 Simulation.

      The revised stimulation protocol allowed us to rerun the parameter sweep under stochastic drive. A direct comparison of alternative wave-generating mechanisms remains an important direction for future work.

      (1.3) The model handles voxel-to-voxel connections crudely when neurons have mixed excitatory/inhibitory properties and varying synaptic strengths. Real connectivity differs dramatically between neuron types (pyramidal cells vs. interneurons, across cortical layers), but the model only distinguishes excitatory and inhibitory neurons. Additionally, uniform synaptic weights ignore natural variations in connection strength based on neuron type, distance, and functional role. Integrating the updated thalamocortical dataset mentioned by the authors, even at regional resolution, would substantially improve the model.

      We thank the reviewer for raising these important points regarding cell-type-specific connectivity and heterogeneity of synaptic weights. We agree that cell-type-specific connectivity and heterogeneous synaptic weights are important limitations. Because the current voxelized projectome is not cell-type specific, we now state these limitations explicitly and outline how future versions of the model could incorporate improved density and synaptic weight assumptions in the Discussion (Construction of the Allen connectivity). Specifically, we now write:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.

      Also, our model uses uniform synaptic weights within each synapse type: every AMPA synapse has conductance g<sub>AMPA</sub> and every GABA synapse has conductance g<sub>GABA</sub>. In reality, synaptic strength varies with presynaptic and postsynaptic cell type, anatomical distance, and the distribution of synaptic weights is typically heavy-tailed (e.g. log-normal [5]). Extending our algorithm to sample synaptic weights from realistic distributions would be a natural next step and is likely necessary for quantitative comparisons with electrophysiological recordings.”

      These additions clarify which aspects of the current model are constrained by the available voxelized projectome and which extensions would require cell-type-resolved connectivity data, region-specific density information, or more detailed synaptic-weight estimates.

      (1.4) While the authors bridge microscopic (single neuron) and mesoscopic (regional connectivity) data to study macroscopic (whole-cortex) waves, they don’t integrate the distinct mechanisms operating at each scale. The framework demonstrates that realistic connectivity enables macroscopic waves but fails to connect how wave dynamics emerge and interact across spatial scales systematically.

      We thank the reviewer for this insightful comment. In the revision, we added the Kuramoto synchrony analysis as a first step toward connecting scales: changes in the microscopic coupling parameter alter network synchrony, and this synchrony measure is closely associated with the macroscopic PGD observable (new Fig. 5). We agree that a full separation of layer-specific microcircuit mechanisms, region-specific connectivity motifs, and whole-cortex wave propagation remains beyond the scope of the present study. The revised text now frames the synchrony-to-wave relationship as one concrete cross-scale link that can be investigated further with this framework.

      (1.5) Claims that Allen connectivity produces higher phase gradient directionality (PGD) than local connectivity appear limited to delta oscillations at very specific coupling strengths and applied currents. Few parameter combinations show significantly higher PGD for Allen connectivity, and these are generally low PGD values overall.

      We agree that the original comparison did not sufficiently establish whether the Allen-connectivity advantage held beyond a small number of parameter choices.

      In the revised manuscript, we expanded the analysis to a 10 × 12 grid of excitatory coupling strengths and Poisson stimulus magnitudes for Allen, local, and uniform connectivity. Across this full (g<sub>AMPA</sub>,V<sub>stim</sub>) grid, Allen connectivity shows higher per-band maximum PGD than local or uniform connectivity, especially in the theta, alpha, and beta bands (Fig. 5b). The delta-band difference is small, consistent with the reviewer’s observation that the original delta-band result did not clearly separate Allen from local connectivity.

      In the revised manuscript, Fig. 5c and Fig. 5d show how PGD varies with stimulus magnitude and coupling strength individually. The full per-band PGD heatmaps for all three connectivities, together with the per-band Allen-minus-opponent gap bars, are provided in Figure 4-figure supplement 6. These additions show that the Allen-connectivity trend is not limited to a single representative operating point.

      (1.6) Broadly, it’s unclear how this computational framework can study memory, learning, sleep, sensory processing, or disease states, given the disconnect between simulated intracellular voltages and the local field potentials or other electrophysiological measurements typically used to study cortical traveling waves. While computationally impressive, the practical research applications remain vague.

      We thank the reviewer for this important point. To bridge the gap between our simulations and experimentally measured data, such as local field potentials (LFP), we used a post-hoc LFP estimation pipeline and added a dedicated Methods subsection describing it (LFP estimation from intracellular voltage). The key idea is that LFP primarily reflects the net transmembrane synaptic current in a local population, which we can reconstruct directly from the voltage traces and the connectivity used in the simulation. Please see Methods section LFP estimation from intracellular voltage.

      This pipeline allows us to test whether the macroscopic traveling-wave structure identified in the voltage traces is also present in an LFP-like signal. In the revised manuscript, we added a new Results section and show, in Fig. 6, side-by-side voltage- and LFP-based snapshots, the voltage-LFP PGD scatter (Pearson r = 0.78-0.89 per band), and per-band PGD comparisons across Allen, local, and uniform connectivity on the LFP signal. These results indicate that our conclusions are not restricted to intracellular voltage and provide a closer bridge to LFP-based experimental measurements.

      (1.7) The paper needs a clearer explanation for why medium coupling (100%) eliminates waves in Allen connectivity (Figure 6) while stronger coupling (150%) restores them.

      Thank you for requesting this clarification. The revised analysis suggests that the non-monotonic relationship between coupling strength and wave activity reflects an interaction between network synchrony and spatial organization. At weak coupling, the network has enough coordination to support propagating waves. At medium coupling, increased synaptic drive pushes the network into an asynchronous irregular state that disrupts coherent wave fronts. At strong coupling, rhythmic synchronization is re-established and again supports wave propagation.

      To substantiate this explanation quantitatively, we measured the Kuramoto order parameter R(t)=|〈e<sup>iϕ(x,t)</sup>〉<sub>x</sub>| from the generalized phase ϕ(x, t) of the band passed voltage field (see updated Quantitative measurement of neuronal activity in Methods) and reduced it to the maximum over each 1-s recording window. We then swept the same ten g<sub>AMPA</sub> values (0.005-0.050 nS) and twelve stimulus magnitudes used for the PGD sweeps, for Allen, local and uniform connectivity, in all five frequency bands. The new analysis is presented in Fig. 5: panel (e) shows synchrony versus g<sub>AMPA</sub> for the three connectivities, panel (d) shows the matching PGD curve, and panel (f) shows the synchrony-PGD scatter with Pearson r per connectivity.

      The synchrony curve captures the main peak-dip-recovery structure of the PGD curve. For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses to a 75 % suppression in the medium-coupling window g<sub>AMPA</sub> = 0.020-0.035 nS, and recovers near unity at g<sub>AMPA</sub> ≥ 0.040 nS. Uniform connectivity follows the same U-shape with a slightly earlier dip. Per-band versions of the PGD- and synchrony-vs-coupling curves are provided in Figure 5-figure supplement 1, showing that the peak-dip-recovery profile holds across all five canonical bands for Allen and uniform connectivity. Across the 120- point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson 0r ranging from ≈ 0.30 to ≈ 0.92 across band-connectivity combinations; per-band scatters in Figure 5-figure supplement 2). Local connectivity follows a different trajectory: its synchrony is moderate at weak coupling and decays monotonically with gAMPA without recovering at strong coupling. This is consistent with local connectivity not supporting large-scale propagating waves at strong coupling, so the peak-dip-recovery interpretation applies mainly to Allen and uniform connectivity. We have added this synchrony analysis to the revised manuscript:

      “To diagnose the mechanism behind this profile, we measured the Kuramoto order parameter R(t) = |⟨e<sup>iϕ(x,t)</sup>⟩<sub>x</sub>| from the generalized phase field ϕ(x, t) of each frequency band and recorded its maximum over each simulation window (Quantitative measurement of neuronal activity). The resulting synchrony curve (Figure 5e) resembles the trend of the maximum PGD well (Figure 5d). For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses in the medium-coupling window, and recovers at strong coupling. Across the full 120-point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson r = 0.30- 0.92, all p < 10−3; Figure 5f, with per-band scatters in figure Supplement 2). Therefore, the PGD trough at medium coupling may be a synchrony trough: increased synaptic drive pushes the network into an asynchronous irregular state, while strong coupling re-establishes rhythmic synchronization that supports wave propagation. This synchrony-to-wave bottleneck seems more significant to the networks with long-range connectivity (Allen, uniform), partly because long-range connections can augment synchrony in the coupled neuronal network [6]. We also note that although uniform connectivity is able to achieve almost an identical level of synchrony to that of Allen connectivity, the PGD remains much lower due to the loss of spatial organization within.”

      (1.8) Does using a single connectivity parameter (ρ = 300) across all regions miss important regional differences in cortical connectivity density?

      We thank the reviewer for raising this important point. We agree that a single global density parameter misses region-to-region variation in local synaptic density, and we have extended the Discussion (Construction of the Allen connectivity) to state this limitation alongside the cell-type-specific connectivity and synaptic-weight limitations discussed in response to Point 1.3. Specifically, we now write:

      “At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      This paragraph is placed directly after the cell-type and weight-heterogeneity limitations added in response to Point 1.3.

      Reviewer #2 (Public review):

      (2.1) This work presents a spiking network model of traveling waves at the whole-brain scale in the mouse neocortex. The authors use data from the Allen Institute to reconstruct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in the neocortex.

      Overall, the results are interesting and shed new light on the dynamic organization of activity across the neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      We thank Reviewer 2 for the positive assessment and accurate summary of our work.

      (2.2) Description of Algorithm 1: While the Methods section clearly explains the density parameter ρ, the statement on line 358 concerning the “ideal” average number of connections is a little unclear. The authors should explicitly clarify that ρ is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity. The ρ parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in the mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      We thank the reviewer for raising this important point about our connectivity algorithm and simulation. We have clarified in the revised Methods (Use of the voxelized connectivity data, Methods 4.2) that ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity. Specifically, we now write:

      “The density parameter ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity: higher values of ρ yield more connections per neuron and thus higher biological realism, at the cost of greater memory and runtime [7].”

      A careful accounting of the biological scales involved (synapse density per neuron, total cortical population) and incorporating region or cell-type-specific connection density is left as a future direction. We have noted this in the Discussion subsection:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      (2.3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      The reviewer is correct, and we are grateful for the prompt to be more precise. The Results phrasing around Fig. 2 has been softened to avoid any implication of narrowband rhythmicity, and now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Moreover, we have revised the Introduction to describe [11] as demonstrating traveling waves in broadband (5-40Hz) activity, making clear that traveling waves can occur without requiring narrowband oscillations (see also our response to your related point below). In the revised manuscript, we also separate the broadband activity into canonical frequency bands and analyze the wave activity in each band independently.

      (2.4) Line 217: The authors should clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013) or differ from them.

      We thank the reviewer for this suggestion. We agree that [12] is an important experimental benchmark. A direct quantitative comparison is difficult because the original data were not aligned to the Allen Brain Atlas CCF used in our simulations. We therefore revised the Discussion to identify this comparison as a future direction, alongside the visual-cortex bidirectional-wave data of [13]:

      “A more detailed quantitative comparison with experimental cortical-wave studies, such as the cortex-wide voltage-imaging data of [12] or the bidirectional visual-evoked waves reported by [13], is left as a future direction.”

      (2.5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      We thank the reviewer for raising this important point. The reviewer is correct that higher-frequency oscillations tend to be more spatially localized, which can inherently reduce PGD when measured globally. We therefore revised this analysis by separating the broadband signals into canonical frequency bands and comparing PGD within each band.

      In the revised manuscript, we no longer interpret cross-band PGD differences as evidence for a frequency-to-spatial-scale relationship. Instead, we report the level of macroscopic wave activity within each canonical band. This per-band comparison is summarized in Fig. 5b, which reports the maximum PGD in each band (mean ± SEM across the entire (g<sub>AMPA</sub>,V<sub>stim</sub>) grid) for the three connectivities. Allen connectivity shows higher per-band PGD than local or uniform connectivity in the theta, alpha, and beta bands, without requiring an interpretation of PGD differences across frequency bands.

      For completeness, we also computed the per-band mean(Allen) − mean(opponent) PGD gap across the full parameter grid (10 g<sub>AMPA</sub> × 12 V<sub>stim</sub> = 120 points per connectivity). The result is presented in Figure 4-figure supplement 6: panel (a) gives the per-band maxPGD heatmaps over the (gAMPA, Vstim) grid for each connectivity, and panel (b) gives the per-band Allen-minus-opponent mean-PGD gap. The mean gap against Local is −0.004 in delta, +0.027 in theta, +0.036 in alpha, +0.036 in beta and +0.004 in gamma, and against Uniform is +0.015, +0.044, +0.052, +0.040 and +0.011, respectively. The gap is largest in alpha and broadly concentrated in theta-alpha-beta, rather than in delta as we had originally written. We report these per-band gaps descriptively because they summarize one parameter sweep per connectivity rather than independent biological or simulation replications, and we do not interpret the differences across bands as a meaningful frequency dependence. We have added this analysis to the revised manuscript.

      (2.6)Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detection. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      We agree with the reviewer that the relevant anatomical blocks are already present in the data. Our intent was to relate the sampling procedure to stochastic block models for network generation, not to community-detection algorithms. We have revised this passage in the Discussion subsection Construction of the Allen connectivity: ”In essence, our algorithm belongs to the family of stochastic block models [14], where the block structure is given by the voxelization of the Allen Brain Atlas and the inter-block connection probabilities are set by the voxelized projection strengths.”

      (2.7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      Thank you for raising this important point. We agree that conduction delay is directly tied to axonal propagation time and therefore to anatomical path length. Our original wording was intended to note that the relevant path length cannot be approximated reliably by Euclidean distance in the 3-D coordinate space. We have revised this passage in the Discussion sub-section Cortical model and simulation to state that conduction delay is linked to the white-matter path length of the connection, while noting that our model lacks the actual axonal path geometry through the cortical manifold:

      “However, we did not include conduction delay in our study. Conduction delay is thought to have a proportional relationship with the white matter path length of the connection between two neurons [15]. In our model, although we have the 3D positions of neurons, we do not have geometric information about the cortical manifold. Two neurons can be very close in the Cartesian coordinates measured by the Euclidean distance, but very far in terms of the length of the actual connection in the brain. Therefore, incorporating accurate conduction delay in the model is an important future direction.”

      (2.8) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Thank you for this reference. We have added a citation to [16] in the Discussion subsection Quantitative measurements of 3-D traveling waves, acknowledging that methods for 3-D wave analysis do exist while noting that most published algorithms are designed for 2-D data:

      “There are many techniques available for identifying and measuring large-scale neuronal spatiotemporal patterns [17, 18]. While methods for detecting wave dynamics in three-dimensional data do exist [16], most published algorithms are designed to analyze 2-D data, so our simulation data, which is intrinsically 3-D, presents new challenges for measurement.”

      (2.9) Line 28: It is important to note that the Davis et al. (2020) reference is not actually in the beta band, but instead in the broadband (5-40 Hz). This distinction is important because it demonstrates that waves can occur in neural data without requiring narrowband oscillations.

      Thank you for this correction. We have moved the [11] citation out of the beta-band group in the Introduction and reframed it as a broadband (5-40Hz) reference, so that the sentence now reads:

      “These waves are observed at different frequencies during various brain activities, ranging from slow-wave activity [8, 9], sleep spindles [19], to faster oscillations in alpha [20, 21], beta [22, 23], and gamma [21, 13] frequency bands, as well as in broadband (5-40Hz) activity [11].”

      This distinction is important because it makes clear that traveling waves do not require narrowband oscillations. The revised manuscript therefore separates the broadband activity into canonical frequency bands and compares wave activity within each band.

      (2.10) Line 46-49: This sentence could be clearer, for example, by specifying “certain dynamics” in more precise terms.

      We apologize for the imprecision in the original manuscript. We have clarified the sentence in the Introduction to specify that the dynamics of interest are coexistence patterns of local and global activity in the network, as described in the cited reference [24]:

      “It has been shown that in a coupled neuronal network, the coexistence of global wave activity with locally asynchronous states only arises when the number of oscillators is high enough, where local and global activities can coexist [24].”

      (2.11) Figure 1b(i): Small typo in the label for this panel.

      This has been corrected. Thank you for catching it.

      (2.12) Lines 121-122: It may be important to note that spiking neural networks can also generate self-sustained activity (Vogels and Abbott, JNeurosci, 2005; Kumar et al., Neural Computation, 2008). This self-sustained activity is a form of internally generated “frozen” noise that is fundamentally different from externally imposed noise sources (such as Poisson external input) (Destexhe and Contreras, Science, 2006). Waves appear in this self-sustained activity, as well (Davis et al., Nature Communications, 2021), supporting the generality of this phenomenon.

      Thank you for pointing out these important references. We have added citations to [25], [26], [27], and [24] in the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity:

      “Spiking networks of this scale can also generate self-sustained activity in similar regimes [25, 26], which differs fundamentally from externally imposed noise [27], and traveling waves have been reported under such conditions [24].”

      In the revised manuscript, we also changed the stimulation protocol from constant current to Poisson input, which is more similar to the stochastic input that cortical neurons receive in vivo. Traveling waves observed in our model under this Poisson drive therefore complement, rather than depend on, the self-sustained-activity regime emphasized by the cited works.

      (2.13) Lines 139-140: “anterior-posterior macroscopic waves in both directions” and ”in the reverse direction right after each other” could be clearer. In addition, the study from Aggarwal et al. (Nature Communications, 2022) could be relevant to note at this point.

      We thank the reviewer for these wording suggestions and the relevant reference. In the revised manuscript, we reran the simulations and updated Figure 2 accordingly. The new representative simulation shown in Fig. 2 emphasizes a coherent wavefront sweeping along the anterior-to-posterior axis, and the original passages describing consecutive opposite-direction waves have been removed from the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity, which now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Bidirectional propagation can still occur in the model, but we have chosen not to present it as a focal result of this revised manuscript; consequently, the specific phrasings flagged by the reviewer no longer appear in the Results text.

      We agree that [13] is an important reference, and we now discuss it in the Comparing with experimental data subsection as a potential benchmark for future quantitative comparison with our model.

      (2.14) Figure 7: The caption for this figure could be clearer.

      Line 281: typo ”Ermentrou”.

      Line 418: typo ”excitatory”.

      The typos (“Ermentrou” → “Ermentrout”, and the “excitatory” typo at line 418) have been corrected. The original Figure 7 has been removed from the revised manuscript; its content (PGD versus stimulus magnitude and coupling strength, and PGD by dominant frequency bucket) has been reorganized across the new Figs. 4-6.

      References

      (1) Steriade M, Mccormick DA, Sejnowski TJ. Thalamocortical Oscillations in the Sleeping and Aroused Brain. Science. 1993;262(5134):679-85. Available from: <GotoISI>:// WOS:A1993MD95200029.

      (2) Ye Z, Bull MS, Li A, Birman D, Daigle TL, Tasic B, et al. Brain-wide topographic coordination of traveling spiral waves. BioRxiv. 2023:2023-12.

      (3) Rockland KS.What do we know about laminar connectivity? Neuroimage. 2019;197:772-84.

      (4) Harris JA, Mihalas S, Hirokawa KE, Whitesell JD, Choi H, Bernard A, et al. Hierarchical organization of cortical and thalamic connectivity. Nature. 2019;575(7781):195+. Available from: <GotoISI>://WOS:000496159900061https://www.nature.com/articles/ s41586-019-1716-z.pdf.

      (5) Buzs´aki G, Mizuseki K. The log-dynamic brain: how skewed distributions affect network operations. Nature Reviews Neuroscience. 2014;15(4):264-78.

      (6) Bazhenov M, Rulkov NF, Timofeev I. Effect of synaptic connectivity on long-range synchronization of fast cortical oscillations. Journal of neurophysiology. 2008;100(3):156275.

      (7) Morrison A, Aertsen A, Diesmann M. Spike-timing-dependent plasticity in balanced random networks. Neural computation. 2007;19(6):1437-67.

      (8) Massimini M. The Sleep Slow Oscillation as a Traveling Wave. Journal of Neuroscience. 2004;24(31):6862-70. Available from: https://dx.doi.org/10.1523/jneurosci. 1318-04.2004 https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729597/pdf/ 0246862.pdf.

      (9) Liang Y, Song C, Liu M, Gong P, Zhou C, Kno¨pfel T.Cortex-Wide Dynamics ofIntrinsic Electrical Activities: Propagating Waves and Their Interactions. The Journal of Neuroscience. 2021;41(16):3665-78. Available from: https://www.jneurosci.org/ content/jneuro/41/16/3665.full.pdf.

      (10) Aggarwal A, Luo J, Chung H, Contreras D, Kelz MB, Proekt A. Neural assemblies coordinated by cortical waves are associated with waking and hallucinatory brain states. Cell Reports. 2024;43(4):114017. Available from: https://www.sciencedirect.com/ science/article/pii/S2211124724003450.

      (11) Davis ZW, Muller L, Martinez-Trujillo J, Sejnowski T, Reynolds JH. Spontaneous travelling cortical waves gate perception in behaving primates. Nature. 2020;587(7834):4326. Available from: https://doi.org/10.1038/s41586-020-2802-y.

      (12) Mohajerani MH, Chan AW, Mohsenvand M, LeDue J, Liu R, McVea DA, et al. Spontaneous cortical activity alternates between motifs defined by regional axonal projections. Nature Neuroscience. 2013;16(10):1426-35. Available from: https://doi.org/ 10.1038/nn.3499https://www.nature.com/articles/nn.3499.pdf.

      (13) Aggarwal A, Brennan C, Luo J, Chung H, Contreras D, Kelz MB, et al. Visual evoked feedforward–feedback traveling waves organize neural activity across the cortical hierarchy in mice. Nature Communications. 2022;13(1):4754. Available from: https://doi.org/10.1038/s41467-022-32378-xhttps://www.nature. com/articles/s41467-022-32378-x.pdf.

      (14) Holland PW, Laskey KB, Leinhardt S. Stochastic blockmodels: First steps. Social Networks. 1983;5(2):109-37. Available from: https://www.sciencedirect.com/science/ article/pii/0378873383900217.

      (15) Lemar´echal JD, Jedynak M, Trebaul L, Boyer A, Tadel F, Bhattacharjee M, et al. A brain atlas of axonal and synaptic delays based on modelling of cortico-cortical evoked potentials. Brain. 2022;145(5):1653-67.

      (16) Budzinski RC, Nguyen TT, Min´o-Calero J, Davis ZW, Muller LE. Analyzing transientevoked neural activity in three-dimensional cortical recordings. Physical Review Research. 2023;5(1):013012.

      (17) Townsend RG, Gong P. Detection and analysis of spatiotemporal patterns in brain activity. PLOS Computational Biology. 2018;14(12):e1006643. Available from: https: //doi.org/10.1371/journal.pcbi.1006643.

      (18) Gutzen R, De Bonis G, De Luca C, Pastorelli E, Capone C, Allegra Mascaro AL, et al. A modular and adaptable analysis pipeline to compare slow cerebral rhythms across heterogeneous datasets. Cell Reports Methods. 2024;4(1). Available from: https: //doi.org/10.1016/j.crmeth.2023.100681.

      (19) Muller L, Piantoni G, Koller D, Cash SS, Halgren E, Sejnowski TJ. Rotating waves during human sleep spindles organize global patterns of activity that repeat precisely through the night. Elife. 2016;5. Available from: https://www.ncbi.nlm.nih.gov/pubmed/27855061https://www.ncbi.nlm. nih.gov/pmc/articles/PMC5114016/pdf/elife-17267.pdf.

      (20) Zhang H, Watrous AJ, Patel A, Jacobs J. Theta and alpha oscillations are traveling waves in the human neocortex. Neuron. 2018;98(6):1269-81. e4.

      (21) van Kerkoerle T, Self MW, Dagnino B, Gariel-Mathis MA, Poort J, van der Togt C, et al. Alpha and gamma oscillations characterize feedback and feedforward processing in monkey visual cortex. Proc Natl Acad Sci U S A. 2014;111(40):14332-41. Available from: https://www.ncbi.nlm.nih.gov/pubmed/25205811.

      (22) Bhattacharya S, Brincat SL, Lundqvist M, Miller EK. Traveling waves in the prefrontal cortex during working memory. PLoS Comput Biol. 2022;18(1):e1009827.

      (23) Rubino D, Robbins KA, Hatsopoulos NG. Propagating waves mediate information transfer in the motor cortex. Nature Neuroscience. 2006;9(12):1549-57. Available from: https://doi.org/10.1038/nn1802.

      (24) Davis ZW, Benigno GB, Fletterman C, Desbordes T, Steward C, Sejnowski TJ, et al. Spontaneous travelling waves naturally emerge from horizontal fiber time delays and travel through locally asynchronous-irregular states. Nature Communications. 2021;12(1):6057.

      (25) Vogels TP, Abbott LF. Signal propagation and logic gating in networks of integrate-and-fire neurons. Journal of Neuroscience. 2005;25(46):10786-95.

      (26) Kumar A, Schrader S, Aertsen A, Rotter S. The high-conductance state of cortical networks. Neural Computation. 2008;20(1):1-43.

      (27) Destexhe A, Contreras D. Neuronal computations with stochastic network states. Science. 2006;314(5796):85-90.

    1. eLife Assessment

      This study shows that combining forced cell cycle re-entry with Rbpj deletion enhances Müller glia dedifferentiation and promotes their conversion into retinal neuron-like cells in the uninjured mouse retina. It provides a valuable strategy for improving Müller glia-mediated neurogenesis and advancing regenerative potential in the mammalian retina. Overall, the data are convincing. The authors have also addressed concerns regarding Müller glia function, cell survival, and the limitations of neuronal maturation and integration, further strengthening the conclusions of the study.

    2. Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits of the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly and through multiple methods, including immunostaining, single nucleus RNA sequencing, and in situ hybridization, that coupling notch inhibition with cell cycle re-activation induces the expression of neuronal markers in mammalian MG. The sn-RNA-seq data is particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Comments on revised version:

      The authors sufficiently addressed all concerns. I particularly appreciate the additional experiments to demonstrate retinal function, and the edits to the text regarding retinal and cell function and retinal organization.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines Müller glia (MG) reprogramming in the uninjured mouse retina through a combination of Notch signaling inhibition and AAV-induced proliferation. Building on their prior work showing that Cyclin D1 overexpression and p27^Kip1^ knockdown (CCA) promotes MG proliferation with very limited neurogenesis, the authors now demonstrate that Rbpj deletion alone induces a modest degree of MG-to-neuron conversion without proliferation, in agreement with recent work in the field. However, combining Rbpj deletion with CCA-mediated proliferation substantially enhances MG dedifferentiation and the generation of retinal neuron-like cells. Through genetic lineage tracing, histological analyses, and single-cell transcriptomics, the authors provide evidence that MG-derived cells acquire molecular features of bipolar (ON, OFF, and rod bipolar) and amacrine neurons. Most MG-derived cells appear to survive long-term (up to 9 months).

      Strengths:

      Overall, the study is carefully designed and executed, and the manuscript is clearly written with well-presented figures. While the work does not significantly expand the repertoire of neuronal types generated from mammalian MG beyond what has been previously reported in the field, it provides a valuable and improved strategy for inducing robust MG proliferation and neurogenesis in the mammalian retina.

      Weaknesses:

      (1) It would be better to include a negative control AAV when evaluating the effect of CCA AAV in the Rbpj KO background. This could help distinguish the specific contribution of the CCA construct from potential effects of intravitreal AAV injection itself, which can induce mild inflammation, known to influence MG reprogramming.

      To address this concern, in the revised manuscript we included the result from Rbpj KO eyes injected with a negative control AAV (AAV<sub>7m8</sub>-GFAP-GFP) (Fig. S14a). MG reprogramming efficiency, quantified as the proportion of tdT<sup>+</sup>Otx2<sup>+</sup> cells among total tdT<sup>+</sup> cells, was then compared between the AAV-GFP–treated and Rbpj KO–only eyes. At 4 months post-injection, the percentage of tdT<sup>+</sup>Otx2<sup>+</sup> cells in the AAV-GFP–treated eyes was comparable to that of Rbpj KO alone (Fig. S14b–c), and substantially lower than in the CCA-treated eyes. Together, these results indicate that the enhanced MG reprogramming observed in the Rbpj KO+CCA group is driven by transgenes expressed rather than by nonspecific effects of AAV or injection.

      (2) The extent of MG transduction by the CCA AAV is not clear. As quantifications are normalized to total MG (GFP^+^ or TdTomato^+^) or retinal length, it would be useful to clarify whether near-complete transduction is assumed, or if additional information on transduction efficiency can be provided.

      In our previous study (Wu, Liao, et al., 2025, eLife), we have demonstrated that high-dose (4E10vg/injection) AAV7m8 effectively transduced the whole retina, with near-complete MG transduction observed in the vicinity of the injection site, as evidenced by virtually all MG expressing GFP in these regions. In the revised manuscript, we clarified the transduction efficiency in Line 108-110 on Page 5 and Line 625-626 on Page 27.

      (3) In Figure S10, the reduced MG proliferation observed in the CCA + Rbpj deletion group could also potentially reflect decreased GFAP promoter activity in dedifferentiated MG following Rbpj deletion. Alternatively, MG-derived cells may be more fragile under these conditions.

      We thank the reviewer for these excellent insights. We agree that a down-regulation of GFAP promoter activity following Rbpj-mediated dedifferentiation is a highly plausible explanation for the moderate reduction in proliferation, as lower promoter activity would diminish AAV transgene expression. We have included this possibility in the data interpretation (Line 176-178, page 8). Regarding the alternative possibility of increased cell fragility, we agree that cell death cannot be ruled out, but occasional apoptotic cells over a long period of time are difficult to capture experimentally.

      (4) In the CCA + Rbpj deletion condition, do MG undergo single or multiple rounds of cell division?

      We have previously demonstrated that MG typically undergo a single round of cell division in wild type mouse retina following CCA treatment (Wu, Liao, et al., 2025, eLife). Given our observation that Rbpj deletion suppresses CCA-induced MG proliferation (Fig. S11), it is unlikely that the addition of Rbpj deletion would trigger multiple or continuous rounds of cell division beyond the single-round baseline established by CCA alone. While we did not re-evaluate cell division kinetics in the current study, we reason that CCA similarly drives MG to undergo a single round of division in the Rbpj KO context.

      (5) What fraction of neuron-like cells (bipolar- and amacrine-like) arises from proliferation versus direct transdifferentiation? Quantification of MG-derived cells expressing neuronal markers (e.g., Otx2, HuC/D), with and without EdU labeling, would help distinguish these mechanisms.

      The percentages of MG-derived cells expressing neuronal markers with and without EdU labeling, were shown in Fig 3d-e and Fig S19d-e. In the Rbpj KO-only group, neuron-like cells arise exclusively through direct transdifferentiation without cell division, as no EdU incorporation was detected in Rbpj-deficient MG. In this group, a small fraction of MG-derived cells expressed the neuronal marker Otx2 or HuC/D (Fig 3e, Fig S19e). In contrast, the Rbpj KO+CCA group achieved a substantially higher neurogenesis rate, with a significant proportion of Otx2+ or HuC/D+ MG-derived cells also being EdU+ (Fig. 3d, Fig. S19d), indicating that they arose through de novo neurogenesis. By subtracting the contribution of direct transdifferentiation observed in the Rbpj KO-only group, we estimate that majority of MG-derived neuron-like cells in the Rbpj KO+CCA group were generated through proliferation-mediated de novo neurogenesis.

      (6) In Figure S18a, the authors state that "while the neuron-like clusters were best classified as BC-like and AC-like based on their distinct marker gene expression, they also exhibited mixed expression of genes associated with other retinal neuronal types, including RGC markers (e.g., Tubb3, Myt1l, Grin1) and photoreceptor markers (e.g., Crx, Prom1, Epha10, Gucy2e, Scg3) (Fig. S18a), suggesting that the regenerated cells exist in a hybrid state" and "MG derived neuron like cells also expressed genes characteristic of RGCs and photoreceptors, indicating enhanced lineage". However, many of these genes are not specific to RGCs or photoreceptors and are instead broadly expressed in retinal neurons or enriched in bipolar/amacrine populations. Therefore, it is unclear whether these cells exhibit hybrid RGC or photoreceptor identity.

      We thank the reviewer for this insightful comment and for pointing out the need for greater precision in our terminology regarding these markers. While individual markers may lack absolute, 100% cell-type exclusivity, genes such as Tubb3 and Gucy2e serve as widely accepted lineage-associated genes that characterize RGC and photoreceptor programs, respectively (Soto et al., 2008; Sato et al., 2018; Sotani et al., 2024). We have revised the manuscript to replace terms "RGC-specific genes" and "photoreceptor-specific genes" with "RGC signature genes" and "photoreceptor signature genes", respectively. Furthermore, these RGC- and photoreceptor-signature genes are co-expressed across the entire Otx2+ MG population rather than being segregated into distinct, specialized subpopulations (Fig. 4d, Fig. S20). This uniform distribution indicates that these cells possess a hybrid transcriptional program that concurrently incorporates elements of both RGC and photoreceptor identities.

      (7) The authors provide a thorough molecular characterization of MG-derived cells through immunostaining and single-cell sequencing. However, their morphological features, synaptic connectivity (e.g., synaptic marker expression), and electrophysiological properties remain largely uncharacterized. While these experiments may be technically challenging, this limitation should be discussed.

      We agree with the reviewer that characterizing the precise morphological features, synaptic connectivity, and electrophysiological properties of MG-derived cells is a crucial step for any neuronal regeneration study, and we acknowledge that this represents an important limitation of our current study.

      As demonstrated by snRNA-seq data, the MG-derived neuron-like cells exhibit an incompletely mature state, characterized by hybrid transcriptomic signatures. By immunostaining, we did not observe any MG-derived cells with photoreceptor outer segment or typical RGC morphology. Therefore, it is highly likely that these cells have not established functional synaptic connectivity or acquired mature electrophysiological properties. Performing functional or circuitry assessments at this stage would be premature.

      We have added a comprehensive discussion regarding this limitation, along with future directions for long-term functional validation, in the revised manuscript (Line 502-515 on Page 22).

      (8) The conclusion that CCA + Rbpj deletion induces neurogenesis without compromising MG supportive functions or retinal homeostasis appears somewhat oversold. This claim is primarily based on gross retinal morphology and ZO-1 staining. Given the extent of MG dedifferentiation and ectopic cell generation in the ONL and INL, it is likely that retinal function is affected. Functional assessments (e.g., ERG) would be required to support this conclusion. The authors should consider tempering this statement.

      To address the concern raised by the reviewer, we performed electroretinography (ERG) to evaluate both scotopic and photopic retinal function in the Rbpj KO+CCA-treated eyes compared to contralateral untreated controls (Supplementary figure S23e-h). In addition, we conducted optomotor response testing to assess whether visual behavior is affected following treatment (Supplementary Figure S23d). The results demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal function.

      (9) Regarding the mechanism by which CCA-induced proliferation enhances MG reprogramming in the Rbpj knockout background, one plausible explanation is that chromatin states (e.g., histone modifications and DNA methylation) are transiently reset during DNA replication and cell division. While this alone may be insufficient to activate neurogenic programs, it could synergize with Rbpj deletion to allow neurogenic transcription factors (such as Ascl1, Otx2, NeuroD1, and NeuroD2) to access previously inaccessible chromatin regions, thereby promoting MG reprogramming.

      We thank the reviewer for the insightful suggestion on the model, which aligns well with our experimental findings. Our snATAC-seq data demonstrate that CCA-induced proliferation broadly increases chromatin accessibility at key neurogenic loci, including Neurod2, Dll1, and Otx2, in active MG compared to resting MG (Figure 6f–h). This chromatin remodeling alone is insufficient to drive neurogenesis, as CCA-only treated MG largely revert to a quiescent glial state. However, when combined with Rbpj deletion, which derepresses downstream neurogenic transcription factors such as Ascl1 and Neurog2 by relieving Notch-mediated transcriptional repression, these newly accessible chromatin regions can be effectively occupied and activated by the available neurogenic factors. The concept that cell division facilitates epigenetic resetting to enhance reprogramming efficiency is well established in the somatic cell reprogramming field, where proliferation rate is directly proportional to reprogramming success by promoting the erasure of lineage-restrictive epigenetic marks and the re-establishment of new transcriptional circuits. In the revised manuscript, we incorporated this mechanistic discussion to provide a more comprehensive interpretation of how proliferation and Notch inhibition converge to promote MG neurogenesis in Line 462-477 on Page 20-21.

      Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able to proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here, they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly - and through multiple methods including immunostaining, single-nucleus RNA sequencing, and in situ hybridization - that coupling notch inhibition with cell cycle reactivation induces the expression of neuronal markers in mammalian MG. The snRNA-seq data are particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Weaknesses:

      Whether the newly generated neurons are functionally integrated remains unclear, and the effect of the manipulation on the function of the retina was not tested. Imaging data suggests that many of the newly generated neurons persist for months, but often appear mislocalized. It is also not clear if the manipulation of MG affects long-term MG function. Cell death was not evaluated, and although the authors evaluated the long-term effect on tight junctions, this data was not quantified, and further analysis on morphology or function was not done. Control eyes were untreated, not vehicle-injected.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The transgenic line may be Glast-CreERT, not Glast-CreERT2.

      We appreciate the reviewer for bringing this to our attention. The formal allele symbol for this transgenic line is Tg(Slc1a3-cre/ERT)1Nat, while this strain is generically classified as "Cre/ERT2" by the Jackson Laboratory. Some published studies referred to this line as Glast-CreERT and others as Glast-CreERT2. To maintain consistency with the formal allele symbol, we have adopted "Glast-CreERT" throughout the revised manuscript.

      (2) For snATAC data in Figure 6 e,f, and Figure 19b. It is most likely gene activity, not gene expression, since these are snATAC, not snRNA data.

      For this inaccurate terminology, we have corrected all relevant figure labels and associated text in the revised manuscript to clearly state "gene activity" instead of "gene expression."

      (3) Some text in Figure 6 is a bit too small to read.

      We have increased the font size of the text elements in Figure 6 to ensure readability and have also reviewed all other figures for consistency. Revised figures with improved legibility have been included in the updated manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) There are multiple instances where further elaboration of methods or tools in the test would improve readability and comprehension by a broader audience. It would be helpful to (early, often, and clearly) explain precisely which cell types are labeled in your mouse line and how. Someone unfamiliar with the mouse line may struggle to understand what is labeled by the tdT or GFP. Likewise, it would help to consistently define what cell types are labeled by tdT+ vs. Sox9+, tdT+, etc.

      We have added a clear and detailed description of the mouse lines and labeling strategy early in the Results section, specifying which cell types are labeled by tdT and GFP and how the labeling is achieved. We have also ensured that the definitions of cell type identifiers (e.g., tdT<sup>+</sup> for MG-derived cells, Sox9<sup>+</sup>/tdT<sup>+</sup> for MG remaining in a glial state) are consistently stated upon first use and maintained throughout the manuscript to improve readability for a broader audience. In addition, we added headings for the quantification graphs to improve readability in all quantification figures.

      (2) It is unclear what the difference is between Figure 1c and S1c, and these should be quantified as the % of positive cells, as described in the text.

      We have removed Figure S1c and moved Figure 1c to supplementary figure 1. The MG labeled by EdU and Sox9 or Otx2 were quantified as % of the EdU+ MG.

      (3) S2e: Clarify what pixel level means, is this pixel intensity?

      Yes, "pixel level" in Figure S2e refers to pixel intensity. We apologize for the ambiguous wording and replaced "pixel level" with "pixel intensity" in the revised figure legend to ensure clarity.

      (4) Figure 2: In the magnified image of the GFP+, Sox9- cell, the GFP is also very faint. Could these cells be dying? Analysis of the expression profile of these cells (or ruling out apoptosis) would better support a dedifferentiation argument.

      The faint GFP signal observed in GFP<sup>+</sup> Sox9<sup>-</sup> cells is a sign of ongoing dedifferentiation rather than cell death. This is likely due to chromatin remodeling during reprogramming. A similar decrease in reporter signal intensity during MG dedifferentiation has been previously reported by Le et al. 2024, 2025, supporting the interpretation that reduced fluorescence is a characteristic feature of this process. It is possible that a small fraction of GFP<sup>+</sup> Sox9<sup>-</sup> cells may undergo cell death over an extended period, which would be difficult to detect using apoptosis assays. Our long-term survival experiments demonstrate that more than 80% of MG-derived neuron-like cells survive for at least 9 months following treatment (Figure 7), indicating that majority of these cells are viable. The discussion is included in line 112-114 on page 5.

      (5) Figure 3: The Crx labeling appears everywhere except the identified cell. This seems the opposite of the point you are making.

      Crx signal of the MG-derived cell (tdT<sup>+</sup> Crx<sup>+</sup>), which is pointed out by arrowhead, is in a ring-like pattern. This pattern is consistent with the euchromatin region in inverted nucleus of rod. Crx labeling appears in other cells in the image as Crx is highly expressed in native photoreceptors.

      (6) I think it would be nice to address, in the discussion, the apparent disorganization and mislocalization of cells in the long-term images.

      We thank the reviewer for highlighting this critical observation. During retinal development, precise laminar positioning of neurons is guided by a coordinated interplay of cell-intrinsic transcriptional programs and extrinsic cues including cell adhesion molecules, guidance factors, and interactions with neighboring cells. In the adult retina, many of these developmental cues are no longer present or active, which likely contributes to the failure of MG-derived neurons to migrate to their appropriate laminar positions. Interestingly, the vast majority of our divided MG cells remained localized within the outer nuclear layer (ONL). Because the ONL is the physiological location of photoreceptors, this preferential position could serve as an advantageous baseline layout for driving targeted photoreceptor differentiation in future work. To address reviewer’s feedback, we have expanded our discussion section (Line 543-559, page 23-34) to cover the mechanisms underlying this structural disorganization and its downstream implications for functional circuit integration.

      (7) I'm not convinced that ZO1 alone is sufficient to suggest MG function normally or that retinal homeostasis is maintained. I suggest tempering that conclusion in the text.

      For the revision, we have performed additional experiments to address this concern. The optical coherence tomography (OCT) images revealed that retinal layer organization and ONL thickness were comparable among the uninjected eyes, GFP AAV-injected control eyes, and CCA-treated eyes, demonstrating that overall retinal architecture was well-preserved (Fig. S23a–c). Optomotor response testing revealed no significant differences in visual acuity across groups, suggesting that visual function remained intact (Fig. S23d). Furthermore, electroretinography (ERG) demonstrated that scotopic and photopic a- and b-wave amplitudes were unaffected by the treatment, confirming that light responses from photoreceptor and inner retinal neuron were preserved (Fig. S23e–h). Taken together, these findings demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal structure and functional visual circuitry.

    1. eLife Assessment

      This study investigates the role of the Z-disc protein Zasp52 in Drosophila flight muscles and provides evidence that an intrinsically disordered region (IDR) helps to stabilize and promote the localization of the protein to the Z-disc. Overall, this represents an important study that provides insights into Z-disc function and maintenance. The data are convincing, supported by strong genetic evidence and behavioral tests, well-controlled experiments, and detailed statistical analyses. Characterization of a new actin-binding motif mutant and FRAP analysis further support the functional importance of the Zasp52 IDR.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

    3. Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Thank you for the helpful comments and criticisms. We provide exciting additional data, in particular a CRISPR actin-binding motif mutant and FRAP analysis of an exon15e-GFP transgene, both further supporting the importance of the IDR in thin filament stability. We believe that these additional experiments provide compelling evidence supporting our conclusion and substantially advance the current limited body of knowledge surrounding the role of IDRs in structural proteins.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      However, some of the phenotypic analysis, in particular the bending of the sarcomere, likely upon mechanical challenge by muscle contractions, needs more detailed investigations to be fully convincing.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

      Weaknesses:

      (1) The western blot quantifications of Zasp isoform expression are weak. No error bars are indicated in the quantifications; the quantifications appear to be more qualitative than quantitative. According to band intensities, the long Zasp isoforms seem to be less present compared to the shorter ones, even in the flight muscles.

      We have now included quantifications with error bars for the Western blots in our resubmission. It is important to keep in mind that the main point in figure 1B is that there are plenty of exon15e-containing isoforms in IFM, in contrast to other tissues with very limited exon15e-containing isoforms. This is confirmed by the analysis of RNA-seq data in figure 1C, and of course, by the flightless phenotype of the exon15e mutant.

      (2) The phenotypic analysis of the sarcomere appears somewhat superficial throughout the paper. Only Zasp52 and phalloidin are shown; no other Z-disc or thick filament proteins. At least myosin stainings and overview images are important to better judge the phenotypic variations. Are the variants between individuals or regional in the same muscle?

      Our images are representative of the observed phenotypes. Phenotypes are consistently present across all individuals, as reflected in our replicates. Interestingly, they appear not to be randomly interspersed among the sarcomeres but concentrated in certain regions of muscle more than others. Full images are available in the online repository FigShare.

      (3) EM images would benefit from better quantification.

      We do not believe that EM images can be meaningfully quantified, because of the many selection steps preceding image acquisition.

      (4) Other proteins were not analysed with the FRAP-based turnover assay for comparison in wild type and mutant. All Z-proteins might turn over faster in the mutant with the defective Z-disc.

      This is the point we are trying to make. The Zasp52 IDR appears to stabilize the Z-disc and is likely involved in fastening a variety of proteins to it.

      Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

      Weaknesses:

      (1) It is not clear in Figure S1A where exon 15e fits within the Zasp52 locus schematic. This is important as a premise of this paper describes this region to be key, and proof from multiple prediction programs would lend more weight to the prediction of the exon being largely disordered. Inclusion of the discussed short linear motifs, comparison with Canoe or LBD3 for similarities and/or an Alphafold structure would help make the authors' point (colorized with known domains).

      We added a bar below figure S2A to show the region corresponding to exon 15e. We used three disorder prediction programs and one structure (order) prediction program. The majority of exon15e is completely disordered and of very low confidence score, and thus uninformative to display as an AlphaFold structure. Likewise, IDR’s are very difficult to classify, therefore we cannot say much more than that LDB3, Zasp52, and Canoe contain IDRs, with Zasp52 and Canoe both having a putative actin-binding domain within the IDR. We now provide data on the function of the ABD in this resubmission.

      (2) Interesting that immobilization rescues the deterioration phenotypes. The authors should explain in more detail how this was done to avoid dehydration/starvation of the flies.

      We provided more details in materials and methods.

      (3) There is a lot of discussion about the potential function of the IDR region, specifically a putative actin binding motif or other 'ordered' regions that may contain short linear motifs. It would strengthen the findings to show which of these may be essential for Zasp52 function in the IFM. The ability to bind actin could be tested biochemically, and/or smaller deletions could be made to unequivocally test the role of the ABD vs other predicted motifs using genetics. If some of these regions are more ordered, where do they lie within, and do they form a predicted fold or structure that gives insight into function?

      We now provide data on the function of the ABD showing that deleting it has almost no phenotypic defects. That means the IDR is largely/entirely responsible for the observed phenotypes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Western blot in Figure 1B needs proper quantification. A ratio between long and short isoforms in the same muscle type might be informative. Is it known which epitope the antibody recognises? Can a GFP insertion that also labels all isoforms be used as verification? Quantifications are also needed in Figure 2A.

      We have added quantifications of all Western blots (Fig. 1B, 2A, and 2A’). The ratio between exon 15e-containing and total Zasp52 in the same muscle type is included in the lowest bar graph in Fig. 1B. The full-length antibody is polyclonal and was raised against Zasp52-PR which contains all ordered domains; the anti-LIM antibody was raised against the last three LIM domains (both are described or referenced in the materials and methods section). Such a GFP insertion cannot exist due to the complex splicing patterns of Zasp52.

      (2) The name of the deletion allele could be specifically indicated in Figure 1A below the red bar.

      Done.

      (3) It would be useful to indicate the order group names in Figure S2B since species names are hard to read.

      For the version of record we provided high-resolution images, where species names can be read. Drosophilids, Ephemeroptera and Odonata are indicated.

      (4) The inverted spelling of the numbers for the control in Figures 2C and 5H is strange.

      Changed to normal spelling.

      (5) The bending of the myofibril at the Z-disc is a really interesting phenotype. However, it seems it is not always visible; at least it is visible in many myofibrils shown in Figure 3B, but in none in Figure 3E, same genotype, just different staining. Hence, I wonder if this bending could be force-induced by the cutting of the thorax during tissue preparation. It would be useful to display some overview images to allow the reader to judge the quality of the tissue preparation, indicating from where the high magnification view shown was taken. The same is true for Figure 5.

      Overview images are available on FigShare. Note that you can see some “H-zone actin” sarcomeres in Fig. 3B, as well as some mildly bent ones in Fig. 3E. We generally selected images that best demonstrated the phenotype described. Furthermore, neither phenotype is fully penetrant so we cannot expect to see it everywhere. Lastly, it is always possible that phenotypes are affected by preparation, since it is impossible to know what the myofibrils look like in situ. However, all samples were prepared using the same protocol with replicates, and since we see a phenotype in our mutants and not in the control, this indicates that something is different between the two.

      (6) The same applies to the visualisation of the "hyper-contracted" phenotype; again, it seems to be an all-or-nothing phenotype in the zoom shown. An overview image should be shown. The zoom in Figure 4E would benefit from displaying phalloidin in a separate channel. Are actin filaments pulled out of the Z-disc? The latter is often seen in non-perfect cuts in wild-type, but the accumulation at the M is curious. It would be informative to locate the ends of the thick filaments in these cases or quantify thick filament lengths; do these invade the Z-discs? This can easily be done by a myosin staining.

      Is this a regional effect or does it depend on the individual or on the preparation? I am surprised to also see the "hyper-contraction" in 10% of wild-type 5-day adults.

      See previous response where we include overview images. Single-channel images are available; it is visible that actin filaments are not pulled out of the Z-disc. Phenotypes are consistent across individuals as evidenced in our replicates but do tend to be concentrated in certain regions of muscle.

      (7) The EM images would benefit from more overview images. At the moment, we only see a single sarcomere from wild type and mutant, with no quantification of the phenotype. Can the authors see the invading thin filaments into the M-band? The disrupted Z-disc phenotypes are impressive. What is the age of the animal shown in Figure 4?

      We have a panel displaying several mutant sarcomeres. Due to the selectivity and challenges of the EM preparation process, we do not believe we can perform meaningful statistics on them. It sometimes looks like myosin heads are visible in the H zone which may support the presence of thin filaments in the H zone (Fig. 4B and C). However, the quality of these particular EM images is not high enough to identify thin filaments. All phenotypes shown are from 3-week-old animals.

      (8) Is UH-3 GAL4 expressed at the adult stage?

      Yes, from 36 h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (9) Figure 6 would strongly benefit from a myosin staining. Do thin and thick filament lengths scale? It seems that overlap is reduced in the double hets. How can this be envisioned with Z-disc stability? Is myofibril diameter reduced?

      We searched for non-additive differences in myofibril diameter but were unable to detect any.

      (10) What is the FRAP turnover rate of a long Zasp-GFP compared to a short one in wild type? A difference would indicate that it is really the IDR domain that keeps Zasp52 longer at the Z-disc, instead of an indirect effect caused by Z-disc morphology

      We have newly added FRAP data of a GFP-tagged exon 15e construct which displays much lower turnover. This indicates that the IDR does indeed retain Zasp52 at the Z-disc.

      Reviewer #2 (Recommendations for the authors):

      (1) The total protein stain should also be included if it is used for quantitation in Figures 1B and 2A-A'.

      These are available on FigShare.

      (2) It is a bit confusing that the Alphafold plot is inversely correlated with the other 3 prediction programs, although this is explained in the legend. Maybe an Alphafold structure would help make the authors' point (colorized with known domains).

      The AlphaFold structure is almost entirely low-confidence disordered region except for the structured domains so we do not believe it would be helpful to include.

      (3) The title of Figure 8 says 'Certain ex15e defects are rescued by immobilization.' What other defects are not rescued? If true, these should be shown.

      There was a full rescue. We deleted the word “certain”

      (4) Please include a brief explanation of the spatiotemporal expression of UH3-Gal4.

      From 36h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (5) Statistics should be added to Figure 8E.

      Figure 8E (now 9E) has statistics.

      (6) The dark blue color used for integrin staining in Figure S3 is difficult to see. Changing this color may help visualize differences. Also, pointing them out with arrows, etc., will help clarify abnormalities.

      We have described these differences in the figure caption. Single-channel images are available for viewing in any color in FigShare.

    1. eLife Assessment

      This study reports important findings regarding social influence on charitable donations, showing that giving is shaped by the statistical properties of others' donations in a manner that can be captured by a reinforcement learning model. The evidence for the conclusions is solid, although the computational modelling could be better motivated and described, the individual differences analyses could be more robust, and some design choices could be better motivated. Overall, the core effect appears robust and is supported by multiple well-designed experiments, but the conclusions that rely upon computational modelling and individual differences may require further support.

    2. Reviewer #1 (Public review):

      This manuscript investigates how people use sequential social information when deciding how much to donate to charity. Across four preregistered experiments, participants first made baseline donations to a set of charities, then observed a sequence of donations from five other people whose mean and variability were experimentally manipulated, and finally made a second donation to the same charities. The authors ask whether the mean and variability of others' donations affect the mean and variability of participants' own donations, and whether individual differences in psychopathy and empathy are associated with responsiveness to social information.

      The main behavioral finding is that participants shifted their second donations toward the mean of the donations they observed: generous social information increased donations, whereas stingy social information decreased donations. In contrast, the variability of observed donations had little effect on the mean donation shift, but did affect the variability of participants' subsequent donations, with more consistent social information producing stronger reductions in variability. The authors also fit several computational models and conclude that a hybrid model, in which second donations reflect both participants' initial donations and learned predictions of others' donations, best accounts for the data. Finally, they report that psychopathic traits are positively associated with donation change and with model-derived social-information use, and that this association generalizes to a perceptual social-influence task in Experiment 4.

      The paper addresses an interesting question and has several strengths, especially the repeated experimental design, the direct manipulation of social-information statistics, and the attempt to connect descriptive behavior with computational modeling and individual-difference measures. However, several aspects of the design and analysis currently block some of the major conclusions. The behavioral results provide convincing evidence that observed donation levels affect later donation decisions. The current evidence is less decisive for the stronger claims that the winning computational model identifies the underlying mechanism, that individual-level model parameters are robust phenotypes, and that psychopathy specifically increases susceptibility to social information.

      Strengths:

      A major strength of the manuscript is that it investigates social influence in charitable giving across four preregistered experiments with relatively large samples. The core mean-effect result is replicated across different donation scales, across hypothetical and incentivized settings, and across student and more general online samples. This gives the descriptive behavioral finding substantially more credibility than would be available from a single experiment.

      The experimental manipulation is also valuable. Rather than presenting only a single prior donation or a simple group average, the authors expose participants to sequences of donations and independently manipulate the mean and variability of this social information. This design allows the authors to ask not only whether social information changes donation levels, but also whether the distributional structure of that information changes the variability of participants' own responses.

      Another strength is the combination of traditional statistical analyses with computational modeling. The hybrid model is a reasonable descriptive candidate because it formalizes the intuitive idea that second donations may depend both on participants' initial preferences and on learned expectations about others' donations. This modeling approach has the potential to clarify mechanisms of social-information use, especially if the validation of the model and its individual-level parameters is strengthened.

      Experiment 4 is a sensible extension because it uses an incentivized design, includes a more diverse sample, examines transfer to novel charities, and adds a perceptual social-influence task. These features broaden the empirical scope of the manuscript and make the psychopathy-related findings more interesting, although the perceptual-task result should still be treated as requiring replication.

      Weaknesses

      The first limitation concerns causal interpretation of the phase effects. Participants always make baseline donations first, then observe social information, and then make second donations to the same charities. There is no non-social repeated-donation control condition. This type of design does support the conclusion that donation changes differ as a function of the observed social-information condition, especially the mean of others' donations. However, it does not by itself fully isolate social influence from other processes that could also occur between a first and second donation to the same item, such as repeated exposure to the charities, slider familiarity, memory of the first donation, regression to the mean, reduced uncertainty, fatigue, or "the experiment clearly wants me to update" demand effects. This issue is especially relevant for the claim that observing others' donations generally reduces the variability of individual donations. The variability effect may well be socially driven, but the absence of a non-social or irrelevant-information repeated-donation control means that this cannot be decisively demonstrated.

      The second limitation concerns the trial-level mixed models. The primary mixed-effects models include random intercepts for participants and items, but do not appear to include random slopes for within-participant or within-item phase effects. Since phase is repeatedly manipulated within participants and items, random-intercept-only models may underestimate uncertainty for some phase interactions, resulting in anti-conservative p-values. The convergent participant-level ANOVA analyses are reassuring, but the trial-level inferential claims would be stronger if the authors reported additional analyses using fuller random-effects structures or other methods that better reflect the repeated-measures structure.

      The third limitation concerns model comparison and model validation. The computational models are fit separately to each participant, and model comparison is based on summed information criteria and protected exceedance probabilities derived from those participant-level fits. This is informative about relative conditional fit within the tested sample and model set. However, the manuscript uses the winning model to support broader claims about latent computational mechanisms, individual computational phenotypes, psychopathy-related susceptibility, and potential intervention relevance. For these claims, the relevant prediction target is generalization to new participants, whose individual parameters are not known in advance. The current model-comparison approach is not well aligned with that target. Additionally, the loss appears to combine prediction trials and donation outcomes, so the selected model may more strongly reflect performance at predicting participants' guesses about others rather than specifically predicting their own donation decisions.

      The fourth limitation concerns the model adequacy checks and recovery analyses. The analyses described as posterior predictive checks do not appear to be posterior predictive checks, because the models are not Bayesian and there consequently isn't a posterior to check. Instead, the analyses appear closer to some sort of in-sample fitted-value reconstruction checks. Such checks provide limited evidence of model adequacy, especially because the same second-donation data used to estimate individual parameters are then used to assess whether the fitted model reproduces the main behavioral patterns. In addition, the reported model and parameter recovery analyses use extremely favorable response-noise assumptions that are not expected to be met in real data. The analyses establish that the models and parameters are mathematically distinguishable in principle, but they do not establish that the individual-level parameters are reliably recoverable under realistic empirical noise levels to the extent required for the analyses performed in the manuscript.

      The fifth limitation concerns the interpretation of the psychopathy results. The association between psychopathic traits and donation change is interesting and appears directionally consistent across experiments. However, the interpretation that psychopathy increases susceptibility to social information is vulnerable to biasing by baseline-distance. The manuscript reports that psychopathy is negatively associated with baseline donations in Experiments 1-3. Participants with lower baseline donations have more room to move toward generous social information, and absolute donation change is partly a function of the distance between the initial donation and the observed social mean for mechanical reasons. Thus, an association between psychopathy and absolute donation change could theoretically arise even if psychopathy does not directly increase social susceptibility.

      A sixth limitation is that we could not find the links to the preregistration. The authors state when preregistered hypotheses were or were not supported, but it is unclear how these hypotheses were phrased. Most notably, it is unclear how variance in the observed donation choices was supposed to influence participants. As a side note, it was not quite clear if the variance in the observations was higher or lower across charities, across observed persons, or across both.

      Several more minor suggestions can also be made regarding the modelling and the presentation of the task, etc.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript examines how the statistical properties of others' charitable donations shape subsequent giving using four preregistered experiments and computational modelling. The authors find that both the average level and variability of observed donations influence donation behaviour, and that individual differences in social information use are associated with psychopathic traits.

      Strengths:

      This is a well-executed paper on the important question of how social information shapes charitable giving. In my view, the combination of preregistered experiments, large sample sizes, computational modelling, and a multi-paradigm approach makes for convincing evidence. The progression across experiments, the use of real donation data rather than deception, the incentivized experiment 4, and the generalization to a second paradigm are all notable strengths. The introduction is clearly written and well-motivated - an enjoyable read. The experimental paradigm is thoughtfully designed, and the methods and supplementary materials are described in considerable detail. The computational modelling provides useful additional insights beyond the behavioural analyses.

      As far as I could tell, the manuscript also adheres closely to the preregistrations. The primary hypotheses, experimental designs, exclusion criteria, and key analyses are all consistent with the preregistered plans. Deviations seem to consist of methodological improvements (e.g., mixed-effects models replacing ANOVAs), additional computational and robustness analyses, and therefore strengthen rather than weaken the manuscript. (NB: for transparency, I would appreciate a clearer distinction between preregistered and post hoc analyses, as well as a brief explanation for why some preregistered secondary analyses are no longer reported; see minor comments below).

      Overall, I enjoyed reading this paper. I believe it will make a valuable contribution. My comments below are intended to further strengthen an already solid manuscript.

      Weaknesses:

      (1) The rationale for the social-information phase could be clarified further. Given the research question, I wondered why participants observed the five donations sequentially (and only briefly) rather than simultaneously. In particular, variance is arguably more difficult than the mean to encode and remember, and a sequential presentation may both obscure distributional differences and introduce primacy or recency effects. It would be helpful if the authors could better motivate this design choice, and indicate whether they examined possible order effects.

      Relatedly, I felt somewhat uncertain about the purpose of asking participants to predict each donation before observing it. The prediction phase appears to play an important role in the computational model, but its theoretical role is not clearly introduced. Is it intended as a measure of participants' evolving beliefs about the descriptive norm, or primarily as a modelling device? Finally, were these predictions incentivized (e.g., for accuracy), and if not, how should readers interpret them?

      (2) I would appreciate having the full experimental materials reproduced in the Supplementary Information. This would make it easier to understand what participants experienced during the task, including what they were told about the "other participants" whose donations they observed.

      Minor points:

      (1) The interpretations around domain-generality would be strengthened by reporting the association between social information use in the charitable giving task and in the BEAST. Currently, both measures are shown to correlate with psychopathy, but it remains unclear whether individuals who rely strongly on social information in one task also do so in the other. Reporting this correlation (or explaining why it cannot be meaningfully computed) would provide a nice and direct test of a domain-general tendency to use social information.

      (2) It would help to explain more explicitly why the standard deviation of donations is theoretically interesting in its own right. The motivation for studying the mean seems immediately intuitive, whereas the motivation for focusing on variability could be elaborated on further in the Introduction.

      (3) As I said above, I think the manuscript follows the preregistrations closely. Maybe I missed it, but it seems that prediction accuracy and reaction-time analyses were omitted. It would improve transparency further if the authors would briefly mention the preregistered secondary analyses that are no longer reported (and explain why they were omitted).

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors aimed to assess the mechanisms of social influence on charitable giving, particularly by separating the role of donation magnitude and variability in others' donations, and by examining the role of incremental social information in a learning framework. They additionally investigated individual differences in the magnitude effects in relation to self-reported psychopathy and empathy. The main findings suggest that magnitude and variability of others' donation impacted the magnitude and variability of the participants' donations, respectively, and that the weight of social information on individual decisions correlates positively with psychopathy, but not with empathy.

      Strengths:

      (1) The findings extend previous evidence for social influence on charitable giving to contexts where social information is provided incrementally, and to effects on the variability in social information (in addition to the mean).

      (2) Individual differences suggest a role for psychopathy, but not empathy.

      (3) Findings are replicated across all 4 (or for some findings 3 out of the 4) experiments, which helps strengthen the claims.

      (4) Multiple experiments are a strength, especially Experiment 4, which helped address concerns/potential confounds in the previous experiments, increase representativeness of the sample, add incentive compatibility, and generalize to another task domain (perceptual).

      (5) For modelling, strong model and parameter recovery was obtained, thus validating the modelling pipelines.

      (6) The experiments were pre-registered, though it's unclear whether only planned analyses were pre-registered, or specific directional hypotheses. It would help if the manuscript took the reader through the pre-registration (and any deviation from it), instead of expecting the reader to do the comparison between the pre-registrations and actual manuscripts.

      (7) The studies are appropriately powered, and power analyses are provided.

      Weaknesses

      (1) Lack of rationale and justification for the between-subjects design.

      While this design may be appropriate in some cases (for example, for the generalization of donation to new charities or as a potential "intervention"), it would have been great to know if the findings related to social influence extend to a within-subjects design, especially given the weak results related to the effects of standard deviation in others' donations. It is possible that variability in others' responses would have a stronger effect if manipulated within individuals, since the same individual exposed to both high-SD and low-SD social information may weight low-SD information more, but this effect may lack when individuals are only exposed to the same variability across trials.

      (2) Motivation for the RL framework.

      The use of reinforcement learning (RL) isn't very well motivated, both in the introduction and methods/results (given the task). In particular, why is RL relevant to studying the problem of social influence, which isn't inherently a learning problem? This should be better motivated in the introduction. Second, when taking the task into account, it's unclear why RL is an appropriate model, given that from the perspective of the participant, the 5 others are different individuals, so the model shouldn't assume that predicting an individual's donation should be related to the previous individual's donation. Unless participants are informed that there is some dependency between the 5 donors they observe on each trial? If so, this should be made clear.

      (3) Specifics of modelling analyses, and separability between prediction and second donation data.

      Does the RL-based model (either prediction-only or hybrid) explain more variance in second donations than a simple linear regression model predicting second donation from initial donation and the mean of others' donations (or each individual other's donation)? It could be helpful to add some models that include social influence (i.e., integration of social and individual information) but no learning mechanisms per se. If this is not done, I do not believe that current results show that participants combine "their initial self-donation tendencies with their predictions of observed others' giving to guide their second individual donations". While participants may update their predictions, the authors should test multiple models of prediction update (fit only on the prediction data to understand the specific mechanisms of prediction update independently of second donation - for example, is it RL, or could it just be a running average, or some other heuristic? In parallel, it would be helpful to test whether it's the learned predictions (or whatever other prediction update mechanism was found to best explain the prediction data) or the actual others' donation information that best explains second donation - when combined with initial donation. These latter models would be fit on second donation data only in order to be comparable. If it's not possible to separate people's predictions from the actual social information (others' donations) then this should be acknowledged as a limitation. Ultimately, separating the modelling by data type (prediction only vs second donation data only) would help provide more insights into the learning mechanisms (if any) and whether it's learned prediction, or just social information, which influences second donation.

      (4) Missing statistics in generalization to novel donation results.

      On page 13, in the generalization effect, the authors mention that "Compared with participants exposed to High-SD social information, those exposed to Low-SD social information exhibited less variability in their novel donations, with this effect being especially pronounced in the Low-Mean condition." Was this supported by a significant interaction between SD and Mean condition? If so, please report the statistics of the interaction; if not, it's probably better to refrain from making this claim.

      (5) Behavioral index of social influence individual differences.

      For the first analysis reported on the association with psychopathy (Figure S9), as well as empathy (Figure S10), the absolute change between first and second donation does not seem like the appropriate marker of social influence. While I understand from Figure 2 that most participants changed their donation in a direction consistent with the social information, it would appear more appropriate to calculate an index of donation change consistent with influence, so calculated as D2 - D1 for the high mean groups and D1 - D2 for the low mean groups. This would be a better measure to interpret high values as an index of social influence.

      (6) Interpretation of psychopathy effects.

      a) The general idea that high psychopathy would be associated with increased social influence seems counterintuitive. While I appreciate that the authors controlled for additional variables such as age, gender, condition, and other model parameters, is it possible that this effect could be instead explained by the availability heuristic (the social information is more readily available to participants than their individual choice from the baseline trials), lower memory for their own choice, or lower IQ/cognitive abilities? These appear to be important confounds to address to be able to interpret the findings.

      b) Related to this, and given that psychopathy/empathy were negatively/positively related to baseline donation amounts, it would be good to account for baseline mean donation amount in the individual difference analyses.

      c) Finally, the authors interpret this association in line with other studies that have shown strategic social blending in psychopathy - while this seems possible in contexts where others are present, it doesn't really seem to be the case in this task. Did participants believe the other donors were watching them somehow? It also appears contradictory for the incentivized experiment, whereby if high psychopathy participants would no longer be able to "maintain a favorable social image while still pursuing their own self-interests" (p.23), since as soon as incentivization is added, participants' own self-interests are directly in conflict with the social image. Was participants' understanding of the incentive compatibility tested in Experiment 4?

      (7) Asymmetry between generous vs stingy social influence and link with psychopathy.

      a) Was such an asymmetry present - in other words, were people more strongly influenced by generous others or stingy others, or were the two comparable? I believe some analyses could be added to test this, and this is also where a within-subject design could help (e.g., different parameters for the two directions of social influence at the individual levels).

      b) Related to that, does the correlation with psychopathy vary between conditions? It appears important to test if the increased social susceptibility is general or specific to increases (~high mean group, generous social influence) or decreases (~low mean group, stingy social influence) in donation. I understand that the main effect of psychopathy survived controlling for conditions, but it would still be interesting to test for an interaction between psychopathy and condition in predicting donation changes (calculated as suggested in point 5 above) or social influence weight.

      (8) Perceptual task in Experiment 4.

      a) While it is good to show that there was no correlation between psychopathy and initial estimate in the perceptual task, were there differences in initial estimate accuracy (i.e., difference between initial estimate and correct answer) along psychopathology? If so, this should be controlled for in the analyses. Given that social influence is always in the direction of the true value, the proportional deviations between initial estimate and social information could yield larger numerical differences and induce larger changes in estimate.

      b) Even if previous studies have excluded rounds in which participants update their estimate in the opposite direction of the social information or move beyond it, I believe analyses that include those rounds should be included, especially in the context of individual difference analyses. Could it be that individuals who are high in psychopathy or low in empathy have a higher proportion of rounds where they go against the social influence? The same question applies to the main 4 experiments (in case this criterion was applied to) as well as the perceptual task.

      c) Because the perceptual task was completed by the same participants as Experiment 4, were the two social influence measures correlated across tasks? Was psychopathy better predicted by a combination of predictors across the two tasks?

      (9) Were individual difference measures examined in relation to the variability effect?

      (10) Discussion.

      The authors argue against a role for opportunistic conformity. While I tend to agree with their interpretation, I believe that it could be strengthened as follows:

      a) First, it relies on a null result (the absence of a difference in decreases between low-mean low-SD and low-mean high-SD groups), which I do not believe was explicitly tested; and even if it was, it should ideally be corroborated by Bayesian statistics to provide strength of evidence for the null effect.

      b) Second, this could be a great opportunity to dive into the mechanisms of social influence in the model, by testing the theory that only the lowest (or highest) donation from the group (rather than the mean, or the learned prediction) influences donation. Could a subset of participants be better fitted by such a model?

      (11) Methods. Maybe I missed it, but it's unclear what participants were told about the other donors they are observing. It is mentioned that they were fully debriefed after the experiment, but what they were told in the instructions appears important. Was believability tested (this also relates to my comment #1 about the rationale for a between-subjects design, which creates fairly biased sets of social information from the perspective of a single participant)? And related to my comment #2, what participants were told about the donors could help justify the rationale for the RL framework.

    5. Author response:

      We thank the editors and reviewers for their thoughtful and constructive comments on our manuscript. We are pleased that they considered the core behavioral findings important and robust, especially the results showing that the magnitude and variability of others’ donations affected the magnitude and variability of participants' donations, respectively. We also appreciate their acknowledgement of the strengths of the experimental design, large sample sizes, the incentive-compatible and across-domain measures included in Experiment 4, and combined behavioral and computational approaches.

      We agree that the manuscript would benefit from greater clarification in several areas, further analyses, and more cautious interpretations. In the revised manuscript, we plan to clarify the rationale for sequentially presenting social information, the role of prediction responses, the theoretical motivation of examining the variability of others’ donation, the use of the between-subjects design, and the motivation for the RL framework. We also agree that the lack of a non-social repeated-donation control condition limits the interpretation of the phase effects. Our design permits strong inferences about differences in donation changes across different conditions, but it cannot establish that the phase-related changes are exclusively attributable to social information exposure. We will revise the wording accordingly, moderate the causal language, and explicitly discuss the limitations of our design.

      To strengthen the behavioral analyses, we plan to supplement the current mixed-effects linear models with models that reflect the repeated-measure structure of the task, including random slopes for the phase. We will also add statistics in the generalization results section and test the asymmetry between generous vs. stingy social influence. Moreover, the reviewers raised an important concern regarding the associations between psychopathy and the donation change. Because psychopathy is negatively correlated with initial donations in several experiments, absolute donation changes may partly reflect the distance between initial donation and the observed donation mean. We therefore plan to reanalyze the psychopathy effects by using signed donation changes and trial-level discrepancies between participants’ initial donations and the observed social information. These additional analyses will enable a more direct and precise assessment of whether psychopathy is associated with greater susceptibility to social influence.

      We further agree that the comparison and validation of the computational models should be strengthened. In the revised manuscript, we plan to clarify that the learning models are intended to describe the updating beliefs about a group-level donation norm from sequential social information, rather than learning about a single donor. We will expand the candidate model set to include non-learning models, such as models based on the actual social mean, a running average. We will also model the prediction phase and the second donation phase separately. This will help identify the models that provide explanatory values for both predictions of others’ donations and individual donation behaviors. In addition, because the models were not estimated via a Bayesian framework, we agree that the term “posterior predictive checks” is inappropriate. We will rename these analyses. We will also rerun the parameter and model recovery analyses using empirically informed noise levels separately for the prediction and donation phases. In addition, we will implement model-evaluation processes, such as cross-validation, that better reflect prediction for new participants.

      In Experiment 4, we plan to directly report the association between social-information-use measures in the perceptual and the donation task to strengthen the domain-generality effect. We will additionally examine whether social susceptibility in the perceptual task is associated with psychopathy by including all trials, including those in which participants moved away from or beyond the social value.

      Finally, we will correct the reporting and presentation issues identified by the reviewers, including the social information use equation in the perceptual task, the pseudo-SD of individual donations formula, supplementary figure captions, and task duration. We will also provide fuller experimental materials and make the preregistration links more prominent. In addition, we intend to make the analysis code, model-fitting scripts, and data available during the revision process.

      We greatly appreciate the editors’ and reviewers’ thoughtful suggestions, which will help us substantially strengthen the manuscript. We are grateful for the opportunity to address these important points and believe that the planned revisions will enhance the manuscript’s clarity, robustness, and its contribution to the understanding of social influence in donation behaviors.

    1. eLife Assessment

      In this important study, Lau et al. identify non-conserved nucleotides within the common binding motifs of BLIMP1 and IRF4 that provide a molecular mechanism for their distinct roles as crucial transcription factors during the antibody-secreting cell differentiation. The major strength of this manuscript is the solid and detailed characterization of human in vitro plasma cell differentiation. However, several overstatements exist, therefore requiring careful revision to improve the manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablst/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. Direct testing of selected regulatory elements would make the causal claim stronger. Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fully demonstrated mechanism of gene regulation.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations. Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study. Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20+ cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation. Have both conditions been tested? This experiment also assumes that all B cells have equal potential to differentiate into antibody-secreting cells. What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20+ fraction would be enriched in these cells. Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138+CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c? What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells? Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

    4. Author response:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablast/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      We agree that CRISPR/Cas9 editing of bulk D7 cultures complicates interpretation because this population contains both activated B cells and PB/prePCs. We will therefore revise the text to distinguish phenotypic effects measured across the bulk D7 culture from the downstream single-cell analysis focused on cells along the prePC-to-PC trajectory. In particular, our interpretation of IRF4 and BLIMP1 function in prePCs is based primarily on the D9 scRNA-seq analysis, in which cells arrested in the activated B cell compartment are not used to define the perturbed PC-trajectory states. We will clarify this analytic design in a future revision and temper language implying that all effects arise exclusively within prePCs.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      We agree that the divergence between IRF4 and PRDM1 perturbations could reflect differences in protein turnover, or hierarchical dominance, in addition to developmental timing. We will revise the Discussion to state that our data support a temporally ordered model in which IRF4 acts early to license the prePC-to-PC transition and BLIMP1 consolidates the terminal state, but that the current experiments do not exclude alternative explanations related to hierarchical dominance or degradation kinetics. We will also note in the revised Discussion that degron-based perturbations, rescue experiments, and gain-of-function analyses would be needed to resolve the functional ordering of IRF4 and BLIMP1 with higher temporal precision.

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. (1) Direct testing of selected regulatory elements would make the causal claim stronger. (2) Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fudlly demonstrated mechanism of gene regulation.

      We agree that the current data support the motif lexicon primarily as a predictive model for differential TF occupancy rather than as a fully causal mechanism of gene regulation. We will therefore revise the relevant text in the Results and Discussion. The EMSA data directly test nucleotide-dependent binding preferences, and the CUT&RUN/multiome analyses show that these motif variants are differentially associated with IRF4- or BLIMP1-bound DEG-linked OCRs. However, direct causal testing of endogenous regulatory elements, for example by base editing of selected ISRE/EICE variants, will be required to determine whether these variants are sufficient to predictably alter gene activity in differentiating plasma cells.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations.

      Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study.

      Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20<sup>+</sup> cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation

      We agree that CD138 alone does not fully define terminal PC maturation. In the revised manuscript, we will clarify that CD138 was interpreted in the context of a broader maturation profile, including CD20 downregulation, ICAM2 upregulation, IRF8 loss, IRF4/BLIMP1 expression, Ki-67 loss, and antibody secretion. The D7 PB population was proliferative and CD138<sup>-</sup>, whereas D21 CD20<sup>-</sup> cells were largely Ki-67<sup>-</sup> and included CD138<sup>+</sup> cells, supporting their progressive maturation. We will include Ki67 analysis in CD138<sup>-</sup> and CD138<sup>+</sup> cells at D21 in the revision.

      We agree that the motivation and culture conditions for this experiment required a clearer explanation. The purpose of the D7 sort-and-reculture experiments was to identify the developmental window enriched for cells competent to generate PCs, thereby defining the stage at which IRF4 and PRDM1 should be perturbed. Sorted D7 CD20<sup>+</sup> actB cells and CD20<sup>-</sup>CD38<sup>+</sup>CD27<sup>+</sup> PBs were both placed into the same D7-D14 differentiation conditions, allowing a direct comparison of their PC-generating competence under identical culture conditions. We will clarify this design in the Results and Methods. We have not tested whether returning D7 CD20<sup>+</sup> cells to D0 conditions involving CD40L stimulation restores PC differentiation, and we will now acknowledge in the revision that the CD20<sup>+</sup> fraction may contain cells with distinct intrinsic differentiation potential, including cells differentiating into non-PC states.

      What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20<sup>+</sup> fraction would be enriched in these cells.

      We agree that the CD20<sup>+</sup> D7 cells may contain anergic or memory B cell precursors. However, this does not alter the interpretation that the CD20<sup>-</sup> (CD38<sup>+</sup>/CD27<sup>+</sup>) PBs contain a PC precursor population. Even if memory B cells are generated in the CD20<sup>+</sup> fraction by D7, based on the sorting experiments, their presence would have little-to-no effect on developing PC precursor populations.

      Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138<sup>+</sup>CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      We agree that the heterogeneous nature of CD20<sup>-</sup> cells complicates the interpretation. However, we would like to emphasize that all D7 CD20<sup>-</sup> cells are Ki67<sup>+</sup> whereas all D21 CD20<sup>-</sup> cells are nearly all Ki67<sup>-</sup>. Though we concede these could include recently proliferated PCs at D21, it is consistent with this population becoming quiescent. To better address this question in a future revised version, we will include direct measurements of Ki67 levels in CD138<sup>+</sup> and CD138<sup>-</sup> cells at D21.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      We agree that the current data do not by themselves prove that IRF4 acts earlier than BLIMP1 in developmental time. We will revise the text to state that the data are consistent with a temporally ordered model, rather than demonstrating strict sequential action. The latter interpretation is based on the distinct IRF4 KO stunted PC state observed by D9 scRNA-seq (Fig. 3C), together with the stronger early phenotypic effect of IRF4 loss (Fig. 3A). However, because both factors are mutually reinforcing and because perturbations were not performed across multiple time points, alternative explanations remain possible, including differences in editing efficiency, protein stability, and feedback-dependent TF decay. We will modify the text to better explain the rationale behind this interpretation while also acknowledging alternative interpretations that do not involve sequential IRF4-BLIMP1 functions (see response to Reviewer 1).

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c?

      We thank the reviewer for identifying these points of confusion. We will revise Fig. 3B to display outlier events and frequencies more clearly and revise the figure legend to clarify the donor origin of the displayed UMAPs and corresponding supplemental analyses. We will also clarify that Fig. 3B and Fig. 3C represent distinct readouts: flow cytometry measures IRF4 and BLIMP1 protein abundance, whereas scRNA-seq resolves transcriptional states after perturbation. Thus, the “stunted PC” state in IRF4 KO cells is defined at a transcriptional level as a population positioned between prePCs and PCs, not as a population retaining normal IRF4 or BLIMP1 protein expression. We further clarify that the apparent reduction in proliferative plasmablast-like cells at D9 likely reflects both the later timepoint relative to D7 and differences between transcriptional cell-cycle gene scores and Ki-67 protein persistence.

      What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells?

      The signature genes defining the scores in Fig. S3D are listed in Table S2. We will clarify the source of these genes in the figure legend of a future revised version.

      Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      We thank the reviewer for this suggestion to highlight IRF4, BLIMP1 and exemplar target genes in the various populations. We will update Fig. S3 to show transcript levels of IRF4, PRDM1 and an example of one of each of their target genes in unperturbed cells to demarcate their normal expression pattern.

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

      We agree that the data do not exclude the possibility that BLIMP1 is required before the perturbation window but is less continuously required once cells have progressed beyond a defined prePC stage. Our statement that BLIMP1 is not required to initiate the prePC-to-PC transition is based on the observation that PRDM1 KO cells did not accumulate in the prePC or stunted PC intermediate states observed after IRF4 loss. We will revise the text to make this interpretation more precise by stating that BLIMP1 appears less important than IRF4 for progression into a PC-like transcriptional state but is required for efficient consolidation of the mature PC program. We will also acknowledge that differences in editing efficiency, protein persistence, and timing of BLIMP1 action could contribute to the observed differences in phenotypes.

    1. eLife Assessment

      This important study shows that cave-adapted Astyanax mexicanus have shifted from avoiding alarm and decay odors to being attracted by them, alongside sex-specific responses to social odors. The evidence for the behavioral and heritability claims is convincing, supported by analyses of different cave populations and F2 hybrids, as well as starvation-induced plasticity experiments. Whole-brain pERK mapping offers suggestive mechanistic insight, although uncertainties in anatomical assignments and aspects of statistical and ethogram reporting temper the strength of the neurobiological conclusions. The work will be of interest to anyone working on the evolution of behavior.

    2. Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

    3. Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    4. Author response:

      The authors thank the reviewers for their thorough and fair assessment of our manuscript. We are currently working to edit the manuscript based on the critiques and guidance offered by the reviewers. This will consist of fixing grammar and typos, expanding the material and methods section to include more information on the behavioral assays, modifying graphs for clarity between visuals and interpretations, and correcting our mistakes in neuroanatomical labeling.

      Public Reviews:

      Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      We agree with the reviewer that our light testing data suggests complexity in the response to indirect white light within certain cave populations. In addition, an expanded sample size would likely provide clarity on whether individuals fall within two groups, behavior that looks like direct light or infra-red light, that could relate to genetic variation within cavefish populations. We are currently working to reassess the current data and determining the best course of action for continued studies related to light exposure.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      We agree with both reviewers that the methodological section on behavior needs to be expanded. We are currently editing our methods section to include more detail on water exchanges, odor preparation, timing, biological replicates, and binning.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      We agree that our annotation was incorrect or more accurately mislabeled in our write-up of the preprint and submitted manuscript. Therefore, we have now re-assessed the regions using the tissue cleared and light sheet collected zebrafish atlas, Adult Zebrafish Brain Atlas (AZBA; doi: 10.7554/eLife.69988). We are now editing the resubmission in the following manner: our initial labeling of the ventromedial thalamus will be changed to the lateral olfactory tract (nLOT) of the pallium and the preoptic region to the ventromedial thalamus (VM). We will provide a comparable z-slice of the AZBA segmentation file to illustrate the similarity in position. This would also support a known continuous circuit of olfactory integration, with information flowing from the lateral olfactory tract-to the piriform cortex-to the thalamus. We also agree that a more accurate assessment in Astyanax would require HCR in situ hybridization of markers for those specific brain regions or a neurocomputational brain atlas for adult Astyanax populations. Finally, we assert that this small dataset is preliminary at best and only provides regions of shared activity that could explain anything from perception related processes to relay of odor signaling unrelated to perception. Further work with larger sample sizes and additional populations are currently underway for a follow-up study on the neurobiological basis of olfactory processing and perception in adult cavefish.

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      We agree with the reviewer that the brain mapping section lacks a sufficient sample size and no control group for comparison (surface fish). However, we found the variation in pERK intensity (notably the putative nLOT) to be informative and decided to include the dataset in the manuscript. We are currently working to fill in these data gaps by sampling all populations and increasing the Pachon cavefish sample size. This will be a follow-up study as mentioned above in the last rebuttal paragraph.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      We agree with the reviewer that the hybrid results setup a promising follow up project to map these traits genetically. We are continuing to test odor perception in other hybrid populations and have started Quantitative Trait Locus mapping experiments.

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

      Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      We agree with the reviewer and direct their attention to the same critique by reviewer 1. We believe this is predominantly preliminary data that was included due to the conspicuous increase in pERK signal from the putative lateral olfactory tract (nLOT) for food and decay exposed cavefish. We are continuing to work on odor stimulated brain mapping and look forward to publishing a comprehensive dataset across wildtype and hybrid populations.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      We agree with the reviewer that the ethograms are challenging to read, especially due to our use of different colors for odor categories. We are currently preparing alternative graphs for displaying ethograms that reduce confusion and make following bout transitions for individual traces and bout probabilities for populations easier on the eyes.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    1. eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRISPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole-genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. However, the evidence supporting the proposed involvement of under-replicated region/replication-termination-zone resolution and TRAIP/URR-like pathways is currently incomplete and could be strengthened with an increased number of reciprocal daughter-cell pairs and by genetic or molecular perturbation, or alternatively, this can be addressed by changing the discussion.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest. The evidence for structural complexity associated with some induced SCEs is intriguing, but the mechanistic interpretation should either be tested directly or presented more cautiously.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9-induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs. Given that potential, the current manuscript would benefit greatly from any experiments characterizing this sub-population: are these cells in a particular cell cycle state, experiencing changes in gene expression, or do they have other unique biological properties?

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

      Weaknesses:

      The number of informative RDCPs is limited, and the mechanistic interpretation of the "WWC-or-WCC/deletion" signature is more suggestive than definitive. In particular, the manuscript invokes (even though only in the Discussion section) URR or replication-termination-zone resolution and discusses TRAIP-dependent CMG unloading, nuclease cleavage, and polymerase theta-mediated joining, but these pathway components are not directly tested herein. A more conservative conclusion that some Cas9-associated SCEs coincide with structural alterations is more appropriate, particularly in the Discussion and Conclusion. For example, the statement that this work provides "direct genetic evidence" for a URR-type mechanism is overstated unless supported by additional experiments or a more extensive analysis of alternative models. Similarly, while the authors explain the limitations of acute Cas9 disruption of LIG3, LIG4, XRCC1, and XRCC4, the manuscript should clarify what biological questions this experiment can and cannot answer.

    3. Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed large-scale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations. The language and logic in the paper can be improved, and some of the claims seem incorrect. For example, the abstract reads "A single Cas9 cut at a unique genomic locus led to strong local enrichment of SCE at the break site, reaching up to 41% in the same cell cycle and 17% in the subsequent division, indicating that DSB repair frequently engages non-local inter-sister repair." The evidence that only a single Cas9 cut was made is lacking (see my earlier comment); it is not clear how local enrichment or non-local inter-sister repair are defined.

    4. Reviewer #3 (Public review):

      Summary:

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are “genetically silent”. Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly “permissive” for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

      Strengths:

      This is an interesting paper that molecularly explores sister chromatid exchanges, which represent an important challenge in molecular biology since they are genetically silent.

      Weaknesses:

      A complexity of the current paper is that it heavily relies on a recently published paper (Chovanec et al 2026, NAR) describing the powerful but complex technique sci-L3-Strand-seq. Knowledge of this paper is a prerequisite to understanding the current manuscript because no reminder is provided. In addition, the current manuscript presents the use of the sci-L3-Strand-seq technique in the study of SCE after Cas9-induced DSBs, while a companion study is referred to several times for containing results about SCE in XRCC1 KO. At some point, one questions the relevance of splitting the use of sci-L3-Strand-seq in different papers instead of making a single integrated one.

    5. Author response:

      We thank the editors and reviewers for their thoughtful evaluation and constructive feedback. We are pleased that all three reviewers recognize the importance of mapping Cas9-induced sister chromatid exchanges (SCE) as a previously invisible repair outcome, and that the RDCP analysis is a notable feature of the study.

      We note that since our manuscript was posted, two companion studies in Science have provided direct biochemical evidence for the TRAIP-dependent pathway we discussed:

      (1) Fujisawa & Labib (Science, 2026; DOI: 10.1126/science.aeh2300) showed that TTF2 bridges CDK1-phosphorylated TRAIP to DNA Polymerase epsilon in the replisome, triggering mitotic CMG helicase disassembly, fork cleavage, and repair via SCE. Loss of this pathway reduced replication stress-induced SCE approximately two-fold in mouse ES cells.

      (2) Can et al. (Science, 2026; DOI: 10.1126/science.aeh1834) independently identified the same CDK1-TTF2-TRAIP axis in Xenopus egg extracts and validated it in HCT116 cells, showing that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions.

      We will incorporate these references in the revised discussion while still framing our RDCP observations as consistent with, rather than definitive proof of, this pathway.

      Below we briefly address the main points raised in the public reviews.

      Reviewer #1:

      We agree that the mechanistic interpretation of the RDCP signature should be presented more cautiously. We will reframe the URR/TRAIP discussion as a model, replacing language such as "direct genetic evidence" with "consistent with." We will add a summary table of RDCP data. We will also expand the description of rescued SCE calls. We will clarify what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) - we think that there is a notable difference at the bulk vs. at the single-cell level depending on the nature of the assay. Fig.1 fonts, labels, and pileup plot descriptions will be improved.

      We agree that the high-SCE subpopulation is particularly interesting and we cannot currently distinguish higher RNP uptake, a permissive cell-cycle state, altered expression, or stochastic variation. This may be better explored by future co-assays with sci-L3-Strand-seq.

      Reviewer #2:

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We will add a brief discussion on this limitation in extrapolating the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements.

      We agree that "a single Cas9 cut" should be revised to "Cas9 targeting of a single genomic locus" to accurately reflect the experimental design. We will clarify the possibility of multiple rounds of cutting at the same sites. We will also clarify "non-local" by modifying Fig.1 - we used this term specifically to include the possibility of inter-sister NHEJ.

      Reviewer #3:

      We will improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

    1. eLife Assessment

      This valuable study provides insights into the developmental regulation of the unusual form of holometabolous metamorphosis that occurs in the black soldier fly, in which a distinct prepupal stage is interposed between the final larval instar and pupation. This type of life history strategy is similar to what is seen in insects at the hemimetabolous-holometabolous boundary, but, given the phylogenetic position of dipterans, it is likely to be a derived trait. The combination of developmental characterization, gene expression profiling, and RNAi-mediated functional analyses provides solid evidence supporting the authors' conclusions and will be of interest to researchers studying insect development and the evolution of metamorphosis.

    2. Joint Public Review

      Summary:

      In this study, the authors investigated the developmental and molecular basis of the unusual metamorphic program of the black soldier fly, Hermetia illucens, which differs from the canonical holometabolous life cycle by inserting a distinct, non-feeding prepupal instar between the final larval stage and pupation. Most insects that undergo complete metamorphosis molt to the final instar and then develop into the prepupal stage without molting. H. illucens, however, undergoes a molt before entering a non-feeding prepupal stage. Thus, it is an unusual, novel developmental strategy, and its regulation has remained a mystery. Through an integrated approach combining detailed morphological characterization, developmental gene expression profiling, and RNAi-mediated functional analyses of the core components of the Metamorphic Gene Network (MGN), the authors examine the developmental identity of this prepupal stage and how the temporal deployment of conserved metamorphic regulators has been reorganized to accommodate this atypical developmental program. In particular, they show that the prepupal stage expresses a distinct combination of the key genes known to regulate life history transitions, including unusually high levels of Br-C expression.

      Strengths:

      The study represents a valuable contribution to insect developmental biology. A major strength is the comprehensive characterization of postembryonic development, which establishes a robust developmental framework for H. illucens. This is complemented by detailed expression profiling and RNAi-based functional analyses of the Metamorphic Gene Network (MGN), comprising the temporal specifier factors, Kr-h1, chinmo, Br-C, and E93. The results show that these conserved regulators are deployed in a modified temporal sequence that accommodates the distinctive prepupal stage while largely preserving their canonical developmental functions. Together, the morphological, molecular, and functional data support the conclusion that the prepupal stage of H. illucens is a distinct developmental transition associated with a characteristic configuration of the metamorphic gene network. The results are supported by solid methodology and approaches and will serve as valuable resources for future investigations into insect development, the evolution of metamorphosis, and the diversification of insect life-history strategies.

      Weaknesses:

      While the study successfully establishes the developmental identity of the prepupal stage and its association with a modified temporal deployment of the MGN, some aspects of the proposed regulatory model are less directly supported by the experimental evidence.

      (1) Several regulatory interactions within the MGN remain inferential rather than experimentally demonstrated in H. illucens. In particular, the proposed relationship between juvenile hormone (JH), Kr-h1, and chinmo is based primarily on expression dynamics and RNAi-induced transcriptional changes. Although these observations are consistent with the proposed model, they do not directly demonstrate that JH induces chinmo expression or establish the regulatory relationship between Kr-h1 and chinmo in this species. As a result, the corresponding regulatory interactions presented in the final model should be regarded as plausible hypotheses rather than experimentally validated mechanisms.

      (2) A second limitation concerns the developmental role assigned to Br-C and E93 during the larval-to-prepupal transition. The authors conclude that sustained Br-C expression is a defining molecular feature of the prepupal stage and discuss its potential role in prepupal specification. However, the functional analyses of both Br-C and E93 were initiated only after larvae had already entered the prepupal stage. Consequently, while the RNAi experiments convincingly demonstrate essential roles for Br-C during the prepupal-to-pupal transition and for E93 during adult differentiation, they do not directly address whether either factor is required to trigger the formation of the prepupal stage itself. Therefore, the molecular mechanisms governing the initiation of this distinctive developmental transition remain unresolved. In particular, the proposed lack of repression of E93 by Br-C is only weakly supported, yet may be an essential feature of the prepupal stage of Hermetia illucens.

      (3) Although knockdowns of Kr-h1 and chinmo knockdowns look superficially similar, it would be good to confirm this with higher-magnification views of the cuticles for all three treatments (control, Kr-h1 RNAi, and chinmo RNAi). In other species, Kr-h1 knockdown leads to premature adult cuticle development, whereas chinmo knockdown typically leads to premature appearance of pupal characteristics. Similarly, in Fig. 4A and 4D, higher-magnification images of the cuticle would be helpful.

      (4) (Relating to Line 336 and Figure 7): "This low but persistent prepupal Kr-h1 expression, together with modest chinmo expression from PPD0 to PPD8, may be correlated to a JH-dependent antimetamorphic effect that maintains the prepupal stage." However, we are not aware of a function of JH in extending the prepupal stage. In addition, in most insects, the prepupal stage expresses high Kr-h1 expression; this peak likely prevents the animal from turning into an adult instead of the pupa. We presume the same holds true for H. illucens (although the lower expression of Kr-h1 during that stage is curious). As a result, we suggest that Fig. 7D be revised as it may be difficult to distinguish between pupal formation and prepupal maintenance given the experimental set-up. Fig. 7E may also need to be modified since the development of the pupa may require Kr-h1. It is worth noting that at the prepupal stage, JH and Br-C are co-expressed in many insects. If the authors think that Kr-h1 expression needs to be low at this time, this would imply a novel interaction between Kr-h1 and Br-C, and should be discussed.

    3. Author response:

      We thank the editors and reviewers for their careful evaluation and constructive suggestions. During the review process, we identified and corrected several presentation, terminology, citation, and figure-legend errors, and these corrections have been incorporated into the current version of preprint. Following the reviewers’ comment, we have also revised the title of the manuscript. We are now preparing a substantive revision that will distinguish more clearly between experimentally supported conclusions and hypothetical regulatory relationships. We plan to examine gene expression at an earlier time point after Br-c knockdown, further investigate the relationships among Kr-h1, chinmo, and E93 using additional RNAi experiments, characterize the cuticular phenotypes of precocious prepupae at higher magnification, and determine whether severe E93-knockdown individuals exhibit evidence of a repeated pupal developmental program. After completing these experiments, we will cautiously revise the proposed regulatory model and moderate conclusions that are not directly supported by the current evidence.

  2. Aug 2026
    1. eLife Assessment

      What can a neural network trained to imitate animal behavior tell us about biology? This valuable work uses deep reinforcement learning to train an artificial neural network to transform the dynamics of a recurrent neural network based on the C. elegans connectome into an adult Drosophila walking program in a physical model of the fly body, demonstrating that achieving plausible output dynamics does not in and of itself imply biologically meaningful simulation. Evidence for this basic claim is solid, but more extensive analyses, better methodological description, and a discussion of deeper challenges in the undertaking of biological brain modeling would strengthen the study. This result demands the attention of the practitioners of the growing field of connectome simulation for the purpose of gaining mechanistic understanding of nervous system function.

    2. Reviewer #1 (Public review):

      Summary

      The authors build a "digital sphinx" by stitching together two neural network models: (i) a recurrent network with fixed parameters derived from the C. elegans connectome and imputed physiological (e.g. neural input/output) functions, and (ii) a feedforward encoder-decoder model with learnable parameters intended to represent a central brain - to - motor interface, then harnessing the combined model to a Drosophila biomechanical model situated in a physics simulator, and finally using deep reinforcement learning (DRL) training to optimize the parameters of the encoder-decoder model to reproduce a set of spatiotemporal patterns of jointed limb activations that together produce the overall organismal behavior of walking, within the physics simulator.

      The primary intent of this paper is to dispel the recent grandiose claims made in the mainstream press by a private company, Eon Systems, to have achieved a major advance in biologically based brain simulation of the production of a set of ethologically relevant motor behaviors by the fly. Representatives of the company referred to this modeling and training process euphemistically and deceptively as "brain uploading". The authors proceed with a reduction-to-triviality exercise by constructing their own high-parameter dynamical brain-plus-body model situated in a physical simulation that produces, after training by reinforcement learning, satisfying ethological behavioral imitation in the same vein as the private company claim, but based on a clearly absurd and biologically unrealistic set of model assumptions.

      Secondarily, the paper provides two overall admonitions that they assert their computational demonstration illustrates: that training high parameter network models to imitate behavior, even if they possess some biological detail, will deliver little or no biological insight, and that models of behavioral generation must be built from detailed biological data and, crucially, developed in a hypothesis generation/falsification loop with experimental validation, in order to be scientifically useful.

      Appraisal

      The authors are well justified in challenging the non-rigorous claims of "uploading" or even the delivery of a neurobehavioral simulation with potential scientific utility, in unison with the vocal criticisms of many other researchers in the fields of AI and neuroscience, and it is an important message to deliver to the world. However, the authors' own modeling counter-exercise, while clever and vivid in imagery, suffers from its own lack of rigor, both in disclosure of implementation and in scientific case-making. Some sacrifice of clarity and thoroughness in the interest of brevity is inevitable under the brief format of this manuscript; however, we suggest that crucial additions and modifications should be made to avoid falling into a similar trap of non-rigorous sensationalism.

      Because the private company claims were not accompanied by a scientific paper, preprint, code repository, or much methodological disclosure of any kind, the authors have the particular challenge of building a refutation case against an undefined target. As a consequence, the authors chose their own task, model structure, and training paradigm.

      The authors argue that brain models need to be built from biological data to be useful for yielding biological insight. We agree with the overall principle; however, in practice, this procedure is fraught with epistemological difficulty. Biological modeling suffers from a unique challenge within the larger endeavor of scientific/physical modeling, which is that it is generally unclear as to precisely what biological quantities should be measured and at what level of detail they should be measured. Additionally, biological data will by necessity be incomplete and noisy, and thus decisions of coarse-graining must be made at the outset of large-scale data collection projects, and some, possibly a substantial, level of data imputation will have to be performed in order to build testable models in our lifetimes. Despite the astonishing success of scaling (in both parameter count and corpus size) in engineered neural networks for certain human-like tasks, it is not at all clear that simply adding more detail to biological models will produce deeper scientific insight, or whether cataloging parameters from snapshot data will yield functional simulations. The failed Blue Brain mega-project should provide a lesson, as well as Marder's longstanding work on parameter variation in neural systems. The coupled, pernicious questions of choosing measurement detail and modeling detail represent a deep, unsolved challenge area for the field, and this context should be raised in the text.

      The message about overinterpreting models trained with deep reinforcement learning, while valid and important, should be broadened to be a message about overinterpreting trained high-parameter models in general, in their ability to fit data or reproduce simple behavior. Other parameter optimization/learning procedures for building underdetermined and/or high-parameter models risk the same misinterpretation. The prescription of building models in conjunction with experimental prediction and validation is an important point.

      The authors leave out an additional important and underappreciated challenge of brain-model-building, which is that imitating a time segment of behavior is a computational task of unspecified, and possibly low complexity. Successful recapitulation of behavioral time series may simply not be considered cognitively interesting, even if the model is built entirely on biological data. While quantifying task complexity is another open area of computational and neuroscientific research, the authors should, at a minimum, describe their particular task data in explicit mathematical terms and preferentially provide some complexity analysis. In the absence of task complexity analysis, at a minimum, computational controls should be applied to demonstrate the necessity of whatever structure or data is being asserted in the model. This epistemological practice is glaringly absent in much, if not most, of the neurobehavioral modeling literature. This paper would be a good opportunity to set an example of rigor.

      Finally, the authors' description of prior work in the field of whole-organism neurobiological simulation feels incomplete and skewed toward work in Drosophila versus other model organisms. An internet search reveals many published efforts to build neurobehavioral models at varying levels of detail in C. elegans, of which only two are referenced.

      We do feel this work constitutes an illustrative scientific exercise and important counterpoint to the sensationalism building around efforts in neurobiological simulation. It should inspire further work in defining a rigorous and scientifically productive epistemological framework for these kinds of brain modeling efforts.

      Further Comments

      (1) The authors oversell the completeness and quality of connectome datasets and what they lack.

      Language such as "complete wiring diagrams," "nearly comprehensive connectomes" neglects the well-appreciated gaps in biological data that most practitioners believe necessary for useful, detailed models to be built. There is a brief mention that biological parameters "remain unknown" and that interfaces are "incompletely characterized", but beyond that, the authors do not explain which parameters are missing, why these parameters might matter, and what still needs to be addressed in order to make any plausible whole-brain emulation claims. This may also inadvertently bolster the sensationalist claims that the manuscript is trying to deflate by giving the impression that neurobiological and physiological data collection is a near-complete exercise.

      (2) Prior work in C. elegans neurobehavioral modeling should be more acknowledged, if nothing else, for why it has been largely unsatisfying.

      C. elegans is rarely discussed, while Drosophila is primarily focused on. The status of C. elegans connectomics, physiological mapping, biomechanics, and neurobehavioral modeling is worth more treatment.

      (3) Critiques of Eon Systems announcements also, by and large, apply to more detailed and disclosed efforts in neurobehavioral modeling using RL for parameter imputation, and this should be recognized.

      By way of reference to a tweet in the first paragraph, the authors are responding to a recent claim made by a startup that they have fully "uploaded" a fly brain, a significant advance vis-à-vis prior work in neurobehavioral modeling in Drosophila, such as references [3 and 9], which are mentioned as background in the paper but left out of the methodological critique. But one of the central warnings of the paper is around the challenge of interpretability when using reinforcement learning to optimize model parameters. The authors also should acknowledge that the use of RL has been justified by building neurobehavioral model builders as a proxy for the learning and tuning processes thought to occur during animal development.

      (4) Substantiate the reservoir computing explanatory claim with appropriate computational controls.

      The reservoir computing idea is the only piece of hypothesizing a necessary function for the central brain component model in the paper. This claim could be substantiated with some basic computational controls rather than just hypothesized. We suggest the following possibilities as additions to the model: (a) replace the connectome with an RRNN, (b) shuffle the connectome, or (c) use other simple dynamical systems in place of the worm brain model.

      Specific Manuscript Comments

      (1) Abstract

      "New connectome datasets and musculoskeletal models now enable integrated, closed-loop simulations of the neural and biomechanical systems of the fruit fly Drosophila, an ideal model organism to investigate embodied intelligence."<br /> This sentence could mislead non-specialists into thinking all current simulations are novel because the connectome datasets are new. In fact, FlyWire (2024), NeuroMechFly (2022), and other connectomes have already been available for some years now. We believe that this sentence is a chance to make the opposite point that these resources have existed for a while, and that many simulations have been built before.

      "However, many biological parameters of the nervous system and the body, as well as how they interface, remain unknown."<br /> Some examples of specific parameter/physiological data types that are missing and thought to be critical, such as neuronal input/output functions, are warranted. See below for a comment on the confusing construct of "interface" as a distinct entity from the neural network.

      (2) Introduction

      "Among animals that walk, the integration of brain wiring and body models is perhaps closest to fruition in Drosophila, due to the recent completion of multiple complete wiring diagrams (known as connectomes) of the fly nervous system." ...and... "The fly is the only animal with legs for which nearly comprehensive connectomes of its brain and nerve cord exist."<br /> The walking qualifier allows the authors to skirt around the substantial and decades-long work on connectomes in C. elegans, which crawls and does not walk. Yet sinusoidal crawling is a multidimensional, adaptive behavior, so it seems this exclusion was for narrative convenience rather than contextual accuracy.

      "Despite this progress, closed-loop integration of biomechanical and neural models remains far from straightforward."<br /> Work (and shortcomings) in C. elegans neurobehavioral modeling should also be stated here alongside the fly.

      "Where interfaces between brains and body models are missing or only partially characterized, one approach is to train an artificial neural network (ANN) to approximate these interfaces with deep reinforcement learning (DRL)."<br /> The choice of "interface" as a distinct, well-defined neurobiological entity is somewhat confusing and may mislead non-practitioner readers. If neuronal and muscular (and their interactions) physiology are incorporated into a neurobehavioral model, then in principle there is nothing left to call an "interface". It would be clearer to explain that prior neurobehavioral models have often inserted a trainable multilayer feedforward network between sensory and central brain and between the central brain and motor effectors in order to have a substrate for learning, and that this insertion may render the entire biological modeling exercise scientifically pointless, or at a minimum require a set of computational controls.

      "In building virtual animal models, a motor policy is commonly learned by DRL so that the integrated, closed-loop virtual body successfully mimics the detailed kinematics of real animal behavior."<br /> The authors could define "motor policy" in simple terms and give a brief example.

      "Additional realism is added when the motor policy network is constrained by a connectome dataset. However, many biophysical parameters for individual neurons and synapses remain un-measured."<br /> "motor policy network" is confusing; this is referring to the entire network model here, presumably.

      (3) Methods

      "We used the adult hermaphrodite C. elegans nematode connectome dataset [15, 16, 5], including the identities of its 302 neurons and their synapses (Fig. 1A)."<br /> We believe the authors should specify the dataset type, which is a structural, unsigned connectome lacking grounding in physiological function.

      "The policy network was trained in closed loop using PPO as implemented by MIMIC-MJX"<br /> The authors should define "PPO" and "MIMIC-MJX" in simple terms and explain why they were used.

      (4) Discussion

      "Its role in the movement policy could be fulfilled equally well by a randomly connected RNN, akin to reservoir computing [20], since all the learning happens in the black-box ANN motor decoder."<br /> See above - this computational exercise should actually be performed.

      "Looking further ahead, swapping brain and body models of related species may one day yield real insights into how their brains and bodies diverged through evolution. However, far more model development and experimental validation is needed before we can learn anything from such a digital sphinx."<br /> These two sentences about future possible cross-species chimeras feel superfluous and unsubstantiated, and weaken the main argument of the paper about whole-brain emulation.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use DRL to train a C. elegans connectome-based ANN to control stepping in a D. melanogaster body model. The resulting system can walk. This shows that one needs further constraints to derive biologically meaningful results from this approach.

      Strengths:

      The authors perform a very simple experiment with a clear outcome. The interpretation (or lack of interpretation) is a striking cautionary tale.

      Weaknesses:

      There is little analysis of precisely how robust this result is to parameter variation and network wiring. The worm also undulates in an oscillatory fashion. Thus, it is possible that the network is tapping into biologically meaningful motifs to generate oscillations for walking. As well, it would be useful to examine which heuristics one can use to determine whether modeling efforts are sufficiently constrained (i.e., how much biological data will be necessary to start obtaining fruitful, interpretable outcomes from DRL task optimization). For example, their "solution" using the worm connectome is not sparse (i.e., it uses many neurons). Perhaps a signature of a biologically-meaningful, interpretable result is one that is sparse?

    4. Reviewer #3 (Public review):

      Summary:

      The authors construct a computational chimera by attaching a C. elegans connectome to a Drosophila body biomechanical model and use deep reinforcement learning to link neural activity to motor output. The model is able to produce walking, but is considered a priori to be scientifically meaningless, and the work is treated as a cautionary tale in complex interpretation layers unconstrained by experiment or data.

      Strengths:

      In a period of increasing excitement about linking AI and neuroscience, I respect very much that the authors work through a nontrivial example of nonsense results, rather than just making a theoretical case. It offers a clear and memorable existence proof that matching outputs of complex trained networks does not mean the internal dynamics are themselves emulated.

      Weaknesses:

      While I understand that the work was a rapidly produced comment on science-by-press-release, the message seems too important to be treated in quite as pithy a manner as it is. In particular, because the computational experiment is so memorable, it is worth getting the message right to avoid a set of readers who take from it that they should dismiss this category of neuroAI wholesale (which the authors absolutely do not imply!).

      One part of me reads this work and thinks that by intentionally wiring up the sensory feedback in a particularly nonsense way, the authors have just made a bad model, and sometimes bad models can still generate sensible outputs, especially when expressive models are optimized to fit those sensible outputs. But I think this work is trying to say something more specific than this, and I would like it to be a bit clearer about that. The authors do a fairly good job of sharing a view about what should have been done instead, but this message would benefit from having some more concrete suggestions to avoid a simplistic interpretation. A few thoughts:

      (1) It's not entirely obvious to me that the model is "scientifically meaningless." As the authors know extremely well, Drosophila walking is thought to be driven by simple central pattern generators coupled to leg-specific implementations. The C. elegans neural circuit is clearly capable of producing rhythmic activity as well. A version of the model they ran could have identified biologically valid rhythmic activity in the C elegans circuit and mapped it via the DRL to the right locomotor behavior in the fly. While this would not be a good emulation of the fly, it's not a concept devoid of scientific meaning. Similarly, if the ANN is converting a rhythmic signal to coordinated walking, it's not obvious to me that there aren't useful principles to identify in how it achieves this - it's basically the equivalent of that post-CPG circuitry, no?

      (2) Similarly, is this outcome going to be relatively specific to rhythmic behaviors? I suspect that it would be harder to push the C. elegans connectome to produce some behaviors than others - for example, adding in visual navigation and other motor patterns, or a ring attractor. Rhythmic circuits arise in many places, and both biology and dynamical systems tell us they can come from numerous configurations of elements and interactions.

      (3) Aside from the nonsense formulation of the problem, I would have liked to know more about what the authors should have done to know their model was useless. Put another way, if the authors hadn't known that their model was bad from the beginning (e.g., if they had stuck a fly brain in the middle of it, gotten the sensory feedback right), would there have been some way to figure out if it was meaningful or meaningless based on the results of the trained model itself?

    1. eLife Assessment

      This study provides valuable insights into the role of the bile acid receptor TGR5 in regulating bone marrow adipose tissue. This revised manuscript has additional characterization of TGR5 expression in hematopoietic and stromal populations, and phenotypic differences observed between sexes, during aging, and under high-fat diet conditions. Although the scope of the study has expanded, the mechanism of how TGR5 in the microenvironment regulates hematopoietic recovery is still incomplete.

    2. Reviewer #1 (Public review):

      This study by Alonso-Calleja and colleagues aimed to determine whether TGR5 regulates hematopoiesis and the bone marrow microenvironment under steady-state conditions and following transplantation. The revised manuscript substantially improves upon the original submission by providing additional characterization of TGR5 expression in hematopoietic and stromal populations, incorporating analyses in female mice, and expanding the investigation of bone marrow adipose tissue under aging and high-fat diet conditions. These additions more convincingly establish TGR5 as a regulator of bone marrow adipose tissue and stromal composition.

      Major strengths of the study include the comprehensive characterization of the bone marrow adipose tissue phenotype across multiple experimental settings and the demonstration that TGR5 deficiency consistently alters the stromal compartment. The strongest and most convincing aspect of the work is the identification of TGR5 as a regulator of bone marrow adipose tissue and the bone marrow microenvironment. These findings provide useful insights into how metabolic signaling pathways influence the hematopoietic niche.

      However, the evidence supporting a direct role for TGR5 in hematopoietic recovery following transplantation remains limited. Although reciprocal transplantation experiments and peripheral blood recovery analyses strengthen the manuscript, the conclusions regarding hematopoietic regeneration continue to rely largely on correlative observations. The study does not directly demonstrate that expansion of adipocyte progenitors is responsible for the enhanced recovery phenotype, nor does it establish improved regeneration of hematopoietic stem or progenitor cells within the bone marrow. Overall, the revised work addresses many of the concerns raised in the original review and provides useful new insights into the regulation of the bone marrow microenvironment by TGR5. Nevertheless, the conclusions regarding hematopoietic recovery should remain appropriately tempered, as the mechanistic basis linking the stromal phenotype to enhanced regeneration has not been directly demonstrated.

    3. Reviewer #2 (Public review):

      Summary:

      The authors showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. Notably, TGR5 knockout significantly decreases BMAT, increase the APC population and accelerate the recovery upon bone marrow transplantation.

      Strengths:

      The role of TGR5 is interesting.

      Weaknesses:

      Additional mechanistic studies would further strengthen the work and provide deeper insight into how TGR5 regulates the bone marrow microenvironment.

    4. Author response:

      The following is the authors’ response to the original reviews

      eLife assessment

      This study investigates the role of the bile acid receptor TGR5 in adult hematopoiesis of the mouse model. The findings are potentially useful because the loss of TGR5 leads to dysregulation of bone marrow adipose tissue (BMAT) that has emerging regulatory functions. However, the study is still incomplete because the mechanism of TGR5 is not clear, the stromal cells expressing TGR5 have not been well defined, and there is not strong evidence for the role of TGR5 in recovery from transplant stress.

      We thank the eLife editorial team for handling our manuscript. In our revised version, we took into consideration the suggestions of the reviewers, which we believe have significantly improved the quality of the current study. In summary, our new data provide further evidence that TGR5 is expressed in both hematopoietic cell lineage and stromal cells of the bone marrow (BM), including subpopulation analyses for both. While steady-state hematopoiesis remains intact in TGR5 knockout mice, we demonstrate that loss of TGR5 significantly impacts progenitor reconstitution under stress conditions, alters the BM adipose tissue (BMAT) homeostasis in a sex-specific manner and influences the balance of stromal cell differentiation. In particular, TGR5 deficiency resulted in reduced regulated BMAT and an accumulation of adipocyte progenitors, correlating with improved hematopoietic recovery following BM transplantation. The BMAT decrease is further observed in physiologically and pathophysiologically relevant contexts such as aging and obesity, where TGR5 deficiency is associated with a decrease in the myeloid bias. Collectively, our findings support a previously unrecognized role for TGR5 in maintaining BM niche integrity and highlight its potential as a modulator of hematopoietic support, although the precise molecular mechanisms still need to be elucidated.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      Alonso-Calleja and colleagues explore the role of TGR5 in adult hematopoiesis at both steady state and post-transplantation. The authors utilize two different mouse models including a TGR5-GFP reporter mouse to analyze the expression of TGR5 in various hematopoietic cell subsets. Using germline Tgr5-/- mice it's reported that loss of Tgr5 has no significant impact on steady-state hematopoiesis, with a small decrease in trabecular bone fraction, associated with a reduction in proximal tibia adipose tissue, and an increase in marrow phenotypic adipocytic precursors. The authors further explored the role of stroma TGR5 expression in the hematopoietic recovery upon bone marrow transplantation of wild-type cells, although the studies supporting this claim are weak. Overall, while most of the hematopoietic phenotypes have negative results or small effects, the role of TGR5 in adipose tissue regulation is interesting to the field.

      Strengths:

      This is the first time the role of TGR5 has been examined in the bone marrow.

      This paper supports further exploration of the role of bile acids in bone marrow transplantation and possible therapeutic strategies.

      We thank the reviewer for pinpointing the strengths of our study.

      Weaknesses:

      (1) The authors fail to describe whether niche stroma cells or adipocyte progenitor cells (APCs) express TGR5.

      Using the TGR5:GFP reporter model, we identified GFP<sup>+</sup> cells in the stroma-enriched CD45-Ter119-CD31- population that contains the adipogenic progenitor cells (APC).

      These data, along with the corresponding gating strategy are outline in Figure 6A and B.

      We found the subpopulation analyses within the stroma gate challenging at the individual mouse level given the limited cell numbers for these progenitor populations. We attempted to circumvent this by concatenating the individual files per genotype (WT and TGR5:GFP) for two independent experiments. We were surprised to find that there is little to no GFP expression in the APC population. Nevertheless, we found that the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup>, multi-potent stem cell-like population (Ambrosi et al., 2017), reproducibly showed a TGR5:GFP positivity comparable to that of Ly6<sup>lo</sup> monocytes. We believe this would be compatible with our results showing an increase in CFU-F in Tgr5<sup>-/-</sup> mice, but we remain cautious about its interpretation. Results are shown in Author response image 1 for the concatenated flow plots and numerically in Error! Reference source not found. considering all individual mice analyzed (i.e., non pooled) for completeness. We kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 1.

      Flow cytometry gating strategy used to identify stroma subpopulations in the stroma enriched CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup> gate and their GFP signal for BM cells in TGR5:GFP mice. Subpopulations were immunophenotypically defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (adipogenic progenitor cells (APC)) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations) (Ambrosi et al., 2017). Results are shown as the concatenated data of all mice in each phenotype for two independent experiments, as indicated.

      Author response table 1.

      Frequencies in the total live cell gate for adipogenic progenitor cells (APC) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations, MPSC-like) (gating as in Ambrosi et al., 2017) in the experiments presented in Author Response Figure 1, expressed as average +/- 95% confidence interval for the two independent experiments. The number of mice per genotype and per experiment is indicated in parenthesis. **Paired, two-tailed Student’s t-test statistical analysis for the combined experiments (n=8 per group) indicates p < 0.01 for GFP expressing cells in APC versus MPSC-like populations.

      (2) Although the authors note a significant reduction in bone marrow adipose tissue in Tgr5-/- mice, they do not address whether this is white or brown adipose tissue especially since BA-TGR5 signaling has been shown to play a role in beiging.

      The nature of BMAT and how it relates to brown, white or brown/beige adipose tissue has been a persisting question in the field. Our understanding is that BMAT is currently considered as a distinct adipose depot that is neither white nor brown/beige (Sebo et al., 2019; Suchacki et al., 2020). BMAT does not express UCP1 to an appreciable extent, with reports showing that detectable expression possibly stems from contamination by tissues surrounding bone (Craft et al., 2019). Beyond this consideration, as the regulated BMAT in Tgr5<sup>-/-</sup> mice is almost absent, determination of the brown/beige vs white nature of the little regulated BMAT that remains would be technically extremely challenging.

      (3) In Figure 1, the authors explore different progenitor subsets but stop short of describing whether TGR5 is expressed in hematopoietic stem cells (HSCs).

      We have added these data to the manuscript as part of Figure 1C.

      We have further expanded our data in Figure 1D and Figure 1–figure supplement 1A with the expression of TGR5:GFP in megakaryocyte progenitors (Lin<sup>-</sup>cKit<sup>+</sup>Sca1CD150<sup>+</sup>CD41<sup>+</sup>) as shown in Author response image 2.

      Author response image 2.

      A, representative flow cytometry gating strategy used to identify megakaryocyte progenitors (MkProg) and GFP positivity in TGR5:GFP mice and their wild-type controls. B, frequencies of GFP<sup>+</sup> cells in MkProg population in the BM of 8-12-week-old male TGR5:GFP mice and their controls (n=3 for wild-type control mice, n=4 for TGR5:GFP mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Two-tailed Student’s t-test (B) was used for statistical analysis. p-values (exact value) are indicated.

      Finally, we have completed the characterization of BM progenitor populations to include the erythroid lineage, thus covering the main hematopoietic populations. We have added these data to Figure 1E and Figure 1–figure supplement 1B.

      (4) Are there more CD45+ cells in the BM because hematopoietic cells are proliferating more due to a direct effect of the loss of Tgr5 or is it because there is just more space due to less trabecular bone?

      We observe an average 20% increase in CD45<sup>+</sup> cell counts in baseline Tgr5<sup>-/-</sup> mice. The absolute volume of bone and BMAT lost in these animals does not account for 20% of the medullary cavity volume, so we speculate that the increase in CD45<sup>+</sup> counts is not solely due to increased available volume for these cells.

      (5) In Figure 4 no absolute cell counts are provided to support the increase in immunophenotypic APCs (CD45-Ter119-CD31-Sca1+CD24-) in the stroma of Tgr5-/- mice. Accordingly, the absolute number of total stromal cells and other stroma niche cells such as MSCs, ECs are missing.

      These data are now included in the manuscript and in Author response image 3. Although we detect an increase in the relative frequency of APCs (Figure 7A), on a per-leg basis we did not detect an increase in APC numbers per leg (Author response image 3). We did however observe a decrease in total CD45<sup>-</sup>Ter119sup>-</sup>CD31sup>-</sup> stromal-enriched cells in Tgr5<sup>-/-</sup> mice. Nevertheless, given that only 2–5% of stromal cells are recovered in single-cell suspensions compared with native tissue (Coutu et al., 2017; Gomariz et al., 2018), we consider results for absolute stromal cell numbers prone to over interpretation. Our conclusion, therefore, remains one of relative enrichment of immunophenotypic APCs, supported by in vitro findings of increased adipogenesis and CFU-F formation after plating equal cell numbers.

      Endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) were also quantified, with no differences observed between groups.

      Author response image 3.

      Absolute number of adipocyte progenitor cells (APC), stroma cells (CD45-Ter119CD31<sup>-</sup>) and endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) (n=5 for both Tgr5<sup>+/+</sup> and Tgr5<sup>-/-</sup> mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Unpaired, two-tailed Student’s t-test was used for statistical analysis. p values (exact values) are shown.

      (6) There are issues with the reciprocal transplantation design in Fig 4. Why did the authors choose such a low dose (250 000) of BM cells to transplant? If the effect is true and relevant, the early recovery would be observed independently of the setup and a more robust engraftment dataset would be observed without having lethality post-transplant. On the same note, it's surprising that the authors report ~70% lethality post-transplant from wild-type control mice (Fig 4E), according to the literature 200 000 BM cells should ensure the survival of the recipient post-TBI. Overall, the results even in such a stringent setup still show minimal differences and the study lacks further in-depth analyses to support the main claim.

      We thank the reviewer for this comment. On the one hand, we respectfully disagree on the relevance of the effect size, as Tgr5<sup>-/-</sup> mice recover from low platelet and leukocyte counts significantly faster than wild-type controls. Tgr5<sup>-/-</sup> recipients recovered their platelet levels faster than the Tgr5<sup>+/+</sup> recipients, showing values consistently above 200.000/µL one week sooner; platelet levels below this threshold are considered a risk factor for bleeding events (Morowski et al., 2013; Vannini et al., 2019). In addition, on day 15 post-irradiation, Tgr5<sup>-/-</sup> recipients had higher neutrophil levels, at over the 500 cells/µL-threshold value for infection risk (Morowski et al., 2013; Vannini et al., 2019), but this difference was no longer statistically significant upon stringent multiple-comparison correction (uncorrected p-value 0.028, multiplecomparison corrected p-value 0.108). Underlining the relevance, in a clinical setting, G-CSF is routinely administered to patients daily to enhance myeloid recovery even if the acceleration of recovery is by 1-2 days (Trivedi et al., 2009).

      From the perspective of mortality, we agree that it is higher than expected and constitutes a limitation of our work. Note that we discovered a mistake in data plotting during the preparation of the revised version of the manuscript which renders the mortality curve no longer statistically significant, with a p value that changed from 0.0254 to 0.0676. We have duly changed our interpretation in the text to that of a trend towards higher survival.

      Regarding the mortality in our experiments being higher than expected, we have unfortunately suffered from cases of “swollen muzzles syndrome” in our facilities that have greatly hampered our ability to perform myeloablation experiments (Garrett et al., 2018), as even sublethal doses have resulted in the appearance of side effects that were reasons for euthanasia under Swiss legislation. For example, a strong reduction in mobility requires immediate euthanasia. All experiments were performed blinded to genotype allocation, so we can reasonably exclude experimenter bias. Finally, it could be argued that mice with more marked symptomatology leading to euthanasia are more likely to have hematopoietic deficits, which in our case was mostly seen for WT animals. We have therefore chosen to report mortality alongside the longitudinal assessment of peripheral blood counts to ensure clarity of the coherent effect and full transparency. We strongly believe that this disclosure, when examined as a limitation in the discussion section, aligns with the open science efforts advocated by eLife. Unfortunately, it is beyond the scope of this manuscript and the authors' timeline to rederive the Tgr5<sup>-/-</sup> colony in our new facility and perform new bone marrow transplantation studies with titrating doses of BM.

      Lastly, the choice of 250,000 BM cells serves as the control dose per recommended standards for competitive repopulation assays (Purton and Scadden, 2007). Quoting from this methodological reference paper: “In our experiments, each recipient mouse receives cell doses … together with 2x10<sup>5</sup> competing congenic bone marrow. We have found these cell doses sufficient to detect both reductions (Purton et al., 2006) and increases (Janzen et al., 2006; Walkley et al., 2005) in HSC numbers in different mutant mice… Furthermore, caution should be used when designing competitive repopulation assays, as it has been shown that the reliability of this assay is critically dependent on the numbers of HSCs present in the populations being assessed: when too few or too many HSCs (recipients of <1 1x10<sup>5</sup> or >2x10<sup>7</sup> bone marrow cells each from donor and competing sources) are present, the data may not be meaningful (Harrison et al., 1993).”

      Moreover, our lethal radiation rescue experiments with 2.5x105 cells are designed to deliver a minimal hematopoietic source known to rescue the great majority of animals in standard conditions, and thus to best mimic the situations when enhancement of hematopoietic recovery would be clinically meaningful (as used by our group members in (Naveiras et al., 2009; Tratwal et al., 2020; Vannini et al., 2019)). In our hands and in the absence of “swollen muzzles syndrome”, this approach leads to 80-95% overall survival, which mirrors the clinical setting of autologous transplantation that our transplants are meant to mimic. In clinical practice strategies, enhancing hematopoietic recovery would be most useful to improve morbimortality in either alternative hematopoietic progenitor allotransplants, known associated with delayed engraftment, or in the case of febrile neutropenia associated to bacteriemia, which affects 11-15% of both hematopoietic autologous or allogeneic transplant patients (Gil et al., 2013; Gooley et al., 2010). In summary, septic neutropenia was unfortunately modelled by the rescue experiments presented in Figure 7 C-J with swollen muzzles syndrome concomitant to the reconstitution, but we strongly believe that this complication enhances the clinical relevance of our data.  

      (7) Mechanistically, how does the loss of Tgr5 impact hematopoietic regeneration following sublethal irradiation?

      As delineated in the previous point, we have been seriously conditioned by cases of “swollen muzzles syndrome” (Garrett et al., 2018), which has stopped us from proceeding with more irradiation experiments for this particular study. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis is now a focus for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      (8) Only male mice were used throughout this study. It would be beneficial to know whether female mice show similar results.

      We thank Reviewer #1 for this question, as it led us to perform the characterization of steady-state hematopoiesis and morphological bone and BMAT evaluation of young female mice, yielding new findings that strengthen our manuscript. In summary, we have found that the decrease in BMAT is sexually dimorphic, with young females not showing reduced BMAT levels. Conversely, we did find a trend towards lower trabecular bone content in females. We present these data as part of Figure 3.

      Reviewer #2 (Public Review):

      Summary:

      In this manuscript, the authors examined the role of the bile acid receptor TGR5 in the bone marrow under steady-state and stress hematopoiesis. They initially showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. They further demonstrated that TGR5 knockout significantly decreases BMAT, increases the APC population, and accelerates the recovery upon bone marrow transplantation.

      Strengths:

      The manuscript is well-structured and well-written.

      We thank Reviewer #2 for this comment.

      Weaknesses:

      The mechanism is not clear, and additional studies need to be performed to support the authors' conclusion.

      We agree with Reviewer #2 that more studies are needed to understand the role of TGR5 in the hematopoietic system. We have been hampered in our studies of stress hematopoiesis because of frequent cases of swollen muzzles syndrome (Garrett et al., 2018), which prevented us from conducting additional experiments involving myelosuppression. Furthermore, the identification of the mechanism that links changes in the adipocyte differentiation axis and hematopoietic support, which is more complex than initially thought, has become a top priority for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      Recommendations For The Authors:

      Reviewer #2 (Recommendations For The Authors):

      (1) Figure 1: the authors showed the presence of TGR5 in hematopoietic cells using a GFP report in mice. What's the expression pattern of TGR5 in the nonhematopoietic cells? For example, adipocytes, stromal cells, endothelial cells, etc. In addition to analyze the percentage of TGR5-GFP+ cells, the authors should also quantify the expression levels of TGR5 in various hematopoietic and niche components.

      We thank Reviewer #2 for this question, which we have addressed as points 1 and 3 from Reviewer #1. The expression of TGR5:GFP signal in stromal cells, HSCs, and the various hematopoietic compartments is now presented respectively as new panels in Figure 6A-B, Figure 1C-D, Figure 1–Figure Supplement 1A-B, and Figure 1figure supplement 2A, as well as Author response image 1 for stromal progenitors. Please note that specifically for the stromal compartment, we kindly requested Reviewer #1’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section, as we are concerned about over interpretation for these rare APC and multi-potent stem cell-like subpopulations.

      Consistent with the typical low-abundance expression of G protein-coupled receptors, where ligand-mediated activation is more relevant than absolute receptor abundance, TGR5 is also lowly expressed in most cell types. Although technical limitations prevent us from directly quantifying TGR5 expression in adipocytes (due to the difficulty of isolating these populations from bone marrow), previous studies have reported TGR5 expression in human BMSC-derived adipocytes (Velazquez-Villegas et al., 2018) . For similar reasons, we could not quantify TGR5 expression levels by RT-qPCR in the bone marrow niche, as isolating enough cells from each compartment to reliably detect a low-abundance receptor is technically challenging.

      Regarding endothelial cells, previous studies have shown TGR5 expression in vascular endothelial cells (Kida et al., 2013) as well. In our dataset the number of endothelial cells recovered was limited, largely due to the lack of a dedicated endothelial isolation protocol for the BM. Even so, the data we collected are shown as Author response image 4 and Author response table 2 and suggest that TGR5:GFP level in endothelial cells is lower than in the stromal and hematopoietic compartments. As for Author response image 1 and Author response table 1, we remain cautious about the interpretation of GFP expression in these low frequency stromal and endothelial populations and we kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 4.

      Representative flow cytometry gating strategy used to identify endothelial cells in BM, defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> and their GFP positivity in TGR5:GFP mice.

      Author response table 2.

      Frequency of GFP-expressing cells in the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> gate for endothelial cells presented in Author Response Figure 4, expressed as average (standard deviation). The number of mice per genotype and per experiment is indicated in parenthesis.

      (2) Figure 2: Regarding the competitive transplantation, the authors should bleed the mice every 4 weeks and show the dynamics of donor chimerism up to 16 or 20 weeks. The difference at the 3-week time point is very tiny and is this change significant? There is no statistical significance shown in Supplemental Figure 2 Panel C at a 3-week time point.

      The full data for repetitive monthly bleedings in primary, secondary and tertiary transplants is now shown in Figure 2J and Figure 2-figure supplement 1I-K. The statistically significant difference on the first month, which represents a 26% loss in short-time progenitor phenotype, is small but potentially clinically relevant as it is these progenitors that sustain early hematopoietic recovery from severe leucopenia and thrombocytopenia. Indeed, the hematopoietic phenotype described in Figure 7 C-J for accelerated rescue of Tgr5<sup>-/-</sup> recipients with wild-type bone marrow would be coherently associated to this short-term progenitor phenotype.  

      (3) Figure 3: in addition to the irradiation/transplantation model, the authors should also use alternative models for stress hematopoietic, for example, 5-FU, etc.

      This is an excellent suggestion, but unfortunately beyond the scope of this manuscript due to the move of one of the co-senior authors to another institution. Follow-up studies are however planned with ablative chemotherapy models relevant to the hematopoietic transplant setting (non 5-FU).

      (4) To get a better understanding of the role of TGR5 in the bone marrow niche, the authors should get and/or generate the TGR5 floxed mice and this will allow the authors to delete it from specific cell populations using cell-type specific Cre. In addition, it would be very interesting to investigate further the molecular mechanisms downstream of TGR5.

      We much appreciate this comment and the interest it reflects in TGR5 BM biology. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis are now a focus for one of the involved laboratories and should produce concrete follow-up manuscripts soon.  

      References

      Ambrosi, T.H., Scialdone, A., Graja, A., Gohlke, S., Jank, A.-M., Bocian, C., Woelk, L., Fan, H., Logan, D.W., Schürmann, A., Saraiva, L.R., Schulz, T.J., 2017. Adipocyte Accumulation in the Bone Marrow during Obesity and Aging Impairs Stem Cell-Based Hematopoietic and Bone Regeneration. Cell Stem Cell 20, 771-784.e6. https://doi.org/10.1016/j.stem.2017.02.009

      Coutu, D.L., Kokkaliaris, K.D., Kunz, L., Schroeder, T., 2017. Three-dimensional map of nonhematopoietic bone and bone-marrow cells and molecules. Nat Biotechnol 35, 1202–1210. https://doi.org/10.1038/nbt.4006

      Craft, C.S., Robles, H., Lorenz, M.R., Hilker, E.D., Magee, K.L., Andersen, T.L., Cawthorn, W.P., MacDougald, O.A., Harris, C.A., Scheller, E.L., 2019. Bone marrow adipose tissue does not express UCP1 during development or adrenergic-induced remodeling. Sci Rep 9, 17427. https://doi.org/10.1038/s41598-019-54036-x

      Garrett, J., Sampson, C.H., Plett, P.A., Crisler, R., Parker, J., Venezia, R., Chua, H.L., Hickman, D.L., Booth, C., MacVittie, T., Orschella, C.M., Dynlachta, J.R., 2018. Characterization and Etiology of Swollen Muzzles in Irradiated Mice. Radiation Research 191, 31. https://doi.org/10.1667/RR14724.1

      Gil, L., Poplawski, D., Mol, A., Nowicki, A., Schneider, A., Komarnicki, M., 2013. Neutropenic enterocolitis after high-dose chemotherapy and autologous stem cell transplantation: incidence, risk factors, and outcome. Transpl Infect Dis 15, 1–7. https://doi.org/10.1111/j.1399-3062.2012.00777.x

      Gomariz, A., Helbling, P.M., Isringhausen, S., Suessbier, U., Becker, A., Boss, A., Nagasawa, T., Paul, G., Goksel, O., Székely, G., Stoma, S., Nørrelykke, S.F., Manz, M.G., Nombela-Arrieta, C., 2018. Quantitative spatial analysis of haematopoiesis-regulating stromal cells in the bone marrow microenvironment by 3D microscopy. Nat Commun 9, 2532. https://doi.org/10.1038/s41467-018-04770-z

      Gooley, T.A., Chien, J.W., Pergam, S.A., Hingorani, S., Sorror, M.L., Boeckh, M., Martin, P.J., Sandmaier, B.M., Marr, K.A., Appelbaum, F.R., Storb, R., McDonald, G.B., 2010. Reduced mortality after allogeneic hematopoietic-cell transplantation. N Engl J Med 363, 2091–2101. https://doi.org/10.1056/NEJMoa1004383

      Harrison, D.E., Jordan, C.T., Zhong, R.K., Astle, C.M., 1993. Primitive hemopoietic stem cells: direct assay of most productive populations by competitive repopulation with simple binomial, correlation and covariance calculations. Exp Hematol 21, 206–219.

      Janzen, V., Forkert, R., Fleming, H.E., Saito, Y., Waring, M.T., Dombkowski, D.M., Cheng, T., DePinho, R.A., Sharpless, N.E., Scadden, D.T., 2006. Stem-cell ageing modified by the cyclin-dependent kinase inhibitor p16INK4a. Nature 443, 421–426. https://doi.org/10.1038/nature05159

      Kida, T., Tsubosaka, Y., Hori, M., Ozaki, H., Murata, T., 2013. Bile acid receptor TGR5 agonism induces NO production and reduces monocyte adhesion in vascular endothelial cells. Arterioscler Thromb Vasc Biol 33, 1663–1669. https://doi.org/10.1161/ATVBAHA.113.301565

      Morowski, M., Vögtle, T., Kraft, P., Kleinschnitz, C., Stoll, G., Nieswandt, B., 2013. Only severe thrombocytopenia results in bleeding and defective thrombus formation in mice. Blood 121, 4938–4947. https://doi.org/10.1182/blood-2012-10-461459

      Naveiras, O., Nardi, V., Wenzel, P.L., Hauschka, P.V., Fahey, F., Daley, G.Q., 2009. Bonemarrow adipocytes as negative regulators of the haematopoietic microenvironment. Nature 460, 259–263. https://doi.org/10.1038/nature08099

      Purton, L.E., Dworkin, S., Olsen, G.H., Walkley, C.R., Fabb, S.A., Collins, S.J., Chambon, P., 2006. RARgamma is critical for maintaining a balance between hematopoietic stem cell self-renewal and differentiation. J Exp Med 203, 1283–1293. https://doi.org/10.1084/jem.20052105

      Purton, L.E., Scadden, D.T., 2007. Limiting factors in murine hematopoietic stem cell assays. Cell Stem Cell 1, 263–270. https://doi.org/10.1016/j.stem.2007.08.016

      Sebo, Z.L., Rendina-Ruedy, E., Ables, G.P., Lindskog, D.M., Rodeheffer, M.S., Fazeli, P.K., Horowitz, M.C., 2019. Bone Marrow Adiposity: Basic and Clinical Implications. Endocr Rev 40, 1187–1206. https://doi.org/10.1210/er.2018-00138

      Suchacki, K.J., Tavares, A.A.S., Mattiucci, D., Scheller, E.L., Papanastasiou, G., Gray, C., Sinton, M.C., Ramage, L.E., McDougald, W.A., Lovdel, A., Sulston, R.J., Thomas, B.J., Nicholson, B.M., Drake, A.J., Alcaide-Corral, C.J., Said, D., Poloni, A., Cinti, S., Macpherson, G.J., Dweck, M.R., Andrews, J.P.M., Williams, M.C., Wallace, R.J., Van Beek, E.J.R., MacDougald, O.A., Morton, N.M., Stimson, R.H., Cawthorn, W.P., 2020. Bone marrow adipose tissue is a unique adipose subtype with distinct roles in glucose homeostasis. Nat Commun 11, 3097. https://doi.org/10.1038/s41467-020-16878-2

      Tratwal, J., Bekri, D., Boussema, C., Sarkis, R., Kunz, N., Koliqi, T., Rojas-Sutterlin, S., Schyrr, F., Tavakol, D.N., Campos, V., Scheller, E.L., Sarro, R., Bárcena, C., Bisig, B., Nardi, V., de Leval, L., Burri, O., Naveiras, O., 2020. MarrowQuant Across Aging and Aplasia: A Digital Pathology Workflow for Quantification of Bone Marrow Compartments in Histological Sections. Front Endocrinol (Lausanne) 11, 480. https://doi.org/10.3389/fendo.2020.00480

      Trivedi, M., Martinez, S., Corringham, S., Medley, K., Ball, E.D., 2009. Optimal use of G-CSF administration after hematopoietic SCT. Bone Marrow Transplant 43, 895–908. https://doi.org/10.1038/bmt.2009.75

      Vannini, N., Campos, V., Girotra, M., Trachsel, V., Rojas-Sutterlin, S., Tratwal, J., Ragusa, S., Stefanidis, E., Ryu, D., Rainer, P.Y., Nikitin, G., Giger, S., Li, T.Y., Semilietof, A., Oggier, A., Yersin, Y., Tauzin, L., Pirinen, E., Cheng, W.-C., Ratajczak, J., Canto, C., Ehrbar, M., Sizzano, F., Petrova, T.V., Vanhecke, D., Zhang, L., Romero, P., Nahimana, A., Cherix, S., Duchosal, M.A., Ho, P.-C., Deplancke, B., Coukos, G., Auwerx, J., Lutolf, M.P., Naveiras, O., 2019. The NAD-Booster Nicotinamide Riboside Potently Stimulates Hematopoiesis through Increased Mitochondrial Clearance. Cell Stem Cell 24, 405-418.e7. https://doi.org/10.1016/j.stem.2019.02.012

      Velazquez-Villegas, L.A., Perino, A., Lemos, V., Zietak, M., Nomura, M., Pols, T.W.H., Schoonjans, K., 2018. TGR5 signalling promotes mitochondrial fission and beige remodelling of white adipose tissue. Nat Commun 9, 245. https://doi.org/10.1038/s41467-017-02068-0

      Walkley, C.R., Fero, M.L., Chien, W.-M., Purton, L.E., McArthur, G.A., 2005. Negative cellcycle regulators cooperatively control self-renewal and differentiation of haematopoietic stem cells. Nat Cell Biol 7, 172–178. https://doi.org/10.1038/ncb1214

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public reviews):

      Weaknesses:

      Reliance on self-reports

      A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.

      We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”

      Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.

      Statistical analysis and interpretation of the dynamics matrix

      Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.

      We thank the reviewer for raising this important methodological point. We address it in three parts.

      (i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.

      (ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T<sup>2</sup>, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.

      (iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”

      Terminology (controllability, stability, sensitivity, Gramian)

      Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.

      This is an excellent and important point. We have made the following changes throughout the manuscript.

      (i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.

      (ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”

      (iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.

      (iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.

      (v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.

      Reviewer #2 (Public reviews):

      Online recruitment, selection effects, and remuneration

      Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).

      We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”

      Intervention ongoing during the second block

      Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.

      This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”

      Observation noise

      As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.

      We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.

      Reliance on formal model comparison

      Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.

      We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.

      We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.

      Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.

      Statistical limitations; Bayesian inference

      The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.

      We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.

      Reviewer #3 (Public reviews):

      Dual meanings of controllability

      An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”

      We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.

      Strategy reminder and possible reappraisal

      As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.

      This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.

      Mechanism of distancing effects (eye movement and oculomotor avoidance)

      Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.

      We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”

      Recommendations for the Authors:

      Reviewer #1 (Recommendations for the Authors):

      (1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?

      Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.

      (2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?

      We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.

      (3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.

      References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.

      (4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.

      The typo has been corrected.

      (5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.

      We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.

      (6) Could the authors develop the rationale behind the bias in Equation 1?

      We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.

      (7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.

      We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.

      (8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).

      We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.

      (9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.

      The sentence has been rephrased: ”Cumulative model weights (w<sub>j</sub>) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”

      (10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.

      We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.

      (11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.

      We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.

      (12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.

      We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.

      <(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.

      Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.

      We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      (14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?

      The regression equation has been corrected to DV = β<sub>0</sub> + β<sub>1</sub>IV + β<sub>2</sub>G + β<sub>3</sub>(IV × G) + ϵ, making the interaction term explicit.

      (15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?

      The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.

      Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.

      Beyond the visual trajectory overlays in Figure 4C, we have now added R<sup>2</sup> and peak cross-correlation as a quantitative measure of individual fit quality. Mean R<sup>2</sup> across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R<sup>2</sup> values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.

      (16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?

      Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.

      (17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.

      Interaction effects are now included in the relevant figure caption (Figure 5).

      (18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?

      Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?

      This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.

      (19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?

      This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.

      (20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?

      Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.

      Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.

      (21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?

      The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.

      (22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).

      Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.

      (23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.

      We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.

      (24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.

      The supplementary tables have been restructured to a hierarchical format.

      (25) Table I9 is not referenced in the manuscript.

      A reference to this table has been added in the appropriate Results section.

      (26) It’s a shame that neither data nor analysis code has been made available.

      Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).

      Reviewer #2 (Recommendations for the Authors):

      (1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.

      The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.

      (2) p 5: ’on [not in] the recruiting platform’.

      Corrected.

      (3) p 8: Clarify notation of x<sub>t</sub>t = 1<sup>T</sup> (and same for u). What does this mean?

      The notation has been corrected.

      (4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.

      The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.

      (5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.

      References have been added at each key definition.

      (6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.

      This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.

      (7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?

      A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.

      (8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?

      Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.

      (9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?

      We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.

      (10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.

      Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.

      (11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].

      This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      Reviewer #3 (Recommendations for the Authors):

      (1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.

      More information about the power analysis have been added to the Participants section of the main text.

      (2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).

      We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.

      (3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”

      Thank you for this comment. Please see our response to the public review comments above, which we hope address this.

      (4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?

      See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.

      (5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.

      We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.

    1. eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination, while other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain. However, the lack of access to raw data hampers the assessment of data quality.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      Comment on revised version.

      Authors made satisfactory revision and I have no further comments.

    3. Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) Neurons in Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.<br /> (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, clear and provide much information with the most appropriate plots and graphics. The study's amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      Weaknesses:

      (1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an 'enormous' field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/days in a raw, or am I mistaken?

      The updated version of the manuscript partially addresses the flaws of the original submission. Here are my general concerns:

      (1) Overall, the revised version comes with a few modifications/additions and no new data. Apart from a new correlation analysis, the improvements are mainly discursive, often non-convincing, justifications of the authors' choices. This may reflect a lack of interest in a publication that, in the meantime, lost its original peer-review value. However, it should be acknowledged that the Authors made an effort to partially address the concerns raised by the reviewers.

      (2) My main concern was the quality of fMRI recordings. In the revised version, the Authors provided a new figure with an example single-mouse fMRI data. However, the depicted regions of interest (ROIs) mostly cover the brain spots that I expected to be the most impacted by the BOLD artifacts caused by the proximity of the air and the big field-of-view. In addition, these ROIs do not appear to match the mouse anatomy shown above the functional data. As an example, the EPI images in the OB are almost entirely covered by the colored mask. The feeling is that the fMRI data was indeed poor, as I worried, and the lack of any public repository of raw data reinforces that feeling. To make this point clear: I do not think the findings are not true, but poor fMRI data quality might have hidden more insightful results and does not foster the use of fMRI to monitor the olfactory pathway, which lowers the impact of this article.