10,000 Matching Annotations
  1. Aug 2026
    1. Reviewer #1 (Public review):

      Summary:

      Despite setting global conservation goals for the conservation of genetic diversity, gathering enough genetic data to assess its current status or monitor change over time for all species would require substantial time and resources. Finding reliable proxies or predictors of genetic diversity change would be a valuable alternative. Selmoni and Schuman test for patterns and trends in coral reef animals' genetic diversity across space and time and evaluate how well environmental variables predict their genetic diversity.

      Strengths:

      The authors have compiled a large genomic dataset from 19 studies, 18 species, and 173 reefs, and use a standardized k-mer-based pipeline to process all datasets. The authors propose that using k-mers may be especially suitable for macrogenomic analyses because they are less computationally intense and do not require a reference genome. This is interesting because there are currently few examples of macrogenetics studies using genomic data, likely due in part to the difficulty associated with processing multiple disparate datasets in an efficient and standardized way.

      Weaknesses:

      I have concerns that the models don't fit the data structure, the data are not suitable to test for temporal change, and that the values being predicted do not reflect meaningful genetic diversity. The manuscript would also benefit from a more cohesive narrative structure, clearly stated research goals, and better engagement with previous literature on this topic.

      Effects of time: The data are not suitable to test for temporal changes across all species. 11 of 19 datasets were only sampled over 1-2 years, and only 5 were sampled across 5 or more years. I don't know generation times for these species offhand, but this is generally too short to draw any conclusions about genetic diversity change over time. While it makes sense to include year in the models as a control variable, in my opinion these patterns should not be interpreted and removed from the discussion, or at least the conclusions should be tempered (for example the section "Are coral reef populations rapidly losing their genetic diversity?").

      Environmental effects models: I'm not convinced that the geographic random intercept term (response variable) represents a local genetic distance that should be related to environments. How many species were sampled at each reef? If only 1 species is sampled, then the reef's average diversity reflects that species, not necessarily environmental effects. If multiple species are sampled at a reef, then I suspect this value is the average diversity of all the species sampled there. Knowing that species have different levels of diversity (e.g., Toczydlowski et al 2025 cited here), what does this metric mean and how is it affected by species composition vs environments? I do not see the value in predicting this variable without considering species identity.

      Introduction: The macrogenetics literature is not well described in the introduction. Most of the cited studies indeed use raw data rather than summary statistics, and most do not aim to predict genetic diversity.

      I wouldn't say that explaining more than 20% of variation in data is a major challenge in macrogenetics. The choice of how many, and which, predictors to use depends on the research question (e.g., on the spectrum of understanding to predicting, see Shmueli 2010 "To Explain or To Predict?"; Arif & MacNeil 2022 "Predictive models aren't for causal inference"). It's proposed here that explained variance can be increased by adding more, or more relevant, environmental predictors-but particularly for macrogenetics studies reusing data from different sources with different study and sampling designs and different species, much of the variation in the data will be due to these factors. That is why conditional R2s are typically quite high (>80%) in mixed models with species/study as random effects, see e.g. Clark & Pinsky 2024 or Karachaliou 2025. Having strong effects of particular variables would also require that all species respond to the predictor in the same way, which is not necessarily expected. Low explanatory capacity alone is not an issue if the goal is not prediction.

    2. Reviewer #2 (Public review):

      Summary:

      The authors measured whether a phylogenetically wide sample of marine species was experiencing declining genetic diversity, as one might expect from widespread habitat threat, over a recent 2-decade time span. They next identified key environmental variables that are impacting genetic diversity within and between reefs. A key insight was to apply k-mer-based genetic distances to massively speed up the reanalysis of genetic data into a common pipeline, which is otherwise onerous. The manuscript ends by highlighting key seascape variables that were associated with increases or decreases with genetic diversity through time. The authors achieved their overall aims, though I remain unsure of how well the identified seascape variable-genetic diversity predictions can be generalized, as implied in the abstract.

      Strengths:

      A key strength of the study was its rigor in vetting the k-mer based distance metrics, checking whether they give population structure patterns (Figure S2) and correlated with nucleotide diversity. This surpasses previous studies that used a similar method. The modeling procedure of genetic diversity predictions from environmental variables is also rigorous and presented with some appropriate nuance.

      Weaknesses:

      I noted five weaknesses, listed below in order of potentially more severe at the top to more minor at the bottom.

      First, only one k-mer distance metric was tested. There are many k-mer distance metrics that will potentially give different weight to different frequencies of polymorphism, like how Watterson's theta and nucleotide diversity weight polymorphisms, depending on their frequency. It would be interesting to see if the seascape variable predictions hold with a Jaccard or cosine distance metric, or if the observed results are purely restricted to the choice of Bray-Curtis distance. Another option would be to use mash, skmer, or (very recent development) re-skmer distances, which attempt to more directly approximate the average nucleotide identity between two sequence sets based on the k-mer sets of their reads while also being faster than traditional alignment. This would potentially give cleaner trends, seeing as the goal is to have a proxy for nucleotide diversity, with the downside of not including the impacts of non-SNP variation.

      Second, it is unclear to me whether the number of species analyzed, 18, can accurately identify important environmental variables in early warning systems, as claimed in the abstract. While this likely represents the best available balance of evidence, it is worth highlighting the manuscript's note that the environment-diversity predictions did not scale across marine realms. This is likely a limitation of data availability, rather than a study design flaw, but is nonetheless important for readers to keep in mind. Would recommend that the abstract acknowledge this limitation.

      Third, it is unclear to me how the included datasets compare in terms of genome-wide coverage. Figure S1 gives sequencing depths in terms of read number, but what is the range normalized for genome size - are the datasets 1x, 5x, 10x, on average, etc? This is probably most key to how the k-mer distance metrics will perform, because low coverage will make two samples appear artificially distant due to rarity of sampling the same k-mer multiple times. However, the correlation between k-mer distance and regular alignment-based SNP distance (Figure 2C) gives some confidence that this effect could be small.

      Fourth, though k-mer distances may in some sense better capture the breadth of DNA sequence diversity, they lack a concrete interpretation of what loci may/may not be under selection as marine environments are increasingly threatened over time. This is a different question than what the manuscript tries to address, but is of interest to the field and is something that would be seemingly difficult to do with k-mers.

      Finally, I had a more minor concern: the k-mer-based distances use only k-mers that are mapped to a reference genome. This is done to thoroughly remove contaminant sequences, which is important, but is a double-edged sword because there is potentially additional pangenomic variation that is real but does not map to a single reference. The authors state in the first paragraph of the discussion that this choice did not affect the results, but I did not find an associated analysis in the supplement or main text. It would be interesting to compare distance metrics based on screening against all known microbial+human genomes vs screening against the reference genomes.

    1. eLife Assessment

      This study presents a valuable analysis of the TARA oceans dataset, advancing genome- and ecosystem-scale metabolic modeling. The authors present convincing evidence for the factors that shape microbial ecology at the molecular scale. The work will be of broad interest to oceanographers and scientists in other areas where metabolic interactions in diverse communities are of importance.

    2. Reviewer #1 (Public review):

      The authors use metabolic reconstruction to build genome-scale models from TARA metagenomes.

      This is not the first paper that attempts a reconstruction of metabolism on this scale. And like those that came before, its weakness is in the analysis of the resulting model.

      I was initially quite excited about this manuscript as the reconstruction itself seems well done. However, the analysis does not deliver major or minor results.

      It could be said that the modelling cuts some corners, for example in the identification of the flux modes. I was not quite able to follow why this is necessary, as numerical linear algebra methods seem to be able to handle problems of this size. But then again, I am quite confident that the sampling approach that the authors use produces a good enough approximation.

      The weak point of the manuscript is the analysis that follows. The conclusions contain very few hard results, and the results that are presented talk mostly about concepts that are somewhat arbitrarily introduced and hard to relate back to nature. The authors construct a mix of their own concepts and others borrowed from cancer research, and neither of the metrics that are used is very convincing.

      For example, the authors say they contribute to the complexity-stability debate, but that debate focuses on a particular notion of stability. What the authors call stability here is an entirely different concept that is completely unrelated to the stability of the complexity-stability debate; I would rather describe it as a resilience metric rather than stability.

      The problem is that the metrics that are used lack an underlying firm grounding. Ecology was for a long time struggling with the same problem (and to some extent still is), but eventually much progress has been made by rigorously deriving metrics that are easy to relate back. The same would be possible here. Using Steuer's structural kinetic method, the reconstructed metabolism could be described dynamically, which would allow genuine stability analysis as well as other crucial metrics such as the observability and controllability, impact and sensitivity, which would make it far easier to relate results back.

    3. Reviewer #2 (Public review):

      This study used publicly available Tara Oceans and Tara Oceans Polar Circle metagenomic and metatranscriptomic datasets, including viromes, to construct sample-specific, gene-scale metabolic models. This reviewer understood that, for each sample, genes or transcripts associated with metabolic pathways were integrated into a single virtual "superorganism" or "community cell." These sample-level models were then compared across global ocean regions to investigate spatial patterns in heterotrophic prokaryotic metabolism, metabolic synergy, and the potential effects of virus-encoded auxiliary metabolic genes. However, it was not clear whether archaeal genes were also included in the heterotrophic prokaryotic fraction.

      The study represents an ambitious and potentially valuable attempt to connect large-scale environmental omics data with constraint-based metabolic modeling. The authors handled a very large dataset and introduced quantitative approaches based on flux sampling, Reaction Cumulative Correlation, and synergy scores to describe community-level metabolic phenotypes using several mathematical formulations. The recovery of previously defined oceanic ecological zones from reaction-based models may provide preliminary support for the ecological relevance of the framework, although only approximately 26% of prokaryotic genes and 7% of viral genes were mapped to known metabolic reactions.

      However, the framework should be understood as gene- or reaction-resolved, sample-level metabolic modeling rather than genome-resolved community modeling. By combining all detected genes within a sample into a single superorganism, the approach likely loses taxon-specific metabolic information and cannot directly distinguish intracellular metabolism from interspecies metabolic exchange. Consequently, several ecological interpretations, including cooperation, stability, reaction essentiality, and viral impacts, remain strongly dependent on the underlying model assumptions.

      Overall, the manuscript presents a novel and scalable framework with considerable potential. However, greater methodological clarification and more cautious interpretation are needed before the ecological and biogeochemical conclusions can be fully supported.

    1. eLife Assessment

      This is a valuable manuscript that leverages information already being collected in mosquito surveillance, but that is currently discarded, building on ideas developed during the SARS-CoV-2 pandemic. The work is mostly rigorous, involving ample data from real-world surveillance, independent laboratory testing, and simulation modeling to validate the inference method. The evidence is solid for demonstrating feasibility and biological plausibility, although some conclusions would be strengthened by additional sensitivity analyses on key assumptions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript seeks to make use of information about Ct values from PCR testing of mosquito pools for West Nile virus infection to make inferences about mosquito prevalence and West Nile risk. It does so through analysis of empirical data and simulated data with a realistic agent-based model.

      Strengths:

      This work is conceptually innovative for mosquito-borne viruses, building on ideas developed primarily during work on SARS-CoV-2. Exploring this topic is worthwhile regardless of the outcome. The use of data, testing in multiple labs, and the complementarity of modeling and empirical data analysis are all strengths of the approach.

      Weaknesses:

      Some of the primary weaknesses include a dependence of the results on relatively narrow model assumptions, and a lack of compelling improvement over existing methods. None of these weaknesses are fatal flaws; they are modest weaknesses that limit the potential of or excitement about the method.

    3. Reviewer #2 (Public review):

      Summary:

      The authors extend their previous population-based Ct-value framework for inferring community epidemic trajectories from human infections to vector infections, using mosquitoes as vectors for West Nile virus. They use agent-based modelling to distinguish virus-positive detections arising from non-active infection states from those reflecting active infections, and then apply this framework to mosquito surveillance data from Colorado and Texas.

      Overall, this is a well-designed and carefully evaluated study. The manuscript proposes a feasible and potentially valuable framework for vector infection surveillance. The findings are supported by both mechanistic agent-based simulations and applications to real-world mosquito surveillance data, which strengthens the biological plausibility and practical relevance of the proposed approach.

      Strengths:

      A major strength of the study is its clear methodological extension from human infection surveillance to vector infection surveillance. The agent-based modelling framework provides a useful basis for distinguishing active infections from virus-positive detections that may reflect non-active infection states. The application to surveillance data from two different geographic settings further supports the feasibility of the framework. Overall, the study is carefully designed, and the model schematic and main analyses are generally clear.

      Weaknesses:

      (1) It would be helpful if the authors could provide plots showing variation across locations and over time. This would further support the claim made in the paragraph at lines 101-107.

      (2) Figure 2: The model schematic is clear in terms of workflow, but it would benefit from more information on model parameterization. In particular, it would be helpful to clarify which parameters or migration rates were estimated from the data and which were assumed based on prior literature.

      (3) Figure 4: I wonder whether the authors examined how changes in the proportion of mosquitoes with static viral-kinetics trajectories would affect the observed bimodal distribution. Relatedly, it would be useful to know whether there is a threshold proportion at which the method becomes less able to distinguish active from static viral-kinetics patterns.

      Conclusion:

      Overall, the evidence is reasonably strong for demonstrating the feasibility and biological plausibility of the proposed framework. Some conclusions would be further strengthened by additional sensitivity analyses on key assumptions, especially the proportion of static viral-kinetics trajectories and spatial-temporal heterogeneity across surveillance sites.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      We thank the reviewer for pointing out that, as E/DXXD has been considered a conserved caspase-3 cleavage site, TEVDA has also been validated as a caspase-6 cleavage site, which we have missed. We will further confirm this using site-mutated expression vectors in S2 cells.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      We thank the reviewer for the suggestion. We will attempt to label Xkr-mCherry under apoptosis with other vesicles and change the ER marker to Cnx99A (Calnexin ortholog in Drosophila).

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      We thank the reviewer for their careful reading, and we will correct the mistakes in the revised manuscript.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      We agree with the reviewer’s opinion. These mutants from the BDSC are not completely deleted but partially deleted; therefore, the RT-qPCR assay may not be very accurate. We will detect the mRNA levels of tm9sf4, dorp9, and sac1 using RT primers from different cDNA regions to make the results more convincing.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      We thank you for raising this intriguing hypothesis. Indeed, PM localization of Lact:C2 was reduced in xkr<sup>ko</sup> cells, and the distribution could not be rescued after apoptosis. Unlike xkr<sup>ko</sup>, tm9sf4, and sac1 RNAi-treated cells displayed weak PS disorder, which may be due to the efficiency of knockdown. However, the statistical results indicated that the ratio of PM/Cyto was reduced in tm9sf4 and sac1 RNAi-treated cells.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      We thank for the reviewer’s suggestion. We will add the statistical results of Fig. 3I in the revised version.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      We thank you for raising this intriguing hypothesis. We will delete the relevant statement to eliminate unnecessary misunderstandings.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      We thank for the reviewer’s suggestion. We will move this to the main text and provide quantification in the revised version.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

      We thank for the reviewer’s suggestion. We will weaken the statement in the Results section of Figure 6 and perform osbpl9 knockdown to observe efferocytosis in Raw264.7 cells, as OSBPL9 interacts with Xkr8 strongly.

      Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

      Thank you for your comprehensive and generous assessment of our work and for the time and expertise you have devoted to reviewing our manuscript. We will revise the manuscript accordingly and provide a point-by-point response in the revised version.

      Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      We sincerely thank you for your careful reading and accurate summary of our manuscript. We appreciate the time, effort, and expertise you have dedicated to evaluating our work, and we will try our best to improve our manuscript according to your suggestions.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      We thank for the reviewer’s suggestion. We used AcD-treated S2 cells and untreated S2 cells to perform mass spectrometry. To clarify this, we will add a detailed method description in the method section.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      We thank for the reviewer’s suggestion. We used anti-HA magnetic beads to perform immunoprecipitation, and single HA-TM9SF4 was used as a negative control.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      We thank for the reviewer’s suggestion, we will add the negative control for BiFC results in the revised version.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      We sincerely thank you for the careful consideration of our statistical results, which were obtained from 3 repeats. We will choose another flow cytometry histogram of tm9sf4 and sac1 to make the data and images more consistent.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      We sincerely thank the reviewer for their careful and thoughtful observations. Given the results we have observed, we will add this to the discussion section in the revised version.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      We thank for the reviewer’s suggestion. We will overexpress Xkr with or without AcD treatment to observe whether the localization or amount of Lact:C2 changes and to re-evaluate the role of Xkr in PS exposure.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      We thank the reviewer for the suggestion, and we believe that the TIRF technique can help us demonstrate the ER-PM signal. Since our college has no TIRF microscope, we will try our best to seek cooperation from other colleges to achieve this experiment.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

      We sincerely thank the reviewer for this careful and thoughtful observation, and we agree that the interpretation in Line 285-288 is overstated. To explain this, we plan to detect the Lact:C2 and MAPPER signals in S2 and xkr<sup>ko</sup> cells with or without AcD to confirm how Xkr regulates PS via ER-PM under apoptotic conditions.

    2. eLife Assessment

      This study presents valuable findings implicating Xkr and its newly identified binding partners in regulating the exposure of apoptotic signal for phagocytosis, suggesting that Xkr might facilitate phosphatidylserine transfer at the ER-plasma membrane contact sites. However, several conclusions, including caspase-independent activity of Xkr and its role at the ER-PM contact sites, are not well supported by experimental data. Furthermore, some key controls are missing, and in numerous places the authors' discussion has extended far beyond what the data support. The findings presented are currently incomplete and do not fully support key mechanistic claims of the study.

    3. Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

    4. Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

    5. Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

    1. eLife Assessment

      An intriguing set of electrophysiological and analytical results suggests that adult rat neurons in the medulla oblongata crucial for mediating descending pain modulation are embedded in a network that generates a specific pattern of oscillatory activity. These valuable findings may broaden our understanding of how descending pain modulation is achieved. The supporting evidence is incomplete, and the authors need to provide several methodological clarifications and additional data interpretations to strengthen the manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      The authors hypothesized that "RVM neurons operate across multiple temporal scales, integrating fast responses associated with reflex-linked control with slower fluctuations reflecting ongoing network or state-dependent modulation". The hypothesis was tested with the established ON/OFF-cell model and probabilistic modeling. The study is conceptually interesting and methodologically sophisticated. The findings build toward the conclusion that pain-control circuits operate across multiple timescales.

      Strengths:

      The use of Bayesian regression and Gaussian process modeling to quantify and characterize recovery dynamics and ongoing oscillatory activity.

      The authors show that slow rhythmic activity appears preferentially in ON- and OFF-cells but not in NEUTRAL-cells, suggesting that the oscillations are related to pain-modulatory circuitry rather than being a generic feature of all recorded neurons.

      The observation that some oscillatory activity is coherent with autonomic measures aligns with broader views of the RVM as a hub integrating nociceptive and homeostatic regulation.

      Some pitfalls are appreciated and discussed by the authors, including the functional significance of slow fluctuations, the influence of anesthetics on global brain-state dynamics, the molecular profiles of the studied ON- and OFF-cells, and the heart rate as a covarying signal of RVM neuronal activity.

      Weaknesses:

      A general weakness is that the work is mostly descriptive and relies on anesthetized preparations. Whether the observed rhythms occur in awake animals and are linked to fluctuations in pain behavior needs to be confirmed in future studies.

      The study measures limited autonomic variables. The causal relationship between "slow fluctuations" and "ongoing physiological state" is unclear and overstated, since the data presented appear correlational.

      ON- and OFF-cells in the RVM are identified by their responses correlated with reflexive activity. The significance of the observed oscillations in spontaneous pain conditions is unclear.

      It is uncertain whether the observed rhythms truly reflect intrinsic RVM organization rather than anesthesia-dependent phenomena; the authors appreciated this pitfall, though.

      The Gaussian process analysis suggests predictability and quasi-periodicity, but predictability alone does not necessarily imply a true biological oscillator.

      Conclusion:

      The results support the authors' hypothesis. The findings provide a compelling conceptual message about the multiscale organization and dynamics of descending pain-control circuits and encourage further studies on the topic.

    3. Reviewer #2 (Public review):

      Using electrophysiological recordings in a well-characterized animal model of acute pain, and analytical and modeling methods, the authors show that descending pain-modulatory neurons in the rostral ventromedial medulla (RVM) operate across various timescales. They have both rapid multi-phase responses to noxious stimuli that unfold over tens of seconds, with distinct fast and slow recovery dynamics. Additionally, they generate slow quasi-periodic oscillations with approximately 5-minute periods during ongoing activity. These oscillations are statistically predictable and cell-type specific, demonstrating that descending pain control is organized through structured temporal dynamics that encompass immediate stimulus-evoked responses and slower fluctuations associated with physiological state.

      A novel discovery is a ~5-minute quasi-periodic oscillation in ongoing ON- and OFF-cell activity. This oscillation, along with its coherence with heart rate, forms the basis for the claim that descending pain circuits exhibit intrinsic multi-timescale organization. However, it's crucial to demonstrate that this periodicity is independent of external experimental cycles such as methohexital infusion pharmacokinetics, servo-controlled temperature regulation, or slow autonomic feedback loops, all of which operate on similar timescales. For instance, the 300-second period closely matches typical drug infusion cycling and thermoregulatory feedback intervals. Therefore, heart-rate coherence peaks at multiples of this period could equally reflect a shared external driver rather than intrinsic RVM organization. Although the absence of this cyclic structure in Neutral cells argues against this possibility, the authors might want to explicitly discuss this potential confound.

      The findings are important and novel in that they characterize an intriguing structure in the activity of ON and OFF neurons in the RVM. However, in the absence of a causal manipulation causality can only be inferred. That there is no phase-dependence of withdrawal latency argues against a causal role. The author are encouraged to qualify their conclusions (and their title) accordingly.

      Because anesthesia can affect global dynamics, this might affect the oscillations reported. Without awake validation, it remains uncertain whether these rhythms reflect an intrinsic property or an anesthesia-induced regime. Again, the absence of oscillations in Neutral cells argues against this possibility, but it is still possible that ON/OFF cells are embedded in different circuits that are affected differently by anesthesia.

      Analyses of many of the ON-cells had longer training windows (>1 sec) compared to those for the NEUTRAL cells. Could this have reduced the ability to fit and validate periodicity for the latter cell type?

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript by Ashworth and colleagues, the authors investigate the temporal dynamics of the rostral ventromedial medulla (RVM), a key output node in a major descending pain-modulation circuit. Using data from extrasellar single-unit recordings of RVM ON, OFF, and NEUTRAL cells in lightly anesthetized rats, the authors' computational modeling yielded two major findings: (1) heat-evoked ON burst and OFF pause, followed by exponential recovery components in10s of seconds; and (2) ON and OFF cells exhibit periodic fluctuations in ~5-minute cycles that are statistically predictable.

      Strengths:

      The manuscript's concept is innovative, offering the first quantitative analysis of multi-timescale dynamics in physiologically characterized RVM pain-modulating neurons. This advances a field that has mostly depended on qualitative or single-timescale descriptions. The authors use contemporary Gaussian process and probabilistic models to capture statistically predictable slow dynamics. The study is further strengthened by identifying ON-, OFF-, and NEUTRAL-type cells using well-established criteria grounded in decades of RVM research. The combination of rapid reflex-related responses and slower ongoing rhythms supports a dual-timescale framework, providing a more integrated understanding of how these neurons may regulate reflex activity and state-dependent processes.

      Weaknesses:

      Several limitations are noted. Incomplete characterization of light anesthesia during recording sessions, such as methohexital stability and clear criteria for identifying "lightly anesthetized" states. While the NEUTRAL cell control is helpful, it does not fully address concerns about circuit specificity or systemic confounds. The findings are male-dominant, which may limit their generalizability. The synchrony between ON and OFF cells was suggested but not directly tested. The heart rate coherence with ON, OFF, and NEUTRAL cell activity results is intriguing but does not fully clarify how these neurons influence heart rate, particularly within the "lightly anesthetized" model.

    5. Author response:

      We are pleased that the reviewers found the study conceptually novel and the analytical framework rigorous. In response we have substantially revised the manuscript to clarify methodological details, temper several interpretations, expand discussion of alternative explanations, and include additional analyses using the existing dataset. We have deliberately revised the manuscript so that our conclusions are limited to those directly supported by the data, namely that physiologically identified RVM pain-modulatory neurons exhibit structured dynamics spanning multiple temporal scales. We do not interpret the slow fluctuations as evidence for a specific intrinsic oscillator or for a causal role in physiological state regulation. We have also expanded the rationale for the lightly anaesthetized preparation, emphasizing that it provides both the recording stability required for prolonged single-unit recordings from sparse neurons in the deep RVM and a controlled physiological setting in which the baseline temporal organization of the circuit can be characterized while minimizing ongoing sensory, motor, and behavioral influences.

      Regarding the rationale for the lightly anaesthetized preparation, these experiments take advantage of the well-validated lightly anaesthetized Sprague-Dawley rat in which much of the foundational data concerning physiology and function of RVM neurons was obtained. This “middle-out” strategy [1] has allowed direct connections between the activity and pharmacology of identified RVM neurons and altered nociceptive behavior. This protocol demonstrably spares the essential links between brainstem pain-modulating neurons and nociceptive transmission pathways. Although the focus here was on ongoing activity, precluding the repeated nociceptive testing needed to link neuronal activity to nociceptive threshold, previous work has demonstrated that ongoing activity of OFF and ON-cells is correlated with nociceptive sensitivity [2] and that alterations in OFF- and ON cell firing in response to pharmacological manipulation and in models of persistent pain states, stress, and sickness have behavioral relevance [3–6,6–24]. Further, conclusions from work in lightly anaesthetized rats have repeatedly been found to be congruent with behavioral observations by other groups in awake rats and mice [19,25–36] and with functional imaging evidence in humans [37–40]. The lightly anaesthetized model has thus established a circuit-level explanatory framework for behavioral findings obtained in several species in multiple laboratories.

      A further consideration for the present study is that the lightly anaesthetized preparation allows us to examine the underlying temporal organization of the RVM under controlled conditions, without the additional factors that would necessarily come into play in an awake animal. Dynamics would inevitably be influenced by ongoing sensory input, behavioral priorities, arousal and other internal state changes. These factors would make it difficult to distinguish the intrinsic dynamics of the descending pain-modulatory system from the effects of the animal’s constantly changing experience.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors hypothesized that “RVM neurons operate across multiple temporal scales, integrating fast responses associated with reflex-linked control with slower fluctuations reflecting ongoing network or state-dependent modulation”. The hypothesis was tested with the established ON/OFF-cell model and probabilistic modeling. The study is conceptually interesting and methodologically sophisticated. The findings build toward the conclusion that pain-control circuits operate across multiple timescales.

      Strengths:

      The use of Bayesian regression and Gaussian process modeling to quantify and characterize recovery dynamics and ongoing oscillatory activity.

      The authors show that slow rhythmic activity appears preferentially in ON- and OFF-cells but not in NEUTRAL-cells, suggesting that the oscillations are related to pain-modulatory circuitry rather than being a generic feature of all recorded neurons.

      The observation that some oscillatory activity is coherent with autonomic measures aligns with broader views of the RVM as a hub integrating nociceptive and homeostatic regulation.

      Some pitfalls are appreciated and discussed by the authors, including the functional significance of slow fluctuations, the influence of anesthetics on global brain-state dynamics, the molecular profiles of the studied ON- and OFF-cells, and the heart rate as a covarying signal of RVM neuronal activity.

      Weaknesses:

      A general weakness is that the work is mostly descriptive and relies on anesthetized preparations. Whether the observed rhythms occur in awake animals and are linked to fluctuations in pain behavior needs to be confirmed in future studies.

      The study measures limited autonomic variables. The causal relationship between “slow fluctuations” and “ongoing physiological state” is unclear and overstated, since the data presented appear correlational.

      ON- and OFF-cells in the RVM are identified by their responses correlated with reflexive activity. The significance of the observed oscillations in spontaneous pain conditions is unclear.

      It is uncertain whether the observed rhythms truly reflect intrinsic RVM organization rather than anesthesia-dependent phenomena; the authors appreciated this pitfall, though.

      The Gaussian process analysis suggests predictability and quasi-periodicity, but predictability alone does not necessarily imply a true biological oscillator.

      Conclusion:

      The results support the authors’ hypothesis. The findings provide a compelling conceptual message about the multiscale organization and dynamics of descending pain-control circuits and encourage further studies on the topic.

      We thank Reviewer R1 for their thoughtful and balanced assessment of our work. We are grateful for the reviewer’s positive evaluation of the conceptual framework, the analytical methodology, and the conclusion that the results support the hypothesis that RVM pain-modulatory neurons operate across multiple temporal scales. We agree with the reviewer’s central assessment that the present study is primarily descriptive and that several important questions regarding the origin and functional significance of the slow dynamics remain unresolved. We also appreciate the reviewer’s emphasis on clearly distinguishing observations directly supported by the data from their mechanistic interpretation. Although the manuscript already acknowledged that the present findings do not establish the mechanistic origin of the slow dynamics, we agree that this distinction could be made more explicit. We have therefore revised the manuscript to clarify that the approximately 5-minute fluctuations represent structured, quasi-periodic activity whose underlying origin cannot be determined from the present experiments. Throughout the manuscript we now explicitly acknowledge that these dynamics may arise from interactions between RVM circuitry and broader physiological or network processes, including anaesthesia-related state modulation, autonomic regulation, or other slow network influences. We also emphasise that the relationship between RVM activity and heart rate is correlational and does not establish a causal interaction. Finally, we now discuss more explicitly the complementary roles of controlled lightly anaesthetized and awake preparations. The present preparation was chosen to characterize baseline RVM dynamics under controlled sensory and behavioral conditions, whereas future awake studies will be important for determining how these dynamics are expressed and modulated during ongoing behavior, sensory experience, and chronic pain.

      (1) In the abstract, the ”timescales” are vaguely stated as ”rapid activation”, ”fast recovery dynamics,” and ”slow dynamics”. Quantifying these expressions with approximate ranges (milliseconds, seconds, tens of seconds, minutes, etc.) whenever possible would benefit readers.

      We thank the reviewer for this helpful suggestion. We have revised the Abstract to provide approximate timescales for the different phases of neuronal activity, distinguishing the rapid stimulus-evoked response (sub-second), recovery dynamics (seconds to hundreds of seconds), and ongoing quasi-periodic fluctuations (approximately 5 minutes). We believe these revisions improve the clarity of the Abstract and better convey the central findings of the study.

      (2) The abstract states that ”Effective pain therapies increasingly target neural circuits...” The connection to therapy is not clarified in the manuscript. A brief statement about how temporal dynamics might influence neuromodulation, analgesic interventions, or chronic pain could strengthen translational impact.

      We appreciate this suggestion. We have revised both the Abstract and Discussion to better explain the potential translational relevance of our findings. Rather than making a broad statement regarding pain therapies, we now briefly discuss how understanding the temporal organisation of descending pain-modulatory circuits may ultimately inform the design and timing of neuromodulatory interventions. We also emphasise that these implications remain speculative and require future investigation.

      (3) To address inter-animal and inter-neuron variability. Are the effects consistent across animals? Are all neurons oscillatory? Are the reported timescales driven by a subset of cells?

      We thank the reviewer for raising this important point. We have expanded the Results and Discussion to clarify the degree of variability observed across neurons and animals. In particular, we now emphasise that the slow quasi-periodic dynamics are not uniformly expressed across all neurons, but rather represent a structured population-level phenomenon with variability in predictability and modulation strength between cells, particularly within the OFF-cell population. We also clarify the consistency of the observed timescales across animals and discuss this variability as an important feature of the underlying circuitry rather than evidence for a single homogeneous oscillatory process.

      (4) Discuss what circuit mechanisms generate the oscillations. Are they driven by inputs from the PAG or intrinsic to RVM?

      We agree that the mechanisms underlying the slow temporal dynamics are an important question. We have expanded the Discussion to consider several possible sources of these dynamics, including intrinsic RVM circuitry, descending inputs from higher-order structures, and broader physiological or brain-state fluctuations. We emphasise that the present experiments cannot distinguish between these possibilities and have revised the manuscript to make this limitation more explicit while highlighting it as an important direction for future work.

      Reviewer #2 (Public review):

      Using electrophysiological recordings in a well-characterized animal model of acute pain, and analytical and modeling methods, the authors show that descending pain-modulatory neurons in the rostral ventromedial medulla (RVM) operate across various timescales. They have both rapid multi-phase responses to noxious stimuli that unfold over tens of seconds, with distinct fast and slow recovery dynamics. Additionally, they generate slow quasi-periodic oscillations with approximately 5-minute periods during ongoing activity. These oscillations are statistically predictable and cell-type specific, demonstrating that descending pain control is organized through structured temporal dynamics that encompass immediate stimulus-evoked responses and slower fluctuations associated with physiological state.

      A novel discovery is a 5-minute quasi-periodic oscillation in ongoing ON- and OFF-cell activity. This oscillation, along with its coherence with heart rate, forms the basis for the claim that descending pain circuits exhibit intrinsic multi-timescale organization. However, it’s crucial to demonstrate that this periodicity is independent of external experimental cycles such as methohexital infusion pharmacokinetics, servo-controlled temperature regulation, or slow autonomic feedback loops, all of which operate on similar timescales. For instance, the 300-second period closely matches typical drug infusion cycling and thermoregulatory feedback intervals. Therefore, heart-rate coherence peaks at multiples of this period could equally reflect a shared external driver rather than intrinsic RVM organization. Although the absence of this cyclic structure in Neutral cells argues against this possibility, the authors might want to explicitly discuss this potential confound.

      The findings are important and novel in that they characterize an intriguing structure in the activity of ON and OFF neurons in the RVM. However, in the absence of a causal manipulation causality can only be inferred. That there is no phase-dependence of withdrawal latency argues against a causal role. The author are encouraged to qualify their conclusions (and their title) accordingly. Because anesthesia can affect global dynamics, this might affect the oscillations reported. Without awake validation, it remains uncertain whether these rhythms reflect an intrinsic property or an anesthesia-induced regime. Again, the absence of oscillations in Neutral cells argues against this possibility, but it is still possible that ON/OFF cells are embedded in different circuits that are affected differently by anesthesia.

      We thank Reviewer R2 for their careful and constructive assessment of our work and for recognising the novelty of identifying structured multi-timescale dynamics in physiologically characterised RVM neurons. We particularly appreciate the reviewer’s thoughtful consideration of alternative explanations for the observed low-frequency temporal structure.

      The reviewer raises an important question regarding the extent to which the approximately 5-minute quasi-periodic dynamics reflect processes generated within descending pain-modulatory circuitry versus broader physiological or experimental influences. As discussed in the original manuscript, the present experiments cannot determine the precise mechanistic origin of these dynamics, and we have revised the Discussion to make this distinction more explicit. We now consider possible contributions from autonomic regulation, thermoregulatory processes, anaesthesia-related state modulation, and other slow physiological influences. We also clarify that the NEUTRAL-cell population argues against a uniform global effect acting similarly across all RVM neurons, but cannot exclude systemic influences that preferentially engage ON- and OFF-cell circuitry.

      We have additionally expanded the rationale for the lightly anaesthetized preparation. This preparation was not used solely for technical convenience. Stable single-unit recordings from physiologically identified ON- and OFF-cells are technically challenging because the RVM is a deep brainstem structure and these functional cell classes are relatively sparse; suppression of spontaneous movement therefore permits substantially greater recording stability over the prolonged epochs required here. Importantly, the preparation also provides a controlled physiological setting in which the underlying temporal organization of RVM activity can be examined while reducing the continuously changing sensory, motor, arousal, and behavioral influences that would necessarily contribute to RVM activity in an awake animal. Awake preparations are essential for determining how RVM neurons respond during ongoing behavior and natural sensory experience, but that is a complementary question to the one addressed here: whether physiologically identified RVM neurons exhibit structured temporal dynamics under controlled conditions.

      We have therefore revised the manuscript to present the lightly anaesthetized preparation as both a methodological choice and an important boundary condition on interpretation. We continue to acknowledge that anaesthesia may influence slow network dynamics, and that future awake recordings will be required to determine how the temporal structure identified here is expressed in the behaving animal. We have also revised the title and several sections of the manuscript to ensure that our conclusions consistently reflect the correlational nature of the data and do not imply mechanistic or causal interpretations beyond those directly supported by the experiments.

      (1) Analyses of many of the ON-cells had longer training windows (> 1 sec) compared to those for the NEUTRAL cells. Could this have reduced the ability to fit and validate periodicity for the latter cell type?

      We thank the reviewer for raising this important point. The difference between ON/OFFand NEUTRAL-cell analyses reflects the available recording durations rather than differences in the Gaussian process fitting procedure. All cell classes were fitted using the same GP model and training strategy; however, some NEUTRAL-cell recordings were shorter (960 s versus 1500 s for ON- and OFF-cells), resulting in correspondingly shorter training segments. Whilst the minimum frequency recoverable from a 960 s training segment is 0.00104 Hz, meaning that the 0.0033 Hz frequency observed in ON- and OFF-cells would be recoverable if present in a 960 s recording. We agree that shorter recordings could, in principle, reduce the ability to estimate slow periodic structure. We have therefore clarified this point in the Methods and Discussion. Importantly, the absence of predictable low-frequency dynamics in NEUTRAL-cells is supported not only by GP prediction performance but also by the independent power spectral analysis, the low latent GP variance, and the near-flat phase-normalised reconstructions, suggesting that the difference between cell classes is not solely attributable to recording duration. To address this concern, we will revise the manuscript to repeat the GP analysis after truncating the ON- and OFF-cell recordings to match the duration of the NEUTRAL-cell recordings. We will also include a 960-second-long simulated NEUTRAL-cell recording with periodic structure, to demonstrate that this would be located by our method if present.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript by Ashworth and colleagues, the authors investigate the temporal dynamics of the rostral ventromedial medulla (RVM), a key output node in a major descending pain-modulation circuit. Using data from extrasellar single-unit recordings of RVM ON, OFF, and NEUTRAL cells in lightly anesthetized rats, the authors’ computational modeling yielded two major findings: (1) heat-evoked ON burst and OFF pause, followed by exponential recovery components in10s of seconds; and (2) ON and OFF cells exhibit periodic fluctuations in 5-minute cycles that are statistically predictable.

      Strengths:

      The manuscript’s concept is innovative, offering the first quantitative analysis of multitimescale dynamics in physiologically characterized RVM pain-modulating neurons. This advances a field that has mostly depended on qualitative or single-timescale descriptions. The authors use contemporary Gaussian process and probabilistic models to capture statistically predictable slow dynamics. The study is further strengthened by identifying ON-, OFF-, and NEUTRAL-type cells using well-established criteria grounded in decades of RVM research. The combination of rapid reflex-related responses and slower ongoing rhythms supports a dual-timescale framework, providing a more integrated understanding of how these neurons may regulate reflex activity and state-dependent processes.

      Weaknesses:

      Several limitations are noted. Incomplete characterization of light anesthesia during recording sessions, such as methohexital stability and clear criteria for identifying “lightly anesthetized” states. While the NEUTRAL cell control is helpful, it does not fully address concerns about circuit specificity or systemic confounds. The findings are male-dominant, which may limit their generalizability. The synchrony between ON and OFF cells was suggested but not directly tested. The heart rate coherence with ON, OFF, and NEUTRAL cell activity results is intriguing but does not fully clarify how these neurons influence heart rate, particularly within the “lightly anesthetized” model.

      We thank Reviewer R3 for their thoughtful and constructive assessment of our work. We appreciate the reviewer’s emphasis on providing additional methodological detail and placing the findings within the context and limitations of the experimental preparation. In response, we have substantially expanded the Methods to provide a more complete description of the lightly anaesthetized preparation, the methohexital infusion protocol, physiological monitoring, and the rationale for the ongoing recording paradigm.

      We have also clarified why this preparation was appropriate for the question addressed here. In addition to enabling stable long-duration single-unit recordings from sparse, physiologically identified neurons in the deep RVM, the lightly anesthetized preparation provides a controlled physiological setting in which baseline temporal dynamics can be characterized while minimizing ongoing sensory, motor, and behavioral influences. We nevertheless acknowledge that anaesthesia may alter slow brain-state dynamics, and we now make this limitation more explicit throughout the manuscript. We have also revised the Discussion to more clearly acknowledge the predominantly male sample, the interpretation of the NEUTRAL-cell population as a comparison group rather than a definitive control for systemic effects, the limitations of inferring synchrony from pseudo-population data, and the correlational nature of the heart-rate coherence analysis.

      (1) How does the lightly anesthetized preparation affect evoked and oscillation activity modeling? Given that cell activities can be highly influenced by the state of sedation and the pharmacology of methohexital, detailing how light anesthesia was achieved and determined can help interpret the limitations of the current model. For example, did the methohexital rate adjustments occur during the ongoing activity period used for GP modeling? What specific criteria defined “lightly anesthetized” beyond stable paw withdrawal latency, such as stable respiratory rate, EMG (reflex vigor?), and core temperature? Given RVM activity coupled to autonomic/thermoregulatory circuits, data on these variables should be reported, or their absence should be acknowledged.

      We thank the reviewer for this important comment. We have substantially expanded the Methods and Discussion to describe both the rationale for the lightly anaesthetized preparation and the criteria used to maintain it.

      The preparation offers both technical and conceptual advantages for the present question. Technically, the RVM is a deep brainstem structure and physiologically identified ON- and OFF-cells are relatively sparse. Prolonged extracellular recordings therefore depend on maintaining stable electrode–neuron contact, which is readily disrupted by spontaneous movement. Light methohexital anaesthesia suppresses spontaneous movement while preserving nocifensive withdrawal responses and the canonical physiological response patterns used to identify ON-, OFF-, and NEUTRAL-cells. Conceptually, the aim of the present study was to characterize the baseline temporal organization of identified RVM neurons rather than to determine which sensory, cognitive, or behavioral events drive their activity in an awake animal. An awake preparation would necessarily introduce continuously changing sensory input, motor activity, arousal, behavioral priorities, and other internal-state variables, all of which are known to influence RVM activity. These are important influences in their own right, but for the present question they would make it more difficult to distinguish underlying temporal structure from activity driven by ongoing experience. We therefore view controlled lightly anaesthetized and awake preparations as complementary: the former is useful for identifying foundational circuit dynamics under controlled conditions, whereas the latter will be essential for determining how those dynamics are modified and expressed during natural behavior.

      This preparation has also been extensively used to establish the canonical relationship between ON-/OFF-cell activity and nocifensive responses, pharmacological modulation of the RVM, and top-down control from structures including the hypothalamus and amygdala, with many of these functional relationships subsequently confirmed in awake behavioral experiments. We have added this context to the revised manuscript.

      With respect to physiological monitoring, core temperature was continuously monitored and maintained at 36–37 °C, heart rate was monitored by EKG, and EMG was recorded to monitor withdrawal responses. Light anesthesia was defined functionally by preservation of a stable nocifensive withdrawal response in the absence of spontaneous movement. Respiratory variables were not recorded, and we now acknowledge this explicitly as a limitation. We have also clarified in Methods that the Methohexital rates were not adjusted during the recording windows used for the gaussian process analysis.

      We therefore agree that the findings must be interpreted within the context of the lightly anaesthetized preparation, but we do not view awake recordings as a direct substitute for the present experiment. Rather, awake studies provide the important next step of determining how the structured dynamics identified under controlled conditions are modulated by sensory experience, behavioral state, and ongoing cognition.

      (2) It is unclear how the absence of slow oscillations in NEUTRAL cells can be used as an internal control for anesthesia and systemic drift. It is unlikely that NEUTRAL cells are identified in every single-cell recording session for them to be used as a consistent internal control. Also, as the authors suggested that the shared modulatory inputs to ON/OFF cells explain the coordinating mechanism for ON/OFF rhythmicity, the lack of rhythmicity or coherence in majority of the NEUTRAL cells may indicate that they do not receive the same modulatory inputs as ON/OFF cells. Would this make NEUTRAL cells insensitive to systemic changes throughout the recording sessions? Do rhythmic vs. non-rhythmic cells differ in location within the RVM?

      We appreciate the reviewer’s important distinction. We agree that NEUTRAL-cells should not be considered a definitive internal control for anaesthesia or systemic physiological drift. NEUTRAL-, ON-, and OFF-cells were not necessarily recorded simultaneously within the same session, and the functional classes may differ in the systemic or modulatory inputs they receive. Our intended inference is therefore narrower: the absence of comparable low-frequency temporal structure in most NEUTRAL-cells argues against a uniform global process that imposes the same temporal pattern on all RVM neurons. It does not exclude anaesthesia-related, autonomic, thermoregulatory, or other systemic processes that preferentially influence ON- and OFF-cell circuitry. We have revised the manuscript throughout to make this distinction explicit and now refer to NEUTRAL-cells as an informative comparison population rather than as a definitive control for systemic influences.

      Indeed, as the reviewer suggests, differential sensitivity to common modulatory inputs could itself contribute to the distinction between ON/OFF- and NEUTRAL-cell dynamics. This interpretation is also compatible with the observation that a subset of NEUTRAL-cells shows low-frequency coherence with heart rate despite lacking the structured approximately 5-minute temporal dynamics observed in the ON/OFF populations.

      We additionally examined the reconstructed recording locations and found no obvious anatomical segregation between neurons showing stronger versus weaker low-frequency structure within the sampled RVM region. We now state this in the revised manuscript. We appreciate the reviewer’s point that there may be locational differences between RVM rhythmic and non-rhythmic cells, which should be addressed in future work; however, determining this would require a substantially larger sample size, for example with multichannel probe recording, for a valid analysis.

      (3) It is important to acknowledge that findings are effectively male-only (77 M and 6 F). Although a recent publication demonstrated that RVM ON and OFF cell activities do not differ substantially on an individual level between male and female rats, sex differences in RVM population dynamics remain unexplored. The current finding may not be generalizable to females.

      We thank the reviewer for highlighting this important limitation. We now explicitly acknowledge in the Discussion that the present dataset is predominantly male (77 males, 6 females) and therefore does not permit meaningful assessment of sex differences in population dynamics. Although previous studies suggest that individual ON- and OFF-cell responses are broadly comparable between sexes, the generalisability of the present findings to female animals remains unknown and should be addressed in future work.

      (4) It was suggested that strong synchrony exists within each functional population (e.g., ON and OFF cells). However, phase-relationship or coherence analyses were lacking. Since there were recoding sessions with > 2 cells/animals, were there enough recordings that contain simultaneous ON/OFF pairs to allow for these analyses?

      We thank the reviewer for this helpful suggestion. We agree that direct analyses of synchrony between simultaneously recorded neurons would provide valuable additional information. However, the number of simultaneous recordings containing identifiable ON/OFF-cell pairs was insufficient to support a robust phase or coherence analysis. We have therefore revised the Discussion to avoid implying that synchrony has been directly demonstrated and instead describe the results as evidence for consistent low-frequency temporal structure across recordings. We also identify direct analysis of synchrony in larger simultaneously recorded neuronal populations as an important direction for future work.

      (5) It was intriguing that the ON-cell population’s ongoing activity shows a predictive structure, while the OFF-cell population does not (Figure 5). However, this interesting asymmetry in ongoing activity between two cell classes was not adequately explained in the discussion. For example, since shared modulatory inputs were proposed as the coordinating mechanism for ON/OFF rhythmicity, how may this difference in ON and OFF rhythm predictivity occur?

      We appreciate the reviewer drawing attention to this interesting observation. We have expanded the Discussion to consider possible explanations for the greater predictability observed in ON-cells relative to OFF-cells. Although both populations exhibited similar dominant timescales, ON-cells displayed larger latent GP variance and more consistent predictive performance, whereas OFF-cells exhibited greater heterogeneity across recordings. We now discuss several possible explanations for this asymmetry, including differences in intrinsic cellular properties, network coupling, or modulation amplitude, while emphasising that the present data do not allow these possibilities to be distinguished.

      (6) The relationship between RVM activity oscillations and cardiac rhythms appears to be covariate but may not support the ”physiologically meaningful” claim with the current analysis. Additional discussion could help clarify the findings of a) how the RVM oscillation period of 300s relates to the heart rate peak/oscillation period of 600s (Figure 6d) and b) how NEUTRAL cells show heart rate coherence but lack rhythmicity.

      We thank the reviewer for this thoughtful comment. We have revised the Discussion to more carefully interpret the heart-rate coherence analysis. In particular, we now emphasise that the observed coherence demonstrates shared low-frequency temporal structure but does not establish a causal relationship between RVM activity and cardiac dynamics. We also discuss the relationship between the approximately 300-s RVM timescale and the broader low-frequency components observed in the heart-rate spectrum, noting the limited frequency resolution available at these timescales. Finally, we expand our discussion of the NEUTRAL-cell results, clarifying that significant coherence in NEUTRAL-cells despite the absence of comparable structured low-frequency firing dynamics is consistent with shared physiological influences acting on multiple cell classes without implying that the slow temporal structure originates within NEUTRAL-cells.

      References

      (1) Noble, D. The Music of Life: Biology beyond the Genome (Oxford University Press, 2006).

      (2) Heinricher, M. M., Barbaro, N. M. & Fields, H. L. Putative Nociceptive Modulating Neurons in the Rostral Ventromedial Medulla of the Rat: Firing of On- and Off-Cells Is Related to Nociceptive Responsiveness. Somatosensory & motor research 6, 427–39 (1989).

      (3) Barbaro, N. M., Heinricher, M. M. & Fields, H. L. Putative Nociceptive Modulatory Neurons in the Rostral Ventromedial Medulla of the Rat Display Highly Correlated Firing Patterns. Somatosensory & Motor Research 6, 413–425 (1989).

      (4) Heinricher, M. M., Haws, C. M. & Fields, H. L. Evidence for GABA-mediated control of putative nociceptive modulating neurons in the rostral ventromedial medulla: Iontophoresis of bicuculline eliminates the off-cell pause. Somatosensory & Motor Research 8, 215–225 (1991).

      (5) Heinricher, M. M. & Kaplan, H. J. GABA-mediated inhibition in rostral ventromedial medulla: Role in nociceptive modulation in the lightly anesthetized rat. Pain 47, 105–113 (1991).

      (6) Heinricher, M. M. & Tortorici, V. Interference with GABA transmission in the rostral ventromedial medulla: Disinhibition of off-cells as a central mechanism in nociceptive modulation. Neuroscience 63, 533–546 (1994).

      (7) Heinricher, M. M., McGaraughty, S. & Grandy, D. K. Circuitry Underlying AntiOpioid Actions of Orphanin FQ in the Rostral Ventromedial Medulla. Journal of Neurophysiology 78, 3351– 3358 (1997).

      (8) Heinricher, M. M., McGaraughty, S. & Farr, D. A. The role of excitatory amino acid transmission within the rostral ventromedial medulla in the antinociceptive actions of systemically administered morphine. Pain 81, 57–65 (1999).

      (9) Heinricher, M. M., McGaraughty, S. & Tortorici, V. Circuitry Underlying Antiopioid Actions of Cholecystokinin Within the Rostral Ventromedial Medulla. Journal of Neurophysiology 85, 280–286 (2001).

      (10) Heinricher, M. M., Schouten, J. C. & Jobst, E. E. Activation of brainstem N-methyl-daspartate receptors is required for the analgesic actions of morphine given systemically. Pain 92, 129–138 (2001).

      (11) McGaraughty, S. & Heinricher, M. M. Microinjection of morphine into various amygdaloid nuclei differentially affects nociceptive responsiveness and RVM neuronal activity. Pain 96, 153–162 (2002).

      (12) Heinricher, M. M. & Neubert, M. J. Neural Basis for the Hyperalgesic Action of Cholecystokinin in the Rostral Ventromedial Medulla. Journal of Neurophysiology 92, 1982–1989 (2004).

      (13) Heinricher, M. M., Martenson, M. E. & Neubert, M. J. Prostaglandin E2 in the midbrain periaqueductal gray produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Pain 110, 419–426 (2004).

      (14) Heinricher, M. M., Neubert, M. J., Martenson, M. E. & Gonc¸alves, L. Prostaglandin E2 in the medial preoptic area produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Neuroscience 128, 389–398 (2004).

      (15) Kincaid, W., Neubert, M. J., Xu, M., Kim, C. J. & Heinricher, M. M. Role for Medullary Pain Facilitating Neurons in Secondary Thermal Hyperalgesia. Journal of Neurophysiology 95, 33–41 (2006).

      (16) Ortiz, J., Heinricher, M. & Selden, N. Noradrenergic agonist administration into the central nucleus of the amygdala increases the tail-flick latency in lightly anesthetized rats. Neuroscience 148, 737–743 (2007).

      (17) Xu, M., Kim, C. J., Neubert, M. J. & Heinricher, M. M. NMDA receptor-mediated activation of medullary pro-nociceptive neurons is required for secondary thermal hyperalgesia. PAIN 127, 253 (2007).

      (18) Ortiz, J. P., Close, L. N., Heinricher, M. M. & Selden, N. R. α2-Noradrenergic antagonist administration into the central nucleus of the amygdala blocks stress-induced hypoalgesia in awake behaving rats. Neuroscience 157, 223–228 (2008).

      (19) Edelmayer, R. M. et al. Medullary pain facilitating neurons mediate allodynia in headache-related pain. Annals of Neurology 65, 184–193 (2009).

      (20) Martenson, M. E., Cetas, J. S. & Heinricher, M. M. A possible neural basis for stress-induced hyperalgesia. Pain 142, 236–244 (2009).

      (21) Heinricher, M. M., Maire, J. J., Lee, D., Nalwalk, J. W. & Hough, L. B. Physiological Basis for Inhibition of Morphine and Improgan Antinociception by CC12, a P450 Epoxygenase Inhibitor. Journal of Neurophysiology 104, 3222–3230 (2010).

      (22) Heinricher, M. M., Martenson, M. E., Nalwalk, J. W. & Hough, L. B. Neural basis for improgan antinociception. Neuroscience 169, 1414–1420 (2010).

      (23) McGaraughty, S., Farr, D. A. & Heinricher, M. M. Lesions of the periaqueductal gray disrupt input to the rostral ventromedial medulla following microinjections of morphine into the medial or basolateral nuclei of the amygdala. Brain Research 1009, 223–227 (2004).

      (24) Rogness, V. M. et al. Descending Facilitation of Nociceptive Transmission From the Rostral Ventromedial Medulla Contributes to Hyperalgesia in Mice with Sickle Cell Disease. Neuroscience 526, 1–12 (2023).

      (25) Smith, D. J. et al. Dose-Dependent Pain-Facilitatory and -Inhibitory Actions of Neurotensin Are Revealed by SR 48692, a Nonpeptide Neurotensin Antagonist: Influence on the Antinociceptive Effect of Morphine1,2. The Journal of Pharmacology and Experimental Therapeutics 282, 899–908 (1997).

      (26) Hurley, R. W. & Hammond, D. L. The Analgesic Effects of Supraspinal µ and δ Opioid Receptor Agonists Are Potentiated during Persistent Inflammation. Journal of Neuroscience 20, 1249–1259 (2000).

      (27) Kovelowski, C. J. et al. Supraspinal cholecystokinin may drive tonic descending facilitation mechanisms to maintain neuropathic pain in the rat. Pain 87, 265–273 (2000).

      (28) Porreca, F. et al. Inhibition of Neuropathic Pain by Selective Ablation of Brainstem Medullary Cells Expressing the µ-Opioid Receptor. Journal of Neuroscience 21, 5281–5288 (2001).

      (29) Zhang, Y. et al. Identifying local and descending inputs for primary sensory neurons. Journal of Clinical Investigation 125, 3782–3795 (2015).

      (30) Franc¸ois, A. et al. A Brainstem-Spinal Cord Inhibitory Circuit for Mechanical Pain Modulation by GABA and Enkephalins. Neuron 93, 822–839.e6 (2017). URL https://www.ncbi.nlm. nih.gov/pmc/articles/PMC7354674/.

      (31) Kim, J.-H. et al. Yin-and-yang bifurcation of opioidergic circuits for descending analgesia at the midbrain of the mouse. Proceedings of the National Academy of Sciences 115, 11078–11083 (2018).

      (32) Nguyen, E. et al. Medullary kappa-opioid receptor neurons inhibit pain and itch through a descending circuit. Brain 145, 2586–2601 (2022).

      (33) Jiao, Y. et al. Molecular identification of bulbospinal ON neurons by GPER, which drives pain and morphine tolerance. The Journal of Clinical Investigation 133, e154588 (2023).

      (34) Nguyen, E., Grajales-Reyes, J. G., Gereau, R. W. & Ross, S. E. Cell type-specific dissection of sensory pathways involved in descending modulation. Trends in Neurosciences 46, 539–550 (2023).

      (35) Fatt, M. P. et al. Morphine-responsive neurons that regulate mechanical antinociception. Science 385, eado6593 (2024).

      (36) Wang, Q. et al. Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain. Nature 1–10 (2026).

      (37) Brooks, J. C., Davies, W.-E. & Pickering, A. E. Resolving the Brainstem Contributions to Attentional Analgesia. The Journal of Neuroscience 37, 2279–2291 (2017).

      (38) Mills, E. P. et al. Brainstem Pain-Control Circuitry Connectivity in Chronic Neuropathic Pain. Journal of Neuroscience 38, 465–473 (2018).

      (39) Mills, E. P., Keay, K. A. & Henderson, L. A. Brainstem Pain-Modulation Circuitry and Its Plasticity in Neuropathic Pain: Insights From Human Brain Imaging Investigations. Frontiers in Pain Research (Lausanne, Switzerland) 2, 705345 (2021).

      (40) Oliva, V., Hartley-Davies, R., Moran, R., Pickering, A. E. & Brooks, J. C. Simultaneous brain, brainstem, and spinal cord pharmacological-fMRI reveals involvement of an endogenous opioid network in attentional analgesia. eLife 11, e71877 (2022).

    1. eLife Assessment

      This important technical study introduces SCOPE, an optics-free spatial reconstruction method based on bidirectional sender and receiver oligonucleotides on barcoded hydrogel beads. By sequencing proximity-encoded chimeric molecules, the authors computationally reconstruct 2D and 3D spatial information at an impressive scale. The technical demonstrations in synthetic bead systems are convincing and establish proof-of-principle that large spatial domains can be reconstructed without microscopy. The methodological advance is clear and the scale is impressive. This work will be of interest to those working on spatial mapping.

    2. Reviewer #1 (Public review):

      Summary:

      Liao et al. present SCOPE (Spatial reConstruction via Oligonucleotide Proximity Encoding), a method for reconstructing spatial organization from diffusion-defined DNA barcode interactions without the use of optical imaging. In SCOPE, hydrogel beads bearing unique DNA barcodes contain both "sender" and "receiver" oligonucleotides. Upon enzymatic release, sender oligos diffuse locally and hybridize to receiver oligos on neighboring beads, forming chimeric molecules that encode spatial proximity. Sequencing these products yields an interaction matrix, which is then used to reconstruct a spatial coordinate map.

      The authors demonstrate reconstruction of synthetic two-dimensional shapes, a large multicolor Snellen eye chart, and the interior surface of three-dimensional molds. The work expands the conceptual and experimental landscape of optics-free spatial sequencing.

      Strengths:

      SCOPE employs bidirectional sender and receiver oligonucleotides on every bead, rather than using asymmetric transmitter-receiver architectures found in other diffusion-based methods. The symmetric design may improve detection sensitivity and reconstruction strategies, and represents a meaningful variation on optics-free spatial encoding.

      A notable strength of this study is the physical scale achieved. The authors reconstruct a Snellen chart spanning approximately 704 mm² and demonstrate molded 3D structures on the order of 75-100 mm³. Although some larger-scale warping is evident, and is discussed as potentially due to non-uniform diffusion, the relative local positioning across these large areas appears impressively accurate.

      The authors extend reconstruction beyond two-dimensional arrays to three-dimensional molded surfaces. This demonstrates that the assay and the computational methods for interpreting proximity graphs can support non-planar spatial relationships, expanding the scope of optics-free spatial inference.

      The revised manuscript also strengthens the interpretation of these three-dimensional experiments by providing additional evidence that the limited recovery of interior regions primarily reflects technical constraints associated with molecular recovery from the hydrogel beads on the interior rather than an inherent limitation of the SCOPE framework.

      Weaknesses:

      Although the method is discussed in the context of spatial genomics and potential tissue applications, it is currently demonstrated only on engineered two-dimensional bead arrays and three-dimensional shapes fabricated in molds. The authors appropriately acknowledge that additional work will be required to establish performance in heterogeneous biological tissues, where diffusion and molecular recovery are likely to be more complex.

      The revised manuscript substantially clarifies the limitations of the current three-dimensional implementation by providing additional discussion and control experiments supporting the interpretation that reduced recovery of interior beads is primarily a technical limitation of the present protocol. Although limited volumetric sampling remains a current constraint of the method, the authors appropriately discuss applications in which surface-resolved reconstruction may still be informative.

      The computational reconstruction workflow is now described in greater detail, including automated parameter selection, fixed versus optimized hyperparameters, and explicit discussion of the manual flattening step required for the largest Snellen reconstruction. The revised manuscript also explains that many anticipated tissue-section applications present a more constrained reconstruction problem than the intentionally challenging proof-of-concept demonstrations presented here, providing additional context for the expected performance of the method in future biological applications.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Liao et al. present SCOPE (Spatial reConstruction via Oligonucleotide Proximity Encoding), a method for reconstructing spatial organization from diffusion-defined DNA barcode interactions without the use of optical imaging. In SCOPE, hydrogel beads bearing unique DNA barcodes contain both "sender" and "receiver" oligonucleotides. Upon enzymatic release, sender oligos diffuse locally and hybridize to receiver oligos on neighboring beads, forming chimeric molecules that encode spatial proximity. Sequencing these products yields an interaction matrix, which is then used to reconstruct a spatial coordinate map.

      The authors demonstrate reconstruction of synthetic two-dimensional shapes, a large multicolor Snellen eye chart, and the interior surface of three-dimensional molds. The work expands the conceptual and experimental landscape of optics-free spatial sequencing.

      Thank you for this accurate summary of the work.

      Strengths:

      SCOPE employs bidirectional sender and receiver oligonucleotides on every bead, rather than using asymmetric transmitter-receiver architectures found in other diffusion-based methods. The symmetric design may improve detection sensitivity and reconstruction strategies, and represents a meaningful variation on optics-free spatial encoding.

      A notable strength of this study is the physical scale achieved. The authors reconstruct a Snellen chart spanning approximately 704 mm² and demonstrate molded 3D structures on the order of 75-100 mm³. Although some larger-scale warping is evident, and is discussed as potentially due to non-uniform diffusion, the relative local positioning across these large areas appears impressively accurate.

      The authors extend reconstruction beyond two-dimensional arrays to three-dimensional molded surfaces. This demonstrates that the assay and the computational methods for interpreting proximity graphs can support nonplanar spatial relationships, expanding the scope of optics-free spatial inference.

      Thank you for highlighting these strengths of SCOPE.

      Weaknesses:

      Although the method is discussed in the context of spatial genomics and potential tissue applications, it is currently demonstrated only on engineered two-dimensional bead arrays and three-dimensional shapes fabricated in molds. It remains unclear how SCOPE would perform in heterogeneous biological environments, where diffusion may exhibit additional non-uniformities. A biological proof-of-concept, even limited in scope, would help define the method's strengths and limitations more clearly.

      We concur with the reviewer that a biological proof-of-concept is a key next step, and that diffusion will be more heterogeneous in this more complex environment. To this end, we are actively working to further develop SCOPE for use in tissue sections, with the goal of capturing transcriptomes, accessible chromatin, and genomes. As part of this work, we also hope to systematically explore a range of tissue permeabilization and tissue clearing approaches to mitigate the impact of heterogeneity on performance.

      The reconstruction of three-dimensional structures lacks strong sampling from volume interiors. This is speculated to be due to several possible factors; however, this limitation constrains the method to reconstruction of volume surfaces rather than comprehensive three-dimensional profiling.

      Thank you for highlighting this important limitation. The 3D reconstructions are indeed constrained by undersampling of volume interiors. We anticipate that this might be addressed via relatively minor adjustments to the protocol, e.g. using light- or base-labile linkers to trigger oligo release, with the expectation that this will improve reaction consistency throughout the volume. However, even if we are unable to resolve this issue, we note that surface-resolved reconstructions may be useful for some goals, e.g. embedding a bead-packed gel within a tissue lumen, such as the gut. This could enable surface beads to capture RNA transcripts from adjacent cells, while bead–bead associations serve to define the surface topology.

      The reconstruction workflow involves multiple preprocessing steps and embedding choices. While these appear to work well for synthetic shapes with known geometry, it is less clear how parameter choices would be made in contexts where ground truth is unknown. Clarifying how reconstruction robustness is assessed without prior knowledge of spatial structure would help readers understand how the method could be practically deployed, particularly in more heterogeneous tissue contexts.

      Thank you for the opportunity to clarify. The computational pipeline used for 2D SCOPE reconstruction is designed to operate on a standardized input format and can be applied to arbitrary datasets without prior knowledge of spatial structure. For example, as shown in Figure 3, both the circle and “swoosh” geometries were reconstructed using the same algorithm and identical initial parameters. While certain hyperparameters are pre-specified (e.g. the number of k-nearest neighbours used to compute the pairwise distance matrix for UMAP), these are fixed across datasets. Other parameters, such as UMAP’s “min_dist,” are selected via an automated heuristic grid search that proceeds without user intervention. The agreement with ground truth in these controlled settings, together with the reproducibility of stochastic reconstructions (see Figure 3E-F), supports the robustness of the approach.

      Importantly, there was one exception. Reconstruction of the Snellen eye chart dataset required a manual step, involving an initial 3D UMAP embedding followed by a 2D projection to “flatten” the result. We suspect this reflects radial non-uniformities in sender/receiver oligo diffusion at larger spatial scales. Addressing such confounders algorithmically by explicitly modelling diffusion heterogeneity represents an important area for future work, with the goal of entirely eliminating the need for manual intervention.

      Finally, we note that these benchmark shapes represent somewhat contrived examples, and the geometries encountered in practice may often be much less complex. For example, in conventional spatial genomics, the geometry consists of a bead monolayer forming a flat, regular surface on a rectangular slide of known dimensions. Regardless of the tissue architecture overlaid on this surface, the reconstruction problem is defined by the bead monolayer itself, inferred through sender-receiver interactions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It would be helpful to further clarify the limitations in interior sampling of three-dimensional structures by providing a more explicit comparison with Qian and Weinstein's volumetric DNA microscopy (UMI-UEI) approach. In particular, do the authors anticipate that the current limitation in SCOPE's interior sampling can be mitigated through experimental optimization, or might this represent an inherent challenge associated with hydrogel-based bead scaffolds relative to substrate-free approaches? A more detailed discussion of this point would help readers understand whether the observed volumetric constraint is technical and potentially solvable, or structural to the platform design.

      We thank the reviewer for this suggestion and agree that a more explicit comparison to volumetric DNA microscopy helps clarify the origin of the current limitation in interior sampling. Based on our experiments to date, we view this constraint as primarily technical and, in principle, addressable.

      A key distinction between SCOPE and volumetric DNA microscopy(Qian and Weinstein, 2025), using the UMI– UEI framework, lies in the recovery of recorded molecules from the interior of the sample. In volumetric DNA microscopy, the hydrogel-embedded specimen can be fully digested and treated with Proteinase K, enabling efficient liberation and recovery of molecules throughout the volume for downstream sequencing(Qian et al., 2026). In contrast, in the current implementation of SCOPE, we do not dissolve the polyacrylamide hydrogel, and recovery therefore relies largely on diffusion of chimeric molecules out of the scaffold. This likely biases against molecules generated in the interior and leads to reduced sampling of internal regions. This interpretation is supported by a control experiment in which barcoded beads were allowed to settle in solution at the bottom of a tube in the absence of a polymerized hydrogel scaffold. In this setting, 3D UMAP reconstruction yielded a solid, non-hollow structure consistent with the expected conical geometry of the tube bottom, indicating that SCOPE is capable of recovering volumetric structure when recovery is not diffusion-limited. Taken together, these observations suggest that the apparent “hollowing” in current 3D reconstructions reflects a limitation in molecule recovery from hydrogel scaffolds, rather than an inherent constraint of the SCOPE framework itself.

      We are currently exploring potential solutions, including the use of reducible crosslinkers to enable hydrogel dissolution and/or mechanical shearing of the gel. If these experiments are successful, we would plan to include the results in revisions to the manuscript, together with an appropriately edited version of the paragraph above. If they are unsuccessful and major experimental effort is going to be required to address this issue, we would likely move forward with textual changes only, incorporating the points made in the paragraph above into the discussion.

      References

      Qian N, Li J, Yasser R, Yu M, Weinstein JA. 2026. Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions. Nat Protoc. doi:10.1038/s41596-025-01329-3

      Qian N, Weinstein JA. 2025. Spatial transcriptomic imaging of an intact organism using volumetric DNA microscopy. Nat Biotechnol 1–11.

    1. eLife Assessment

      This important study investigates how sleep loss and circadian disruption affect whole-organ metabolism in flies (Drosophila melanogaster) and reports that wild-type flies align metabolism in anticipation of diurnal rhythm, while mutant flies with impaired sleep or circadian function shift to reactive or misaligned metabolism. The integration of chamber-based flow-through respirometry with LC-MS metabolomics is innovative, and the significance of the findings is significant. The strength of evidence needed to support the conclusions is generally convincing, although the absence of direct measures of locomotor activity makes it difficult to separate intrinsic metabolic changes from potential behavioral differences in mutants.

    2. Reviewer #1 (Public review):

      Summary:

      This study by Akhtar et al. aims to investigate the link between systemic metabolism and respiratory demands, and how sleep and circadian clock regulate metabolic states and respiratory dynamics. The authors leverage genetic mutants that are defective in sleep and circadian behavior in combination with indirect respirometry and steady-state LC-MS-based metabolomics to address this question in the Drosophila model.

      First, the authors performed respirometry (on groups of 25 flies) to measure oxygen consumption (VO2) and carbon dioxide production (VCO2) to calculate the respiratory quotient (RQ) across the 24-hour day (12h:12h light-dark cycle) and assess metabolic fuel utilization. They observed that among all the genotypes tested, wild type (WT) flies and per0 flies in LD and WT flies in DD exhibit RQ >1. They concluded the >1 RQ is consistent with active lipogenesis. In contrast, the short-sleep mutants fumin (fmn) and sleepless (sss) showed significantly different RQ; the fmn exhibits a slight reduction in RQ values, suggesting increased reliance on carbohydrate metabolism, while sss exhibit even lower RQ (0.94), consistent with a shift toward lipid and protein catabolism.

      The authors then proceeded to bin these measurements in 12-hour partitions, ZT0-12 and ZT12-24, to assess diurnal differences in average values of VO2, VCO2, and RQ. They observed significant day-night differences in metabolic rates in WT-LD flies, with higher rates during the day. The diurnal differences remain in the short-sleep mutants, but the overall metabolic rates are higher. WT-DD flies exhibit the lowest respiratory activity although the day-night differences remain in free-running conditions. Finally, per01 mutants exhibit no significant change in day-night respiratory rates, suggesting that a functional circadian clock is necessary for diurnal differences in metabolic rates.

      They then performed finer resolution 24-hours rhythmic analysis (RAIN and JTK) to determine if VO2, VCO2, and RQ exhibit 24-hour rhythmic and if there are genotype-specific differences. Based on their criteria, VCO2 is rhythmic in all conditions tested while VO2 is rhythmic in all conditions except in fmn-LD. Finally, RQ is rhythmic in all 3 mutants but not in WT-LD and WT-DD. Peak phases for the rhythms were deduced using JTK lag values.

      The authors proceeded to leverage a previously published steady-state metabolite datasets to investigate potential association of RQ with metabolite profiles. Spearman correlation was performed to identify metabolites that exhibit coupling to respiratory output. Positive and negative lag analysis were subsequently performed to further characterize these associations based on the timing of the metabolite peak changes relative to RQ fluctuations. The authors suggest that a positive lag indicates that metabolite changes occur after shifts in RQ, and a negative lag signifies that metabolite changes precede RQ changes. To visualize metabolic pathways that exhibit these temporal relationships, clustered heatmap and enrichment analysis were performed. Through these analyses, they concluded that both sleep and circadian systems are essential for aligning metabolic substrate selection with energy demands, and different metabolic pathways are misregulated in the different mutants with sleep and circadian defects.

      Strength:

      The research questions this study explore are significant given metabolism and respiratory demand are central to animal biology. The experimental methods used, including the well characterized fly genetic mutants, the newly developed method for indirect calorimetry measurements, and LC-MS based metabolomics, are all appropriate. This study provides insights into the impact of sleep and circadian rhythm disruption on metabolism and respiratory demand and serves as a foundation for future mechanistic investigations.

      Comments on revised version.

      The authors have thoughtfully revised the manuscript. They have now provided clarifications regarding perceived conceptual flaws in the original version. They have also performed additional data analysis to support their conclusions and provide clarifications on statistical methods when appropriate. Overall, the revised manuscript is much improved and the conclusions are generally well supported by their results.

    3. Reviewer #2 (Public review):

      This is an innovative and technically strong study that integrates dual-gas respirometry with LC-MS metabolomics to examine how sleep and circadian disruption shape metabolism in Drosophila. The combination of continuous O₂/CO₂ measurements with high-temporal-resolution metabolite profiling is novel and provides fresh insight into how wild-type flies maintain anticipatory fuel alignment, while mutants shift to reactive or misaligned metabolism. The use of lag-shift correlation analysis is particularly clever, as it highlights temporal coordination rather than static associations. Together, the findings advance our understanding of how circadian clocks and sleep contribute to metabolic efficiency and redox balance.

      However, there are several areas where the manuscript could be strengthened. The authors should acknowledge that their findings may be gene-specific. Because sleep deprivation was not performed, it remains uncertain whether the observed metabolic shifts generalize to sleep loss broadly or are restricted to the fmn and sss mutants. This concern also connects to the finding of metabolic misalignment under constant darkness despite an intact clock. The conclusion that external entrainment is essential for maintaining energy homeostasis in flies may not translate to mammals. It would help to reference supporting data for the finding and discuss differences across species. Ideally, complementary circadian (light-dark cycle disruption) or sleep deprivation (for several hours) experiments, or citation of comparable studies, would strengthen the generality of the findings. Figures 1-4 are straightforward and clear, but when the manuscript transitions to the metabolite-respiration correlations, there is little description of the metabolomics methods or datasets, which should be clarified. The Discussion is at times repetitive and could be tightened, with the main message (i.e., wild-type flies align metabolism in advance, while mutants do not) kept front and center. Terms such as "anticipatory" and "reactive" should be defined early and used consistently throughout.

      Overall, this is a strong and novel contribution. With clarification of scope, refinement of presentation, and a more focused Discussion, the paper will make a significant impact.

      Comments on revised version.

      The authors have satisfactorily addressed my concerns in the revised manuscript

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate how sleep loss and circadian disruption affect whole-organism metabolism in Drosophila melanogaster. They used chamber-based flow-through respirometry to measure oxygen consumption, carbon dioxide production, in wild-type flies and in mutants with impaired sleep or circadian function. These measurements were then integrated with a previously published metabolomics dataset to explore how respiratory dynamics align with metabolic pathways. The central claim is that wild-type flies display anticipatory coordination of metabolic processes with circadian time, while mutants exhibit reactive shifts in substrate use, redox imbalance, and signs of mitochondrial stress.

      Strengths:

      The study has several strengths. Continuous high-resolution respirometry in flies is challenging, and its application across multiple genotypes provides good comparative insight. The conceptual framework distinguishing anticipatory from reactive metabolic regulation is interesting. The translational framing helps place the work in a broader context of sleep, circadian biology, and metabolic health.

      Weaknesses:

      At the same time, the evidence supporting the conclusions is somewhat limited. The metabolomics data were not newly generated but repurposed from prior work, reducing novelty. The biological replication in the respirometry assays is low, with only a small number of chambers per genotype. Importantly, respiratory parameters in flies are strongly influenced by locomotor activity, yet no direct measurements of activity were included, making it difficult to separate intrinsic metabolic changes from behavioral differences in mutants. In addition, repeated claims of "mitochondrial stress" are not directly substantiated by assays of mitochondrial function. The study also excluded female flies entirely, despite well-documented sex differences in metabolism, which narrows the generality of the findings.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study by Akhtar et al. aims to investigate the link between systemic metabolism and respiratory demands, and how sleep and the circadian clock regulate metabolic states and respiratory dynamics. The authors leverage genetic mutants that are defective in sleep and circadian behavior in combination with indirect respirometry and steady-state LC-MS-based metabolomics to address this question in the Drosophila model.

      First, the authors performed respirometry (on groups of 25 flies) to measure oxygen consumption (VO2) and carbon dioxide production (VCO2) to calculate the respiratory quotient (RQ) across the 24-hour day (12h:12h light-dark cycle) and assess metabolic fuel utilization. They observed that among all the genotypes tested, wild type (WT) flies and per0 flies in LD and WT flies in DD exhibit RQ >1. They concluded the >1 RQ is consistent with active lipogenesis. In contrast, the short-sleep mutants fumin (fmn) and sleepless (sss) showed significantly different RQ; the fmn exhibits a slight reduction in RQ values, suggesting increased reliance on carbohydrate metabolism, while sss exhibits even lower RQ (0.94), consistent with a shift toward lipid and protein catabolism.

      The authors then proceeded to bin these measurements in 12-hour partitions, ZT0-12 and ZT12-24, to assess diurnal differences in average values of VO2, VCO2, and RQ. They observed significant day-night differences in metabolic rates in WT-LD flies, with higher rates during the day. The diurnal differences remain in the short-sleep mutants, but the overall metabolic rates are higher. WT-DD flies exhibit the lowest respiratory activity, although the day-night differences remain in free-running conditions. Finally, per01 mutants exhibit no significant change in day-night respiratory rates, suggesting that a functional circadian clock is necessary for diurnal differences in metabolic rates.

      They then performed finer-resolution 24-hour rhythmic analysis (RAIN and JTK) to determine if VO2, VCO2, and RQ exhibit 24-hour rhythmic and if there are genotypespecific differences. Based on their criteria, VCO2 is rhythmic in all conditions tested, while VO2 is rhythmic in all conditions except in fmn-LD. Finally, RQ is rhythmic in all 3 mutants but not in WT-LD and WT-DD. Peak phases for the rhythms were deduced using JTK lag values.

      The authors proceeded to leverage a previously published steady-state metabolite dataset to investigate the potential association of RQ with metabolite profiles. Spearman correlation was performed to identify metabolites that exhibit coupling to respiratory output. Positive and negative lag analysis were subsequently performed to further characterize these associations based on the timing of the metabolite peak changes relative to RQ fluctuations. The authors suggest that a positive lag indicates that metabolite changes occur after shifts in RQ, and a negative lag signifies that metabolite changes precede RQ changes. To visualize metabolic pathways that exhibit these temporal relationships, a clustered heatmap and enrichment analysis were performed. Through these analyses, they concluded that both sleep and circadian systems are essential for aligning metabolic substrate selection with energy demands, and different metabolic pathways are mis regulated in the different mutants with sleep and circadian defects.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strength:

      The research questions this study explores are significant, given that metabolism and respiratory demand are central to animal biology. The experimental methods used, including the well-characterized fly genetic mutants, the newly developed method for indirect calorimetry measurements, and LC-MS-based metabolomics, are all appropriate. This study provides insights into the impact of sleep and circadian rhythm disruption on metabolism and respiratory demand and serves as a foundation for future mechanistic investigations.

      We thank the reviewer for the positive comments.

      Weaknesses:

      There are some conceptual flaws that the authors need to address regarding circadian biology, and some of the conclusions can be better supported by additional analysis to provide a stronger foundation for future functional investigation.

      At times, the methods, especially the statistical analysis, are not well articulated; they need to be better explained.

      Thank you for this suggestion and have revised and we expanded the Methods and figure legends to improve transparency and reproducibility.

      Specifically, we have:

      (i) Strengthened the rhythmicity description by specifying that rhythmicity was assessed in Nitecap using the RAIN algorithm with FDR-adjusted p-values (significant p ≤ 0.05; trending 0.05–0.1), and that period and peak phase were estimated with JTK_CYCLE (JTK lag = peak phase), with per-genotype period, phase, and p-values reported in Table 1;

      (ii) Clarified sample size and replication by adding these lines to the methods section and indicating sample sizes in figure legends. Figure legends now report n (chambers) and SEM.

      “Each genotype measurement represents an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). The chamber was treated as the experimental unit for all analyses.”

      In addition, we have expanded the description of metabolomics-respirometry correlation analyses to include dataset structure, time-matching across ZT, normalization steps, the use of Spearman correlations, and interpretation of lagged associations.

      (iii) We have added these lines to the methods section:

      “To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite time points (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Reviewer #2 (Public review):

      This is an innovative and technically strong study that integrates dual-gas respirometry with LC-MS metabolomics to examine how sleep and circadian disruption shape metabolism in Drosophila. The combination of continuous O<sub>2</sub>/CO<sub>2</sub> measurements with high-temporal-resolution metabolite profiling is novel and provides fresh insight into how wild-type flies maintain anticipatory fuel alignment, while mutants shift to reactive or misaligned metabolism. The use of lag-shift correlation analysis is particularly clever, as it highlights temporal coordination rather than static associations. Together, the findings advance our understanding of how circadian clocks and sleep contribute to metabolic efficiency and redox balance.

      We thank the reviewer for the positive comments.

      However, there are several areas where the manuscript could be strengthened.

      The authors should acknowledge that their findings may be gene specific. Because sleep deprivation was not performed, it remains uncertain whether the observed metabolic shifts generalize to sleep loss broadly or are restricted to the fmn and sss mutants. This concern also connects to the finding of metabolic misalignment under constant darkness despite an intact clock.

      We agree that our findings should be framed as genotype- and condition-specific. The phenotypes we report arise from chronic, genetically encoded sleep loss (fmn, sss) and clock loss (per01); because acute sleep deprivation was not performed, we do not claim these effects generalize to sleep loss broadly. This also bears on the reviewer's point about constant darkness: the metabolic misalignment we observe in WT-DD occurs despite an intact clock, and we therefore interpret it as a consequence of removing external light-dark cues under our conditions. We have scoped the claims accordingly in the Abstract, Results, and Discussion (subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness, DD”).

      The text now reads as follows:

      “We restrict our conclusions to the genotypes and conditions tested (fmn, sss, and per<sup>01</sup>), and we do not generalize these effects to acute sleep deprivation because sleep deprivation was not performed in this study. Accordingly, the ‘metabolic misalignment’ observed in constant darkness (DD) likely results from a decrease in synchrony due to the removal of external light:dark cues.”

      The conclusion that external entrainment is essential for maintaining energy homeostasis in flies may not translate to mammals. It would help to reference supporting data for the finding and discuss differences across species. Ideally, complementary circadian (lightdark cycle disruption) or sleep deprivation (for several hours) experiments, or citation of comparable studies, would strengthen the generality of the findings.

      Thank you. We have tempered the interpretation and expanded both the discussion and its citations. We now (i) avoid stating that external entrainment is universally “essential” for energy homeostasis, (ii) explicitly discuss fly-mammal differences (sleep architecture, thermoregulation, feeding control, and entrainment mechanisms), and (iii) anchor the translational comparison to the mammalian circadian-misalignment and sleep-loss literature already integrated in our Discussion (refs [3, 43-46]), noting that establishing cross-species generality will require additional paradigms (constant light, acute sleep deprivation).

      The text now reads as follows:

      “These phenotypes parallel mammalian systems, where sleep loss and circadian misalignment are linked to elevated basal metabolic rate, a shift toward carbohydrate oxidation and lipid/protein catabolism, and blunted, phase-shifted respiratory oscillations [3, 43-46]; physiological differences between flies and mammals nonetheless caution against direct mechanistic extrapolation. In constant darkness, our DD data show that endogenous free-running regulation persists but that removing external light-dark cues degrades temporal coordination between respiration and metabolism; however, we acknowledge that the coupling may be different in mammals.”

      Figures 1-4 are straightforward and clear, but when the manuscript transitions to the metabolite-respiration correlations, there is little description of the metabolomics methods or datasets, which should be clarified.

      Thank you for noting this. We agree that the transition to the metabolite–respiration correlation analyses required clearer description of the metabolomics datasets and processing. We have revised the Methods and the corresponding Results text to briefly summarize the metabolomics dataset parameters and workflow, including how metabolomics and respirometry measurements were time-matched across ZT, the normalization procedures applied prior to analysis, the use of Spearman rank correlations, and how we interpret lagged relationships between metabolite abundance and respiratory outputs.

      The text now reads as follows:

      Methods:

      “Metabolomics-respirometry integration and lag analysis

      To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite timepoints (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Results:

      “Temporal Profiling of Respiratory Quotient in Wild-Type Flies Under Light-Dark Conditions

      RQ values corresponding to each 2-hour Zeitgeber Time (ZT) point were extracted from the 5-minute-binned dataset. Building on this alignment, we explored temporal relationships by systematically shifting the RQ time series by −120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes relative to the metabolite dataset. The continuous respirometry time series showed an oscillatory day-night pattern in RQ; metabolomics was then used to relate time-matched and lagged metabolite dynamics to RQ patterns (Figure 4).”

      The Discussion is at times repetitive and could be tightened, with the main message (i.e., wild-type flies align metabolism in advance, while mutants do not) kept front and center.

      Thank you for this helpful suggestion. We have revised the Discussion to reduce repetition and improve focus by keeping the central takeaway explicit throughout, and by consolidating overlapping paragraphs into a more streamlined narrative.

      We added this revision at the start of the Discussion, in the opening subsection “Temporal Misalignment Alters Fuel Utilization and Respiratory Rhythms.”

      The Discussion now reads as follows:

      “Across the manuscript, the central takeaway is that wild-type flies under LD exhibit anticipatory alignment of fuel selection with time of day, whereas short-sleep mutants (fmn, sss) and clock-disrupted flies (per01) show reactive or misaligned metabolism under our conditions. We therefore focus the Discussion on loss of temporal coordination between respiratory output and pathway-level metabolism, rather than reiterating rate changes alone.”

      Terms such as "anticipatory" and "reactive" should be defined early and used consistently throughout.

      Thank you for this suggestion. We agree and have revised the manuscript to define these terms early (at first use) and apply them consistently throughout. We added this definition in two places:

      (i) In the Results, at the start of the metabolomics-respirometry integration section where we first introduce the lag analysis, and

      (ii) In the Methods, within the paragraph describing the lag analysis workflow, using identical wording.

      The text now reads as follows:

      “We define ‘anticipatory’ as metabolite changes that precede the associated respiratory shift (negative lag) and ‘reactive’ as changes that follow or coincide with the respiratory shift (positive lag), and we use these terms consistently throughout.”

      Overall, this is a strong and novel contribution. With clarification of scope, refinement of presentation, and a more focused Discussion, the paper will make a significant impact.

      We again thank the reviewer for the positive comments.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate how sleep loss and circadian disruption affect whole-organism metabolism in Drosophila melanogaster. They used chamber-based flow-through respirometry to measure oxygen consumption and carbon dioxide production in wild-type flies and in mutants with impaired sleep or circadian function. These measurements were then integrated with a previously published metabolomics dataset to explore how respiratory dynamics align with metabolic pathways. The central claim is that wild-type flies display anticipatory coordination of metabolic processes with circadian time, while mutants exhibit reactive shifts in substrate use, redox imbalance, and signs of mitochondrial stress.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strengths:

      The study has several strengths. Continuous high-resolution respirometry in flies is challenging, and its application across multiple genotypes provides good comparative insight. The conceptual framework distinguishing anticipatory from reactive metabolic regulation is interesting. The translational framing helps place the work in a broader context of sleep, circadian biology, and metabolic health.

      We thank the reviewer for the positive comments.

      Weaknesses:

      At the same time, the evidence supporting the conclusions is somewhat limited. The metabolomics data were not newly generated but repurposed from prior work, reducing novelty.

      Thank you for raising this point. We now make the provenance of the metabolomics dataset explicit in the manuscript. Importantly, the current study uses this dataset in a new analytical context by integrating it with continuous VCO<sub>2</sub>/VO<sub>2</sub> respirometry through timematched and lag-aware analyses. This approach allows us to evaluate dynamic relationships between respiratory output and metabolite profiles across circadian time, which was not addressed in the original metabolomics study. We have clarified this point in the Introduction and Methods.

      The text now reads as follows:

      In the Introduction:

      “To provide a more comprehensive perspective on metabolic regulation, we complemented newly generated respiratory measurements with steady-state metabolomic profiling using liquid chromatography-mass spectrometry (LC-MS) data previously published from our group [27]. This integrative framework enabled timematched and lag-aware analysis of respiratory output and metabolite profiles across Zeitgeber time in the LD cycle….”

      In the Methods:

      “The metabolomics dataset analyzed in this study was previously published and is publicly available, as described in detail in [27, 31]. In the present study, these data were integrated with respirometry measurements to assess temporal relationships between metabolite abundance and respiratory output.”

      The biological replication in the respirometry assays is low, with only a small number of chambers per genotype.

      Thank you for highlighting this concern. We suggest that this is a lack of clarity in our initial description of the design and that the replication structure should be stated more explicitly. We have revised the Methods and all relevant figure legends to clearly report biological replication using the chamber as the experimental unit, including n (number of chambers) per genotype and the associated error structure. We also clarify sampling depth by stating that each genotype measurement reflects an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). This information is now reported consistently to make the unit of analysis transparent.

      We added this clarification in the Methods under “Respirometry Setup” (where chamber loading and experimental design are described) and ensured that each relevant figure legend explicitly reports n (chambers) and SEM.

      The text now reads as follows:

      “Each genotype measurement represents an average of ~300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), with the chamber as the experimental unit; n (chambers) and SEM are reported in each figure legend.”

      Importantly, respiratory parameters in flies are strongly influenced by locomotor activity, yet no direct measurements of activity were included, making it difficult to separate intrinsic metabolic changes from behavioral differences in mutants.

      A detailed timing comparison between behavior (feeding and locomotion) compared to respirometry is given in our master response to Reviewer 1, Major comment 4.

      In addition, repeated claims of "mitochondrial stress" are not directly substantiated by assays of mitochondrial function.

      Thank you. We agree that “mitochondrial stress” requires direct functional evidence. We therefore directly assayed mitochondrial respiration (baseline gut-tissue OCR in fmn and per01 versus iso31 controls), added a new Methods subsection and Figure 9, and reframed our wording from “mitochondrial stress/impairment” to altered (elevated) baseline mitochondrial respiration. The experimental details are now in the Methods and the result in the Results (both quoted below). sss was not assayed, so we removed the functional mitochondrial-stress claim for sss; we retain Vaccaro et al. (2020) as prior support for fmn.

      Methods — new subsection “Gut Tissue Respirometry” now reads: “Oxygen consumption rate (OCR) was measured in dissected gut tissue from iso31, fmn, and per01 flies using the Resipher System (Lucid Scientific, GA, USA). Baseline OCR (fmol/mm<sup>2</sup>/s) was averaged over a 24-hour window following a 12-hour acclimation; values from 3 independent runs were pooled, median-normalized to iso31 within each run, and log2(x+10)-transformed. A single iso31 outlier (third per01 run) was excluded; no other values were removed. Each mutant was compared with iso31 using the MannWhitney test (GraphPad Prism 10), with significance at p < 0.05 (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 after excluding one iso31 outlier).”

      The text has been added to Results:

      “Because pathway-level metabolomics implicated mitochondrial pathways in the sleep and circadian mutants, we directly assayed mitochondrial respiration by measuring baseline oxygen consumption rate (OCR) in dissected gut tissue from fmn and per01 relative to iso31 controls. Both fmn and per01 guts showed significantly elevated baseline OCR (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 across 3 runs; Mann-Whitney test, p<0.05; Figure 9A,B). Because only baseline OCR was measured, we interpret this as altered (elevated) baseline mitochondrial respiration rather than reduced capacity or a specific coupling defect (Figure 9).”

      Discussion- fmn:

      “These interpretations are further supported by gut-tissue respirometry showing elevated baseline mitochondrial respiration in fmn relative to iso31 controls. Together with prior evidence of ROS accumulation (oxidative stress) in fmn (Vaccaro et al., Cell, 2020), these functional data indicate that chronic sleep loss in fmn is associated with altered mitochondrial respiration.”

      Discussion- per01:

      “Gut-tissue respirometry in per01 likewise showed elevated baseline mitochondrial respiration, providing functional evidence consistent with the metabolomic signatures of disrupted redox balance and mitochondrial metabolism.”

      The study also excluded female flies entirely, despite well-documented sex differences in metabolism, which narrows the generality of the findings.

      Thank you for raising this point. We agree that sex is an important biological variable in metabolic regulation. While our Methods state that male flies were collected, we have now made this explicit and unambiguous by stating that only males were used for respirometry and metabolomics integration, and we have added this as a limitation in the Discussion, noting that sex-specific physiology could influence the magnitude and/or timing of the effects we report. We also highlight inclusion of females as an important future direction.

      The text now reads as follows (Methods):

      “Only male flies were used for all respirometry experiments and for integration with the metabolomics dataset. This was done to reduce variability introduced by female reproductive status (e.g., mating/egg production) and associated metabolic differences, enabling a clearer comparison across genotypes and lighting conditions.”

      The text now reads as follows (Discussion):

      “Because only males were analyzed, our conclusions may not generalize to females, which can show sex-specific metabolic physiology. Inclusion of female flies and direct sex comparisons across LD and DD conditions will be an important future direction. More specifically, females carry a higher reproductive and biosynthetic load (egg production) that typically raises metabolic rate and can shift RQ toward lipogenesis and alter the amplitude and phase of diurnal respiratory rhythms; females might therefore show larger or differently-timed effects than the males studied here.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors appear to be using WT-DD as a condition to disrupt circadian rhythm (line 216). Although rhythmicity is often dampened in DD compared to in LD, circadian rhythm is defined as rhythm in a constant condition (e.g., DD) after entrainment. WT-DD is not a condition that the authors should use if they want to disrupt circadian rhythms in flies. WTLL would be a better condition to use since flies become arrhythmic in LL, not DD.

      Thank you for this clarification. We agree that DD is not a circadian-disrupting condition; circadian rhythmicity is defined by rhythms that persist under constant conditions after entrainment. Our intent was not to treat WT-DD as arrhythmic, but to use DD to assess free-running circadian regulation in the absence of external light–dark cues. We have revised the manuscript to clearly distinguish diurnal rhythms under LD from free-running circadian rhythms under DD, and to avoid implying that DD abolishes rhythmicity.

      The Results text now reads as follows:

      “We used constant darkness (DD) to assess free-running circadian regulation in the absence of external light–dark cues.”

      In the Discussion, in the subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness,” we revised the interpretation of WT-DD to clarify that the observed metabolic effects reflect removal of external light–dark cues under our experimental conditions, while avoiding overgeneralization.

      The Discussion text now reads as follows:

      Our DD data show that endogenous free-running circadian regulation persists but that removing external light–dark cues degrades the temporal coordination between respiration and metabolism, indicating that external entrainment normally strengthens this coordination under our conditions; however, we acknowledge that the coupling may be different in mammals.

      (2) The authors need to be more precise in the use of the term "circadian" throughout the manuscript. When describing 24h rhythmicity in the LD condition in flies, they can use the diurnal rhythm. Circadian rhythm is an endogenous rhythm without external time cues (e.g., DD rhythm).

      Thank you for this point. We agree and have revised the manuscript to use terminology consistently: rhythms measured under LD are now referred to as diurnal (LD) rhythms/patterns, and the term circadian is reserved for endogenous free-running rhythms in DD. We updated the Methods, Results, Discussion and figure legends throughout to correct instances where LD rhythmicity was previously labeled as “circadian.”

      We added this clarification in the Methods (“Drosophila Strains, Entrainment and Collection”) and in the Results/figure legends where LD time courses are described.

      The text now reads as follows (Methods):

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).”

      We revised the Results text where WT-DD and LD time courses are described, replacing imprecise references to ‘circadian disruption’ or ‘circadian cycle’ with ‘free-running conditions in DD,’ ‘24-hour cycle,’ or ‘diurnal pattern under LD,’ as appropriate. We also revised the Discussion to avoid describing LD patterns as circadian and to avoid implying that DD disrupts circadian rhythmicity.

      (3) Lines 253-256: The authors' interpretation of this data is not accurate. The authors observed significant day-night differences in VO2 and VCO2 in WT-DD. This suggests there is circadian control over metabolism rhythms. The authors noted there is "limited circadian control".

      Thank you for pointing this out. We agree that the significant day-night differences in VO<sub>2</sub> and VCO<sub>2</sub> in WT-DD support persistent endogenous (circadian) control of respiratory rhythms under constant darkness. We have revised the Results text (lines 253-256) to remove the statement implying “limited circadian control” and instead describe the WTDD effect as maintained rhythmicity with altered amplitude and/or phase relative to LD, rather than loss of rhythmic regulation.

      We added this revision in the Results section under “Diurnal Variation in CO<sub>2</sub> Production, O<sub>2</sub> Consumption, and Respiratory Quotient Across Genotypes” (WT-DD description; lines 253-256).

      The text now reads as follows:

      WT-DD flies, maintained in constant darkness, exhibited the lowest overall respiratory activity. Despite the absence of environmental light cues, VCO<sub>2</sub> and VO<sub>2</sub> retained significant day-night differences (Figure S4), consistent with persistent free-running circadian control in constant darkness (with altered amplitude and/or phase relative to LD).

      (4) Besides sleep, metabolic rates are known to be affected by food consumption. Measuring food consumption of the sleep and circadian mutants might provide insights into whether the metabolic rates are more affected by changes in sleep profile or food consumption. This might also be important given fumin displays impaired dopamine transport function and defective dopamine reuptake, and dopamine is known to affect eating behavior. This was not considered and/or discussed.

      It is technically challenging to measure feeding during respirometry, so we acknowledge it as a limitation. To test whether behavioral timing could explain our results, we compared our respiratory rhythms with the feeding and activity rhythms reported in Malik et al. (2026); the comparison and its interpretation are now in the Discussion (quoted below). This is our master response to the activity/feeding concern and is crossreferenced from Reviewer 3’s public review and Recommendation 1; feeding and activity were not measured for sss.

      We added this clarification in the Discussion (Limitations/confounds) where we address potential behavioral contributors (activity/feeding) to respirometry outcomes.

      The text now reads as follows:

      “Locomotor activity and feeding could not be measured during the respirometry recordings, so genotype differences in respiratory parameters should be interpreted with caution. To assess whether behavioral timing could account for these differences, we compared our respiratory rhythms with the feeding and activity rhythms reported for these genotypes in Malik et al. (2026): wild-type feeding peaked at ZT ~3.25 and fmn feeding was phase-delayed to ZT ~4.5, while fmn also showed elevated locomotor activity, particularly during the dark period. In our data the fmn VCO2 rhythm peaks at ZT ~3.25 (RQ at ZT ~4.25), so the respiratory peak slightly precedes the feeding peak; the phase of the fmn respiratory rhythm is therefore not driven by feeding, although the elevated activity of fmn may contribute to its higher overall metabolic rate. Feeding and activity were not measured for sss.”

      (5) It is not clear whether the authors are simply analyzing the SAME dataset in Figure 1, Figure S4, and Figure 2-3, but just with different resolutions. They need to better articulate this point.

      We thank the reviewer for pointing this out. While the source data for Figures 1, 2-3, and Figure S4 use the same underlying respirometry dataset, they present different analyses to address specific questions. Figure 1 shows the full time-course traces (fine time bins), Figures 2-3 extract rhythmicity metrics (e.g., period/phase) from those same time-series, and Figure S4 collapses the same data into simple day (ZT0-12) vs night (ZT12-24) averages.

      We added this clarification in the Figure legends for Figure 1, Figures 2-3, and Figure S4, and also noted it in the Methods where the respirometry analysis outputs (time-series binning, rhythmicity analysis, and day/night averaging) are described.

      (6) Figure 3 and page 14: What are their criteria for differentiating "conserved" vs "distinct" phase alignment? It is not clear whether this conclusion is supported by any statistical analysis.

      We now specify that the “conserved” vs “distinct” phase descriptions refer to early- vs late-peaking rhythms, and we state the criterion explicitly: a phase was called “conserved” when it fell within ±3 h of the WT-LD peak. The revised text reads: “VO<sub>2</sub> peaked … suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”; “RQ peaked … indicating distinct phase alignment, with the sleep-mutant phases ~8 h apart (Figure 3, Table 1).”

      The text now reads as follows:

      “VO<sub>2</sub> peaked …suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”

      “RQ peaked … indicating distinct phase alignment compared to respiratory output, with phases of the sleep mutants 8 h apart (Figure 3, Table 1).”

      (7) Figure 4: The authors need to provide more details as to how the metabolite dataset was utilized to generate this figure and how they made the conclusion that their analysis "revealed a distinct circadian rhythmicity in RQ, characterized by oscillatory patterns indicative of coordinated substrate utilization across the day-night cycle".

      We have clarified how the respirometry and metabolomics data are used for Figure 4 and corrected the overstated rhythmicity claim:

      (1) The RQ patterning in Figure 4 is derived from the continuous respirometry time series, not from the metabolomics dataset, which is used only for the time-matched and lagged correlation analyses that relate metabolite dynamics to RQ. (2) We removed the statement that the analysis “revealed a distinct circadian rhythmicity in RQ”: RQ was not statistically rhythmic in WT-LD or WT-DD (Table 1), and Figure 4 instead shows the day-night RQ pattern that serves as the reference for the lag-based metabolite correlations. We revised the Methods, the Results paragraph introducing Figure 4, and the Figure 4 legend accordingly.

      The text now reads as follows:

      “Respiratory quotient (RQ) was recorded continuously and averaged into 5-minute bins. To integrate with metabolomics collected every 2 hours, we extracted the corresponding 2-hour ZT RQ values and performed a lag analysis (−120 to +120 min) to relate metabolite dynamics to RQ patterns; rhythmicity of RQ itself was assessed from the respirometry time series.”

      (8) It is unclear why the examples of hydroxyhexadecenoylcarnitine and quinolinate were chosen to be presented in Figure 5a. The authors should clarify their choice of these two examples. Also, the authors should generate a supplemental table with the "several metabolites demonstrating strong correlations (line 294).

      Thank you for this suggestion. Hydroxyhexadecenoylcarnitine and quinolinate were selected as representative examples, and we agree that the metabolites supporting the strong RQ-associated correlations should be provided more explicitly. We have now added a Supplementary Table listing the metabolites demonstrating strong correlations with RQ across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions, using the same selection criterion applied in the heatmap analyses (|ρ| ≥ 0.7, p < 0.05).

      We also revised the Results text near the statement describing strong metabolite-RQ correlations to direct readers to this new table.

      The text now reads as follows:

      Hydroxyhexadecenoylcarnitine and quinolinate are among the strongest positively- and negatively-lagged RQ-correlated metabolites in WT-LD (ρ = +0.78 at +120 min and ρ = −0.77 at −120 min; Supplementary Table 1), illustrating the two opposite lag directions of the workflow. The full set of metabolites showing strong RQ-associated correlations across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions is provided in Supplementary Table 1.

      (9) Although Spearman correlation analysis suggests some correlation between RQ and the two metabolites shown in Figure 5, the correlation shown in Figure 5b does not appear to be compelling. Results shown in Figure 5b do not provide confidence that conclusions based on clustered heatmap analysis shown in Figures 6 to 8 are meaningful. In addition to Spearman correlation, the authors might consider performing additional statistical methods to provide further support.

      Thank you for this comment. We suggest that Fig. 5b was not explained clearly and have clarified. Figure 5b is a lag analysis, not a separate correlation result: it shows how the Spearman correlation changes when the RQ time series is shifted forward or backward in time relative to the metabolite timepoints. The goal is to illustrate lead-lag timing (which shift gives the strongest association), rather than to present a single “strong” correlation as standalone proof.

      To address the concern about confidence in the heatmap-based results (Figs. 6-8), we have strengthened the reporting by providing effect sizes (Spearman ρ) and lag for the metabolite-respirometry associations (now included as Supplementary Table 1). This allows readers to evaluate the statistical support underlying the clustering, beyond the visual patterns in the heatmaps.

      We additionally report multiple-testing–corrected significance for the metabolite–RQ correlations (Benjamini–Hochberg FDR) alongside nominal p in Supplementary Table 1, using the same correction already applied to the pathway enrichment in Supplementary Table 2.

      Minor comments:

      (1) Line 82: The authors should clarify what they mean by "circadian collection". Except for WT-DD, my interpretation is that they collected their samples in LD, so that would not be "circadian collection".

      Thank you for catching this. We agree that “circadian collection” was imprecise. We have revised line 82 to clarify that samples collected under LD were collected across diurnal (LD) time (ZT), and we now reserve “circadian” specifically for collections under constant conditions (DD). We also updated the wording throughout the manuscript to maintain this distinction consistently. We added this clarification in the Methods section “Drosophila Strains, Entrainment and Collection” (line 82).

      The text now reads as follows:

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).

      (2) The authors cited Frayn 1983 to indicate how the RQ value can be used to reflect metabolic fuel utilization. Is this interpretation accepted for all animals, including flies?

      RQ is widely used in indirect calorimetry as an index of relative substrate utilization, including in small model organisms, but we agree that it should be interpreted with appropriate caveats. We have revised the manuscript to clarify that we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.

      We made this change in two places: in the Introduction, where RQ is introduced and Frayn is cited, and in the Methods, under “Carbon Dioxide and Oxygen Analysis and Calculations,” where RQ is defined.

      The text now reads as follows:

      Introduction: “These measurements allow for the estimation of energy expenditure and respiratory quotient (RQ), which can provide an index of relative substrate utilization, with appropriate caveats[13, 14]. In this study, we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.”

      Methods: “RQ was used as an index of relative shifts in substrate utilization over time and between genotypes, interpreted conservatively rather than as a precise measure of absolute carbohydrate versus lipid oxidation as this has not been directly characterized in flies.

      (3) Figure S4: Since the authors are comparing day-night differences, they should plot them in the same graph to make it easier to compare.

      We have revised Figure S4 to plot day and night within the same graph/panel for each metric (VCO<sub>2</sub>, VO<sub>2</sub>, RQ), using side-by-side day vs night groupings per genotype to facilitate direct visual comparison, and we updated the legend accordingly.

      (4) Line 252: When comparing diurnal differences, the authors should not use the word "rhythm". Pattern or profile might be a better word to use.

      We appreciate the suggestion. Where we use “rhythm” we refer specifically to 24-hour oscillations established statistically by JTK_CYCLE and RAIN (Figures 2-3, Table 1); for the coarser day-versus-night comparisons we agree “pattern” or “profile” is preferable and have adopted it there. We have gone through the manuscript to apply this distinction consistently.

      (5) Figure 6a: larger font labels are necessary for the metabolites.

      Thank you. We have revised Figure 6a to increase the metabolite label font size for improved readability.

      Reviewer #3 (Recommendations for the authors):

      (1) Activity controls: To strengthen the paper, include or reference direct measures of locomotor activity (e.g., DAM system). This would allow the separation of metabolic changes from behavioral differences and would enable better analysis of circadian patterns.

      This activity/feeding confound is addressed in full in our response to Reviewer 1, Major comment 4; we cross-reference it here to avoid repetition.

      (2) Mitochondrial function: "mitochondrial stress" should be supported by additional assays such as mitochondrial enzyme activities, high-resolution respirometry, or reactive oxygen species measurements.

      See our full response to the mitochondrial point in Reviewer #3's public review above, including new Figure 9 and the Gut Tissue Respirometry Methods.

      (3) Sex differences: Provide a clear justification for excluding female flies. If feasible, incorporate female data or explicitly discuss how sex differences could alter metabolic outcomes.

      This is addressed in the public review for reviewer 3.

      (4) Statistical presentation: Increase n. Clarify in figure legends whether error bars represent SD or SEM and ensure consistency across all figures (Figure legends).

      We have revised all figure legends to clearly state that error bars represent SEM and ensured this is applied consistently across all figures. We also explicitly report the n for each genotype (with chambers as the experimental unit) in the relevant legends.

      The text now reads as follows:

      “Error bars represent SEM, and n denotes the number of chambers (experimental units) per genotype.”

      (5) Sample size and replication: Indicate more clearly that the chamber, not the individual fly, is the experimental unit. Discuss limitations of replication and statistical power in the text.

      Thank you for this comment. We have clarified throughout the Methods and figure legends that the chamber (not the individual fly) is the experimental unit. We now state that each genotype measurement reflects an average of 300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), and we report n (number of chambers) and the error structure (SEM) for each genotype.

      Added to Discussion:

      “We also acknowledge the limits of this replication: with n = 12 chambers per genotype, statistical power to detect small-magnitude differences and subtle phase shifts is limited, and negative calls (e.g., arrhythmicity) should be interpreted with this caveat.”

      (6) Writing and clarity: (a) Streamline the Discussion to focus on mechanistic themes (anticipatory vs reactive alignment, substrate shifts, redox imbalance).

      (b) The opening phrase ("Precise temporal regulation of metabolism by sleep and circadian rhythms is essential for dynamic energy homeostasis") is dense, vague, and difficult to interpret. Consider rephrasing to something more concrete, for example: "Sleep and circadian rhythms tightly control when and how the body uses energy, but we do not yet know exactly how this timing connects to oxygen use and breathing needs."

      Thank you for these helpful suggestions. We have revised the Discussion to reduce repetition and improve focus by organizing it around the key mechanistic themes raised by our data, including anticipatory vs reactive metabolic alignment, substrate-use shifts, and redox/mitochondrial imbalance. We revised the opening sentence of the Abstract to make the biological question more concrete and accessible, following the reviewer’s suggestion.

    1. eLife Assessment

      This study presents fundamental results on questions related to the presence/absence of the Entner-Doudoroff pathway in cyanobacteria. In contrast to an earlier study, compelling evidence is given that Synechocystis PCC 6803 lacks both an Entner-Doudoroff pathway and a related bypass but contains a promiscuous aldolase. This study successfully reconciles data from different studies and lessons learned from a previous misconception.

    2. Reviewer #1 (Public review):

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED, actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations under the light of more recent findings, and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause for the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; given that currently, genome sequencing is an affordable technique, this kind of mistakes can now be easily prevented by confirming the correct generation of the mutant by DNA sequencing, as proposed by the authors in a recently published preprint (Theune et al, bioRxiv 10.64898/2026.04.08.717167).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but certainly would confer additional robustness to the enzymatic results.

      Comments on revised version.

      The authors have provided satisfactory replies to all my suggestions and corrections, and I have no further changes to suggest.

    3. Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data of different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data is presented to reconcile contradicting data of earlier studies and to serve as basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetitions, and nevertheless misses some specifications.

      Comments on revised version.

      The materials and methods section has been significantly improved. The revised manuscript is now recommended for publication as it is.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Some of the authors proposed in a PNAS paper in 2016 the occurrence of the EntnerDoudoroff (ED) pathway in cyanobacteria and plants, on the basis of several lines of biochemical and genetic evidence. However, more recent results indicated that one of the two specific enzymes of the ED pathway (EDD) is missing in Synechocystis PCC 6803. The authors carried out additional experiments, which demonstrated that EDD is missing, and one of the enzymes (ED aldolase) is a promiscuous enzyme which seems to be involved in proline metabolism and is not actually participating in the ED pathway as initially believed. The results described in this paper are strong evidence that this new interpretation is appropriate, and therefore, it corrects the previous proposal, providing an honest description of the reasons why the authors had reached the wrong conclusion about the existence of the ED pathway in cyanobacteria and plants.

      We thank Reviewer 1 for the summary and comments. Based on our finding that EDA is a promiscuous aldolase that, in addition to the cleavage of KDPG to GAP and pyruvate (a reaction of the ED pathway) catalyzes other reactions in vitro, we proposed potential in vivo functions of EDA, including its involvement in proline metabolism. However, these assumptions require further experimental testing. We do not yet have definitive findings regarding the function of the promiscuous aldolase EDA in Synechocystis in vivo, but respective studies are currently underway.

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations in light of more recent findings and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides, they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause of the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; this could have been easily prevented by confirming the correct generation of the mutant by DNA sequencing.

      We thank the reviewer for this kind comment. We agree that biochemical characterization of enzymes as well as DNA sequencing to check deletion mutants, are important and valuable tools. As outlined in the manuscript and additionally in more detail in a recently submitted article, which is available at bioRxiv (Theune et al. 2026, doi: https://doi.org/10.64898/2026.04.08.717167) and is currently under review at PLOS One, we suggest that genome sequencing of deletion mutants in combination with complemented strains as controls are required to minimize the risk of misinterpretation based on secondary mutations (1). During the early stages of our research on the ED pathway, and later as well when we were already trying to resolve the conflicting results that had accumulated concerning the ED pathway, genome sequencing for Synechocystis mutants was not affordable as a routine procedure (2-4). Therefore, we could not have easily prevented this misconception based on this technique at that time. However, we strongly encourage genome sequencing of deletion mutants (in combination with complemented strains) as routine procedures these days (1).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      As stated above and in the manuscript, we agree that the in vivo role of EDA requires further experimental work which is in progress. However, our results demonstrate that EDA splits KDPG into GAP and pyruvate in vitro, but we assume that this reaction does not play a role in vivo due to the absence of its substrate.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but it certainly would confer additional robustness to the enzymatic results.

      Synechocystis is easily amenable to genetic manipulation, and we agree that expression and purification of all enzymes from this host would have been ideal. However, the first characterization of Synechocystis EDA was performed with proteins that were purified from Synechocystis and showed activity on KDPG at comparable rates as proteins that were purified from E. coli in this study (2). Moreover, most biochemical characterizations of EDAs from archaea, bacteria and plants were performed after recombinant expression in E. coli and yielded highly active enzyme as in the case of Synechocystis is this study (5-7). Therefore, we currently have no reason to worry that the expression in E. coli might affect the enzymatic activity of EDA. The main reason for utilizing E. coli as an expression strain in this study was to gain higher yields of protein for in-depth analyses.

      Bibliography:

      I think the list of papers used in this manuscript is complete and up to date. However, I do miss recent papers which addressed one aspect that was proposed in the original 2016 PNAS paper: the authors wrote, "We therefore suggest that Prochlorococcus might oxidize glucose via the ED pathway under mixotrophic conditions, as shown for Synechocystis." Recent studies checked this hypothesis and have shown that the ED pathway seems to be also missing in Prochlorococcus and marine Synechococcus, and I think this manuscript is a good place to cite them, since these results are consistent with the findings of this paper.

      We will include a references from Moreno-Cabezuelo et a. 2023 (DOI: 10.1128/spectrum.03275-22) in which the proteomes of three marine Prochlorococcus and three marine Synechococcus strains were investigated upon exposure to glucose (8). Protein levels of EDA were either downregulated or not affected while proteins involved in OPP pathway and CBB cycle were upregulated. The authors of this study conclude that this indicates that the latter processes rather than the ED pathway are involved in photomixotrophy in these strains. However, flux analyses are still missing.

      Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner-Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data from different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data are presented to reconcile contradicting data of earlier studies and to serve as a basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetition, and nevertheless misses some specifications.

      We thank Reviewer 2 for the kind summary and comments and will improve the materials and methods part accordingly in a revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Some additional aspects that could be improved:

      (1) L182-184: Are there any known in vitro attempts to determine whether some DHADs can accept 6PG as substrate? If not, did the authors check this possibility in the lab?

      To the best of our knowledge, as mentioned in the manuscript, some DHADs are tested for their activity towards gluconate and some other substrates but not 6PG. It was suggested that since gluconate is smaller it might fit in the catalytic site of DHADs normally occupied by DHIV (7). We discussed this point in the manuscript (Line 177-185).

      In this study we tested the DHAD Slr0452 from Synechocystis and DHAD from Synechococcus with 6PG as substrate but the enzyme did not catalyze 6PG dehydration. The result for Slr0452 is shown in Figure S3.

      (2) L234-241: This paragraph shows how important it is to avoid relying only on sequence alignments to assign functions to proteins (and enzymes in particular). The authors stress this idea elsewhere in the paper, but I think it should receive even more attention in a manuscript like this. Physiological characterization of the protein function is paramount, especially in the case of enzymes. Also, the GDH1 overexpression mutant was done in E. coli; this adds an additional layer of uncertainty, since the protein processing in E. coli might not be entirely identical to that carried out in cyanobacteria, as I mentioned above. This, in turn, could be one of the reasons for not finding the expected enzymatic activity. This comment is also relevant for the results shown in Table 1 (page 12).

      As outlined following lines L234-241 we tested crude cell extracts from Synechocystis WT, a Synechocystis strain overexpressing putative GDH1 (Sll1709) and a Dzwf deletion mutant which can be assumed to upregulate a GDH/GK bypass if present in Synechocystis for GDH activity but did not find any. This strongly indicates that sll1709 does not code for an active GDH in Synechocystis. We thereafter also overexpressed GDH1 in E. coli and again could not detect any GDH activity. In case of GDH1, enzyme activity was therefore tested both in Synechocystis and E. coli and yielded similar results.

      (3) L309-310: "we currently have no explanation for the gluconate that was detected in previous IC-ESI-MSMS measurements in Synechocystis". This is a serious issue, given how important this evidence was for the conclusions of the 2016 PNAS paper and the hypothesis of the ED pathway in cyanobacteria. I would suggest that the authors provide some possible explanations, even if it is based on studies from other teams.

      It would be very speculative and therefore in our eyes not helpful to search for explanations in this case as indeed the measurements were done in another lab. One explanation might be that 6P gluconate got dephosphorylated and yielded gluconate as an artifact. However, as we have no experimental validation for this idea and furthermore cannot test it. We therefore prefer to not further comment on this aspect.

      (4) L347: The presence of an insertion in the sequence of zwf in the ∆gnd mutant is a welcome explanation for the observed results: abolition of 6PG production in this mutant. This is another important message to stress in this manuscript: construction of mutants should always be double checked by DNA sequencing of the relevant genomic regions, to ensure that these kinds of side problems do not appear, leading to confusing results.

      We agree with the reviewer and would get even one step further and rather suggest combining whole-genome sequencing with complemented mutants as it is difficult to know all relevant genomic regions that should be sequenced. We discuss this issue in even more detail in another manuscript that is currently available as online preprint in bioRxiv and is under review (9). This work was now added as a citation in line 566 and the end of the following statement (line 564-566): Routine complementation of deletion mutants and sequencing of selected genes or the entire genome are effective means of identifying secondary mutations that can lead to misleading phenotypes (9).

      (5) L418-419 and Table 2: The results observed for the authors (i.e., that EDA could also catalyze reactions with OAA and KHG, albeit with substantially lower catalytic efficiency than with KDPG), is a matter for concern: if their hypothesis is correct (meaning that this EDA is fundamentally involved in the PEP-pyruvate-OAA node and/or proline metabolism, and not with the ED pathway), then why should it keep in the evolution of these organisms such a strong preference for KPDG, when it is not being used physiologically for the ED pathway)? Furthermore, the Km for KDPG is lower than for OAA or KHG, leading to a very big difference in the Kcat/Km values.

      To solve these questions further respective studies are underway. As EDD is absent from Synechocystis no KDPG should be available in the cells so that catalytic activity on KDPG should be irrelevant in vivo. The in vivo role of Eda requires further clarification.

      Hereafter, I will mention some aspects, following the instructions of eLife, which are related to suggestions for improved experiments/data/analyses, improvements of writing and presentation, and minor corrections to text/figures.

      (1) L81: I would modify the text to "Accordingly, this raises further questions...".

      Thanks for this hint. We modified the text accordingly.

      (2) L189: Add "pages" after "following".

      Thanks for pointing this out. We replaced “following” by “below”.

      (3) Page 7: The whole beginning of the Results section is actually more discussion than description of results, and the first mention of figures appears in L207 of page 8. Given the content of the paper, I think the authors might reconsider using "Results and Discussion" rather than different, specific sections for Results and Discussion. This is one of the papers where I think the combined use of both makes sense and will allow an easier understanding of the message.

      We thank the reviewer for this valuable suggestion and changed the heading to Results and Discussion.

      (4) L181: Add reference regarding the llvD/EDD superfamily.

      We added the following references in lines 172-176 and 185-189:

      (1) Melse, O., Sutiono, S., Haslbeck, M., Schenk, G., Antes, I., & Sieber, V. (2022). Structure Guided Modulation of the Catalytic Properties of [2Fe− 2S]-Dependent Dehydratases. ChemBioChem, 23(10), e202200088.

      (2) Ren, Y., Vettenranta, E., Penttinen, L., Jänis, J., Rouvinen, J., & Hakulinen, N. (2025). The engineered dimer of L-arabinonate dehydratase from Rhizobium leguminosarum bv. trifolii: The role of intersubunit interactions in IlvD/EDD family. Biochemical and Biophysical Research Communications, 757, 151610.

      (3) Ahmed, H., Ettema, T. J., Tjaden, B., Geerling, A. C., Van Der Oost, J., & Siebers, B. (2005). The semi-phosphorylative Entner–Doudoroff pathway in hyperthermophilic archaea: a reevaluation. Biochemical Journal, 390(2), 529-540. ff

      (4) Bräsen, C., Esser, D., Rauch, B., & Siebers, B. (2014). Carbohydrate metabolism in Archaea: current insights into unusual enzymes and pathways and their regulation. Microbiology and Molecular Biology Reviews, 78(1), 89-175.

      (5) Figure 2: The data shown in column plots in Fig 2A, B and C, and 4B, could be presented in tables, which would save space while providing the same information: basically, very little/no activity in some cases vs high levels of activity in others.

      We would like to keep the column plots as we still think that they visualize our data well.

      (6) L255 "Unfortunately, we were not able to overexpress putative GDH2". It would be interesting to give more details about the possible reasons for this fact.

      We tested different growth conditions for recombinant GDH2 expression. The expression culture was incubated at 37 °C for 3 hours as well as overnight at 18 °C for overnight after induction. Both experiments did not resolve the expression problem.

      (7) L409-410: I think this sentence should include a brief part explaining the kind of essay used to test this activity.

      We added now that the LDH-coupled continuous assay was used (see line 403).

      (8) Figure 6E: Please give the specific activity in U/mg, as in Fig 6F, instead of percents.

      100% is given in U/mg units in the figure legend as “control without effector (100 %; specific activity of 4.3 U/mg)”. For easy comparison of effectors, the relative activity (%) is often used. We would therefore prefer to keep the current data presentation.

      (9) L576: Provide the origin of the utilized PCC 6803 strain, given there is a certain level of variability in this strain (glucose tolerance, etc). Also, even if there are some methods which are very widely used, I think the Materials and Methods section should either properly describe them or else cite the source. For instance, BG11 medium is mentioned, but no further information is given.

      We included the information that the glucose-tolerant Synechocystis strain was utilized and added the receipt of and a citation for BG11 medium (10).

      (10) L582 Generation of mutants: This section mentions the Gibson Assembly cloning method, but I miss further information to allow the reader to reproduce the methodology with as many details as possible, or at least cite papers which do so.

      We added a reference in which Gibson Assembly is described (11). Together with the primers listed in Table S3 the generation of mutants is now reproducible.

      (11) L609 Please give information in g, not rpm, for centrifugation. Also, mention the model and brand of the centrifuge and rotors used. Also, immunoblotting is very loosely described. This is also valid for other sections, for instance, L618, L636.

      We now added the following information: Cells were harvested by centrifugation at an RCF (relative centrifugal force) of 3,992 x g in a Beckman Coulter with a JLA-8.1000 Rotor for 20 minutes at 4°C. We now added a reference (12) in which immunoblotting is described in more detail.

      (12) L613: Describe the "small scale purification".

      We now added the information that the small-scale purification was performed using a 50-ml aliquot of the large culture which was treated as described below for the remaining sample.

      (13) L619-620: Describe the composition of the lysis buffer.

      The composition of the lysis buffer is already described as follows: lysis buffer (50 mM NaPO<sub>4</sub> pH=7.0; 250 mM NaCl; 1 tablet complete protease inhibitor EDTA-free

      (Roche) per 50 mL)

      (14) L691: Specify which amounts of auxiliary enzymes in coupled enzymatic assays were used.

      Thanks for pointing this out. We have now integrated the information that 1U of each of the auxiliary enzymes was utilized in the coupled enzymatic assays.

      (15) L716: The authors mention several times using a "double beam spectrophotometer". Please provide the model and brand.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany).

      (16) L717 and 718: define "∆absorption".

      In line 715, the information is given that absorption was measured at 340 nm, "∆absorption" is accordingly the ∆absorption at 340 nm. This information was added.

      (17) L723: "Synechocystis" should be in italics.

      Synechocystis is now written italics.

      (18) L749-759: This section should be described in more detail: preparation of protein extracts, SDS, immunoblotting, etc, or cite references of the same team where these methods were properly described.

      In this section the listed methods are already described in detail.

      (19) 798-799: "frozen cell pellets". Please provide numbers to specify the amount of material used.

      Thanks for pointing this out. We now added the information that frozen cell pellets with a wet weight of 2.4 g wet weight were resuspended.

      (20) L871-872: "It was ensured that auxiliary enzymes were not rate-limiting. One unit (1 U) of enzyme activity is defined as 1 µmol substrate consumed or product formed per minute" is repeated several times in the manuscript (L 907-909, L936-938). I would advise using it the first time, and on other occasions, refer to the same conditions as described above.

      We have accordingly circumvented the repetition of 1 U definition from the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Interpunctuation, especially comma placement, should be improved.

      We improved interpunctuation, especially comma placement to the best of our knowledge.

      (2) Line 63: delete the first "which".

      “Which” was deleted.

      (3) Lines 81/82: revise sentence.

      We revised the sentence to: Accordingly, this raises further questions about the presence of the ED pathway in cyanobacteria and plants.

      (4) Line 164: "presumed" instead of "presumes".

      The word was changed accordingly.

      (5) Line 228: by others? especially in reference 1?

      We deleted by others as the reference is given.

      (6) Line 240: "or" instead of "no".

      “no” was replaced by “or”

      (7) Figure 2: The axes are not well visible, and the explanation for the positive control in panel B is missing in the legend.

      Axes from figures 2, 3 and 4 were enlarged. For Figure 2B the following information was added in the legend: As a positive control, 0.05 U glucose dehydrogenase from Pseudomonas sp. was added to Δzwf cultures and to purified putative GDH1 (Sll1709). Axes from figures 2, 3 and 4 were enlarged.

      (8) The investigation on the general absence/presence of the GDH/GK bypass in cyanobacteria may not be necessary for this study.

      We included this data in this manuscript as the mistaken assumption that the GDH/GK bypass exist in Synechocystis lead among other observations to the misinterpretation of an existing ED pathway in Synechocystis. We would therefore prefer to keep these data in the manuscript.

      (9) Line 316: delete "on".

      “on” was deleted.

      (10) Line 352: delete "or".

      “or” was deleted.

      (11) Lines 352/353: ZWF expression level appears to be reduced accordingly. This should be stated.

      We agree that Zwf expression might be lower, however, we are not entirely sure if this is truly valid and would rather test this assumption further as described in the following lines.

      (12) Figure 4: The axes are not well visible.

      Axes from figures 2, 3 and 4 were enlarged.

      (13) Line 363: values are not only normalized to protein content, but also to activity found for the WT.

      We now added: The values are normalized to Zwf enzyme activity found in the WT based on protein content.

      (14) Line 408: no separate subsection required.

      The title for a new subsection was deleted.

      (15) Lines 437-439: These are results descriptions, which should not be placed in the legend, but in the main text, as is partially done.

      We deleted these result descriptions in the legend.

      (16) Lines 441/442: formatting: one or no bracket pair.

      The brackets were corrected.

      (17) Lines 443/444: refer to Table 2 instead of giving the values in the legend to avoid duplication.

      We deleted the values and now refer to Table 2.

      (18) Line 503: delete "identified".

      We deleted the second “identified” in the sentence and changed the wording to: Apart from four identified cyanobacteria that possess potential EDDs. In addition, we also added the names of the four cyanobacteria that were found including the sequence IDs of the putative EDDs.

      (19) Lines 529/530: revise sentence and format.

      We added one sentence and revised the following sentence: In contrast to Synechocystis EDA, EDA from Synechococcus prefers OAA over KDPG. The catalytic efficiency of Synechococcus EDA on oxaloacetate is rather low (OAA 0.437 s<sup>-1</sup> mM<sup>-1</sup>), however, its activity can be enhanced by NADP<sup>+</sup>(13).

      (20) Line 542: revise sentence.

      We revised the sentence to: It remains to be investigated whether this reaction could play a role in vivo, with KDPG potentially acting as a regulatory metabolite at low concentrations.

      (21) Line 577: The glass tubes used for cultivation should be specified.

      We now added the following information: Custom-made glass tubes were placed in a photobioreactor (manufactured by Willi Hilke, Uslar, Germany).

      (22) Lines 584-585: unclear, was the resistance cassette not placed in the gene to be deleted?

      Yes, the resistance cassette was placed in the gene to be deleted and was fused for homologous recombination to two DNA fragments approximately 200 bp directly upstream and downstream of the gene. This information was now added.

      (23) Line 607: cultivation equipment to be specified.

      The following information was now added: For the purification of GDH1 from Synechocystis, a 6 L photoautotrophic culture of the P3-His-GDH1 overexpression strain was grown in a 10 L glass flask at 28°C, illuminated with constant light (50 µmol m<sup>-2</sup> s<sup>-1</sup>) and gassed with filter sterilized ambient air to an OD<sub>750</sub> of about 1.

      (24) Line 670: GTS should be specified.

      Thank you for this hint. This was a typo. GST was meant not GTS. This was now corrected and GST was specified as Glutathione S-Transferase.

      (25) Line 679: delete "gluconate kinase (GK) and".

      The second gluconate kinase (GK) was deleted and sentence was revised to:

      For gluconate kinase (GK) activity measurements in Synechocystis crude cell extracts the GK reaction was enzymatically coupled to 6-phosphogluconate dehydrogenase (GND) reaction, the latter providing NADP<sup>+</sup> reduction, which was monitored photometrically at 340 nm.

      (26) Line 686: again GK activity? Difference unclear. Was the previously described procedure for GND activity determination?

      GK activity measurements in Synechocystis crude cell extracts and GK activity measurements with recombinant enzyme that was expressed in E. coli were done in two different labs with different protocols. Therefore, the first description refers to measurements with Synechocystis while the second measurement refers to measurements with E.coli. This is now specified more clearly.

      (27) Type/supplier of spectrophotometers and centrifuges used should be given.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany). As this study was performed in two different labs over a period of 8 years including one lab moving to a new location, it is now impossible to specify all centrifuges that were utilized. Even though we agree that it would be good to provide this information, we now would have difficulties to be specific.

      (28) Consider the referencing of published methods to streamline the materials and methods section.

      We now streamlined the materials and methods section by deleting repetitions as outlined below. However, as protein expression, protein purification and enzymatic tests were performed in different labs, in some cases several protocols are given.

      (29) Line 757: give specifics of anti-rabbit antibody and define PBS-T and PBS-T Cytiva.

      Specifics were added to the text.

      (30) Lines 761ff: It is not given for all genes used how they were derived. All synthesized?

      We now added detailed information for all genes.

      (31) Line 762: codon-optimized for? E. coli?

      The information was added that genes that were expressed in E. coli were codon-optimized for E. coli.

      (32) Lines 782-786: Rationals for experimental strategies do not belong to materials and methods sections, but to results sections.

      The part was deleted in the materials and methods section and transferred to the results section.

      (33) Lines 818/819: repetitive.

      We removed the repetition and refer to the purification method as stated above in the materials and methods section.

      (34) Lines 847-851: True for all EDA-type assays? Kinetic parameters are shown in Table 2 rather than Table 1.

      Yes, true for all EDA-type assays. We changed the Table number to 2.

      (35) Line 863: delete "in".

      “in” was deleted

      (36) Lines 888-893: sounds repetitive.

      The lines were modified accordingly.

      (37) Lines 908/909: repetitive.

      The repetitive comment on the definition of 1U was deleted.

      (38) Lines 928-938: repetition

      The repetition was deleted.

      References

      (1) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in <em> Synechocystis </em> PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (2) X. Chen et al., The Entner–Doudoroff pathway is an overlooked glycolytic route in cyanobacteria and plants. Proceedings of the National Academy of Sciences 113, 5441-5446 (2016).

      (3) D. Schulze et al., GC/MS-based 13C metabolic flux analysis resolves the parallel and cyclic photomixotrophic metabolism of Synechocystis sp. PCC 6803 and selected deletion mutants including the Entner-Doudoroff and phosphoketolase pathways. Microbial Cell Factories 21, 69 (2022).

      (4) A. Makowka et al., Glycolytic Shunts Replenish the Calvin–Benson–Bassham Cycle as Anaplerotic Reactions in Cyanobacteria. Molecular Plant 13, 471-482 (2020).

      (5) V. Zaitsev et al., Insights into the Substrate Specificity of Archaeal Entner–Doudoroff Aldolases: The Structures of Picrophilus torridus 2-Keto-3-deoxygluconate Aldolase and Sulfolobus solfataricus 2-Keto-3-deoxy-6-phosphogluconate Aldolase in Complex with 2-Keto-3-deoxy-6-phosphogluconate. Biochemistry 57, 3797-3806 (2018).

      (6) J. S. Griffiths et al., Cloning, isolation and characterization of the Thermotoga maritima KDPG aldolase. Bioorg Med Chem 10, 545-550 (2002).

      (7) S. E. Evans et al., Plastid ancestors lacked a complete Entner-Doudoroff pathway, limiting plants to glycolysis and the pentose phosphate pathway. Nature Communications 15, 1102 (2024).

      (8) J. Moreno-Cabezuelo, G. Gómez-Baena, J. Díez, J. M. García-Fernández, Integrated Proteomic and Metabolomic Analyses Show Differential Effects of Glucose Availability in Marine Synechococcus and Prochlorococcus. Microbiol Spectr 11, e0327522 (2023).

      (9) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in Synechocystis sp. PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (10) R. Y. Stanier, R. Kunisawa, M. Mandel, G. Cohen-Bazire, Purification and properties of unicellular blue-green algae (order Chroococcales). Bacteriol Rev 35, 171-205 (1971).

      (11) D. G. Gibson et al., Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature Methods 6, 343-345 (2009).

      (12) M. Boehm et al., Comprehensive study on ferredoxin isoforms in the cyanobacterium Synechocystis sp. PCC 6803. bioRxiv 10.64898/2026.04.08.717189, 2026.2004.2008.717189 (2026).

      (13) N. Xie, C. Sharma, K. Rusche, X. Wang, Phosphoketolase and KDPG aldolase metabolisms modulate photosynthetic carbon yield in cyanobacteria. The Plant cell 10.1093/plcell/koae291 (2024).

    1. eLife Assessment

      This study applies a relatively novel method, ABR (Automated Behavioural Response system), to an arboreal small mammal, providing some valuable results on the responses of these small primates to both natural predators and human stimuli. However, the study in its current form is incomplete in its framing and methodology and could be improved to appeal more broadly to behavioural biologists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

      Strengths:

      The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

      Weaknesses:

      The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

      (1) Improve the framing of the paper:

      The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

      Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

      (2) Methods:

      The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

      Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

      Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

      Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys? Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

      (3) Analysis:

      A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

      For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

      Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

      Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

    3. Reviewer #2 (Public review):

      Summary:

      The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

      Strengths:

      This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

      Weaknesses:

      Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

    4. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

      Strengths:

      The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

      Weaknesses:

      The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

      (1) Improve the framing of the paper:

      The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

      Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

      We agree that the reviewer makes a valid point that while we were focusing on explicit hypothesis testing in our statements that the other papers mentioned are making predictions based off ecological theory. We will remove line 87 and make sure to limit these comments in the manuscript. We will reframe the introduction and discussion to better reflect this and change the verbiage throughout to focus on behavioural ecology, the impacts of anthropogenic noise and the conservation implications of the study as well as the novel arboreal aspect of the work. We will also update the title to better fit this framing of the study.

      (2) Methods:

      The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

      We do touch on the limitation of not identifying individuals in the discussion (lines 252-258), but we will draw this into a specific section which will address this and the other limitations mentioned here.

      Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

      For one group, there were only 13 videos without a playback where a marmoset was present. So, we used these 13 videos for this group, and selected 13 videos at random from the other groups to match sample size. For these groups, we assigned each video without a playback but with a pygmy marmoset a sequential number, and then used a random number generator to select 13 videos for analysis. We refer to these videos as controls as they are negative controls, and the reviewer is correct – they do measure basal levels of behaviour. In contrast, the playback of control sounds is a procedural control (Bui et al., 2022). We will change how we refer to these controls in the manuscript to reflect the types of control they are. We use the comparison between negative controls and videos with playbacks in experiment 1 to assess the impact of the playback procedure itself, though the reviewer is correct that our conclusions would be better supported if we explicitly compared the intervention playbacks with the procedural control. We will add post-hoc tests to make this comparison explicit.

      Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

      We edited the audios so that they were at similar volume levels. We will update the methods with this information and touch on this in the updated limitations section discussed above.

      Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys?

      Although we do not have information about exposure rates to predators, we do mention levels of exposure of human noise in our methods section on lines 318-320 and we also touch on this in our discussion lines 206-212.

      We will update the methods section to be clearer that all groups are exposed to high levels of anthropogenic noise due to their proximity to the ecotourism lodge and community. We will also expand on this in the discussion as a limitation as having a broader array of groups with varying levels of exposure to humans would allow us to see the broader behavioural reactions to these playback stimuli.

      For the predators we mention in the methods section line 382 “All four species have been found in the study area (Barker and Papworth, 2024)” but we will further expand on this to provide information on the density of raptors in the area based on our previous study.

      Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

      The number of videos was the total number of times the camera trap triggered. These included videos triggered by another animal or foliage movement where there was no marmoset present, and includes both videos with and without a playback.

      We will make this clearer in our description of the results and will report the number of unsuccessful playbacks.

      (3) Analysis:

      A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

      For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

      We used various tests from the performance package (e.g. check_overdispersion) to test the fit of different distributional models (e.g. Poisson, negative binomial) to the same data. Although some models were not overdispersed and did not show evidence of zero-inflation, they did have singular fits, and/or a conditional R<sup>2</sup> of 1.0 suggesting overfitting. We therefore did not choose these models. We will clarify this and provide further details of our approach in the manuscript.

      The models for the behaviours not reported did not fit well using any distributional model. These behaviours had very low occurrence (0 seconds in most videos), and so there was very sparse data for generating estimates, and very low variation within / between groups and conditions. Therefore, these models generated the errors ‘Model nearly unidentifiable’ and warnings about singular boundaries. These suggest we would not be able to reliably generate model estimates, so we do not report the results. We will change the manuscript to make this reasoning more explicit.

      Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

      We will update the results reporting to be standardised as the reviewer has suggested.

      Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

      We will include video number in a reanalysis of the data.

      Reviewer #2 (Public review):

      Summary:

      The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

      Strengths:

      This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

      Weaknesses:

      Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

      We thank the reviewer for their comments and we will update the discussion to add more detail on the ecological relevance of the behaviours that we observed.

      References

      Bui, S., Madaro, A., Nilsson, J., Fjelldal, P.G., Iversen, M.H., Brinchmann, M.F., Venås, B., Schrøder, M.B. and Stien, L.H. 2022. Warm water treatment increased mortality risk in salmon. Veterinary and Animal Science, 17, 100265. https://doi.org/10.1016/j.vas.2022.100265

    1. eLife Assessment

      This valuable study identifies an upstream repressive region in the Saccharomyces cerevisiae DDI2/3 promoter and shows that chemical stress induced by cyanamide or MMS is accompanied by reduced histone abundance and decreased nucleosome occupancy across the DDI2/3 locus. Genetic histone depletion is sufficient to enhance DDI2/3 expression, while MNase-seq and genetic analyses support a model in which the transcription factor Fzf1 both activates transcription through its established promoter-binding function and promotes chromatin changes that relieve nucleosome-mediated repression, helping to explain the unusually strong induction of DDI2/3 relative to other Fzf1 targets. The study is significant, as it provides a conceptually interesting extension of Fzf1-mediated stress regulation. The strength of the evidence is solid, although additional controls and a more cautious treatment of whether Fzf1 directly drives nucleosome eviction versus acting indirectly through transcriptional recruitment or chromatin-remodeling factors would further strengthen the mechanistic conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Du et al. identify a putative URS (nucleotides −709 to −229) that contributes to CY/MMS-induced DDI2/3 transcription in the promoter of S. cerevisiae DDI2 through promoter mapping. They further showed that CY/MMS leads to histone loss, and that genetic depletion of histone triggers DDI2/3 expression in an Fzf1-dependent manner. Using MNase-seq, the authors demonstrate that CY treatment leads to nucleosome loss at the DDI2/3 promoter and coding sequences, and that this chromatin remodeling process requires Fzf1. Based on these results, the authors propose that Fzf1 promotes CY-induced DDI2/3 expression through two distinct mechanisms: as a transcription factor and as a regulator of nucleosome occupancy. This dual mode of regulation may contribute to the exceptionally high induction of DDI2/3 relative to other Fzf1 target genes in response to CY.

      Strengths:

      This manuscript identifies the URS region at the DDI2 promoter that regulates CY/MMS-induced DDI2 expression. In addition, the authors revealed an important role of nucleosome occupancy in regulating DDI2/3 transcription, proposing the intriguing dual regulation model of Fzf1. Overall, this work furthers our understanding of how Fzf1 mediates the increase of DDI2/3 expression in response to CY/MMS treatment.

      Weaknesses:

      Overall, the study is of interest, and the data are generally convincing; however, several conclusions would benefit from further experimental validation. Certain controls are necessary for several experiments to improve the strength of evidence. Several major points are listed below:

      (1) For Figure 6A and Figure 7A, the authors concluded that there are 'joint effects' of CY treatment and histone depletion. However, it is unclear whether CY treatment acts dependently or independently of histone depletion. As shown in Figure 2, both CY and MMS can cause histone reduction. In addition, the depletion system in the RMY102 strain only depletes about half of the H3 (based on the western blots in Figure 5D, E). It would be necessary to test whether H3 abundance is further depleted in CY-treated RMY102 by Western blot.

      (2) Proper controls are missing in Figure 6A and Figure 7A. The authors compare gene expression levels in RMY102 + histone depletion + CY/MMS treatment with RMY102 + non-histone depletion. There are two variables here: histone depletion and CY/MMS treatment. It would be more convincing to include RMY102 YPGal+ CY/MMS treatment (5-40 mM), so that the impact of histone depletion and CY/MMS treatment on Fzf1 target gene expression levels would be clearer.

      (3) In lines 242-249 and 269-272, the authors compared RMY102 versus BY4741 to conclude that histone depletion affects dose dependency of CY/MMS treatment. However, RMY102 and BY4741 might have different responses to CY/MMS due to strain background differences. Thus, in line with point 2, showing the expression level curves for non-histone depleting RMY102 treated with different doses of CY/MMS would be necessary.

      (4) Figure 4 shows that CY and MMS have differential impacts on cell growth, which is intriguing. However, the rest of the data did not provide further insights in regard to this observation. It might be helpful to speculate possible underlying mechanisms in the Discussion session.

    3. Reviewer #2 (Public review):

      Summary:

      Previous work established that FZF1 is both necessary and sufficient for activation of FZF1 target genes through the CS2 sequence motif, which is present upstream of FZF1-responsive targets. This study extends that model by demonstrating that, in addition to direct binding of FZF1 to CS2 elements, FZF1 can also promote reduced nucleosome-mediated repression, thereby contributing an additional layer of transcriptional regulation.

      The authors investigate why FZF1-dependent transcriptional responses exhibit different magnitudes despite FZF1 binding to CS2 elements with similar affinity. Using promoter constructs derived from the DDI2-3 gene, the authors identify a region upstream of the CS2 element that functions as a repressive regulatory element. Based on this observation and publicly available datasets, the authors propose that this repression may be mediated through nucleosome occupancy.

      Strengths:

      The authors demonstrate that the DDI2-3 promoter contains positioned nucleosomes and show that chemical stress results in decreased histone protein levels and reduced histone-associated transcripts. They further examine whether histone depletion alone is sufficient to activate the DDI2-3 response and find that reduced histone levels increase expression, although chemical treatment produces an additional increase that remains dependent on FZF1. These findings suggest that FZF1 contributes to reductions in nucleosome occupancy at DDI2-3 and SSU1, revealing a second, potentially independent mechanism by which FZF1 regulates transcriptional responses to chemical stress.

      Overall, the authors provide strong evidence that nucleosome occupancy influences the magnitude of FZF1-mediated DDI2-3 responses to chemical stress. This work has important implications for understanding how transcriptional networks evolve to generate highly tuned responses by combining multiple regulatory mechanisms acting on shared molecular components.

      Weaknesses:

      However, several additional considerations should be addressed. While histone depletion may contribute to differential FZF1-mediated responses, alternative mechanisms may also influence the observed transcriptional differences. For example, YHB1 exhibits basal expression that is independent of FZF1, and SSU1 contains the CS1 regulatory element, which can promote increased expression independently of FZF1 responsiveness. Therefore, differences in promoter architecture and the presence of alternative regulatory sequences may also contribute to differential FZF1 responses and should be discussed.

      Additionally, the authors should clarify whether nucleosome depletion is directly mediated by the FZF1 ZF5 domain or occurs indirectly as a consequence of RNA polymerase II (Pol II) recruitment. Although the data presented in Figure 9 are consistent with a direct interaction model, the current evidence does not fully exclude the possibility that Pol II recruitment contributes to subsequent nucleosome/histone depletion. Unless there is direct experimental evidence demonstrating that FZF1 ZF5 independently promotes nucleosome remodeling, this alternative mechanism should be acknowledged and considered in the discussion.

    4. Reviewer #3 (Public review):

      Summary:

      In the manuscript titled "Dual regulation of chemical stress-induced DDI2/1 3 expression by a transcription factor Fzf1 and nucleosome in Saccharomyces cerevisiae" Du et al have discovered a dual role of Fzf1 in transcriptional control of DDI2/3 during cyanamide (CY) or MMS treatment. While previous literature established that Fzf1 regulates multiple targets (DDI2/3, SSU1, YHB1, and YNR064C) by binding the CS2 consensus sequence, it remained unclear why DDI2/3 uniquely undergoes a massive 1,000-fold induction under cyanamide (CY) stress, whereas the others show only a 20- to 40-fold induction. In this work, the authors showed that Fzf1 functions beyond standard transcriptional activation. Using MNase-seq and a series of promoter truncation mutants, the authors mapped Upstream Repressing Sequences (URS) in the DDI2/3 promoter that are heavily occupied by nucleosomes. The authors showed that Fzf1 is essential for chromatin remodelling and nucleosome eviction (specifically at the -2 nucleosome position) to de-repress the DDI2/DDI3 expression.

      Strengths:

      The two-tier mechanism of action of Fzf1 in controlling the DDI2/3 expression during CY/MMS stress is compelling and novel.

      Weaknesses:

      While the authors presented the in vivo MNase-seq data that show Fzf1 is necessary for nucleosome displacement, the current study lacks any in vitro mechanistic proof. As the authors acknowledge, it remains unknown whether Fzf1 directly displaces nucleosomes on its own (perhaps through unmapped post-translational modifications induced by chemical stress) or whether its ZF5 activation domain merely acts as a scaffold to recruit separate chromatin remodelling complexes.

      To test nucleosome repression, the authors utilised extreme global interventions, such as deleting the SPT10 gene or halting de novo histone synthesis using a galactose-to-glucose medium shift in the RMY102 strain. While these methods effectively deplete histones and induce DDI2/3 up to 30- to 350-fold, completely depleting cellular histones causes massive, pleiotropic secondary effects across the entire genome, which can obscure specific regulatory relationships.

    1. eLife Assessment

      The study addresses a timely problem that is of importance in the cfRNA field, and it provides a potentially valuable resource. The use of a common processing pipeline, publicly available code, and variance-partition analyses are notable strengths. However, several central conclusions, particularly those concerning protocol-level effects, the dominance of technical over biological variation, and the proposed universal QC criteria remain more general than the study design and available data can support. Therefore the evidence is currently incomplete and needs to be improved through the revision process.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Tuñí-Domínguez et al. present a large-scale, cross-study meta-analysis of 2,356 cell-free RNA sequencing (cfRNA-Seq) samples. The authors aimed to systematically evaluate the impact of pre-analytical variables and library preparation protocols on biological interpretation. By harmonizing publicly available and internally generated datasets through a uniform bioinformatics pipeline, the study seeks to establish standard quality control (QC) metrics and provide evidence-based guidelines for protocol selection in cfRNA-Seq biomarker discovery.

      Strengths:

      A major strength of the study is the scale and breadth of the harmonized dataset. The use of a common computational pipeline reduces variation arising from differences in bioinformatic processing and enables more direct comparisons among published datasets. The authors examine multiple complementary dimensions of data quality rather than relying on a single sequencing metric. The inclusion of variance-partition analyses, healthy-control-only analyses, and analyses restricted to samples with low gDNA contamination strengthens the evaluation of technical heterogeneity. The public availability of the analysis code and processing configurations further increases the reproducibility and potential utility of this work. The study convincingly demonstrates that technical and pre-analytical factors are major sources of variation across existing plasma cfRNA-seq datasets.

      Weaknesses:

      Several central conclusions are broader than the current cross-study design can fully support. Protocol category is strongly associated with dataset, laboratory, sample source, and collection procedure, making intrinsic protocol effects difficult to separate from study-specific effects, particularly for categories represented by only one or a few studies. The conclusion that technical variation generally overwhelms biological variation also requires qualification because diverse diseases and cancer types are combined into broad phenotype categories, and many phenotypes are concentrated within individual datasets. The interpretation and additional value of the proposed NG80 and NP80/NG80 metrics require further support, especially given their variable behavior across datasets. Most importantly, the thresholds used in the proposed universal QC framework are not independently validated and are strongly influenced by library-preparation strategy. The current findings therefore support protocol-dependent comparisons and identification of technical trade-offs more strongly than they support a universal definition of library quality.

      Major concerns:

      (1) Protocol effects remain difficult to distinguish from study-specific effects.

      The manuscript interprets Broad Protocol Category as a major determinant of transcriptomic variation. However, protocol, dataset, laboratory, sample source, and collection procedure are strongly interconnected. Although both Dataset and BPC are included in the variance-partition model, several protocol categories are represented by only a limited number of studies.

      This is particularly problematic for WRO, which is represented by a single study. Its apparent characteristics therefore cannot be separated from Reggiardo-specific laboratory, cohort, provider, or sample-processing effects. The authors should clarify the stability of the variance-partition results and, where feasible, provide sensitivity analyses. Alternatively, conclusions based on protocol categories represented by one or very few studies should be explicitly presented as study-specific observations rather than broadly validated protocol properties.

      (2) The interpretation and robustness of the proposed library-diversity metrics require further support.

      The manuscript attributes the high NG80 values in the Block and Sun datasets to gDNA contamination. However, other datasets with similarly low FSR values, including Wang and Giráldez, do not show comparably high NG80 values. This suggests that gDNA contamination alone is insufficient to explain library diversity and that sequencing depth, fragment length, mapping behavior, or library construction may also contribute.

      The NP80/NG80 ratio is conceptually reasonable, but its additional value beyond the RNA-biotype composition shown in Figure 3C is unclear, particularly given the large variability observed in WRR datasets. The authors should further examine the relationships among FSR, NG80, sequencing depth, and protocol characteristics and assess the stability of NP80/NG80. If additional validation is not feasible, the interpretation and generality of these metrics should be moderated.

      (3) The conclusion that technical variation generally overwhelms biological variation requires qualification.

      The manuscript provides convincing evidence that technical heterogeneity is a major source of variation in the combined cross-study dataset. However, the conclusion that donor phenotype contributes only negligible variation may be broader than the analysis supports.

      Phenotypes are reduced to healthy, cancer, and non-cancer disease categories despite substantial biological heterogeneity, and many disease groups are concentrated within particular studies. Consequently, disease-specific biological variation may partly be assigned to the Dataset term. A similar issue affects the cellular-origin analysis, where collection center and phenotype are substantially associated in the Chen dataset, making their individual contributions difficult to distinguish.

      Where sample sizes permit, the authors should examine more specific disease categories or perform within-dataset analyses. Otherwise, the conclusions should be narrowed to state that technical variation dominates the present heterogeneous cross-study aggregation, rather than implying that phenotype-associated cfRNA signals are generally negligible.

      (4) The proposed universal QC framework is insufficiently justified and protocol-dependent.

      Figure 6 defines high-quality libraries using NG80 >1,000 together with FSR >20% or FER >75%. However, the manuscript does not explain how these thresholds were selected or validate them against an independent measure of reproducibility, analytical performance, or biomarker utility.

      These criteria are also strongly affected by library-preparation strategy. FER and FSR favor libraries enriched for exonic or spliced RNA, whereas the NG80 cutoff disadvantages WRR libraries containing abundant noncoding transcripts. This is difficult to reconcile with the recommendation of WRR for exploratory transcriptomic and microbial analyses.

      The authors should justify the threshold selection and assess its sensitivity and protocol dependence. Ideally, the criteria should be validated against an independent performance endpoint. Otherwise, Figure 6 should be reframed as a descriptive comparison, and protocol- or application-specific guidance should replace a universal binary definition of library quality.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript systematically evaluates the impact of experimental workflows on plasma cell-free transcriptome sequencing (cfRNA-seq) data. The authors integrate a large number of cfRNA sequencing datasets from multiple publicly available studies and establish a unified bioinformatics framework to systematically assess the effects of experimental workflows, genomic DNA contamination, library diversity, and preanalytical factors on cfRNA transcriptomic profiles.

      Strengths:

      This study addresses the technical heterogeneity that may hinder cfRNA biomarker discovery and clinical translation and is of substantial value.

      Weaknesses:

      The effects of disease phenotype, technical confounding, criteria for library quality, and conclusions regarding DNase treatment require further clarification and validation.

      Major Points:

      (1) The conclusion that donor phenotype explains only a small fraction of transcriptomic variation requires further support from within-study analyses.

      The authors conclude from variance partitioning across all studies that phenotype explains only a small fraction of cfRNA transcriptomic variation. However, the included studies encompass different diseases, while phenotype is simplified into healthy, cancer, and non-cancer disease categories, and the overall transcriptomic variation is strongly influenced by study-specific and experimental workflow batch effects. Therefore, the cross-study pooled analysis may underestimate genuine disease-associated cfRNA differences within individual studies conducted under the same experimental workflow.

      Recommendation: The authors are encouraged to perform within-study phenotype analyses in datasets that include both healthy controls and disease samples and have sufficient sample size, and to quantitatively estimate the proportion of transcriptomic variance explained by phenotype. For example, the Zhu dataset includes healthy controls and liver cancer samples, and the authors have already observed relatively clear phenotype-associated clustering between the two groups. The contribution of healthy-versus-liver-cancer phenotype to transcriptomic variation could therefore be quantified within this dataset. If similar results are obtained across multiple independent cohorts, the findings could then be summarized across studies. The authors should also note that a low contribution of phenotype to global transcriptomic variance does not necessarily imply that disease-associated cfRNA signals lack biological or clinical relevance.

      (2) Comparison of the relative contributions of technical factors and disease phenotype may be affected by confounding.

      Figure 1C shows that phenotype is strongly or even completely confounded with technical variables such as collection center and centrifugation protocol in some cohorts. Nevertheless, the variance partitioning analysis across all samples is used to conclude that technical factors are the primary sources of variation, whereas phenotype contributes little. In the presence of such confounding, technical effects and disease-associated biological effects may not be reliably estimated independently, and this conclusion therefore requires more direct validation.

      Recommendation: The authors are encouraged to perform an independent within-study variance analysis in cohorts in which technical variables and phenotype are relatively balanced. For example, in the Moufarrej cohort, phenotype is essentially unconfounded with collection center/centrifugation protocol (Cramer's V = 0). Phenotype and relevant technical variables could be included simultaneously in a within-cohort model to quantify their respective contributions to cfRNA transcriptomic variation. If technical factors still explain a larger fraction of variance in such relatively unconfounded cohorts, this would provide stronger support for the central conclusion of the manuscript.

      (3) The use of NG80 as a criterion for defining "high-quality libraries" requires further validation.

      The authors use NG80 as a metric of library diversity and further apply NG80 > 1,000 as one criterion for defining high-quality libraries in Figure 6. However, because NG80 is based on gene counts, it may be affected by sequencing depth. In addition, gDNA contamination can artificially increase NG80, whereas genuinely abundant non-coding RNAs in WRR libraries can lower NG80. Therefore, a higher NG80 does not necessarily indicate better overall library quality, and the metric may reflect both technical quality and genuine RNA composition.

      Recommendation: The authors are encouraged to re-evaluate NG80 after downsampling samples to a common number of mapped fragments and to examine the relationship between NG80 and sequencing depth. The rationale for the NG80 > 1,000 threshold should also be further justified, and the impact of alternative NG80 thresholds on high-quality library classification and the main conclusions should be assessed. Unless there is evidence that this threshold reliably predicts library reproducibility or biomarker-related information content, library diversity and overall library quality should be clearly distinguished, and NG80 should not be presented as a universal criterion for high-quality libraries.

      (4) Conclusions regarding DNase treatment should be interpreted more cautiously.

      The Toden study did not explicitly report DNase treatment. The manuscript infers that DNase digestion was performed based on the fact that this study originated from the same laboratory as other studies and used a similar workflow; this inference should not be treated as an established experimental fact. In addition, the authors state that double DNase treatment is the most effective approach among non-EB workflows, but this conclusion is mainly based on cross-study comparisons, in which DNase strategy varies together with laboratory, sample handling, and cohort-specific factors. The current evidence is therefore insufficient to establish that double DNase treatment itself is superior.

      Recommendation: The authors are encouraged to label the DNase status of the Toden study as "not reported" or "inferred", unless confirmation can be obtained from the original authors. The conclusion that double DNase treatment is the most effective approach should also be tempered, with explicit acknowledgment that it requires direct parallel validation using the same samples under different DNase treatment strategies.

    1. eLife Assessment

      This study presents a valuable patient-derived human cell platform for investigating LMNA-associated cardiomyopathy and identifies cyproheptadine as a compound that improves abnormal calcium handling across multiple LMNA variants in vitro. The experimental approaches are generally rigorous and establish reproducible disease-associated cellular phenotypes; however, the evidence remains incomplete because the mechanistic basis of the observed rescue is not established, and the translational conclusions substantially exceed the presented data.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors use patient-specific induced pluripotent stem cell-derived cardiomyocytes (hiPSC-CMs) from six patients carrying three distinct pathogenic LMNA variants to investigate disease mechanisms underlying LMNA-associated dilated cardiomyopathy (DCM). The authors report shared abnormalities in nuclear morphology, electrophysiology, calcium handling, and contractility across all patient-derived lines and demonstrate that correction of representative LMNA variants by CRISPR/Cas9 rescues several of these phenotypes. They further develop a high-throughput phenotypic drug screening platform using calcium transient measurements and identify cyproheptadine as the only compound among 1,280 FDA-approved drugs that consistently improves calcium transient abnormalities across all patient-derived lines.

      The study addresses an important clinical problem and establishes a technically sophisticated patient-derived screening platform. However, while the experimental work is generally well executed, several of the major biological and translational conclusions are not sufficiently supported by the presented data.

      Strengths:

      The major strength of this study is the development of a patient-derived functional screening platform using multiple LMNA mutations rather than focusing on a single pathogenic variant. The inclusion of six patient-derived iPSC lines representing three distinct mutations increases the generalizability of the observations and allows identification of disease features that appear reproducible across different genetic backgrounds.

      The phenotypic characterization is comprehensive and includes nuclear morphology, transcriptomics, manual and automated electrophysiology, calcium imaging, and impedance-based contractility measurements. Importantly, CRISPR-mediated correction of representative LMNA variants provides convincing evidence that the observed cellular abnormalities are directly attributable to the pathogenic variants.

      Finally, the implementation of an unbiased high-throughput drug screen using patient-derived cardiomyocytes represents a valuable technical advance that could facilitate therapeutic discovery in inherited cardiomyopathies.

      Weaknesses:

      The principal weakness of the manuscript is that the central conclusions substantially exceed what is directly demonstrated by the data.

      The manuscript repeatedly concludes that dysregulated calcium handling represents a shared pathogenic mechanism underlying LMNA-associated cardiomyopathy. However, the presented experiments establish only that abnormal calcium handling is a shared cellular phenotype across the studied variants. The data do not distinguish whether calcium dysregulation is a primary disease mechanism or whether it is secondary to the numerous upstream abnormalities already known to result from LMNA dysfunction, including altered nuclear architecture, defective mechanotransduction, chromatin remodeling, and transcriptional dysregulation. This distinction is critical because the manuscript repeatedly interprets correction of calcium handling as correction of the underlying disease process without directly demonstrating this relationship.

      A related concern is the interpretation of the drug screening results. The primary screen is entirely based on normalization of calcium transient parameters (CTD75, FWHM, and T75-25). Consequently, the screen identifies compounds capable of correcting calcium cycling rather than compounds that necessarily modify disease biology. Although cyproheptadine subsequently improves impedance-derived contractile parameters, it remains unknown whether treatment rescues other defining features of LMNA cardiomyopathy, including abnormal electrophysiology, nuclear defects, transcriptional alterations, or broader cellular stress responses. Thus, the conclusion that cyproheptadine represents a "novel treatment" for LMNA-associated cardiomyopathy is considerably stronger than the evidence presented.

      The mechanistic studies are also insufficient to support the proposed mode of action. The manuscript proposes CHRM2 as the likely mediator of cyproheptadine activity primarily because it is the only appreciably expressed known target in the transcriptomic dataset. However, no functional experiments test this hypothesis. Without genetic or pharmacological interrogation of CHRM2, the proposed mechanism remains speculative. Likewise, alternative mechanisms of cyproheptadine action, including serotonergic signaling, histamine receptor antagonism, direct calcium channel modulation, or antioxidant effects, are not investigated.

      Another conceptual weakness is that the manuscript promises to distinguish both shared and variant-specific disease mechanisms but ultimately focuses almost exclusively on shared phenotypes. The transcriptomic analyses remain largely descriptive and are not leveraged to identify mutation-specific biological pathways or explain differences among the three LMNA variants. Given the unique cohort assembled in this study, this represents a missed opportunity to generate broader biological insight into LMNA-associated disease.

      The transcriptomic analyses themselves would also benefit from more rigorous interpretation. RNA from multiple independent differentiations was pooled prior to sequencing, limiting assessment of biological variability and reducing confidence in statistical inference. Similarly, many functional analyses emphasize the number of wells, recording sweeps, or individual cells while the number of independent biological differentiations is less prominently presented. Greater emphasis on biological replication would strengthen confidence in the robustness of the findings.

      Finally, the translational implications of the work should be interpreted more cautiously. All therapeutic studies are performed in relatively immature two-dimensional iPSC-derived cardiomyocytes. No validation is presented in engineered heart tissues, multicellular cardiac organoids, animal models, or human tissue. The current evidence supports the conclusion that cyproheptadine is a promising in vitro phenotypic modifier rather than an established therapeutic candidate for LMNA-associated cardiomyopathy.

      Overall, this study establishes a valuable patient-derived platform for investigating LMNA-associated cardiomyopathy and demonstrates the utility of functional high-throughput screening in identifying compounds that improve disease-associated cellular phenotypes. However, the manuscript currently overstates both the mechanistic significance of calcium dysregulation and the therapeutic implications of cyproheptadine. A more restrained interpretation of the findings together with additional mechanistic validation would substantially strengthen the impact of the work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors conducted a functional high-throughput drug screening using hiPSC-CMs derived from patients with LMNA-DCM. Cyproheptadine emerged as a therapeutic candidate.

      Strengths:

      The screen appears well designed.

      Weaknesses:

      The single candidate that emerged from the screen, cyproheptadine, raises issues with potency. In addition, validation studies that are both expected and necessary for a drug proposed as a novel therapeutic for human cardiomyopathy have not yet been performed. Rigor could be improved once the basic mechanistic and validation studies discussed below have been performed.

    1. eLife Assessment

      This study presents a valuable finding on impacts of the common driver mutations APC, KRAS G12D, and TP53 in murine colorectal organoids. In particular, the authors examine how the order of APC and TP53 acquisition influences tumor phenotype. The evidence supporting the claims of the authors is solid. The work will be of interest to scientists working in the field of breast cancer.

    2. Reviewer #2 (Public review):

      Summary:

      This study addresses an important and timely question in colorectal cancer biology by systematically examining the effects of the common driver mutations APC, KRAS G12D, and TP53 in murine colorectal organoids, with particular emphasis on how the order of APC and TP53 acquisition influences tumor phenotype. These mutations are well known to be frequent, truncal, and often co-occurring in colorectal cancer. While it is increasingly appreciated that mutational order can shape tumor behavior, studies directly comparing the phenotypic consequences of alternative APC-TP53 mutation orders remain rare. This work therefore addresses a relevant and timely question.

      Strengths:

      A major strength of the study is its focus on previously unexplored biology, combined with the generation of multiple isogenic murine organoid models with controlled mutational sequences. The authors employ careful and robust quality control of the CRISPR-mediated alterations, and the inclusion of both in vitro and in vivo experiments strengthens the relevance of the work.

      Weaknesses:

      There are, however, several limitations that should be considered when interpreting the findings. First, KRAS G12D activation is used as the initiating alteration, whereas APC loss is generally believed to be the initiating event in most human colorectal cancers. Second, the analysis is restricted to comparing only two mutation orders (KAT versus KTA), which limits the breadth of conclusions that can be drawn about mutation ordering more generally. Finally, key RNA-sequencing and in vivo experiments rely on a limited number of isogenic lines, which constrains interpretability.

      The study aimed to systematically investigate how the accumulation and sequence of driver mutations influence colorectal cancer initiation. The data provide intriguing evidence that the relative timing of APC and TP53 loss may impact tumor initiation and survival in a hostile microenvironment. However, given the limited number of biological replicates, these observations should be interpreted with caution and would benefit from further validation.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Li et al. used genetically engineered murine intestinal organoids to investigate how the temporal order of oncogenic mutations influences cell state and tumourigenicity of colorectal epithelial cells. By sequentially introducing Apc and Trp53 loss-of-function mutations in alternate orders within a Kras^G12D background, the authors generated isogenic organoid lines for both in vitro and in vivo characterisation. Bulk RNA-seq reveals expected transcriptional changes with relatively modest differences between the two triple-mutant configurations (KAT vs KTA). The key finding emerges from transplantation assays: while KAT and KTA organoids show equivalent tumourigenic potential in immunodeficient mice, only KAT organoids form tumours in immunocompetent hosts (5/10 vs 0/10), suggesting that mutation order shapes susceptibility to immune-mediated clearance. The experiments are well-executed, and the conclusions are generally supported by the data.

      Strengths:

      The experimental system is well-designed for the question. By combining a Kras^G12D transgenic background with sequential CRISPR-mediated knockout of Apc and Trp53 in alternate orders, the authors generated truly isogenic organoid lines that differ only in mutational sequence. This is technically non-trivial and provides a clean platform for dissecting order effects, a question otherwise difficult to address experimentally.

      The authors performed comprehensive baseline characterisation of these organoids, including morphological and histological assessment, quantification of organoid-forming efficiency and proliferation, and bulk RNA-seq profiling. While these analyses revealed no major differences between KAT and KTA organoids, and the observed enhancement of epithelial stemness upon Apc loss and proliferative advantage conferred by Trp53 loss are largely expected, the systematic nature of this characterisation establishes a useful methodological template for future organoid-based studies.

      The authors further investigated the functional impact of mutational order using subcutaneous transplantation assays. By comparing tumour formation in immunodeficient versus immunocompetent hosts, the authors uncover a genuinely unexpected finding: KAT and KTA organoids behave equivalently in the absence of adaptive immunity, but diverge dramatically when immune pressure is applied (KAT: 5/10; KTA: 0/10). This observation is arguably the most compelling aspect of the study and opens an interesting line of inquiry.

      We greatly appreciate your comments on this study.

      Weaknesses:

      The authors acknowledge that initiating with Kras^G12D does not reflect the typical human sporadic CRC trajectory, where APC loss is usually the first event. While this design choice was pragmatic, it means the observed order effects are contextualised within an artificial starting point. It remains unclear whether the Apc/Trp53 order would matter in a Kras-wild-type background, or whether the Kras-driven cellular state is a prerequisite for these phenotypes to emerge.

      We agree with the reviewer that initiating tumorigenesis with Kras<sup>G12D</sup> does not fully recapitulate the most common trajectory of sporadic human CRC, where APC loss typically occurs first. We had noted this point in the original Discussion and further clarified it more explicitly in the Introduction part of the revised manuscript as shown in Line 97–103.

      Our experimental design was intended to establish a controlled and genetically tractable system to interrogate the principle of mutation order effects. In this context, Kras<sup>G12D</sup> activation provides a stable oncogenic baseline that facilitates sequential genome engineering and comparison of isogenic lines.

      Although APC loss is frequently the initiation event, a recent study has suggested that Kras<sup>G12D</sup> priming can reshape the selective landscape for subsequent driver events, including Apc alterations (PMID: 41339549). Consistent with this notion, our data indicate that Kras<sup>G12D</sup> activation induces a permissive oncogenic cellular state that may influence the phenotypic consequences of later mutations. We therefore speculate that the Kras<sup>G12D</sup>-primed context may contribute to the observed order-dependent effects.

      We agree that testing Apc Trp53 order in a Kras-wild-type background would be an important future direction, and we have pointed this out explicitly in the revised Discussion as shown in Line 549–554.

      Subcutaneous implantation provides a tractable readout of tumourigenicity, but the cutaneous immune microenvironment differs substantially from that of the intestinal mucosa. Given that the central claim concerns immune-mediated selection, orthotopic transplantation would more directly test whether the observed order effects hold in a physiologically relevant context.

      In the present study, we employed subcutaneous transplantation as a widely used platform to assess tumorigenic potential under controlled immune conditions. This approach offers high reproducibility, straightforward tumour monitoring, and has been broadly applied in organoid-based cancer studies in both immunodeficient (PMID: 23273993, 23776211, 32209571, 33055221) and immunocompetent (PMID: 32209571, 33055221, 41672595) settings.

      Importantly, our primary goal was to determine whether mutation order influences susceptibility to immune-mediated clearance, rather than to model the full complexity of the intestinal niche. The clear divergence between KAT and KTA specifically in immunocompetent hosts supports the existence of intrinsic mutation order-dependent immune vulnerability.

      Nevertheless, we fully agree with the reviewer that orthotopic transplantation would provide a more physiologically relevant immune microenvironment and represents also an important direction for future investigation. We have explicitly discussed this limitation and highlight orthotopic validation as an important future direction in the revised Discussion as shown in Line 563–571.

      The ssGSEA comparison involves only 14 ATK tumours, and the key comparisons (Figure 6E) yield borderline significance (p=0.052). More fundamentally, since mutation order cannot be inferred from the clinical samples, the authors are correlating organoid-derived IFN signatures with tumour immunophenotypes without direct evidence that these patients' tumours followed a KAT-like trajectory. The reasoning becomes circular: KAT organoids define the signature used to identify KAT-like clinical tumours.

      We thank the reviewer for raising this important point. We would like to clarify that our intention was not to infer the actual mutation order in clinical samples, which indeed cannot be reliably reconstructed from bulk tumour RNA-seq data.

      Instead, our goal was to determine whether the transcriptional programs distinguishing KAT and KTA organoids could be observed in human CRC cohorts. In this context, the organoid-derived IFN-related signature was used as a molecular reference to assess potential clinical correlation, rather than to classify tumours by evolutionary trajectory.

      We agree that the statistical significance in Figure 6E is modest (p = 0.052), and we have revised the text (Line 478–480) to present this analysis more cautiously as a suggestive trend rather than definitive evidence. We also clarified this limitation explicitly in the revised manuscript (Line 537–542) to avoid overinterpretation.

      Furthermore, the most striking finding of the study, that KTA organoids fail to form tumours in immunocompetent hosts while KAT organoids can, lacks a mechanistic follow-up. The transcriptomic differences between KAT and KTA are modest when cultured as monocultures, yet their in vivo fates diverge dramatically. The authors do not address why these subtle intrinsic differences translate into such divergent immune susceptibility, nor do they characterise the immune response adequately (beyond limited CD4/CD8 IHC at tumour peripheries).

      We thank the reviewer for this important point. We agree that the mechanistic basis underlying the differential immune susceptibility between KAT and KTA remains incompletely resolved.

      A practical limitation of the current study is that KTA grafts failed to establish tumours in immunocompetent hosts, which precluded downstream histological and immune profiling of established lesions. As a result, our in vivo immune characterization of KTA grafts is nearly impossible.

      Nevertheless, our transcriptomic analyses indicate that KAT and KTA organoids differ in interferon-response and immune-related programs prior to transplantation, and those differentially expressed genes were consistently preserved in tumour cells derived from immunodeficient hosts. These results suggest the presence of intrinsic tumour-cell-autonomous differences may influence immune recognition or clearance.

      We have expanded the Discussion to outline several non-mutually exclusive mechanisms that could account for this phenotype, including altered interferon responsiveness, differential antigen presentation capacity, and changes in tumour cell-intrinsic immune escape programs (Line 527–533). These hypotheses are consistent with the transcriptional differences observed prior to transplantation and provide a framework for future mechanistic investigation. We agree that deeper immune profiling (e.g., immune infiltrate composition, antigen presentation status, and functional immune assays) will be important to fully elucidate the mechanism and represents a key direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study addresses an important and timely question in colorectal cancer biology by systematically examining the effects of the common driver mutations APC, KRAS G12D, and TP53 in murine colorectal organoids, with particular emphasis on how the order of APC and TP53 acquisition influences tumor phenotype. These mutations are well known to be frequent, truncal, and often co-occurring in colorectal cancer. While it is increasingly appreciated that mutational order can shape tumor behavior, studies directly comparing the phenotypic consequences of alternative APC-TP53 mutation orders remain rare. This work, therefore, addresses a relevant and timely question.

      Strengths:

      A major strength of the study is its focus on previously unexplored biology, combined with the generation of multiple isogenic murine organoid models with controlled mutational sequences. The authors employ careful and robust quality control of the CRISPR-mediated alterations, and the inclusion of both in vitro and in vivo experiments strengthens the relevance of the work.

      We greatly appreciate your comments on this study.

      Weaknesses:

      There are, however, several limitations that should be considered when interpreting the findings. First, KRAS G12D activation is used as the initiating alteration, whereas APC loss is generally believed to be the initiating event in most human colorectal cancers.

      We sincerely thank the reviewer for their insightful comments regarding the initiation of tumorigenesis with a Kras mutation rather than the more canonical Apc loss, which was also raised by the reviewer #1. We fully agree that the Apc-first represents the most prevalent sequence in human colorectal cancer (CRC), We have more clearly explained the rationale for our experimental design in the revised Introduction part as outlined in our response to reviewer #1.

      Second, the analysis is restricted to comparing only two mutation orders (KAT versus KTA), which limits the breadth of conclusions that can be drawn about mutation ordering more generally.

      We thank the reviewer for this critical concern, which we agree is essential for strengthening the robustness and generality of our findings. However, as a proof-of-concept study of Apc and Trp53 loss, two major oncogenic events in CRC, serves as a biologically meaningful starting point for dissecting order-dependent effects. Although it is of great significance to compare all six possible mutation orders of these three driver genes, generating and thoroughly characterizing all genotypes (with identical replicates) represents a substantial undertaking beyond the scope of this initial study.

      Finally, key RNA-sequencing and in vivo experiments rely on a single isogenic line, which substantially constrains interpretability.

      The aim of the study was to systematically investigate how mutation accumulation and order influence colorectal cancer initiation. While the data suggest that the relative timing of APC and TP53 loss may be particularly important for tumor initiation, the absence of biological replication makes it difficult to draw robust conclusions. Engraftment efficiency and tumor behavior can be influenced by many factors for a single clone, including additional passenger mutations acquired during culturing, as well as epigenetic differences that are independent of the engineered mutations.

      We thank the reviewer for this concern. We apologize that we have not made a clear presentation of our data source. Indeed, for all major in vitro and in vivo assays of double and triple mutants, we analyzed at least two independently derived clones per genotype. These independent clones harbour distinct mutations in target genes and were treated as biological replicates throughout the study.

      To improve clarity and transparency, we have revised the relevant figure legends and further provided a Table S5 to explicitly indicate the clonal origin of each data point throughout the study.

    1. eLife Assessment

      This important study shows that prenatal alcohol exposure produces lasting changes in amyloid precursor protein processing, including altered APP C-terminal fragments and Notch intracellular domain levels, that persist into adulthood and track with progressive spatial learning and memory deficits, in both wild-type and 3xTg-AD mice. The convincing study was carefully designed and combined biochemical, histological, and behavioral approaches across multiple ages. The work will interest researchers studying fetal alcohol spectrum disorders, Alzheimer's disease risk, and the developmental origins of neurodegeneration. Nonetheless, reviewers identified opportunities to strengthen the conclusions with additional samples and controls.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Montenegro and colleagues reports a uniquely significant set of compelling results from a carefully designed study. The findings are fundamental and should substantially advance our understanding of whether prenatal exposure to high levels of alcohol produces neural changes that are precursors to the development of Alzheimer-like pathology in an animal model of Fetal Alcohol Spectrum Disorder (FASD). The quest was to test whether prenatal exposure to high-dose alcohol in the mouse would result in selective damage that would result in Alzheimer disease-like cellular disorder and mnemonic impairment. Both outcomes emerged and were exacerbated in relevant transgenic mice. The untoward effect on memory endured and even worsened with age.

      Strengths:

      The authors noted the importance of using a validated animal model to test their hypotheses related to AD-like outcomes because human postmortem data are unavailable and even in vivo data are limited to younger FASD cohorts. The authors further noted limitations, which also anticipate experiments that could chart the temporal course of the effect and windows of prenatal alcohol exposure that result in damage or resilience. In short, as the authors state on lines 435-7, "These data provide the first experimental evidence that developmental alcohol exposure impacts core proteolytic pathways central to AD/ADRD pathogenesis."

      Weaknesses:

      Addressing the following would clarify several points in an already well-written paper:

      (1) It would be useful to have a timeline of the study, much like the one provided for the water maze protocol. On that timeline, please include the sample sizes examined, ages at exposure, and other pertinent procedures, indicating which animals remained alive for testing, etc.

      (2) What were the attrition rates for each study group?

      (3) What are the human age equivalents of the maternal mice?

      (4) It would be helpful to see a graph of the BECs of each animal relative to the doses given. That would clarify how the alcohol exposure amount and timing are the same and where they are different for all exposed mice, given that the alcohol levels were somewhat different by group, as noted in the Methods. Were these BEC differences at all related to group differences in outcome measures or memory performance?

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the impact of prenatal alcohol (PAE) on amyloid precursor protein (APP) C-terminal fragments and notch intracellular domain (NICD) levels in adulthood is measured in 3xTg-AD mice.

      Prenatal alcohol alters gamma secretase activity with development and aging. This could have implications for Alzheimer's disease risk in populations without inherited Alzheimer's risk genetics.

      Strengths:

      Strengths include the model, the use of orthogonal approaches, and the rigorous, high-quality data.

      Weaknesses:

      Some figures lack prenatal alcohol treatment in the 3xTg-AD mice.

      Some overstatements should be tempered. For instance, one cannot conclude that the changes in CTFs are driving the changes in learning and memory (as suggested in the last line of the abstract) without a direct intervention testing this. For instance, though PAE caused a more robust learning deficit at 6 mo in WT, the impact on CTFs was less than it was at 3 mo. PAE did not significantly change CTFs or learning/memory in 3xTg-AD mice at 4 months, suggesting the genotype effect takes over at this point. The text should be adjusted to reflect this.

      Conclusion:

      In summary, this is a rigorous assessment of the long-term impacts of PAE on CTFs and learning/memory in adult WT and 3xTg-AD mice.

    4. Reviewer #3 (Public review):

      Summary:

      The goal of this study was to test the hypothesis that prenatal alcohol exposure (PAE) can affect Alzheimer's disease (AD) pathogenesis using biochemical proxies, histology, and a behavioral paradigm sensitive to AD-related memory decline. Major strengths include the breadth of techniques used, consideration of different AD-related molecular markers, and the use of different ages as well as appropriate controls. The authors largely achieved their aims to show that PAE does affect amyloid precursor fragments (CTFs) and notch signaling very early on as well as long-term effects in adulthood, which may uncover a previously underappreciated mechanism that may contribute to AD-related neuropathology and behavioral outcomes during lifespan.

      Strengths:

      (1) Several techniques are used to address molecular, behavioral, and histological PAE-related changes.

      (2) There is use of appropriate controls and an AD-relevant mouse model.

      (3) Different ages are used to address age-related and long-term effects in AD and control mice.

      (4) The novelty of results shows early changes in amyloid-related processes, affected by PAE.

      Weaknesses:

      (1) It is unclear as to whether there are sex differences, particularly in the adult cohort.

      (2) More clarity is needed on sample size per cohort and whether mice that were used for anatomy and biochemical analyses were previously used for behavior. Including a table and noting any overlap would be useful.

      (3) In many instances, two-way ANOVAs with treatment (PAE vs vehicle) and genotype as factors will be useful to report (e.g., Figure 1).

      (4) In Figure 5 and line 253, it is stated that older mice have more severe deficits, but there are no direct statistical comparisons with younger AD mice.

      (5 Lines 270-271 refer to mice as "presymptomatic", but these mice do have behavioral symptoms. Do the authors mean no neuropathology yet? Any data showing lack of robust neuropathology would be useful.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) It would be useful to have a timeline of the study, much like the one provided for the water maze protocol. On that timeline, please include the sample sizes examined, ages at exposure, and other pertinent procedures, indicating which animals remained alive for testing, etc.

      We agree with the reviewer that a timeline would be helpful for clarifying the PAE exposure paradigm and the subsequent sample collection and processing. Sample sizes vary across the different analyses; therefore, we have indicated the sample size for each experiment in the corresponding figure. To further improve clarity, we will add Author response image 1, which will include a schematic of the overall experimental timeline, including the PAE regimen, collection time points, and sample processing, as shown below.

      Author response image 1.

      (2) What were the attrition rates for each study group?

      We thank the reviewer for raising this important point. There was no attrition of animals within the experimental groups; the number of animals included at the beginning and end of the study remained the same. However, we did observe a reduction in litter size following prenatal alcohol exposure (PAE). In our 3xTg-AD colony, litters typically consisted of approximately 8 pups under control conditions, whereas PAE litters occasionally contained 4–6 pups. Thus, the reduction in animal numbers reflects decreased litter size associated with PAE rather than attrition during the study. We would also like to clarify that the primary scope of this study was not to provide a terminal/end-point analysis of disease progression, but rather to investigate the emergence of Alzheimer’s disease (AD)-related phenotypes during early adulthood following PAE. Accordingly, our longitudinal experimental design focused on identifying the earliest molecular, synaptic, behavioral, and neuropathological alterations that emerge during this period. This approach allowed us to examine whether PAE accelerates or precipitates the onset of AD-related symptomatology in the 3xTg-AD model, rather than following the animals through advanced disease stages. We will clarify this rationale in the revised manuscript.

      (3) What are the human age equivalents of the maternal mice?

      We appreciate the reviewer’s question regarding the age of the maternal mice. The dams used in our study were young adult females (2 to 3 months of age) at the time of breeding. Because chronological age does not translate linearly between mice and humans, particularly during development and reproductive maturation, we have avoided assigning a precise human-age equivalent. Based on established comparative developmental frameworks and calculations, these animals represent a 20 years old young-adult in the reproductive stage, rather than an advanced maternal-age condition [1]. We will clarify this point in the revised manuscript.

      (4) It would be helpful to see a graph of the BECs of each animal relative to the doses given. That would clarify how the alcohol exposure amount and timing are the same and where they are different for all exposed mice, given that the alcohol levels were somewhat different by group, as noted in the Methods. Were these BEC differences at all related to group differences in outcome measures or memory performance?

      We appreciate the reviewer’s suggestion to provide a more detailed representation of the BEC data. We agree that displaying the BECs for individual animals would provide additional clarity regarding the consistency of alcohol exposure across groups. We have therefore included the individual BEC values in the new Figure 1. We observed some variability in BECs between the 3xTg-AD and B6129 groups, as noted in the Methods. Importantly, however, the average alcohol consumption was comparable between the two genotypes, indicating that the difference in BECs was not due to differences in the amount of alcohol consumed. All dams in the PAE groups reached BECs above 0.08 g/dL, the commonly used legal blood alcohol concentration limit in the United States, supporting the use of our paradigm as a binge-like alcohol exposure model. We further examined whether the variability in BECs was associated with the differences observed in outcome measures, including memory performance. We did not find evidence that the differences in BECs accounted for the group differences in behavioral or molecular outcomes. Thus, although some intergroup variability in BECs was present, the overall alcohol exposure was comparable, and the observed phenotypic differences were not attributable to differences in alcohol consumption.

      Reviewer #2 (Public review):

      (1) Some figures lack prenatal alcohol treatment in the 3xTg-AD mice.

      We appreciate the reviewer’s observation and agree that the rationale for the different experimental groups across the figures should be clarified. The primary focus of this study is to characterize the effects of prenatal alcohol exposure (PAE) in wild-type B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference to determine whether the effects observed following PAE in wild-type animals overlap with or resemble features of AD pathology. Accordingly, the initial figures focus on the effects of PAE in B6129 mice and include the non-exposed 3xTg-AD group as a reference for the AD phenotype. The last two figures specifically address the effects of PAE in the 3xTgAD model, with the 3xTg-AD mice becoming the experimental subject of interest rather than serving solely as a disease reference. For this reason, the PAE-3xTg-AD group is not included in the earlier figures, whereas it is included in the final two figures where the effect of PAE on the AD model is directly evaluated. We will clarify this experimental rationale in the revised manuscript and figure legends.

      (2) Some overstatements should be tempered. For instance, one cannot conclude that the changes in CTFs are driving the changes in learning and memory (as suggested in the last line of the abstract) without a direct intervention testing this. For instance, though PAE caused a more robust learning deficit at 6 mo in WT, the impact on CTFs was less than it was at 3 mo. PAE did not significantly change CTFs or learning/memory in 3xTg-AD mice at 4 months, suggesting the genotype effect takes over at this point. The text should be adjusted to reflect this.

      We appreciate the reviewer’s careful consideration of this point. We agree that the relationship between APP CTF accumulation and learning and memory deficits should not be interpreted as causal in the absence of a direct intervention experiment. We were careful in choosing the wording throughout the manuscript to describe these findings as associated changes rather than evidence of a causal interaction. Our data demonstrate the presence of APP CTF accumulation and learning and memory deficits following PAE, but they do not establish that CTF accumulation directly drives the behavioral phenotype. We therefore will temper the language in the Abstract and throughout the manuscript to avoid overstatement. We also acknowledge that the relationship between these phenotypes is not necessarily linear across age: although PAE produced a more pronounced learning deficit at 6 months in B6129 mice, the magnitude of APP CTF accumulation was greater at the earlier time point. Importantly, we consider the possibility that the greater APP CTF accumulation observed at earlier ages may represent an early molecular insult whose functional consequences become evident later in life. In this context, the temporal dissociation between the molecular and behavioral phenotypes could be consistent with a “two-hit” model, in which an early-life insult induced by PAE creates or primes a pathological vulnerability that subsequently manifests as cognitive dysfunction with ageing [2]. We recognize, however, that this interpretation remains a hypothesis and would require longitudinal mechanistic studies to establish. This temporal relationship may also contribute to the broader concept of early-life origins of AD/ADRD, suggesting that prenatal environmental exposures may initiate molecular alterations during neurodevelopment that remain detectable or predispose the brain to later-life dysfunction. Similarly, the absence of significant changes in APP CTFs or learning and memory in 4-month-old 3xTg-AD mice suggests that the effects of the AD genotype may become dominant at this stage. Consistent with this interpretation, we state in the Discussion that future studies are needed to identify and experimentally test the direct molecular pathways affected by PAE that ultimately contribute to learning and memory impairment. We will revise the text accordingly to make this distinction clear while highlighting the potential significance of an early molecular insult preceding the later emergence of behavioral phenotypes.

      Reviewer #3 (Public review):

      (1) It is unclear as to whether there are sex differences, particularly in the adult cohort.

      We appreciate the reviewer’s comment regarding potential sex differences. We agree that considering sex as a biological variable is important, particularly for the adult cohorts. To address this point, we will include identifying marks for male and female animals separately in our plots where sample size permits. This will allow the reader to better evaluate potential sex-dependent effects of PAE and to determine whether the observed phenotypes are consistent across sexes. We will also clarify this approach in the revised manuscript.

      (2) More clarity is needed on sample size per cohort and whether mice that were used for anatomy and biochemical analyses were previously used for behavior. Including a table and noting any overlap would be useful.

      We appreciate the reviewer’s suggestion and agree that greater clarity regarding the sample sizes and use of animals across analyses is important. The sample size for each cohort and experimental group is indicated in the corresponding figures, and we will make this information more explicit in each figure legend to facilitate interpretation. Animals that underwent behavioral testing were subsequently used for biochemical analyses, allowing us to examine molecular changes in the same animals in which behavioral phenotypes were characterized. In contrast, for the neonatal cohort, we performed both anatomical and biochemical analyses, to assess the distribution and extent of APP CTF accumulation across the brain during this early developmental period. This approach was selected to provide a broader assessment of the spatial distribution of APP CTF accumulation at birth. We will clarify these experimental details in the revised Methods and figure legends.

      (3) In many instances, two-way ANOVAs with treatment (PAE vs vehicle) and genotype as factors will be useful to report (e.g., Figure 1).

      We appreciate the reviewer’s suggestion regarding the use of two-way ANOVA with treatment and genotype as factors. However, we respectfully disagree that this approach is appropriate for all of the analyses presented in this manuscript. Our experimental design and the specific biological questions addressed in each experiment were not uniform across cohorts. In particular, the primary objective of the study was to characterize the effects of PAE in B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference, while the effects of PAE in the 3xTg-AD model were specifically examined in the later experiments. Therefore, combining genotype and treatment as factors across all datasets would not always reflect the experimental questions or the structure of the cohorts. In addition, some experiments did not include all four groups, making a two-way ANOVA inappropriate for those analyses. Nevertheless, we agree that a two-way ANOVA may be informative for experiments in which both genotype and treatment are fully represented and the experimental design supports this analysis. We will therefore consider and apply two-way ANOVA, where appropriate, to those datasets, including evaluation of the main effects of genotype and treatment and their interaction. We will clarify the statistical approach and its rationale in the revised Methods and figure legends.

      (4) In Figure 5 and line 253, it is stated that older mice have more severe deficits, but there are no direct statistical comparisons with younger AD mice.

      We appreciate the reviewer’s observation. We agree that, in the absence of a direct statistical comparison between age groups, the statement that older mice have “more severe deficits” may be too strong. Our intention was to describe the apparent progression of the phenotype across age rather than to imply that we had statistically demonstrated an age-dependent increase in severity. We have therefore revised the text in Figure 5 and at line 253 to use more cautious language, describing the greater magnitude of the observed deficits in older mice without implying a direct statistical comparison between age groups. We agree that a formal conclusion regarding age-dependent progression would require a statistical analysis, which we will include in the revised manuscript.

      (5) Lines 270-271 refer to mice as "presymptomatic", but these mice do have behavioral symptoms. Do the authors mean no neuropathology yet? Any data showing lack of robust neuropathology would be useful.

      We appreciate the reviewer’s careful observation. We agree that the term “presymptomatic” was not sufficiently precise, particularly because the mice already exhibit measurable behavioral alterations at this age. Our intention was not to suggest that these animals were free of phenotypic abnormalities, but rather that they were at an early stage of disease progression, before the emergence of robust neuropathological features. We have therefore revised the terminology to avoid referring to these mice as “presymptomatic.” Instead, we describe them as being in an early stage of disease development, characterized by emerging behavioral and molecular alterations but without the extensive cardinal neuropathology typically associated with later stages of the 3xTg-AD phenotype. We agree that the distinction between behavioral symptoms and neuropathological progression is important. In this study, our focus was on the emergence of early AD-related phenotypes during young adulthood, rather than on establishing the absence of neuropathology. We have therefore avoided making a definitive claim regarding the lack of neuropathology and have revised the text to more accurately reflect the scope of our data.

      Additional References

      (1) Dutta, S. & Sengupta, P. Men and mice: Relating their ages. Life Sci. 152, 244–248 (2016).

      (2) Gunn, J. S. et al. Exploring the ‘Multiple-Hit Hypothesis’ of Neurodegenerative Disease: Bacterial Infection Comes Up to Bat’. Frontiers in Cellular and Infection Microbiology | www.frontiersin.org 1, 138 (2019).

    1. eLife Assessment

      This fundamental work significantly advances our understanding of the circuit-level implementation of predictive processing by elucidating the functional influence between putative prediction error neurons in layer 2/3 and putative internal representation neurons in layer 5. The evidence demonstrating that neither the hierarchical nor the non-hierarchical variant of predictive processing fully accounts for the presented data is convincing. Moving forward, this line of work would benefit from explicitly comparing different theories, thereby clearly articulating the points raised in this paper.

    2. Reviewer #1 (Public review):

      Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closed-loop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

      The authors conclude that both versions of predictive processing architectures they analyze are likely invalid and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction errors neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

      Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

      In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

      The authors find that the functional impact of L5 neurons to L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

      They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Fig.7).

      All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

      Comments on revised version

      I commend the authors for nicely provided answers and addressing my concerns and clarifying their points on the relationship of their findings to JEPA and current state of understanding.

      In particular for Q5 -I meant if there is also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that the authors identify as PE-?<br /> The authors provided the answer on point.

      Overall, I think this is a fundamental study and the strength of evidence for the claims is exceptional as an exemplary use of existing approaches.

    3. Reviewer #2 (Public review):

      This manuscript reveals functional connectivity of two different classed of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

      Comments on revised version.

      I thank the Authors for answering my questions and updating the manuscript.

      Congratulations on this important work!

    4. Reviewer #3 (Public review):

      Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

      The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

      This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

      Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

      While the work is overall convincing and provides important insights into the circuit-level implementation of predictive processing, I think the connection to JEPA networks would benefit from a more in-depth discussion of its relationship to recently proposed and implemented models. Below, I address the specific points raised by the authors (flanked by ' ... ' to make the author's statements stand out), in particular in relation to the model proposed by Nejad et al. (2025):

      - 'The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

      (1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.'

      I think there is some confusion here. In Nejad et al., both L2/3 and L5 function as encoder networks. Each receives sensory input (with L2/3 receiving this input delayed via L4) and computes a latent representation of that input, denoted z_{L2/3} and z_{L5}, respectively. The prediction is obtained by comparing the output of L2/3 with L5 latent representations (via W_{L2/3->L5} * z_{L2/3}). In other words, L5 is the target in the learning objective.

      Although, they did not explicitly state the mapping with JEPA, their model has the same fundamental property - predictive learning happens in the latent space. Also, this appears very similar to the roles assigned to L2/3 and L5 in your Figure 9B. The fact that you made this more explicit and the new data included, is in my view, a very interesting contribution. However, from an architectural perspective, it appears that your proposal and the model of Nejad et al. are conceptually very similar, and I do not see a substantial difference between the two. This should be made more clear in the Discussion.

      - '2. Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 ('the learning-driving error signal originates in L5'). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.'

      Indeed, in Nejad et al., the layer-dependent mismatch responses are modeled as gradients with respect to neuronal activity, and the model does not explicitly include prediction error neurons. However, this appears to be a modeling choice rather than a fundamental aspect of the proposal, and it does not preclude an implementation with explicit prediction error neurons. In fact, the authors explicitly acknowledge this possibility in the Discussion:

      "The second approach would be to recast our model within a predictive coding framework... Predictive coding jointly optimizes both model parameters and neuronal activities, which could naturally lead to prediction errors observable in the activity of both L2/3 and L5 neurons. Note that these two views are not mutually exclusive."

      While Nejad et al. hypothesize that the learning-driving error signals (that are distinct from their mismatch responses) originate in L5, the abstract loss function itself does not uniquely specify where the underlying comparison between the predicted representation (W_{L2/3 -> L5} z_{L2/3}) and the target representation (z_{L5}) must be implemented. The proposed biological implementation places this computation in L5, but from my understanding, the computational objective itself does not require this specific localization.

      That said, I agree that your proposed model introduces a genuine difference. The functional roles assigned to the vertical projections are effectively reversed: in Nejad et al., the L2/3->L5 projection carries the prediction, whereas the L5->L2/3 projection conveys the error/gradient. In your architecture, by contrast, the L5->L2/3 projection carries the teaching signal (target). This is, in my view, a real and testable interpretational divergence that is worth stating clearly.

      Therefore, I think the novelty lies less in the computational architecture itself and more in committing to a particular biological implementation-one that adds cell-type-specific detail to an implementation that Nejad et al. hypothesised as being compatible with their framework.

      - '3. The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework - silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.'

      My reading of Nejad et al. is consistent with your interpretation of the equations. Specifically, Eq. 3 makes L5 activity depend on L2/3 activity (albeit weakly, with a = 0.3), whereas Eq. 2 contains no L5 term, so L2/3 activity does not depend directly on L5 activity. In that model, the L5->L2/3 pathway carries the learning gradient rather than contributing to the activity dynamics. By contrast, in your proposed model, L2/3 activity depends on L5 activity, whereas L5 activity is not immediately dependent on L2/3 activity (except indirectly through learning/plasticity). So, you state that "silencing L2/3 in a familiar setting should result in no immediate changes in L5 activity.<br /> [...].

      However, I am unsure how to reconcile this prediction with the results shown in Fig. 6. If I understand the figure correctly, optogenetic activation of Rrad-positive (positive prediction error) L2/3 neurons produces a small increase in L5 activity, whereas activation of Adamts2-positive (negative prediction error) L2/3 neurons produces a decrease in L5 activity. Although these experiments involve activation rather than silencing, they nevertheless suggest that perturbing L2/3 activity can have an immediate effect on L5 activity. Could you clarify how this is consistent with the proposed model? In other words, what aspect of the proposed circuitry makes activation effective while silencing is predicted to have no immediate consequence?

      For comparison, Nejad et al. performed a related perturbation analysis in Fig. S15 by scaling the output of L2/3 neurons exhibiting positive mismatch signals (defined through the activity gradients), which increased L5 activity, whereas scaling neurons with negative mismatch signals produced the opposite effect. I am not entirely sure how directly these simulations map onto the experiments shown in your Fig. 6, since the Nejad simulations were performed during mismatch conditions, if I have understood them correctly.

      - '4. The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.'

      Nejad et al. only require two encoders (like in JEPA), how these two are learnt can be done in several ways. While Nejad et al. focus on using a reconstruction loss to learn the L5 target, they also show that it works equally well with non-reconstruction objectives. Therefore, I do not think the presence or absence of a reconstruction objective constitutes a fundamental distinction between the two proposals.

      As your current work presents a conceptual architecture rather than a fully implemented learning algorithm (in a model), the mechanism that would prevent representational collapse has not yet been defined. From my understanding, every joint-embedding approach must address this issue, whether through reconstruction, variance/covariance regularization, stop-gradient or EMA mechanisms, or other approaches. Thus, the absence of a reconstruction objective (or another anti-collapse mechanism) is not, in itself, a distinguishing feature of the proposed architecture, but rather an as-yet unspecified design choice within the learning objective.

      - '5. Lastly, there is time-delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.'

      I agree that the temporal delay introduced by L4 is a key component of the Nejad et al. model and is currently absent from your proposal. However, I would expect temporal delays to emerge naturally in your framework as well, given the multisynaptic and highly parallel organization of cortical circuits. More generally, implementing predictive learning over time (as in JEPA) requires comparing representations at times t and t+1, which seems to require some form of temporal delay. How else would you suggest this is done?

      In general, I think the manuscript would benefit from a clearer discussion of its relationship to the model proposed by Nejad et al. (perhaps following the discussion above), including both the shared conceptual claims and the aspects that genuinely differ between the two frameworks. At present, some of the claims are presented as novel, although at least some of these core ideas have already been proposed in Nejad et al.

      For example, the authors state: "Thus, we propose that layer 2/3 functions to predict layer 5 activity, not sensory input per se, hence making predictions in the internal representation space, not input space." This appears to be exactly what Nejad et al. proposed as discussed above - in their model L2/3 predicts L5 activity (purely in the latent space), as they state in the abstract.

      That said, there are some interesting differences, and I think the community would greatly benefit from making these clear, including the roles assigned to interlaminar connections and the interpretation of the signals carried by these pathways. These differences are interesting and potentially testable, and I think the manuscript would be strengthened by explicitly distinguishing which aspects are in line with the ideas already present in Nejad et al. and which aspects represent new contributions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank you for the time you took to review our work and for your feedback! We have performed the following additional analyses:

      (1) We analyzed the optomotor mismatch response as a function of time spent in the optogenetic closed-loop session.

      (2) We show the correlation between optomotor mismatch response and visuomotor mismatch response for functionally identified PE neurons.

      The main changes to the manuscript are:

      (3) Clarification of our terminology (moved and refined definition of the teaching signals).

      (4) A new figure panel summarizing the functional influence patterns we identified.

      (5) Clarification of our argumentation for the JEPA-inspired proposal.

      All comments are addressed individually in the following.

      Public Reviews:

      Reviewer #1 (Public review):

      Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closedloop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two-photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

      The authors conclude that both versions of predictive processing architectures they analyze are likely invalid, and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing, engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction error neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

      Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

      In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

      The authors find that the functional impact of L5 neurons on L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

      They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Figure 7).

      All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new, exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

      I have several questions about the interpretations of some of the claims and suggestions for potential additional experiments and analyses.

      We thank the reviewer for their comments. We address them below.

      (1) Some of the pieces of the puzzle remain to be identified and demonstrated: the existence of internal representation neurons in L2/3 and ascertaining that the L5/6 neurons analyzed function indeed as internal representation neurons. The authors find that stimulation of L2/3 positive prediction error neurons enhances activity of L5 neurons...If L5 neurons hold a latent representation that serves as a teaching signal for L2/3 neurons (as the authors posit), wouldn't one expect that the input they receive from the positive prediction neurons be suppressive, such that the error is further minimized?

      Not necessarily - this depends on the model one has for how cortex works. In the hierarchical predictive processing model, PE+ neurons are expected to suppress the local internal representation neurons in L5. Our data, however, are not consistent with this model. This is one of the key arguments we build the idea on that JEPA is a better model for cortex than hierarchical predictive processing. In JEPA we would not expect the L2/3 prediction error to update L5 directly, but instead update the prediction of L5 activity (the source of this signal remains to be identified; see our speculations on the origin of prediction signals in the comments to question 3), and drive plasticity in the local L5 encoder network. That something acts like a teaching signal for the L2/3 comparator does not, by itself, determine how the outcome of the comparison influences the source of the teaching signal.

      (2) Do the authors envision any specific differences between the representations of the two encoder networks posited to exist in L2/3 and L5 in the JEPA-like implementation? Are they synchronous/offset in their temporal representations, or any other features?

      Given theoretical work (Mohammadi et al., 2025), one would expect to find differences in learning rates between the predictor and the encoder networks. Assuming the predictor network is also implemented in L2/3, we would expect to see faster learning rates in L2/3 compared to L5. Implementations inspired by related theoretical works also predict differences in learning rate between the two encoder networks (Grill et al., 2020), again with higher learning rates expected in L2/3 compared to L5. Beyond that, however, we are not aware of any experimentally observable differences one might expect to find. We are hoping the computational community will remedy this soon.

      (3) Where is the prediction coming from onto L2/3 neurons? Is it emerging locally in L2/3 from the putative internal representation neurons, or is it long-range - as work from the authors previously proposed? Or a mix of both?

      We expect the predictions to come from long-range inputs. In classical (hierarchical) JEPA one would expect these to be the lateral communication within L2/3. In a non-hierarchical implementation that is capable of operating on arbitrary graphs (as would be necessary for it to work in cortex), we suspect that the L2/3 network will use both long-range L2/3 and long-range L5 input. In JEPA terminology: the encoder A network uses information from non-local sources of both networks to predict local activity of encoder B network (assuming that information is statistically useful in predicting that activity).

      (4) What is the role of the indiscriminate L4 input that appears to enhance activity of both positive and negative prediction error neurons in L2/3?

      The short answer is, we don’t know. We would have indeed expected to find some asymmetry of influence. There are a few options: A) We might be failing to activate a specific interneuron that mediates feedforward inhibition on this pathway by the artificial stimulation of Scnn1a neurons. B) We know that Scnn1a neurons are only a subset of L4 neurons – there are other populations of L4 neurons that might exhibit the opposing influence. C) The Scnn1a population is more than 1 layer of a JEPA away from the comparator and not part of the predictor that forms the actual representation compared against L5 (e.g. L4 could provide an input to L2/3 internal representation neurons, and thus only indirectly influence L2/3 PE neurons). D) The JEPA analogy is wrong.

      (5) Does Figure 7D change in a meaningful manner if the authors plot the correlation between optomotor mismatch response and visuomotor mismatch response specifically for the negative prediction error neurons in L2/3 (Adamts-2) rather than for all L2/3 cells sampled?

      We might be misunderstanding. If the reviewer means genetically identified negative prediction error neurons (Adamts2), we do not have the data to address this question, as we did not perform any recordings of molecularly defined Adamts2 population in L2/3. If the reviewer means functionally identified negative prediction neurons, this would be the two rightmost data points in Figure 7D (the x-axis in this panel is the visuomotor mismatch response strength we use to functionally identify PE- neurons). These two data points on the right-hand side of the plot correspond exactly to what we classify as negative prediction error neurons throughout the rest of the manuscript (15% of the most responsive neurons to visuomotor mismatch).

      Does the reviewer mean, is there also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that we identify as PE-? If so, the answer is yes (Author response image 1). Interestingly, while there is a positive correlation, there is also nonuniformity in response patterns. We think that this is expected, given that our visuomotor coupling paradigm captures only a very small subspace of stimuli that prediction error neurons are tuned to. Bulk stimulation of L5, in contrast, might work better to separate all PE neurons. In other words, we speculate that L5 stimulation is a better predictor of the functional role of an L2/3 neuron.

      Author response image 1.

      Optomotor mismatch response as a function of visuomotor mismatch response for neurons that are functionally classified as PE. Red line shows a linear fit estimated with a bootstrap approach.

      (6) Do the optomotor mismatch responses in L2/3 neurons depend on how long the closed-loop coupling of optogenetic stimulation of Tlx3 L5 neurons and locomotion speed has been in place for?

      No, not that we can measure. We performed an analysis in which we split the optogenetic closed-loop session into two equal parts, early and late. We then quantified the average optomotor mismatch response independently for early and late parts of the session (Author response image 2A). Based on this quantification, we find no evidence of a change as a function of time in the closed-loop session. We also performed a sliding window analysis on a shorter timescale. While it did look like the responses may be smaller in the first few minutes, we do not have sufficient data to address this. None of the differences in response size were significant (Author response image 2B).

      Author response image 2.

      Optomotor mismatch response as a function of experience with artificial closed-loop coupling. (A) Mean L2/3 population response to optomotor mismatch in the first half (Early MM) and in the second half (Late MM) of the optogenetic closed-loop session. (B) Mean L2/3 population response to optomotor mismatch as a function of time in the optogenetic closed-loop session. Error bars indicate SEM. Differences between the mean values were estimated by hierarchical bootstrap and are not significant.

      Reviewer #2 (Public review):

      This manuscript reveals the functional connectivity of two different classes of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

      We thank the reviewer for their comments. We address them below.

      General comments:

      (1) The methods of statistical testing are insufficiently described. I did not understand the description in lines 1105-1119. The authors should provide sufficient details so the reader can reproduce their analyses. For example, it may be helpful to provide specific details of the testing procedure for one of the comparisons (e.g. the first comparison in Table S1).

      We assume the reviewer is not familiar with hierarchical bootstrapping in general. If so, the explanation below would summarize the procedure. This is the procedure with particular emphasis on its application to neuroscience data is described in the paper we reference in that part of the methods (Saravanan et al., 2020). Given that the analysis has become relatively standard (and is described in the reference provided), we think it might be an overkill to add the full explanation below to the manuscript. In addition to the general procedure, there are only 2 pieces of information relevant to fully reconstructing the analysis:

      (1) What are the “levels” (mice, recording sites, neurons, trials)?

      (2) What is the number of bootstrap samples used.

      Thus, we think all information is already provided in the methods. We now also explicitly added the levels when describing the nested structure of the data in the manuscript to increase clarity.

      Hierarchical bootstrap analysis:

      Our data are naturally nested: multiple neurons are recorded within a single mouse, and multiple mice are tested within an experimental group. Standard bootstrapping (sampling with replacement from the entire pool of neurons) fails because it assumes all observations are independent. In reality, neurons from the same mouse are more similar to each other than to neurons from a different mouse. The hierarchical bootstrap (or multi-level bootstrap) preserves this nested structure, ensuring your confidence intervals are not artificially narrow due to pseudoreplication.

      To illustrate the problem, assume you have 100 neurons from Mouse A and 10 neurons from Mouse B, a simple bootstrap will be heavily biased toward Mouse A. Furthermore, the simple bootstrap ignores the fact that the true variance in your population comes from two sources:

      (1) Between-mouse variance (differences in surgery, genetics, or behavior).

      (2) Within-mouse variance (differences in tuning or activity between individual cells).

      Hierarchical bootstrap addresses this problem, and is implemented as follows: To estimate the mean response while accounting for different sample sizes per mouse, a two-level resampling scheme is used

      (1) Resample the higher level (Mice)

      First, you account for the variability between animals.

      Suppose you have N mice in total.

      Randomly draw N mice with replacement from your original pool.

      Note: Because this is with replacement, a single mouse’s data might be included multiple times in one bootstrap iteration, while another mouse might be left out entirely.

      (2) Resample the lower level (Neurons)

      For each mouse selected in Step 1, you must now account for the variability within that specific animal.

      Look at the number of neurons actually recorded from that mouse (let’s call it k<sub>i</sub>).

      Randomly draw k<sub>i</sub> neurons with replacement from that mouse’s specific pool of recorded cells.

      This step is crucial: you always resample the same number of neurons that were originally recorded for that specific mouse. This maintains the "weight" or "influence" that animal had in the original dataset.

      (3) Calculate the resampled statistic

      Calculate the mean of all neurons collected in this “bootstrap sample”.

      (4) Iterate

      Repeat Steps 1–3 many times (typically B = 1,000 or 10,000 iterations).

      The distribution of these B bootstrap means represents your sampling distribution and is used to calculate confidence intervals and p-values. 

      (2) The authors should clarify how the problem of multiple comparisons was addressed for comparisons performed in multiple moments of time, where significance is indicated by a black bar (e.g. in Figure 2F).

      There is no family-wise error correction in cases of comparing response time courses implemented in our analysis. If the reviewer has a good suggestion for how to implement family wise error correction, we would be happy to implement it. We are not aware of anything that is better than what we currently do (the time bin-wise comparison). To briefly explain the problem: In most neuroscience papers, response curves are compared by choosing a time window (e.g. 0.5 to 1 s following a trigger) and calculating mean values of the curves in these windows. This hides the problem of multiple comparisons that arises from the fact that the experimenter is free to choose a response window used for analysis. Note, this also creates a strong incentive to “optimize” choice of an analysis window – a part of the analysis that a reader is typically completely blind to. We could of course also choose an analysis window that sounds reasonable and yields significant differences for our analyses. However, to provide a more unbiased image of the data, we have come to do bin-wise comparisons with a fixed p-value (typically 0.05). This is used in all of our papers at the moment. Given that samples from neighboring timepoints are correlated via a combination of actual responses and a subset of noise sources, the samples are not independent. We now also implemented a correction for spurious positive values by requiring at least 2 neighboring bins to have a p-value below 0.05 to be shown. If the samples were independent, this would mean a false positive rate of 0.0025. Given that they are not, this is a lower bound only. Additionally, any family-wise error correction would be a function of the number of time bins we show in the plot. This would mean that our choice of the time window shown in a plot (-1s to +4s, or +5s, etc.) would change the p-value we consider significant. Thus, there is no explicit family-wise error correction, and we have come to the conclusion that the bin-wise comparison with a fixed p-value is the most unbiased representation of the data we can provide. The alternative would be to additionally plot z-scores or the p-values as a function of time, but in our experience these types of plots are even harder to read for readers not used to it.

      (3) It would be helpful to add a figure in the Discussion summarising the functional connectivity suggested by all experiments.

      We now added a panel that summarizes the functional connectivity observed in our experiments to Figure S10. 

      (4) Throughout the manuscript, the authors use the term "teaching signals", but I am unclear what they mean by it: after reading the definition in lines 45-46, I thought that they corresponded to values (as they are compared to sensory signals). Later (428-430), the text suggests that they correspond to error neurons. But then lines 605-607 say it is not an error signal. The authors should define teaching signals very precisely or remove this term.

      The formal definition of the teaching signal was in footnote 1 of the manuscript. We assume the reviewer may have missed this. We now moved this definition into the main text to increase clarity.

      We use the term teaching signal to mean exactly this definition throughout the manuscript. We suspect, a second source of confusion may come from ambiguity in regards to anatomical vs. functional definitions. We have attempted to emphasize that the definition of teaching signal is a functional one, not an anatomical one (as is the case for ‘prediction’ – predictions are functionally defined, not anatomically – hence it makes sense to ask questions of the form “what are the potential sources of predictions” etc.). A teaching signal is a signal that is compared against a prediction (the ‘ground truth’ the prediction is compared and trained against). In different circuit implementations of predictive processing, different inputs function as predictions and teaching signals. In the hierarchical implementation, activity of the internal representation neuron at the lower level serves as a teaching signal for a prediction signal that is formed by the internal representation neuron from the higher level. In the non-hierarchical implementation, external inputs from the thalamus or other cortical areas serve as teaching signals for the respective prediction signals that are formed by internal representation neurons. This terminology is most intuitive when thinking from the perspective of a prediction error neuron, since a prediction error neuron computes the difference between two inputs signals – one of which functions as a teaching input, and the other one as a respective prediction. Hence, lines 428-430 specify that layer 5 input onto prediction error neurons of layer 2/3 serves as a teaching input.

      Reviewer #2 (Public review):

      Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

      The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

      This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

      Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

      We thank the reviewer for their comments. We address them below.

      While the work is overall convincing and significantly advances our understanding of the circuit-level implementation of predictive processing, there are a few weaknesses that should be addressed or discussed:

      (1) The authors define putative positive prediction error neurons as the 15% of neurons most responsive to grating onset and putative negative prediction error neurons as the 15% most responsive to visuomotor mismatch. While this selection would be expected to overlap with negative and positive prediction error neurons, the criterion is not sufficiently stringent (independent of the exact percentage chosen). In particular, classification of a neuron as a prediction error neuron should ideally be accompanied by evidence that it does not exhibit a significant increase in activity when the prediction matches the sensory input or teaching signal.

      We understand the reviewer’s intuition. We can indeed use other stimuli to identify prediction error neurons, like the relative suppression of responses in closed-loop running onset vs open-loop running onsets. This was the reason behind including Figure S1, to show that our selection results in expected pattern of running onset responses. We don’t typically use the running onset responses as they have an additional confound we have not fully understood. This is that running onset always tends to result in an increase of calcium activity in all neurons. This could have a variety of reasons: A) Contamination of hemodynamic occlusion signals (blood vessels tend to constrict at running onset, making it appear like an increase in calcium activity) – see Yogesh et al., 2025. B) The virtual coupling in our VR is not good enough to provide a true “closed loop” experience. Humans typically notice lags of larger than 30ms – in our VR it is approximately 100 ms. C). Running onset in head-fixed animals is not accompanied by a vestibular input. D) Predictive processing is wrong. We tend to think it is a combination of the three.

      We can also use combinations of the two criteria to select neurons - if the reviewer has a specific selection criteria in mind (top XXX% MM responsive AND top XXX% closed-loop suppressed, etc.) we are happy to repeat the analysis for that specific set of criteria, but the fundamental problem that we are using a functional response to select these neurons does not go away. We know that our functional selection criteria mean we select a population of neurons that is enriched for prediction error neurons. If the enrichment is too weak, we would expect to find no effects in terms of functional influence. It is hard to explain, however, how a weak enrichment could result in a strong effect on functional influence. Our arguments in more lengthy form, for why the visuomotor mismatch is a good stimulus to identify negative prediction error neurons can be found here: Attinger et al., 2017; Jordan and Keller, 2020; Leinweber et al., 2017; O’Toole et al., 2023; Vasilevskaya et al., 2022; Zmarz and Keller, 2016.

      (2) The authors "speculate that the prediction error responses in layer 2/3 may not be computed with respect to sensory input, but with respect to layer 5 activity as a teaching signal." However, it is unclear how this perspective differs from earlier statements in the manuscript. In the Introduction, the authors note that "these signals, typically referred to as sensory signals, we will refer to as teaching signals," and later describe the hierarchical variant as one "in which internal representation neurons act as a source of the teaching signal." Given this framing, it is difficult to identify what is conceptually novel in the updated view. Is the key distinction that layer 2/3 neurons are now proposed to generate predictions in an internal representation space rather than in sensory input space, as briefly suggested in the Discussion? Or are the authors introducing a distinction between an external (sensory) and an internal (cortical) teaching signal? If so, this distinction should be made explicit. Clarifying this point would considerably strengthen the manuscript.

      There might be a misunderstanding regarding our usage of the term teaching signal. In hierarchical predictive processing the teaching signal is typically referred to as a sensory signal, as e.g. in: “prediction error neurons compare predictions to sensory input”. In non-hierarchical predictive processing, or far away from the sensory input (think prefrontal cortex), or for cross-modal interactions “sensory” input is misleading. Also, in non-predictive-processing type models (like JEPA), sensory input has a different functional role. Thus, we operationally define teaching signal as the signal that is compared against the prediction by prediction error neurons.

      The two primary options we are comparing are:

      (1) Is the teaching signal to the L2/3 comparator a bottom-up input to V1 (as one would expect in predictive processing). 

      (2) Is the teaching signal to the L2/3 comparator L5 input (as one would expect in JEPA).

      Our data argue in favor of option 2. We have rephrased parts of the manuscript to try to make this clearer. 

      (3) The authors propose that "L2/3 neurons predict L5 activity, hence making predictions in the internal representation space rather than the input space," and further suggest that, since both deep and superficial cortical layers receive thalamic input, the cortex may function like a JEPA. This idea appears closely related to the model introduced by Nejad et al. (2025), which effectively implements a JEPA-like architecture: L5 activity serves as a target against which L2/3 predictions are compared in a selfsupervised manner, with both L5 and L2/3 (via L4) receiving thalamic input. It would be helpful for the authors to clarify how their framework differs from that model, and to specify the key conceptual or mechanistic distinctions between the present proposal and the approach described by Nejad et al.

      The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

      (1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.

      (2) Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 (‘the learning-driving error signal originates in L5’). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.

      (3) The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework – silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.

      (4) The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.

      (5) Lastly, there is a time delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.

      We expect that the most useful future models should move beyond JEPA, with the emphasis on nonhierarchical models capable of operating on arbitrary graphs. We think cortex functions according to principles of a JEPA (predictions in latent space), and that the role of cell types and the exact computational organization remain to be constrained.

      REFERENCES

      Attinger, A., Wang, B., Keller, G.B., 2017. Visuomotor Coupling Shapes the Functional Development of Mouse Visual Cortex. Cell 169, 1291-1302.e14. https://doi.org/10.1016/j.cell.2017.05.023

      Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020. Bootstrap your own latent: A new approach to self-supervised Learning. https://doi.org/10.48550/arXiv.2006.07733

      Jordan, R., Keller, G.B., 2020. Opposing Influence of Top-down and Bottom-up Input on Excitatory Layer 2/3 Neurons in Mouse Primary Visual Cortex. Neuron 108, 1194-1206.e5. https://doi.org/10.1016/j.neuron.2020.09.024

      Leinweber, M., Ward, D.R., Sobczak, J.M., Attinger, A., Keller, G.B., 2017. A Sensorimotor Circuit in Mouse Cortex for Visual Flow Predictions. Neuron 95, 1420-1432.e5. https://doi.org/10.1016/j.neuron.2017.08.036

      Mohammadi, A.G., Halvagal, M.S., Zenke, F., 2025. Understanding cortical computation through the lens of joint-embedding predictive architectures. https://doi.org/10.1101/2025.11.25.690220

      O’Toole, S.M., Oyibo, H.K., Keller, G.B., 2023. Molecularly targetable cell types in mouse visual cortex have distinguishable prediction error responses. Neuron 111, 2918-2928.e8. https://doi.org/10.1016/j.neuron.2023.08.015

      Saravanan, V., Berman, G.J., Sober, S.J., 2020. Application of the hierarchical bootstrap to multi-level data in neuroscience. Neurons Behav. Data Anal. Theory 3, https://nbdt.scholasticahq.com/article/13927-application-of-the-hierarchical-bootstrap-tomulti-level-data-in-neuroscience.

      Vasilevskaya, A., Widmer, F.C., Keller, G.B., Jordan, R., 2022. Locomotion-induced gain of visual responses cannot explain visuomotor mismatch responses in layer 2/3 of primary visual cortex. https://doi.org/10.1101/2022.02.11.479795

      Yogesh, B., Heindorf, M., Jordan, R., Keller, G.B., 2025. Quantification of the effect of hemodynamic occlusion in two-photon imaging of mouse cortex. eLife 14, RP104914. https://doi.org/10.7554/eLife.104914

      Zmarz, P., Keller, G.B., 2016. Mismatch Receptive Fields in Mouse Visual Cortex. Neuron 92, 766–772. https://doi.org/10.1016/j.neuron.2016.09.057

    1. eLife Assessment

      This important study provides new insights into the patterns of organelle inheritance in the protozoan parasite Toxoplasma gondii. The authors introduce an innovative dual-labeling approach to distinguish maternally inherited from de novo synthesized organelles, representing convincing evidence that different organelles follow distinct inheritance fates during parasite replication. Future studies will be needed to determine whether the residual body functions as a central recycling hub, as the current data are also consistent with alternative models.

    2. Reviewer #1 (Public review):

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting and metabolic compartments) are divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to better understanding the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites. In particular, this study shows that the residue body is a region of the cell syncytium that organelles can be actively transported from. Therefore, it is a space that can actively contribute to the segregation of the late segregating micronemes and rhoptries.

      Weaknesses:

      In addressing the question of residual body participation in sorting of organelles, a clear definition of this structure is required including when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. The authors' definition is as follows: 'The RB originates from the collapse of the maternal parasite during daughter cell budding and occupies the space previously occupied by the mother cell.' As such, a clear marker of the mother cell 'collapse' is required, but such a marker is not identified or used in the study to separate what might be considered an active part of the mother cell during early daughter formation, and the residual body. This might seem like moot a point, but it would help to give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? The authors elegantly show that MyoF is necessary for segregation of micronemes and rhoptries into daughters, and that MyoF depletion leads to accumulation of these organelles within the residual body. Moreover, restored expression of MyoF can then recover these organelles. This clearly demonstrates the activity of the residual body as part of the syncytium space that participates in the maintenance of the vacuole. But does it imply that this space necessarily handles all inherited micronemes and rhoptries as a 'trafficking hub'? My concern with the lack of a clear definition could provide some misinterpretation or overinterpretation of the contribution residual body.

      A further, remarkable conclusion is that maternal micronemes are evenly segregated into daughters through an active process for 'balanced microneme inheritance'. The proportion of maternal micronemes is quantified up to the 8-cell stage and shown to be not significantly different between cells. But would this result be expected with random assortment at this stage? The authors model the probability of a 32-cell stage vacuole occurring with each daughter having within 0-3 maternal micronemes and this is considered unlikely. However, the authors neither present the modelling for the 8-cell stage or show quantification of 32-cell vacuoles. They do show some images of large vacuoles, but it is not possible to determine the distribution of maternal micronemes in these images. A regulated process of segregation would require a complex mechanism where some form of microneme counting would be required to create the proposed balance. It is, therefore, important to have strong data supporting such a hypothesis, but this is not currently presented.

    3. Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of toxoplasmosis. Parasite invasion of host cells, intracellular replication, and subsequent egress, which results in destruction of the infected cell, are central to pathogenicity. This manuscript focuses on understanding how maternal resources, specifically cellular organelles, are shared between daughter parasites during cell division. Many organelles are present as a single copy, making their division and inheritance essential for successful replication. In T. gondii, our understanding of how organelles are divided during cell division remains limited, and this study helps address this important knowledge gap.

      Strengths:

      The major strength of this study is the use of a Halo-based pulse-chase assay to characterize patterns of organelle inheritance and to monitor protein synthesis, turnover, and movement. This approach will be of considerable interest to the field. Using this method, the authors identify three major modes of organelle inheritance:

      (1) Organelles present in multiple copies (such as micronemes and rhoptries) are partitioned between daughter parasites, with additional contributions from newly formed vesicles. Newly synthesized and pre-existing material remain as distinct populations within the cell.

      (2) Single-copy organelles, such as the Golgi and apicoplast, are expanded through the incorporation of newly synthesized material before division.

      (3) Cytoskeletal structures are synthesized de novo during each round of cell division.

      These findings provide a more refined understanding of organelle inheritance and demonstrate that secretory organelles are not generated entirely de novo during each round of division, as was previously thought.

      The paper places particular emphasis on the fate of maternal micronemes and rhoptries during division. The data show that (1) during division in wild-type cells, maternal micronemes and rhoptries are detectable in the residual body (RB); however, the majority of these organelles are localized within the parasite body, either at the apical or basal ends of the daughter parasites (Fig. 6). (2) In the absence of the myosin motor MyoF, micronemes and rhoptries accumulate in the residual body and are not properly trafficked to the daughter cells. Upon restoration of MyoF protein levels, these organelles redistribute to the daughter cells, although in an uneven manner.

      Weaknesses:

      The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. The authors strongly argue that the RB is a central hub for recycling micronemes and rhoptries; however, this conclusion is not fully supported by the data. For example, the authors state:

      Line 227:<br /> "Notably, after endodyogeny was completed, M-MIC2 was redistributed from the RB to the apical tip of the daughter cells (Figure 6A, 11:30, 16:00 h), confirming that the RB serves as a temporary reservoir during microneme recycling (Periz et al., 2019)."

      Line 231:<br /> "In approximately 90% of parasites undergoing replication, M-RON2 was integrated into daughter rhoptries prior to mother cell collapse and formation of the RB (Figure 6B, 4:15-4:30 h and 10:30-10:45 h). Like M-MIC2, M-RON2 was occasionally detected in the RB, though less prominently, suggesting more rapid, tightly regulated, or more efficient recycling due to their lower number."

      Line 326:<br /> "However, we show that the RB temporarily stores maternal secretory organelles, such as micronemes and rhoptries, which are later redistributed to daughter cells in a MyoF-dependent manner (Figure 9B)."

      Thus, the model that all microneme and rhoptry trafficking is RB-dependent is based primarily on the MyoF depletion phenotype (which results in RB accumulation) together with the observation that a relatively small amount of maternal microneme and rhoptry material is detectable in the RB of wild-type parasites. Although the authors' interpretation-that recycling is RB-dependent-is one possible explanation, alternative models are not discussed.<br /> For example, an alternative possibility is that the majority of micronemes and rhoptries are trafficked directly from the apical end of the mother parasite to the daughter cells without passing through the RB. In this scenario, only a subset of the organelles would enter the residual body, perhaps reflecting imperfect trafficking efficiency rather than an obligatory recycling step. Loss of MyoF would impair this trafficking pathway, resulting in the accumulation of secretory organelles within the RB. In other words, RB accumulation could be a consequence of MyoF depletion rather than evidence that all trafficking in wild-type parasites normally proceeds through the RB.

      This alternative interpretation seems particularly relevant for the rhoptries, given that the authors themselves state that "M-RON2 was integrated into daughter rhoptries prior to mother cell collapse and formation of the RB."

      Other comments:

      Figure S10C<br /> To determine whether microneme degradation occurs in the RB, the authors quantified the fluorescence intensity of individual micronemes in control parasites and following auxin washout, showing that after redistribution the fluorescence intensity of individual vesicles is unchanged. However, this is not the appropriate analysis to address the question being asked. To conclude that micronemes are not degraded, the authors would need to quantify the total fluorescence intensity within the entire vacuole. For example, if half of the micronemes were degraded, the remaining micronemes would be expected to retain the same fluorescence intensity as those in the control parasites. Thus, unchanged fluorescence intensity of individual vesicles does not exclude the possibility that degradation has occurred.

    4. Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughter-derived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strength:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are convincing. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the Reviewing Editor for their careful evaluation of our manuscript and for their constructive and insightful comments. Their suggestions have helped us to improve the clarity, rigor, and presentation of our work. In response to these comments, we have substantially revised the manuscript and performed several additional analyses and experiments, as summarized below.

      Major additions and modifications made during revision

      In response to the reviewers' comments, we have substantially revised the manuscript and performed several additional analyses and experiments:

      New analyses

      - Quantification of MyoF recovery following auxin washout using MyoF-mAID-HA immunofluorescence (Figure 8B, Figure S10D).

      - Quantification of maternal MIC2 fluorescence intensity following 24 h MyoF depletion and subsequent redistribution after auxin washout (150 micronemes per condition; Figure S10A,B).

      - Pearson correlation analysis of ANKER1-Halo and HDEL-GFP localization (Pearson's R = 0.92 ± 0.04; n = 30 parasites).

      - Additional probability-based analysis supporting regulated microneme inheritance.

      - Expanded analysis of microneme redistribution across larger replication stages (Figure S5D).

      New figures

      - Figure S6: Dual-labelling analysis of additional Group 2 organelles (ER, apicoplast, glideosome).

      - Figure S10A, B: MIC2 fluorescence intensity analysis following RB retention and redistribution.

      - Figure S10D: Correlation between MyoF recovery and phenotype rescue.

      - Figure S11: Schematic overview of quantification and analysis workflow.

      Additional experimental efforts

      - Generation of a MIC2-Halo / IMC1-mKATE / Cb-Emerald parasite line to improve visualization of RB-associated trafficking.

      - Multiple attempts to perform higher-temporal-resolution live-cell imaging. However, prolonged acquisition resulted in severe phototoxicity, replication arrest, and parasite death, preventing reliable long-term recordings.

      Textual and methodological revisions

      - Expanded Materials and Methods section with detailed descriptions of fluorescence quantification, colocalization analyses, and statistical procedures.

      - Re-evaluation of statistical analyses using two-tailed tests throughout.

      - Revision of manuscript text to clarify the evidence supporting RB-associated trafficking and to better acknowledge current limitations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags, this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting, and metabolic compartments) is divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to a better understanding of the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites.

      Weaknesses:

      In addressing the question of residual body participation in sorting of organelles, it would be useful to clearly define this structure and when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. This might seem like a moot point, but it would give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation, and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? In subsequent cell replications (4, 8, 16... stages), is there a separation between the residual body that links them all and the posterior of each new 'mother cell', and if so, when is this distinction lost? This is important because without a definition, we might be confusing different processes.

      We thank the reviewer for this excellent and thoughtful question. The residual body (RB) emerges at the end of the first replication cycle, where it is delineated by the basal complex and persists as an IVN-associated compartment connecting all daughter parasites through both plasma membrane and cytoplasm. Previous EM and live-cell studies, including ours, have shown that the RB is not a passive remnant but a dynamic structure dependent on F-actin and unconventional myosins, supporting recycling, inter-parasite connectivity, and synchronous growth (Delbac et al., 2001; Muñiz-Hernández et al., 2011; Frénal et al., 2017; Periz et al., 2017).

      In the present study, the MyoF reversibility experiment provides strong support for a model in which RB functions as an active recycling hub. Upon MyoF depletion, maternal microneme and rhoptry proteins accumulate within the RB. Following restoration of MyoF expression, this material is redistributed to daughter organelles. We interpret this reversible phenotype as evidence that the RB represents a distinct and regulated trafficking intermediate rather than simply a by-product of late daughter cell formation.

      We agree with the reviewer that mother cells retain a functional invasion apparatus until very late during daughter formation, and that the delayed release of micronemes and rhoptries likely contributes to their late trafficking toward the cell posterior. However, our data indicate that once released, these organelles transit through a defined RB compartment that actively participates in their recycling rather than merely reflecting positional constraints. This has been previously well illustrated for micronemes, which are trafficked along F-actin filaments within the residual body (Periz et al., 2019).

      At later rounds of replication (4, 8, 16 parasites), previous studies have demonstrated the presence of multiple residual body centres within the same vacuole. However, the precise temporal and structural distinction between the RB linking parasites within the vacuole and the posterior of newly formed mother cells remains insufficiently resolved and is beyond the scope of the present study. Importantly, available ultrastructural and live-cell imaging supports the persistence of shared RB compartments connecting parasites within a vacuole, arguing against a simple conflation of posterior membranes and residual body material.

      While the primary aim of the current work was to investigate the RB's role in organelle recycling, we fully agree that a more precise definition of when and how the RB is formed, remodelled, and ultimately resolved during successive replication cycles will be essential to distinguish recycling from spatial constraints. We have revised the Discussion to better acknowledge this limitation and to avoid overinterpreting the role of the RB in organelle inheritance.

      Are rhoptries/micronemes that originate in one 'mother' able to be sorted to the 'daughters' from a distinct mother in this syncytium? If so, this would make it a sorting centre, but otherwise we could be just capturing the activities at the posterior of any given cell during replication. The authors' further thoughts on this would be very interesting.

      We agree with the reviewer that our current data do not definitively demonstrate whether rhoptries or micronemes originating from one “mother” parasite can be redistributed to daughters derived from another mother within the same syncytial vacuole. Nevertheless, our MyoF chase experiments are consistent with a model in which the RB/IVN functions as an active recycling and sorting hub rather than simply representing posterior trafficking events associated with individual parasites.

      Upon MyoF depletion, maternal micronemes accumulated within the RB. Following restoration of MyoF expression, these accumulated micronemes were subsequently redistributed to daughter parasites. This reversible redistribution is more consistent with an active recycling process than with passive accumulation alone.

      To further support this interpretation, we expanded the analysis presented in Figure S5 by including additional vacuoles and larger replication stages (new panel D). These analyses show that maternal micronemes are redistributed broadly and relatively evenly among daughter parasites. We additionally performed a probability-based analysis demonstrating that the recurrent and homogeneous redistribution patterns observed are highly unlikely to arise from stochastic capture events occurring independently at the posterior end of each parasite during replication. Together, these analyses support the interpretation that microneme redistribution is a regulated process.

      Direct demonstration of recycling between all parasites within a vacuole would require a system allowing simultaneous differential labeling of (i) daughter parasites derived from a specific mother cell and (ii) the maternal organelles originating from that same mother during a subsequent replication cycle. To our knowledge, such an approach is not currently technically feasible. Nevertheless, our live-cell imaging experiments provide additional support for communal redistribution, as microneme material accumulated within the RB was subsequently observed redistributing, albeit unevenly, across multiple tachyzoites within the same vacuole.

      The Group 2 structures are described as those that are divided between daughters and receive newly synthesised proteins that add to the maternal protein of these compartments. While this is a logical conclusion for several that are mentioned, where the maternal protein signal is seen to be depleted with replication (including for the apicoplast, ER, glideosome, and Golgi). Data for the addition of new proteins to these existing structures is actually only presented in direct support of this for the Golgi.

      We thank the reviewer for this important clarification. We initially selected the Golgi as a representative example because its morphology and restricted localization provide the clearest visualization of the dual-labeling dynamics. However, the same experimental approach was applied to all Group 2 organelles analyzed in this study. To address the reviewer's concern more directly, we have now included a new supplementary figure (Figure S6) showing that the same pattern is also observed for the apicoplast, ER, and glideosome.

      We would also like to clarify that the maternal protein signal is not lost during replication but instead becomes progressively diluted as these organelles expand, are partitioned into daughter parasites, and incorporate newly synthesized proteins. The Golgi was originally highlighted because these dynamics are most readily visualized in this compartment; however, the same principle applies to all Group 2 organelles analyzed in this study, as now illustrated in Figure S6.

      Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of Toxoplasmosis. Parasite invasion into host cells, intracellular replication, and then egress, which results in the destruction of the infected cell, is central to pathogenicity. This manuscript focuses on understanding how maternal resources (in this case, cellular organelles) are shared between daughter parasites during cell division. Many organelles are single copy, meaning that division and inheritance by the daughters is crucial for successful replication. The major strength of this study was the use of a Halobased pulse chase assay to characterize patterns of organelle inheritance. The results show that both microneme and rhoptries (secretory vesicles) previously thought to be synthesized de novo are inherited by daughter parasites. Thus, this paper adds new insight to our understanding of cell division in this important parasite.

      Strengths:

      This study demonstrated that pulse labeling of proteins can be used to monitor protein synthesis, turnover, and movement. This approach will be of great interest to the field. Using this method, the authors demonstrate three main modes of organelle inheritance.

      (1) Organelles, where there are multiple copies (such as secretory vesicles, micronemes, and rhoptries), are divided between the daughter parasites, with additional contribution of newly formed vesicles. New and old material remain as separate entities in the cell.

      (2) Single-copy organelles, which are expanded to include newly synthesized material prior to division, such as the Golgi and apicoplast.

      (3) Cytoskeletal structures that are synthesized anew during each round of division. These studies provide more refined insight into patterns or organelle inheritance and demonstrate that secretory organelles are not made de novo during each round of division as previously thought. The paper has a logical flow, and overall, the data is presented in a clear and organized fashion.

      Weaknesses:

      (1) Descriptions of methodology and statistical analysis were incomplete.

      We agree with the reviewer that the description of the methodology and statistical analyses required further clarification. To address this, we have added a new supplementary figure (Figure S11) illustrating the experimental workflow, quantification strategy, and analysis pipeline. We have also expanded the Materials and Methods section to provide detailed descriptions of the experimental design, fluorescence quantification procedures, statistical analyses, and the number of biological replicates. These revisions provide a clearer and more comprehensive description of the methodology and data analysis.

      (2) There are inconsistencies between the data in Figures 1 and 5. In Figure 1, a small amount of maternal IMC is visible in stage 2 parasites. Although this is a ~90% reduction, these parasites should be quantified as parasites with material IMC. However, the graph in Figure 5C indicates that no material parasites have GAPM1a, given that graph 5C is a binary measure (present vs. absent), one would expect a non-zero percent of parasites to have maternal material.

      We agree with Reviewer 2 that, based on the raw fluorescence signal, one might expect a non-zero percentage of parasites to retain maternal IMC material after the first replication. The apparent discrepancy between Figures 1 and 5 reflects our thresholding strategy rather than inconsistent data.

      Figure 5C presents a binary analysis (presence versus absence) using a threshold calibrated from stage 1 parasites and applied uniformly across all markers. Under this criterion, the residual GAPM1a signal after the first replication falls below the detection threshold, resulting in 0% positive vacuoles. Although normalization to stage 2 parasites would detect this weak residual signal, such a protein-specific threshold would compromise direct comparison across the dataset.

      To clarify this point, we have updated the Figure 5C legend to explain the analytical approach and the asterisk associated with GAPM1a. The residual maternal IMC signal visible in Figure 1 represents a rare example selected to illustrate the remaining ~10% signal and is consistent with the absence of detectable maternal IMC1 after replication in Figures 2C and 5E.

      (3) The conclusion from Figure 6 was not justified based on the data. I agree with the author's conclusion that the accumulation of micronemes and rhoptries in the residual body was timedependent. In Figure 6A, the signal observed in the residual body at times 6:30, 13, and 14 hours is not observed in subsequent time points. However, the fate of these micronemes and rhoptries is unclear. It cannot be concluded that these vesicles are recycled back to the mother. They could also have been degraded. In fact, the graphs of microneme inheritance in Figure 2B show a decrease in maternal signal from 100% to 80% between stages 1 and 2, indicating that some microneme degradation is taking place.

      We agree with the reviewer that Figure 6 alone does not definitively establish the fate of micronemes and rhoptries accumulating within the residual body (RB), and that both recycling and degradation remain possible interpretations. Our conclusion that maternal micronemes are predominantly recycled is therefore based on the integration of Figure 6 with our MyoF depletion and recovery experiments, additional quantitative analyses, and previous work demonstrating F-actin-dependent microneme trafficking through the RB (Periz et al., 2019).

      Consistent with this model, MyoF depletion results in the accumulation of maternal micronemes within the RB, whereas restoration of MyoF expression following auxin washout leads to their redistribution across multiple tachyzoites within the same vacuole (Figure 8). Furthermore, maternal microneme signal remains detectable even after prolonged MyoF depletion (up to 48 h) and multiple rounds of replication (Figures 7 and 8), arguing against extensive degradation.

      To further address this possibility, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes retained within the RB after 24 h of MyoF depletion and following redistribution after auxin washout (150 micronemes per condition). No significant difference in fluorescence intensity was observed compared with control maternal micronemes (Figure S10A,B), indicating that maternal microneme signal is preserved during RB retention and redistribution.

      We therefore interpret the decrease in maternal microneme signal observed between stages 1 and 2 in Figure 2B primarily as a consequence of redistribution and dilution rather than degradation, consistent with the stable fluorescence intensity of individual micronemes (Figure 3). Regarding rhoptries, we note that the majority (~90%) are incorporated into daughter parasites before budding is complete, limiting their accumulation within the RB and suggesting that RB-associated trafficking primarily reflects redistribution rather than bulk degradation.

      (4) To convincingly demonstrate that the redistribution of micronemes and rhoptries was due to recovery of MyoF protein levels after auxin washout, a Western blot should be performed to show MyoF protein levels over time. In addition, the decrease in mMIC2 protein levels in the residual body in Figure 8F should be measured and normalized for photobleaching. Both apical and basal signals appear to be reduced over the time course of imaging.

      We agree with the reviewer that demonstrating MyoF recovery following auxin washout is important. Rather than performing a Western blot, we monitored MyoF recovery by immunofluorescence using the HA tag in the MyoF-mAID-HA strain, allowing direct correlation between MyoF reappearance and microneme redistribution at the single-vacuole level. These data are now included in Figure 8B, with the corresponding MyoF presence–phenotype association analysis presented in Figure S10D.

      Regarding photobleaching, we agree that fluorescence loss during time-lapse imaging is an important consideration. However, in this experiment, changes in fluorescence intensity reflect not only photobleaching but also biological redistribution of micronemes and movement of parasites in and out of the imaging plane. In the absence of a stable internal reference fluorophore, applying a standard photobleaching correction could therefore introduce additional inaccuracies. For this reason, we did not quantify fluorescence intensity during the redistribution phase.

      Instead, to assess whether maternal microneme signal is lost during RB retention and redistribution, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes following 24 h of MyoF depletion and subsequent auxin washout (Figure S10A,B). No significant difference was observed compared with control maternal micronemes, supporting the conclusion that redistribution occurs without substantial loss of the maternal microneme pool.

      Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughterderived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns, which they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strengths:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are compelling. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weaknesses:

      (1) The evidence provided to support the claim that microneme and rhoptry inheritance specifically traffics through the residual body does not sufficiently substantiate the claim. The temporal resolution of the imaging is inadequate to precisely trace the path of microneme and rhoptry inheritance. From the data shown in the manuscript, it can be concluded that at least some of the micronemes and rhoptries might be recycled through the residual body, but it is unclear whether many or most of these organelles do so.

      We thank the reviewer for this important comment and refer also to our response to Reviewer 1 above.

      Previous work has demonstrated F-actin-dependent trafficking of micronemes within the residual body (RB) (Periz et al., 2019). Consistent with these findings, our data support a model in which RB-mediated trafficking contributes to maternal microneme recycling. We acknowledge, however, that the temporal resolution of our imaging does not allow continuous tracking of every individual organelle throughout the entire replication process.

      In contrast, our observations indicate that the majority of maternal rhoptry material is incorporated into daughter cells before replication is complete and therefore does not necessarily transit through the RB under normal conditions (Figure 6). Nevertheless, rhoptry inheritance remains dependent on the actin–MyoF trafficking machinery, as MyoF depletion results in the accumulation of rhoptry material within the RB (Figures 6 and 7).

      Taken together, our data support a model in which the RB serves as an important recycling hub for maternal micronemes and can, under conditions of impaired trafficking, also transiently accommodate rhoptry material. However, our imaging resolution does not allow us to conclude that all microneme or rhoptry inheritance obligatorily transits through the RB, and we have revised the manuscript to reflect this limitation more explicitly.

      (2) The absence of specific markers for the residual body brings into question whether microneme inheritance occurs through a discrete residual body or simply via the basal end of the maternal parasite. The authors need a robust way to visualize and define the residual body to claim that micronemes and rhoptries are specifically transported through this structure.

      We agree with the reviewer that the absence of a dedicated residual body (RB) marker remains a limitation and that such a tool would improve the precision of our analyses. To date, no specific RB marker has been identified (see also our response to Reviewer 1). The most reliable proxy currently available is the F-actin chromobody, which labels the dense F-actin network associated with the RB. Using this approach, previous work demonstrated F-actin-dependent trafficking of micronemes within the RB (Periz et al., 2019).

      Building on these findings, our data support a model in which RB-associated trafficking contributes to maternal microneme recycling, whereas rhoptries are more frequently incorporated directly into daughter cells without obvious RB transit. In addition, functional perturbation of the actin–MyoF transport machinery, through MyoF depletion and subsequent recovery, supports the interpretation that the RB represents a discrete actin-associated compartment involved in organelle redistribution. Nevertheless, we acknowledge that our current imaging resolution does not allow us to determine the extent to which all microneme or rhoptry inheritance occurs through the RB, and we have revised the manuscript accordingly.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Comments for revision where either the clarity or accuracy could be improved:

      (1) The methods could do with some further detail with respect to the fluorescence intensity measurement. For example, where the Z-series was taken, and for measurements, were maximum projects taken or single Z-planes? Were all measurements made on unprocessed images, or was any deconvolution, etc, undertaken?

      We updated the Materials and Methods for better clarity and now read as: “Parasites were labeled as described above and allowed to replicate for 24 h on HFF-coated Ibidi live-cell dishes. Approximately 15 fields of view were imaged using Z-stacks spanning 3 μm centred on the vacuoles. For each replication stage (1, 2, 4, and 8 parasites per vacuole), individual tachyzoites were sampled across multiple vacuoles.

      Maximum-intensity projections were generated from non-deconvolved images. Fluorescence intensity (FI) was quantified as the maximum grey value measured within regions of interest (ROIs) drawn on individual tachyzoites from each vacuole stage. ROIs excluded overlapping parasites, neighbouring vacuoles, and regions with atypical signal intensity. For each biological replicate, up to 25 tachyzoites per replication stage were analyzed, and the mean FI value was calculated for each stage. The highest mean FI observed among the stages within a replicate was defined as 100%, and FI values for the other stages were expressed relative to this maximum. Relative FI values were then averaged across three independent biological replicates. Data are presented as mean ± SD.”

      (2) Line 93: Is more JF549 added at each of stages 2, 4, and 8? I assume so to get the progressive increase, but it would help to clarify this here.

      As illustrated in Figure 1a, the second dye is added only once, after 24 h of replication, and is then washed off before imaging. It is not maintained throughout the replication steps. The observed increase in signal is related to the presence of newly synthesized (de novo) proteins generated during replication. The Halo tags of these new proteins are initially free of ligand, as no ligand is present during replication. During the second labelling step, all free Halo tags can bind the dye. The signal intensity at each replication step therefore reflects the amount of de novo material present, which is higher at step 4 than at step 2, because the proteins generated de novo during stage 2 are also present in stage 4. The signal will be determined by both the number of newly produced molecules and their concentration at the same localization.

      (3) Line 103: Check that the Carruthers and Sibley, 1997, ref for Tic20 in the apicoplast is correct. I don't think this could be correct given the date.

      The reference it will be corrected to G van Dooren et al. 2008.

      (4) Figure legends: It would be useful to state what form of microscopy was used in each figure.

      Following the reviewer's advice the legends has been updated.

      (5) Figure 2, S2: How are single organelles tracked, such as rhoptries? I'd assume that a cell will gain new de novo organelles as well, and that this would reduce the signal per cell. Stage 2 has a rhoptry signal in both daughters, so I'd expect the signal to be roughly half for the whole cell, unless the authors can resolve individual rhoptries (which would surprise me with this microscopy). If individual rhoptries were resolved, how was this done, and what was the confidence in this (were there controls?)

      We thank the reviewer for this important point. We do not resolve individual rhoptries with the imaging conditions used in this study. Instead, fluorescence intensity was measured as the maximum grey value within a representative region of interest (ROI) encompassing the apical rhoptry signal while excluding overlapping parasites and regions with atypical fluorescence intensity. The same ROI selection strategy was applied consistently across all replication stages, allowing direct comparison with stage 1 parasites.

      The analyses presented in Figures 3 and S2 show an increasing proportion of tachyzoites lacking detectable maternal rhoptry signal as replication progresses, while the fluorescence intensity of the remaining maternal signal remains relatively stable. Together, these observations are consistent with the redistribution of intact maternal rhoptries rather than a progressive loss of rhoptry fluorescence. 

      (6) Line 150: it is unclear what is meant by 'regulated partitioning'. The more equal inheritance of micronemes versus rhoptries might not indicate a 'regulated partitioning' but just a more uniform distribution, given the larger number of micronemes versus rhoptries.

      We agree with the reviewer that this statement required clarification, and we have revised the text accordingly. Our intention was to emphasize that microneme inheritance appears to be a regulated process, rather than to directly compare it with rhoptry inheritance. The observed differences between these organelles are likely influenced, at least in part, by their different abundances.

      Across successive rounds of replication, daughter parasites consistently inherit comparable amounts of maternal micronemes, even at later replication stages (Figure S4). Given that a single mother parasite contains approximately 30–40 micronemes, whereas successive rounds of endodyogeny can generate up to 32 daughter parasites, a purely stochastic segregation would be unlikely to produce the relatively uniform distribution observed (~1–2 maternal micronemes per tachyzoite). To support this interpretation, we performed an additional probability-based analysis, which indicates that the observed redistribution patterns are unlikely to arise by chance alone. We therefore interpret these findings as supporting the existence of mechanisms that promote balanced microneme inheritance during parasite replication.

      (7) Line 184: How is the maternal signal measured without detecting the internal daughter signal? Is this an average signal for the full parasite, or just for a cross-section of the IMC? And if the latter, how are the different profiles of mother and daughter accounted for? Also, the abbreviation in the brackets doesn't make sense here.

      We thank the reviewer for this important point. We have revised the Materials and Methods section to provide a clearer description of the fluorescence intensity (FI) measurements and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, FI measurements were performed on maximum-intensity projections generated from Z-stack images without deconvolution. A representative region of interest (ROI) was selected, and the maximum grey value was used for quantification. This approach minimizes variability arising from differences in ROI size and provides a robust metric for comparison across replication stages.

      Daughter cell fluorescence was measured using the same approach while excluding overlapping signals from neighboring daughter cells and the maternal IMC. Maternal and daughter signals were distinguished based on their spatial localization and fluorescence labeling. Finally, the abbreviation in brackets has been corrected for clarity.

      (8) Line 191: Subheading a bit unclear. Distinct from other organelles, or are miconeme and rhoptry pathways distinct from each other?

      We agree with the reviewer and have updated the subheading to “Whole-organelle inheritance of micronemes and rhoptries occurs via distinct recycling pathways”

      (9) Line 193: The site of disassembly of the IMC (suggested RB here) might not be the same as the site of degradation. I suggest using 'disassembly' instead here.

      We agree that “disassembly” is an appropriate term to describe the breakdown of the IMC at the residual body (RB). However, we also believe that the RB represents the primary site of IMC degradation, for two reasons. First, if IMC material were not degraded at this site, we would expect to detect Halo-positive signal elsewhere following IMC collapse, which we do not observe. Second, transport of IMC material to an alternative degradation site would be required, but no IMC-positive vesicles are observed, arguing against significant redistribution. Together, these observations support the conclusion that the RB is both the site of disassembly and degradation of maternal IMC.

      (10) Line 215: The conclusion for a difference in timing of microneme and rhoptry segregation is not clearly supported by the data presented. Also, if there are more micronemes than rhoptries, then the frequency of observing a microneme being trafficked through the RB would need to be higher than for rhoptries if the mechanisms were the same. So, a difference in frequency here cannot be used to argue for a different mechanism.

      We agree with the reviewer and the text have been edited to soften our conclusion. 

      (11) Line 234: 'segregation' might be a better term than 'recycling' here because it is actually the sorting into daughter cells that is the important process.

      The text have been edited

      (12) Line 235: I don't think this can be what the authors intend to say. If the maternally-inherited rhoptries are not trafficked through the RB (every time), then how do they get into the daughters? Perhaps this is a case where a clear definition of the RB is required.

      Our observations indicate that maternally inherited rhoptries are frequently incorporated into daughter cells before collapse of the mother cell and establishment of the residual body (RB). Although F-actin is enriched within the RB, an actin network is also present throughout the parasite cytoplasm, where MyoF is likewise localized. We therefore propose that, unlike micronemes, rhoptries do not necessarily transit through the RB during every replication cycle but can be incorporated directly into developing daughter cells while still relying on the same actin–MyoF-dependent trafficking machinery.

      (13) The MyoF Rescue, the experimental plan is not fully described in order to be clear. If the endomembrane architecture was disrupted by MyoF depletion, and this secondary effect caused the segregation phenotype, restoration of MyoF might also simply restore the endomembrane system. So a direct role for MyoF doesn't seem to have been tested in this case.

      We appreciate the reviewer's concern that the segregation phenotype could, in principle, arise indirectly from disruption of endomembrane architecture following MyoF depletion. However, although Golgi morphology is altered in MyoF-depleted parasites, its core functions appear largely preserved. This is supported by the normal biogenesis of de novo micronemes, their correct targeting to the apical pole, their efficient secretion, and the previously reported preservation of parasite invasion. In addition, Golgi markers are not detected in the residual body, where maternally inherited micronemes accumulate, arguing against Golgi-mediated trafficking as the primary cause of the segregation phenotype.

      Taken together, these observations support the interpretation that the segregation defects are more likely to reflect a direct role of MyoF in organelle trafficking and inheritance than a secondary consequence of generalized disruption of endomembrane organization.

      (14) Line 279: Why call it a checkpoint? What is the evidence for its presence here being sensed before a further process is activated, which is what a checkpoint does?

      We agree the reviewer that checkpoint is a misleading term and have been updated to trafficking hub. 

      (15) Line 287 confuses replication of the daughters from cytokinesis, which only happens when each cell loses cytoplasmic connectivity with the other.

      We will clarify this point. In Toxoplasma gondii, cytokinesis represents the final step of daughter cell formation, during which the two fully assembled daughter parasites separate from the mother cell following collapse of the maternal cytoplasm. Historically, the residual body was proposed to arise simply as leftover material from this process. However, multiple studies have now shown that residual body formation is an active and regulated process, dependent on specific cytoskeletal and trafficking factors. Importantly, although cytokinesis marks the physical separation of daughter cells from the mother, parasites within a vacuole remain connected via the residual body and continue to share cytoplasmic and plasma membrane components until egress. 

      (16) Line 298: I don't think there is direct evidence of degradation in the RB. There might be disassembly, but degradation implies proteolysis, which hasn't been tested for.

      We agree with the reviewer that our data do not provide direct biochemical evidence of proteolysis within the residual body (RB) and primarily demonstrate disassembly of the maternal IMC at this site. However, several observations are consistent with local degradation. Following IMC collapse, we do not detect Halo-positive signal elsewhere in the parasite, nor do we observe IMC-positive vesicles or other structures that would suggest transport to a distinct degradation compartment.

      In addition, previous work identified the E3 ubiquitin ligase CSAR1 as a mediator of protein turnover within the RB, supporting the idea that this compartment is associated with degradation-related processes (O'Shaughnessy et al., 2023). While we cannot formally demonstrate proteolysis, these observations support a model in which IMC disassembly is closely coupled to local degradation within the RB.

      (17) The paragraph structure gets a bit confusing at times. See single sentence paragraph, Line 224. Does this sentence justify its own paragraph?

      The text has been edited.

      (18) Make sure Toxoplasma gondii is in italics throughout.

      The text has been edited

      (19) Line 279 cites Figure 10. But there is none.

      The text has been edited

      (20) I advocate introducing a few new acronyms, like DCs. I find that this ultimately reduces the ease with which readers read the work if they don't learn them all quickly.

      We agree that excessive use of acronyms can negatively impact readability. In the present manuscript, all abbreviations used in the text are introduced at their first occurrence in the Introduction, including DCs (line 32), IMC (line 41), PV (line 35), ER (lines 38–39), RB (line 50), and IVN (line 49). We have carefully limited the use of abbreviations to commonly used terms in the field and to those that recur frequently throughout the manuscript, with the aim of balancing clarity and readability. Nevertheless, we are happy to reduce or remove specific abbreviations if the reviewer feels this would further improve clarity.

      Reviewer #2 (Recommendations for the authors):

      (1) Descriptions of methodology and statistical analysis were incomplete as follows:

      (1a) It was unclear how the fluorescence intensity measurements (used to evaluate inheritance vs. new synthesis) were carried out. The y-axis on the graph is labeled average fluorescence intensity (% of max intensity). However, it does not state what was averaged (average fluorescence per vacuole?) and what was max intensity (max pixel intensity in each image or time point with the highest average intensity, relative to the other time points?)

      We agree with the reviewer that the original description of the fluorescence intensity (FI) measurements lacked clarity. We have therefore revised the Materials and Methods section and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, vacuoles were imaged as Z-stacks, and maximum-intensity projections were used for analysis. FI was quantified as the maximum grey value measured within representative regions of interest (ROIs) drawn on individual tachyzoites, rather than as an integrated fluorescence signal across the vacuole. This approach minimizes variability arising from differences in ROI size and allows direct comparison between replication stages.

      For each biological replicate, up to 25 tachyzoites per replication stage were analyzed. The mean FI for each stage was normalized to the highest mean value within that replicate, and data from three independent biological replicates were subsequently averaged.

      (1b) Given the uncertainties with how these measurements were performed, it is difficult to interpret the data. For example, one would expect that the fluorescence intensity of newly synthesized IMC1 in 8-parasite vacuoles would be 4 times higher than that of a 2-parasite vacuole; however, based on the graph in Figure 1B, the measured increase was only 30%.

      We agree that the original description of the fluorescence intensity (FI) measurements required further clarification and have revised the Materials and Methods accordingly. As the reviewer correctly notes, a fourfold increase in FI between 2- and 8-parasite vacuoles would be expected if total IMC fluorescence across the entire vacuole had been measured. However, this was not the parameter quantified.

      Instead, FI was measured as the maximum grey value within representative regions of the daughter IMC, providing a per-cell rather than a whole-vacuole measurement. Using this approach, FI increases between the 2- and 4-parasite stages and then reaches a plateau.

      This behavior is consistent with the biology of IMC biogenesis. Although the total amount of IMC per vacuole increases with parasite number, the amount of IMC protein incorporated into each daughter parasite remains relatively constant. Consequently, once daughter IMCs are fully assembled from de novo-synthesised material, additional rounds of replication increase the total IMC content per vacuole but not the fluorescence intensity measured for individual parasites.

      (1c) T. gondii replicates in an asynchronous manner, so that at the 24-hour time point, a single dish can contain vacuoles containing 2, 4, and 8 parasites. This should be stated explicitly so readers unfamiliar with T. gondii's growth patterns can understand how the experiment was performed.

      We agree with the reviewer and have updated the text line 93. “As Toxoplasma gondii replicates in an asynchronous manner, after 24 of replication, vacuoles containing 1, 2, 4, and 8 parasites can be observed in a single dish.”

      (1d) Colocalization package in Fiji used for ANKER1-Halo with HDEL-GFP and MIC2/RON2 with CbEmeraldFP should be specified.

      We thank the reviewer for this suggestion. Following this recommendation, we performed Pearson correlation analysis for the ANKER1–HDEL-GFP experiment using the Coloc 2 plugins of FiJi. ANKER1Halo and HDEL-GFP showed a strong spatial correlation (Pearson's R = 0.92 ± 0.04, n=30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER.

      We note, however, that this analysis should be interpreted as evidence for co-distribution within the same organelle rather than direct molecular colocalization, as ANKER1 is a transmembrane protein whereas HDEL-GFP labels the ER lumen.

      For all the rest of our analysis, no automated colocalization package or plugin in Fiji was used for the analyses involving ANKER1-Halo with HDEL-GFP or MIC2/RON2 with Cb-EmeraldFP. Colocalization was assessed manually across all experiments by inspecting both full Z-stacks and maximum-intensity projections to ensure robust spatial overlap.

      For MIC2 and RON2, the presence of signal within the Cb-Emerald–positive filament was scored as either cytoplasmic, on the residual body or absence of colocalisation.

      In total, more than 300 and 500 vacuoles were analyzed for MIC2 and RON2–Cb-Emerald colocalization respectively (stable expression of both markers), and more than 150 vacuoles were analyzed for ANKER1-Halo and HDEL-GFP colocalization (transient expression of HDEL-GFP).

      This information has now been added to the Methods section.

      (1e) Statistical methods should be described on an experiment-by-experiment basis. The authors should justify why a one-tailed t-test was conducted. A two-tailed t-test seems more appropriate.

      We agree that statistical methods should be clearly justified on an experiment-by-experiment basis. All the statistical analysis have been performed using two tails and updated in the figures.

      (1f) In Figures 5E and 5F, using boxes to indicate the exact areas of the cell that were used in the fluorescence intensity measurements, rather than arrows, would make this data easier to interpret.

      The figure has been updated.

      Reviewer #3 (Recommendations for the authors):

      The current time-lapse images and videos do not clearly demonstrate microneme movement from the maternal parasite apical end to the residual body and back to the apical end of daughter parasites. As such, the route by which micronemes enter daughter parasites remains inconclusive. To strengthen their claims, the authors should employ higher temporal resolution imaging to definitively capture the movement of micronemes from the maternal apical region into the daughters. From the current data, it also seems plausible that the micronemes may be trafficked into the daughters through the conoid as well, as there is no evidence provided showing a movement of micronemes away from the apical end of the maternal parasite before being present in the daughter parasites.

      We agree with the reviewer that higher temporal resolution imaging would provide a more definitive view of microneme trafficking. However, long-term live imaging of replicating Toxoplasma gondii requires a compromise between temporal resolution and parasite viability. In our experiments, images were acquired every 15–30 min over periods of up to 16 h, as more frequent acquisition consistently induced phototoxicity and prevented completion of parasite replication.

      Despite this limitation, our imaging reliably tracked maternal micronemes over successive rounds of endodyogeny and consistently showed microneme signal associated with the residual body. These observations are in agreement with previous high-temporal-resolution studies, which demonstrated F-actin-dependent microneme trafficking within the residual body over shorter imaging periods (Periz et al., 2019).

      We cannot formally exclude the possibility that some micronemes are transferred directly to daughter parasites through the apical end. However, together with previous studies showing enrichment of F-actin at the basal region of developing daughter cells rather than at the apical tip (Periz et al., 2017), our observations support a model in which RB-mediated trafficking contributes to maternal microneme inheritance.

      The lack of a clear residual body marker needs to be addressed, as the distinction between the basal end of the maternal cell and a bona fide residual body must be explicitly defined to substantiate the major claim of the study. As it stands, it remains unclear whether micronemes and rhoptries as a whole travel through the residual body to be transported into the daughter parasites.

      We agree with the reviewer that a marker specific to the residual body would strengthen this study. Unfortunately, no such marker has been identified to date. The F-actin chromobody is currently the best available proxy, as previous studies have shown that the F-actin network is enriched within the residual body (Periz et al., 2017; Kellermeier et al., 2024). Moreover, high-resolution live-cell imaging has previously demonstrated F-actin-dependent microneme trafficking within this compartment (Periz et al., 2019). We have revised the manuscript to more clearly acknowledge this limitation.

      The conclusions drawn from the actin colocalization data in Figure 6C are based entirely on fixed samples, despite all experimental tools being compatible with live-cell imaging. Supplementing the fixed imaging with live cell data would increase its biological relevance. Published studies have shown that fixation of the actin chromobody results in the loss of resolution of an appreciable amount of the cytosolic F-actin network, and while the localizations analyzed here are primarily along the periphery, since the quantification and text make claims about the colocalization within the cytosol, this potential loss of cytosolic F-actin becomes an issue as there may be more actin available for analysis that is lost due to fixation within the cytosol of the parasites.

      We thank the reviewer for raising this important point and agree that conventional fixation can compromise preservation of the F-actin network. However, we used the same fixation protocol described by Periz et al. (2019), which allows reliable visualization of the RB-associated F-actin network. We have also corrected the description of the fixation protocol in the Materials and Methods.

      Fixation was necessary to image entire vacuoles with sufficient spatial resolution and signal-to-noise ratio for the volumetric analyses presented in Figure 6C. Although some loss of cytosolic F-actin cannot be excluded, this would be expected to reduce, rather than artificially increase, the detection of organelle–actin associations.

      Importantly, previous live-cell imaging studies demonstrated F-actin-dependent microneme trafficking (Periz et al., 2019), and our observations are consistent with these findings. Moreover, the defects observed following MyoF depletion provide independent functional evidence that the trafficking events described here rely on the actin–MyoF transport machinery.

      The statement of colocalization should be backed up by quantitative coefficients like Pearson's coefficient.

      We thank the reviewer for this helpful suggestion. Following this recommendation, we performed a Pearson correlation analysis of ANKER1-Halo and HDEL-GFP using the Coloc 2 plugin in Fiji. ANKER1-Halo showed a strong spatial correlation with HDEL-GFP (Pearson's R = 0.92 ± 0.04, n = 30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER. As ANKER1 is a transmembrane protein and HDEL-GFP labels the ER lumen, this analysis should be interpreted as evidence of co-distribution within the same organelle rather than direct molecular colocalization.

      In contrast, we do not consider Pearson's coefficient appropriate for evaluating the association of micronemes or rhoptries with F-actin. These organelles are predominantly concentrated at the apical pole and, when associated with F-actin, are typically positioned along rather than directly overlapping the filaments. Consequently, Pearson's coefficient would underestimate these biologically relevant associations. We therefore relied on morphological and spatial criteria, which we consider more appropriate for assessing organelle–cytoskeleton interactions.

      In addition, from the methods and presented figure images, specifically in Figure 6C for RON2, how the cytosolic and residual body actin is separated is difficult to discern, as there is a clear residual body actin signal overlapping a parasite. The methods for how this was separated and analyzed should be clearer to remove doubts about how this area was measured, as the current description raises concerns about the counting of the residual body actin within the cytosol.

      We agree with the reviewer that the distinction between cytosolic and residual body (RB)-associated F-actin required further clarification. All analyses were performed manually, as described in our response to Reviewer 2 (comment 1d). The RB was identified by the presence of thick, bundled F-actin filaments at the basal pole that formed a continuous structure connecting parasites within the vacuole, whereas cytosolic F-actin was defined as the thinner filamentous network within the parasite body.

      No automated or threshold-based segmentation was used because the marked differences in filament morphology and fluorescence intensity make reliable thresholding difficult and prone to misclassification. Manual annotation based on spatial localization and filament morphology was therefore considered the most appropriate approach. We have clarified these criteria in the Materials and Methods section.

      Line comments:

      (1) 113: round to rounds - "did not obtain maternal organelles after successive rounds of replication...".

      The text has been updated

      (2) 146: grammatical, "As consequence a progressive" -> "As a consequence", or "Consequently".

      The text has been updated

      (3) 164-166: "Autonomous duplication" implies the separation and duplication of the Golgi occurs on its own, i.e., without any outside intervention, when we know from Carmeille et al. 2021 and Figure 7C here that the Golgi becomes fragmented over rounds of division in the absence of MyoF. I think this is primarily a word choice error with "autonomous".

      We agree with the reviewer and the word autonomous has been removed

      (4) 184: The wording suggests that DC's refers to daughter IMC's, when DC has already been given as an abbreviation for daughter cells previously.

      The text has been updated to correct this error

      (5) 189: de novo is not italicized.

      The text has been updated

      (6) 192: The data shown so far do not show that the RB plays a selective role in organelle recycling.

      The text has been edited to fit better our results “Our data suggest that the organelles trafficking through the residual body (RB) have different fate”

      (7) 208: State that it depends on F-actin, but never show that it is dependent on F-actin through actin disruption, such as cytochalasin D treatment or a specific conditional disruption of F-actin.

      We agree with the reviewer that we did not repeat F-actin disruption experiments (e.g., cytochalasin D or jasplakinolide treatments) in this study. These experiments were performed in our previous work, where pharmacological disruption of F-actin was shown to impair microneme trafficking (Periz et al., 2019). We therefore chose not to repeat these assays.

      Instead, the present study provides complementary evidence by demonstrating that depletion of Myosin F (MyoF), a motor that uses F-actin as a transport track (Kellermeier et al., 2024), disrupts the trafficking of both maternal micronemes and rhoptries. Together, our previous F-actin perturbation experiments and the MyoF depletion data presented here support the interpretation that these trafficking events depend on the actin–MyoF transport machinery. 

      (8) 209: This suggests that the chromobody was transiently expressed in the RON2-Halo line, but the methods suggest MIC2-Halo and RON2-Halo were integrated into a parasite line stably expressing Cb-Emerald.

      The text has been edited.

      (9) 229: The section is confusing with the mention of (now maternal). If I understand correctly, the point being made is that the de novo synthesized MIC2 at stage 2 is now the maternal MIC2 for stage 4, but coloring-wise within the figure, the now maternal MIC2 at stage 4 from stage 2 would still be green. The methods suggest these images were all taken simultaneously, and not at specific timepoints of the same vacuole, so the now maternal line remains confusing.

      The text has been revised for clarity and now reads: “Because Toxoplasma gondii replicates asynchronously, vacuoles at different replication stages coexist within the same culture after 24 h. The second labeling step marks all proteins synthesized since the beginning of the experiment, allowing discrimination between proteins present in the original mother parasite and those synthesized during subsequent replication cycles. As daughter parasites form, they inherit material from their mother, such that proteins synthesized during one replication cycle become maternal proteins in the next. Under MyoF depletion, these newly synthesized protein pools accumulate within the residual body instead of being redistributed to daughter parasites during subsequent rounds of replication (Figure 7, stage 4).”

      (10) 238: The sentences here indicate that Golgi inheritance occurs without issue: "golgi inheritance remained unaffected by MyoF depletion". But, it is evident from the images shown that the Golgi is extremely fragmented, with many more Golgi fragments by stage 8 than there are parasites. This could be solved by rewording and including a line along the lines of "in accordance with the results found in Carmeille et al. 2021".

      The reference to this article is already stated later in the text now line 299-303 “This active role is further supported by the dependence of RB-mediated recycling on F-actin and the class XXII myosin MyoF, which we show to be essential for retrieval of maternal MIC2 and RON2 but dispensable for Golgi inheritance although we noticed a fragmentation of the Golgi, which has been described to depend on MyoF (Carmeille et al., 2021).”

      (11) 319: missing a comma, "Many of these, particularly...".

      The text has been edited.

      (12) 437: Images -> imaged.

      The text has been edited.

      (13) 439: a fresh media -> and fresh media.

      The text has been edited.

      (14) 439: Replication -> replicate.

      The text has been edited.

      (15) 440: images -> imaged.

      The text has been edited.

      Other grammatical issues within the methods:

      (1) Figure 5C: Within the y-axis label of Figure 5C, there is an asterisk with no asterisk explanation within the legend.

      We thank the reviewer for pointing out this oversight. The figure legend has been updated to include an explanation of the asterisk, providing a clearer understanding of the results.

      (2) Figure 5E: The addition of an IMC1 label to match the magenta color would be helpful to readers.

      The figure has been updated

      (3) Figure 6D: No statistics showing significance of colocalization.

      We thank the reviewer for this comment. In Figure 6D, we report the frequency of observed colocalization between organelles (MIC2 and RON2) and F-actin filaments, based on manual analysis across >300 vacuoles for MIC2 and >500 vacuoles for RON2. Specifically, we observed cytoplasmic F-actin colocalization in 94% of vacuoles for MIC2 and 79% for RON2, and residual body F-actin colocalization in 85% of vacuoles for MIC2 and 11% for RON2. This analysis is descriptive and is not intended to compare MIC2 versus RON2 quantitatively; rather, it illustrates the general association of these organelles with F-actin filaments. Standard statistical measures such as Pearson correlation are not meaningful in this context because the organelles are punctate and primarily localized along filaments rather than overlapping continuously. Importantly, the percentages reported represent the fraction of vacuoles in which colocalization can be observed, not the percentage of colocalization between Cb-Emerald and MIC2/RON2 within individual vacuoles.

      (4) Figure 7: The legend title of Figure 7 is at the end of the legend of Figure 6.

      The text has been edited.

    1. eLife Assessment

      This study presents fundamental and significant findings that bovine mammary epithelial cells can be infected with both avian and human influenza A viruses, providing a potential site for viral reassortment. The evidence to support these claims is compelling. The work will be of interest to virologists and evolutionary biologists working on cross-species transmission of viruses and pandemic preparedness.

    2. Reviewer #2 (Public review):

      The authors use a library of influenza A viruses from different strains, classified in lab-adapted, human, avian, and swine according to the animal from which they were isolated. They propose that the cow mammary gland serves as a mixing vessel for influenza A viruses. As a first approach, the authors assess susceptibility to infection across different cell types, including continuous and primary cell lines, bovine mammary cells, and mammary explants. All these cells support polymerase activity. Then, they analyzed changes in the bovine virus's viral fitness relative to an avian precursor. The authors use single-gene replacement to study whether and which RNP segments improve viral transcription. As part of this section, they also test IFN-specific antagonism by NS1 to assess the input of segment 8. Quantitative glycomic analysis was performed on the continuous bovine mammary cell line to demonstrate the presence of both a2,3 and a2,6, which is consistent with their observation that these cells can be co-infected with human and avian IAVs simultaneously. The main question, however, is: what is the glycome in the explants, or directly from tissues?

      Overall, the manuscript is clearly written and provides new insights into the behaviour of the cattle isolate, now compared with a representative group of model or precursor HAs of different origins.

      It would be great if a consistent nomenclature for the IAV strains could be used in the study. There is a mix of origin (Texas), animal from which the virus was isolated (mallard), or abbreviations that do not follow guidelines (IAV07). Are the USSR and Udorn not lab-adapted?

      The experimental setup includes bovine mammary primary and continuous cells, as well as mammary explants. Some of the most significant differences, for example, in viral fitness studies and co-infection experiments, are observed in these explants. Perhaps there could be some additional focus on this observation. The implications in comparison to the results obtained in cultured cells could be described. How will the human and other HA subtype viruses fare in the explants?

      Comments on revised version.

      The authors have satisfactorily addressed the reviewers' comments.

    3. Reviewer #3 (Public review):

      Summary:

      This excellent manuscript by Pinto, Sharp, and colleagues examines bovine tissue tropism for influenza viruses. They find that bovine flu, as well as other strains, have strong replication in mammary tissue. They also map the genetic changes to influenza that improve replication in bovine cells. Overall, the study is well designed and executed and the results are very timely.

      Strengths:

      (1) The experiments are well-controlled.

      (2) The figures are well-constructed and easy to follow.

      (3) The Methods and legends are detailed, with sufficient information.

      Comment on revised version.

      The authors have strengthened the manuscript by addressing comments from the three reviewers and I have no additional concerns/suggestions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, Pinto and colleagues set out to investigate whether the cow udder is a potential mixing site for the influenza virus. The authors have demonstrated that bovine mammary epithelial cells can be infected with both avian and human influenza A viruses, supporting the idea that the cow udder may be a potential site for reassortment. Furthermore, they demonstrate that the bovine-adapted IAV replicates to similar titers in avian epithelial cells when compared to an AIV precursor virus. Thus, suggesting there is no fitness trade-off, and confirms the potential for spill-back of the cattle B3.13 into poultry, which has already been observed. Overall, I believe the authors achieved their aims. However, there are instances in which the results do not entirely support the conclusions (noted in weaknesses). Given the ongoing questions surrounding highly pathogenic avian influenza A virus in dairy cows, this work provides valuable evidence for the potential of the cow udder as a site of reassortment. These findings highlight the need for surveillance of influenza A virus incursions into livestock species, particularly cows. Some specific strengths and questions regarding weaknesses have been outlined below.

      Strengths:

      (1) The authors use a diverse range of cell types and influenza A virus strains, as well as a wide range of techniques to address the questions at hand.

      (2) The use of cells from multiple bovine breeds for the MAC-T, bMEC and explants suggests the phenomenon is not unique to a single breed.

      (3) The results suggesting there is no fitness trade-off for Cattle Texas in an avian host are interesting, and confirm the potential for spill-back of the cattle B3.13 into poultry, which has been observed.

      Weaknesses:

      I have listed my complete questions/concerns below. However, there are two main weaknesses of the article in its current state. Firstly, there is no apples-to-apples comparison in terms of determining a preference for IAV to infect the cow udder over other organs (Q4). The mammary gland and respiratory tract are represented by epithelial cells, but for other organs, fibroblasts were chosen. I think the fairer comparison would be to compare epithelial cells from different organs to demonstrate a preference for the mammary gland. Secondly, the main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. Yet throughout the article, a 10x higher dose of IAV is used in the bMEC cells compared to everything else (Q6). This raises the question of how much of the results are due to a preference for the mammary epithelial cells, and how much is simply due to the increased dose.

      (Q4) When we set out to test if cow mammary gland cells were particularly susceptible to IAV infection compared to other bovine cell types, we used what was available in the Roslin Institute – a mix of primary and continuous cells from various anatomical sites: three epithelial cell types (two mammary, one respiratory tract) two immune cell types and four sets of fibroblasts from various organs. Given the representation of different anatomical sites, cell types and differentiation statuses, we considered this a suitably diverse panel with which to characterise infection dynamics of a broad range of IAVs, before more focussed investigations using the bMEC and explant tissues. Both mammary epithelial cell types grew our library of influenza challenge strains significantly better than the BAT-II respiratory epithelial cells, as well as the two immune cell types and all four fibroblast populations. Of the fibroblast cells, those derived from the brain grew IAV significantly better than the skin and turbinate fibroblasts, while blood-derived macrophages grew virus significantly better than the lymphocytes and non-brain fibroblasts. So there are “apple-to-apple” comparisons as well as apple-to-pear comparisons that give significant differences. We therefore think that our conclusions (in the abstract) that mammary cells are particularly replication competent for IAV, (at the end of the introduction) that “a wide range of cow-derived cells are susceptible” and that (in the results section) that “mammary cells showed the highest susceptibility” are justifiable. However, we agree that testing a wider variety of epithelial cells would be useful and have added text to the Discussion (lines 224-228) to acknowledge this.

      (Q6) We used a higher MOI for bMECs because test experiments with WT PR8 and the Cattle Texas 6:2 reassortant virus showed that MOI 0.01 infections gave more variable results than those run at MOI 0.1, perhaps because of the intrinsic variability of mixed primary cell populations. However, the end-point titres between the two conditions were not significantly different, so we therefore chose to go with the higher MOI. Accordingly, we do not think this choice is a confounding issue. This explanation (line numbers 340-345) and a new Supplementary Figure 11 showing the results of the two MOI tests have been added to the manuscript.

      Reviewer #2 (Public review):

      The authors use a library of influenza A viruses from different strains, classified in lab-adapted, human, avian, and swine according to the animal from which they were isolated. They propose that the cow mammary gland serves as a mixing vessel for influenza A viruses. As a first approach, the authors assess susceptibility to infection across different cell types, including continuous and primary cell lines, bovine mammary cells, and mammary explants. All these cells support polymerase activity. Then, they analyzed changes in the bovine virus's viral fitness relative to an avian precursor. The authors use single-gene replacement to study whether and which RNP segments improve viral transcription. As part of this section, they also test IFN-specific antagonism by NS1 to assess the input of segment 8. Quantitative glycomic analysis was performed on the continuous bovine mammary cell line to demonstrate the presence of both a2,3 and a2,6, which is consistent with their observation that these cells can be co-infected with human and avian IAVs simultaneously. The main question, however, is: what is the glycome in the explants, or directly from tissues?

      We report quantitative glycomics for the primary bovine mammary epithelial cells as well as the continuous line the referee highlights. However, we agree with R2 that a detailed glycomic analysis of primary bovine mammary tissue would allow a better understanding of the actual glycosylation status in vivo. This has been undertaken by the authors and is available as a bioRxiv preprint. This is now cited (ref 25) in the relevant part of the results (line 184-185)

      Overall, the manuscript is clearly written and provides new insights into the behaviour of the cattle isolate, now compared with a representative group of model or precursor HAs of different origins.

      It would be great if a consistent nomenclature for the IAV strains could be used in the study. There is a mix of origin (Texas), animal from which the virus was isolated (mallard), or abbreviations that do not follow guidelines (IAV07). Are the USSR and Udorn not lab-adapted?

      We chose the abbreviated names for a variety of reasons. Partly from common usage (e.g. PR8, Udorn), partly for consistency with other already published papers from the FluTrailMap consortia (e.g. Cattle Texas; Dholakia et al 2026), partly to make diversity obvious in certain figures (e.g. H3N1, H5N2 etc) and partly to avoid confusion between viruses that originate from the same geographic area (e.g. AIV07, AIV09, H5N8-20 etc which are all A/Ck/England/isolate numbers). Overall, we found it more confusing to use the expanded nomenclature. Re AIV07 which the referee criticises for not following naming guidelines – if this is a reference to the EURL nomenclature, AIV07 is the abbreviation for the specific virus A/Chicken/England/053052/2021, our representative virus for EURL genotype EA-2020-C, as we say in the text. This nomenclature has now been added to Table 1, to provide a fuller cross-reference for all the names.

      As to whether USSR and Udorn are lab-adapted – that depends on definitions. There is a continuum of adaptive changes and/or sequence drift starting from the very first growth cycle of an isolate in the laboratory. The viruses we define here as lab adapted are ones that have been deliberately adapted to other host species or which have very long passage histories in multiple laboratory systems resulting in known functionally significant changes; for example, one lineage of PR8 was passaged 77 times in mice, 717 times in cell culture, 30 times in chick embryos, 5 times in ferrets and a further 50 times in chick embryos (https://www.medscape.com/viewarticle/812621_3?form=fpf), rendering it unarguably lab-adapted. We admit that A/USSR/77 and A/Udorn/307/1972 are probably further along this adaptive pathway than more recent isolates such as A/Norway/3433/2018, but are unaware of any specific reason that would put them into our lab-adapted category.

      The experimental setup includes bovine mammary primary and continuous cells, as well as mammary explants. Some of the most significant differences, for example, in viral fitness studies and co-infection experiments, are observed in these explants. Perhaps there could be some additional focus on this observation. The implications in comparison to the results obtained in cultured cells could be described. How will the human and other HA subtype viruses fare in the explants?

      We agree that this is an important and interesting question, and had already tested the strains we used for co-infections: human seasonal pdm09 H1N1 “Norway” and low pathogenic avian influenza “H3N1”, in the mammary explants. Both replicate the avian virus to 20-fold higher titres. We have added this information to the revised manuscript as new Figure panels S6E-H, called out on line 200-201 of the results.

      Reviewer #3 (Public review):

      Summary:

      This excellent manuscript by Pinto, Sharp, and colleagues examines bovine tissue tropism for influenza viruses. They find that bovine flu, as well as other strains, has strong replication in mammary tissue. They also map the genetic changes to influenza that improve replication in bovine cells. Overall, the study is well designed and executed, and the results are very timely.

      Strengths:

      (1) The experiments are well-controlled.

      (2) The figures are well-constructed and easy to follow.

      (3) The Methods and legends are detailed, with sufficient information.

      Weaknesses:

      (1) A comparison to human cells would strengthen the overall impact of the results. Are human mammary cells also uniquely susceptible to influenza? Are bovine mammary cells special in some way?

      This is an interesting question, but we have not tested mammary gland cells from humans (or any other species of mammal). We have however reported elsewhere (Dholakia et al., Nat Commun. 2026 Jan 16;17(1):1603. doi: 10.1038/s41467-026-68306-6.) that Cattle Texas grows well in a variety of human respiratory cells. Here, we are considering the bovine mammary organ as a potential reassortment site for IAVs because of the ongoing viral mastitis epidemic in US dairy cattle; human mammary organs seem unlikely to create a similar opportunity.

      (2) For the virus infection studies with segment 8 swaps, it should at least be noted that some of the phenotypes could be driven by NEP.

      We agree; as Table S1 indicates, NEP has two changes (one shared with NS1) between AIV07 and our B3.13 isolate, so we should not have conflated segment and NS1. We have changed the text to acknowledge this throughout the results (lines 127, 137 and 149) and in the discussion (line 243-244).

      (3) The data demonstrating that bMEC can support co-infection are compelling and important, but would be strengthened with a comparison from a different cell type or species. Do mammary cells uniquely support higher co-infection?

      We have data showing that co-infection also occurs in the continuous MAC-T udder cell line and have now included these data in a revised Figure 4D (described/called out on lines 198-206). We have not tested bovine cells from other organs for co-infection potential as they do not seem to be significant sites of infection in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) How nasal turbinate and cardiac fibroblasts are acquired/cultures is missing from the methods.

      Apologies for the omissions and thank you, because rectifying this brought to light an error in cell naming. The cells originally called bovine cardiac fibroblasts were in fact a second independent preparation of skin fibroblasts. We have corrected the labelling in Figs 1, 2 and S2. The nasal turbinate cells were bought in from ATCC (code CRL-1390). This information has now been added to the methods (line 310).

      (2) Please specify what cell types make up the 2D enteroids, the mammary explants and the nasal turbinates.

      The composition of the 2D enteroids is described in detail in reference [60]. In precis, they are comprised predominantly of epithelial cells, including Paneth, goblet and enteroendocrine cells, as well as stem cells. This information has been added to the methods (lines 374-376) The mammary explants include duct epithelium, connective and muscle tissue, defined by H&E staining of cut sections (see new Figure S12 and text added to the methods on lines 384-386). We have also added the person who did the histology (Rebecca Ross) as an author and to the credit taxonomy (line 877). The nasal turbinate cells appear to be predominantly fibroblast morphology (Methods line 310-311).

      (3) Epithelial cells are the main target for IAV, so why were fibroblasts chosen as a comparison to mammary epithelial cells? This needs justification in text. Without justification, how much can be attributed to the results being mammary-specific, rather than epithelial-specific? The brain (choroid plexus epithelium), heart (epicardium) and skin all contain epithelial cells.

      We think this query is what the referee calls “Q4’ in the public part of their review. Please see our answer above.

      (4) Figure 1, S1, S2 and S3 seem to suggest that mammary cells are more susceptible to IAV infection than cells from other organs. But Figure 2 demonstrates that when it comes to Cattle Texas and AIV07, most of the cell types show high viral titers. If the question is about whether the cow udder is the primary mixing site, would it not be more relevant to investigate which cell type facilitates the best growth of the potential precursor viruses, similar to Figure 4A?

      Figs 1, S1, 2 and 3 all use “full” viruses whereas Fig 2 uses 6:2 reassortants between PR8 and the HPAIVs for biosafety reasons. WT PR8 replicates well in most of the bovine cells tested (Figs S1-3) so we do not see any contradiction. Figure 2 examines the contributions the internal genes make to replication in bovine cells. Re the question over the udder being a potential mixing site – this is where the virus is replicating in the real world; at least in part because of the transmission mechanism, but also because the mammary gland epithelium is highly susceptible to infection, as we show.

      (5) It is misleading to compare the bMEC infection of MOI 0.1, to all other infections of MOI 0.01. This is a consistent problem throughout the article - Figures 1, S1, 2, 4. Why was the bMEC infection at a 10x greater dose? The main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. If we compare the MOI 0.01 experiments alone, then the evidence relies on the immortalised MAC-T cells, compared to primary cell types. In this case, how much can be said about it being mammary specific, rather than immortalised vs primary? I do wonder, for example, how the epithelial nasal turbinates or type II pneumocytes would compare to the primary mammary epithelial cells if they were at the same MOI.

      Please see our answer to this query earlier in the rebuttal.

      (6) Why are the 2D enteroids excluded from Figure 1?

      We had limited supplies of a difficult-to-grow cell model, so we only used them to test the 6:2 viruses (Fig 2).

      (7) The colour scheme for Figure 3 is confusing. In Figure 3A, blue indicates European ancestry, and yellow represents North American ancestry. However, in Figure 3B, these colours now mean something different. To a reader, when there is a colour-coded schematic, it is instinctual to think that this then corresponds to the following panel(s). Since consistently throughout the article, yellow has been used for Cattle Texas, and blue has been used for AIV07, I would suggest choosing different colours to represent European and North American ancestry in Figure 3A

      We’ve changed the figure as the referee suggests and modified the Fig 3 legend accordingly (line 895)

      (8) I am unsure about the conclusions drawn from the results of Figure 3B. In the results, it is framed as trying to determine which segments contributed to the improved activity of Cattle Texas compared to AIV07. In lines 110-111, "... PB2 or PA from AIV07 significantly decreased Cattle Texas minireplicon activity". If PB2 is indeed significant, there is a missing yellow asterisk in Figure 3B.

      Apologies, there was indeed a missing asterisk on the figure; now added.

      Given the significance of PA, why was it not investigated in terms of growth kinetics similar to Figure 3C? Was it overlooked because it doesn't have a North American ancestry? The results of Figure 3B suggest that the 4 amino acid mutation in PA has significantly contributed to changes in polymerase activity.

      The PA changes do indeed matter for minireplicon activity – the key change is K497R, as detailed in our related publication in Nat Comms (citation 17). However, it is less important than changes in PB2, and the PA segment swap by itself has little effect on overall virus replication.

      (9) Similarly, in Figure 3C and lines 116-117, the error bars on the graph are overlapping at 48 hours, suggesting no difference in overall replication. The kinetics are slowed for AIV07 seg1-3, but not for AIV07 seg 1, indicating PB1 does not have an effect. This would then suggest that something in segment 2 or 3 contributes to the slowed kinetics in Figure 3C, which, from the Figure 3B results, is unlikely to be due to PB2. While it was not reassorted, the PA segment is potentially the driver, with its 4 amino acid mutations. I think it is worth performing growth kinetics with and without these 4 amino acid changes in PA.

      We agree that visually on a log<sub>10</sub> scale, the titres of the “WT” 6:2 Cattle Texas and 5:2:1 segment 1 reassortment appear close, but the average titres are 5 and 7-fold different at 24 and 48h respectively, while a 2-way ANOVA with Dunnet’s multiple comparison post-test gives statistical significance at 48h. We have added this information to the figure and its legend (lines 905-906).

      (10) In Figure 3C and E, why did you choose to perform the growth kinetics in the immortalised cell line, when you have access to primary cells? The primary cells would be a more accurate representation of what happens in situ.

      The primary cells were difficult to work with and only available intermittently, so we used what was available at the time.

      (11) In lines 145-147, "thus overall, the reassortment event that replaced segments 1, 2 and 8 alongside drift adaptations in segment 3 may have contributed to the ability of the B3.13 genotype virus to infect cattle". This is not clearly supported by the evidence presented. In terms of segment 1/PB2, the growth kinetics of Figure 3C have overlapping error bars at 48 hours. Where is any evidence presented for the role of segment 2/PB1? There is no change in Figure 3B.

      The referee is correct, calling out seg2 here was an error; we have revised the text (line 144).

      Segment 3 is overlooked in Figure 3 (as highlighted in Q9 and 10), and shows no difference in Figure S4.

      Please see response to Q8; we think segment 3 contributes via PA adaptation, not via PA-X.

      (12) In line 146 "... drift adaptation in segment 3". Make it clear here that you are talking about genetic drift. However, is this likely to be genetic drift? The 4 amino acid mutations are shown to have a significant impact on polymerase activity in Figure 3B, and in Figure 3C, PA potentially contributes to the reduced kinetics. When there are amino acid mutations that correspond to a beneficial phenotypic change, attributing this to drift alone rather than host adaptation is strange.

      Yes, wording clarified (line 145). “Drift” was used to distinguish it from reassortment but we agree this was an incorrect term in the context.

      (13) Figure 4A MAT-C cells: this is ostensibly the same experiment as Figure 3C in terms of the Cattle Texas and AIV07 viruses. If this is the case, how can you explain the difference in kinetics and overall titer? In Figure 3C, Cattle Texas reaches 10^6, and in Figure 4A it reaches almost 10^9. That's almost 3 log difference. Similarly, in Figure 3C Cattle Texas reaches 10^3, but in Figure 4A it reaches 10^6, a 3-log difference. At 24 hrs, they have roughly a 3-log difference between them in Figure 3C, but in Figure 4A this difference is much smaller. As far as I can tell, these are the same viruses, same dose and same cell model. The t0 titer is also vastly different between the two experiments.

      The experiments were done at different times (several months apart, so different cell passage numbers and/or serum batches) and by different people. We have no explanation other than biological variability. However, both groups of experiments include genuine biological replicates done over the course of 2-3 weeks, so in our view represent coherent tests within themselves.

      (14) In Figure 4, why weren't the a-2,6 a-2,3 proportions analysed for the explants and/or used for the co-infection experiment? It showed the greatest difference between the Cattle Texas and precursor viruses in Figure 4A.

      Our data on the proportion of 2,6 and 2,3 SA in bovine udder tissue are now available in a separate preprint (now cited as [25] in our MS). We did not use the explants for co-infection experiments because it would have been technically difficult to read out the outcome by flow cytometry.

      (15) Figure 4D requires a supplementary figure demonstrating the gating strategy, including one of the samples as an example.

      We have compiled a figure of this and added it as new Figure S13 (called out line 556).

      (16) In Figure 4D, why was an MOI of 5 chosen instead of the MOI of 0.01 used throughout the article for MAC-T cell infection? An MOI of 5 (so in a co-infection, a total of 10 virus particles per cell) is completely overloading the cells. At this dose, 10% of the cells were able to be co-infected, but how representative is this of a real co-infection scenario? While it demonstrates it is possible, it potentially remains highly unlikely, similar to the discussion around the swine respiratory tract in lines 235-240.

      We had also performed the co-infections at lower MOI (1) with very similar results – this is now included in Figure 4D. Furthermore, we redid the experiments at a lower MOI of 0.05 and still see co-infection; this now replaces the MOI 5 data in Figure 4D. The text has been revised accordingly (lines 202-205)

      (17) In Figure 4D, why were immortalised cells used when primary bMEC and mammary explants are available? Primary cells would provide more convincing evidence for the potential of the cow udder to be a mixing vessel. Considering that throughout the paper, a 10x higher viral dose is used in the bMEC culture, I wonder if you would need a significantly higher MOI than 5 to produce similar results in a co-infection experiment. The bMEC also has a more even a-2,3 to a-2,6 ratio compared to MAC-T in Figure 4C.

      The bMECs in the original figure are primary cells. In response to other queries, we now include data from the immortalised MAC-T cell line as well.

      Reviewer #2 (Recommendations for the authors):

      Figure 1A, the coloured underline to discriminate continuous and primary cells is lost upon printing... perhaps another way is better?

      We have changed the primary cells to italic text to make the distinction clearer.

      We have made some other minor changes to wording to correct grammatical errors or improve clarity as we went through.

    1. eLife Assessment

      This article provides a computational study to understand how evolutionary selection could introduce structural priors into neural networks to improve learning. This useful investigation provides incomplete evidence that such joint learning is possible and has some counterintuitive features. However, the extent to which the models studied are distinct from those studied in the literature previously remains unclear.

    2. Reviewer #1 (Public review):

      In this article, the authors set out to understand how evolutionary selection could introduce structural priors into neural networks that act as an inductive bias to accelerate learning. To do so, the authors propose an evolutionary conditioning (ED) algorithm and analyse its properties. I find the conceptual framing of the paper very interesting, and it addresses an important question. Since the paper adopts a mostly theoretical approach with no comparison to empirical biological data, I do have a couple of concerns regarding the setup of the computational framework/method, which I think is incomplete and limits how much we can conclude from the current results.

      Major Concerns:

      (1) The evolutionary conditioning (EC) algorithm proposed works as fine-tuning training, plus propagation of the best parent network's weights to the next generation with added Gaussian noise. Conceptually, I find this to be a fairly implausible mechanism for evolution, since it requires carrying the entire set of network weights at some precision. The authors themselves point out that direct weight transfer could be problematic in the introduction.

      (2) More importantly, I would like the authors to conduct a baseline / null model comparison, in which the "evolution" process consists simply of training a neural network for a small number of iterations, adding Gaussian noise, and repeating. The resulting network at each step serves as the "generations", which is then trained further. The same learning speed and dynamics analyses should be applied to this null model. What I am getting at is that I am not sure to what extent the EC algorithm can be thought of as "running a few iterations of SGD" and chaining them together; how much work is the selection process in the GA actually doing?

      (3) I am also unclear on why the EC algorithm does not improve throughout learning. Is this behavior the result of applying only a small number of fine-tuning steps? Presumably, with longer fine-tuning, the individual networks in the middle generations would also improve in performance?

      (4) The EC algorithm applied to a single problem seems somewhat artificial in its setup. I would conceptualize evolution as learning a prior that conditions the network for a range of survival-related tasks. A more realistic setup would apply EC to an ensemble of tasks and then examine its impact on learning a specific task afterwards.

    3. Reviewer #2 (Public review):

      Summary:

      This paper studies the interplay of evolutionary and in-lifetime learning. The authors develop a neural network model in which initial weight configurations evolve under selective pressure, while fitness is determined by the network's performance after a learning period. They show that such a network displays very distinct learning dynamics from those trained by either gradient descent or genetic algorithms alone: in particular, they do not learn the task, but they show evidence of learning-to-learn and unusual representational structure.

      Strengths:

      (1) The writing, figures, and presentation of ideas were clear.

      (2) The question of how evolution on initial weights combines with learning from within-lifetime experience to structure a learning trajectory seems interesting.

      (3) The analysis of existing experiments was well-done, highlighting that though these networks did not really learn, they show latent learning structure that makes the network perform better from less data.

      (4) The interplay between Baldwin & learning dynamics seemed novel and interesting, presenting many attractive puzzles.

      Weaknesses:

      First, the authors point to an important distinction between performing evolutionary selection on the weights pre- or post- lifetime training, the latter of which is Lamarckian. They argue, correctly, that their model is interesting because it selects on the weight initialisation, unlike, for example, Shuvaev et al. However, my understanding is that a long line of papers beginning perhaps with Hinton & Nowlan also do non-Lamarckian evolution: Hinton & Nowlan have unspecified weights (denoted '?' in the paper) that can be inherited and then learnt. Is this not exactly inheritance of initial conditions (in this case, whether learnable or not)? This novelty is a primary motivation of the paper, whereas to me it seems it was already apparent in Hinton & Nowlan, and developed further in what seems to be a long line of uncited literature (see next paragraph). As such, this paper's conclusions seem poorly positioned within the existing state of knowledge/literature.

      Second, the algorithm is framed as novel, but I think it is a rediscovery. This framework is very close to MAML, in which an initial weight configuration is optimised by gradient descent to be good after a few steps of fine-tuning (Finn et al., 2017). The authors' approach differs in using a genetic algorithm to perform the training of the initial weights, avoiding some of the computational complexities of MAML, especially after long fine-tuning. In this, the authors have, I think, rediscovered ES-MAML, MAML where the inner optimisation loop is gradient descent, while the outer is genetic (Song et al., 2020). Other similar work is "Meta-Learning by the Baldwin Effect" (Fernando et al., 2018).

      Further, within Fernando et al. there is a rich literature review, almost none of which are cited by the authors. I point especially to Keesing & Stork, 1990, which appears to show a strong dependence of Baldwin-like improvements on the amount of data, something this paper also shows but explores less thoroughly.

      To summarise my critique thus far: I think the literature already answers the main concern of the motivation (i.e. non-Lamarckian neural network evolution and learning), I think it has already discovered this particular algorithm, and I think past work has more thoroughly analysed behaviours similar to those presented in this paper. Without positioning correctly within this literature, the more general contribution of the paper is hard to establish. The true novelty of the authors' analysis seems to be the emphasis on Saxe et al.-like learning dynamics and its interplay with the Baldwin effect, but I am not certain of this without knowing the literature better.

      Regarding experiments, there was an interesting effect where the EC networks didn't learn but did show latent learning (Figure 2, Figure 3), which sped up later learning (Figure 4). There were a few details I was surprised by on which I would appreciate clarity:

      (1) The main result has basically no headline learning under EC. This will clearly be very dependent on parameters (e.g. if you add or remove enough training steps, the algorithm becomes SGD/GA, which both show learning). It seems like a natural analysis would examine this (e.g. a plot of final performance of EC after 4000 generations with different per-generation learning budgets).

      (2) It is then shown that after 4000 generations EC can learn very quickly to perform the semantic task perfectly, at least within 200 generations (Figure 4C, and perhaps much sooner, Figure 4D, Figure 4F last panel; it was hard to say. This and the previous point seem somewhat inconsistent; was it just that Figure 3 used only 100 fine-tuning steps while the perfect-task-performing networks in Figure 4 required somewhere between 100 and 200? This seems to point to extreme parameter dependence. More broadly, how should I square this inconsistency/near-inconsistency?

      (3) Figure 3k, and especially Figure 4f bottom right panel, seem to show networks that correctly separate all stimuli but cannot classify them. Should I understand this as networks learning to just push apart all pairs of datapoints without structure?

      (4) If I understood the genetic algorithm correctly, only three individuals from each population seeded the next generation. This seems another important parameter to tune, since I think it is far lower than standard evolutionary work, but I am not sure.

      Finally, the paper most interested me as a neural network learning puzzle: how can the network perform so badly, yet lead to such different post-fine-tuning results? The paper pointed to these as 'distinct' learning phenomena without explaining what was causing those differences. The only way I could square these results in my head was as above: that the 4000 generations pushed the initial representation to represent all datapoints differntly, effectively changing the learning problem gradient descent faces from one with a lot of structure (the semantic task) that leads to stepwise learning, to one in which it was basically linear regression on a set of well separated stimuli without the structure necessary for stepwise learning. Since the paper focuses so much on learning dynamics, and studies a task where such things can be precisely probed, it would have been nice to pin down exactly what was happening slightly more.

    1. eLife Assessment

      This valuable study examines whether neurally-derived non-decision time estimates can guide decision models to more reliable parameter estimates. Convincing evidence for their value is obtained from state-of-the-art model-fitting methods applied to two empirical datasets. However, the premise relies on very strong assumptions regarding the degree to which the neural metrics directly index non-decision time and comparisons across models that appear to differ in more ways than just the non-decision time constraint.

    2. Reviewer #1 (Public review):

      Summary:

      This paper proposes a non-decision time (NDT)-informed approach to estimating time-varying decision thresholds in diffusion models of decision making. The manuscript motivates the method well, outlines the identifiability issues it is intended to address, and evaluates it using simulations and two empirical datasets. The aim is clear, the scope is deliberately focused, and the manuscript is well written. The core idea is interesting, technically grounded, and a meaningful contribution to ongoing work on collapsing thresholds.

      Strengths:

      The manuscript is logically structured and easy to follow. The emphasis on parameter recovery is appropriate and appreciated. The finding that the exponential NDT-informed function produces substantially better recovery than the hyperbolic form is useful, given the importance placed on identifiability earlier in the paper. The threshold visualisations are also helpful for interpreting what the models are doing. Overall, the work offers a well-defined, methodologically oriented contribution that will interest researchers working on time-varying thresholds.

      Weaknesses / Areas for Clarification:

      A few points would benefit from additional clarification following the previous revision:

      Returning to one point from my original review. The applications to empirical data describe 6 models, including the FT-DDM with across-trial variability described as a benchmark. Yet the modelling results for both studies (Tables 2 and 3) show only 5 models, omitting the benchmark model. I acknowledge the comment about this in the response letter (no other FT-DDM models in the main text have across-trial variability). Nevertheless, for this to serve as a benchmark, the modelling results really should be reported in Tables 2 and 3 with the other 5 models, and the model goodness of fit figures currently shown in Appendix 7 incorporated into Figures 10 and 13 of the main text. Unfortunately, as it currently reads, it looks as if something is being obscured, which I don't believe is the intention. This won't change the primary conclusions, but it will increase transparency and ease of interpretation.

      Many thanks for introducing Appendix 1 to the revised manuscript. Additional motivation for the central issue is always helpful. However, I'm not (yet) convinced the simulation study in Appendix 1 achieves the intended aim. The simulation study shows that NDT misspecification leads to poorer recovery of other CT parameters (i.e., if a generating value is perturbed and then held fixed at that perturbed value during estimation, recovery of the remaining parameters deteriorates). This demonstrates the consequences of fixing NDT incorrectly, but it does not seem to address the central claim of the manuscript: that NDT and the CT threshold parameters trade off when they are estimated simultaneously. I had expected the simulation study to estimate NDT alongside the remaining model parameters, mirroring the estimation procedure used in the empirical analyses. Such a simulation would directly test whether the two sets of parameters compensate for one another during estimation, and whether this results in poor recovery of the NDT parameters.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model using noisy single-trial estimates of an underlying fixed non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      Weaknesses:

      This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model?

      Comment on revised version.

      Thanks to the authors for their responses and inclusion of additional analyses and simulations. Thanks, in particular for clarifying a crucial detail, that the ndt-informed model actually assumes, like the uninformed model, that there is no variability in the non-decision time, and the idea is that the variable single-trial measurements of non-decision time are noisy estimates of an underlying, constant ndt. The revised paper itself has not made this clear - for example, throughout the Intro, there is no statement that the behavioural model assumes a trial-invariant ndt, and line 231 still calls tau the 'mean' non decision time, implying there is a distribution rather than an invariant single value in the behavioural model.

      One implication of the above is that if the lognormal sigma is purely measurement noise that does not relate to actual variation in the underlying decision process generating behaviour, then the HMP latencies should not relate to behaviour, e.g. shorter latencies predicting shorter RT. I assume that even if the authors did find such a relationship, the principle still stands that a model with fixed ndt is more accurately fit when there are single-trial ndt estimates whose mean provides a constraint on that ndt value, than without such measurements. Still, given ndt variability is a core feature of many decision models, the authors could comment on whether the strategy would work in theory for a model with ndt variability (in the behavioural part), where the single trial estimates would then presumably reflect a mix of measurement noise and genuine ndt variability.

      Another more important implication is that since it is in fact the same model being compared with and without the HMP data guiding the fixed ndt estimate, the reason the fit quality improves with HMP-information is not because it is a better model per se (it is the same model) but because without the HMP guidance, the search algorithm somehow gets lost and fails to find the 'optimal' parameter vector. That is, the parameter vector (just the parameters that relate to the behavioural model itself, not the HMP lognormally-distributed noise associated with VEP measurements) identified as optimal in the HMP-informed version of the model exists in the parameter space of the model without HMP information, but it is just not found? I raised this implication before, and it is still not clear whether it applies. I'm sorry to press on it, but it is critical for readers to understand why it is that neural information can improve overall model fit. Again, the enhancement of parameter recovery (like in Nunez 2025) makes sense, but the enhancement of the "model's fit to behavioural data" does not, without pointing to a deficiency in the search algorithm / fitting procedure.

      The authors state in their replies that the onset of bound collapse is set at accumulation onset and imply that setting it instead at stimulus onset could "mathematically resolve the issue" but they don't do it because it is implausible. It is in fact not only plausible but clearly evidenced in empirical data - collapsing bounds are implemented neurally through urgency signals, and these can begin to dynamically build toward threshold well before, let alone at, stimulus onset. There is nothing bizarre about this - we can prepare movements without sensory input, and indeed even if choosing actions based on a sensory discrimination, motor preparation can launch well before the sensory evidence (e.g. Stanford, Salinas et al 2010) and this in effect collapses the bound on cumulative evidence for triggering action before any evidence actually arrives. So, Urgency/bound-collapse does not need to be triggered by a stimulus; it can start in anticipation of the stimulus. It seems critical, therefore, for the authors to clarify this point - does re-defining the onset of the collapse at stimulus onset remove the trade-off and render unnecessary the neurally-informed ndt estimation?

      Related to this, it is still not clear how bias in the estimation of nondecision time would not be a problem. What if, for example, it is the end of the N2 rather than the peak of the N2 that marks accumulation onset, and/or there is an additional fixed motor time that adds to the N2-based marker to make the full nondecision time that applies in the underlying decision process. By definition (and I think this is essentially what the authors' new simulations verify), because of the trade-offs, this bias would simply be absorbed in shifted estimates of theta and lambda describing the bound collapse function. But wasn't the whole point of the exercise to more accurately estimate those parameters? The obvious implication is that the parameter-estimation accuracy of the ERP-informed model is determined by the accuracy with which the proposed ERP marker directly pinpoints the full nondecision time without bias, but this is not at all obvious in the paper as written. Importantly, in the example scenario I describe above where accumulation onsets when N2 ends, there may still be a perfect correlation of N2 peak latency with underlying ndt across trials - they could still be very strongly "linked" statistically, but we can't know what size offset might be involved.

    4. Reviewer #3 (Public review):

      The current paper addresses an important issue in evidence accumulation models: many modelers implement flat decision boundaries because the collapsing alternatives are hard to reliably estimate. Here, using simulations the authors demonstrate that parameter recovery can be drastically improved by providing the model with additional data (specifically, an EEG-informed estimate of non-decision time). Moreover, in two empirical datasets it is shown that those EEG-informed models provide a better fit to the data. The method seems sound and promising and might inform future work on the debate regarding flat vs collapsing choice boundaries. As an evidence-accumulation enthusiast, I am quite excited about this work, although for the more broader audience the immediate applicability of this approach seems limited because it does require EEG data (i.e. limiting widespread use of the method or e.g. answering questions about individual differences that require very large N).

      Comments on revised version.

      I thought the authors carefully addressed my concerns.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper proposes a non-decision time (NDT)-informed approach to estimating timevarying decision thresholds in diffusion models of decision making. The manuscript motivates the method well, outlines the identifiability issues it is intended to address, and evaluates it using simulations and two empirical datasets. The aim is clear, the scope is deliberately focused, and the manuscript is well written. The core idea is interesting, technically grounded, and a meaningful contribution to ongoing work on collapsing thresholds.

      Strengths:

      The manuscript is logically structured and easy to follow. The emphasis on parameter recovery is appropriate and appreciated. The finding that the exponential NDT-informed function produces substantially better recovery than the hyperbolic form is useful, given the importance placed on identifiability earlier in the paper. The threshold visualisations are also helpful for interpreting what the models are doing. Overall, the work offers a well-defined, methodologically oriented contribution that will interest researchers working on time-varying thresholds.

      We appreciate the positive and constructive feedback. We have addressed your comments in the revised manuscript, as detailed in our responses to the specific comments below.

      Weaknesses / Areas for Clarification:

      A few points would benefit from clarification, additional analysis, or revised presentation:

      (1) It would help readers to see a concrete demonstration of the trade-off between NDT and collapsing thresholds, to give a sense of the scale of the identifiability problem motivating the work.

      Thank you for this constructive suggestion. We conducted a new simulation study in which we considered the non-decision time as a fixed parameter and estimated the remaining parameters of the collapsing threshold model. In this simulation study, we contaminated the non-decision time with noise at six levels (i.e., 0%, 2%, 4%, 6%, 8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters. The results of this simulation are presented in Appendix 1 in the new version of the manuscript.

      (2) Before moving to the empirical datasets, the manuscript really needs a simulation-based model recovery comparison, since all major conclusions of the empirical applications rely on model comparison. One approach might be to simulate from (a) an FT model with across-trial drift variability and (b) one of the CT models, then fit both models to each of the simulated data sets. This would address a longstanding issue: sometimes CT models are preferred even when the estimated collapse in the thresholds is close to zero. A recovery study would confirm that model selection behaves sensibly in the new framework.

      We are grateful for this constructive comment. The revised manuscript includes a model recovery study (see Simulation study 2, in particular Table 1 and the accompanying text). In this simulation, as the reviewer suggested, we generated data from FT-DDM with across-trial variability in drift rate and CT-DDM with hyperbolic and exponential collapsing thresholds. Then, for each generated dataset, we fitted two models and classified the datasets based on goodness-of-fit and estimated decay rate. The results show that incorporating the decay rate, which can be reliably estimated by the NDT-informed modeling framework, significantly improves the precision of model recovery.

      (3) An additional subtle point is that BIC is defined in terms of the maximised log-likelihood of the model for the data being modelled. In the joint model, the parameter estimates maximise the combined likelihood of behavioural and non-decision-time data. This means the behavioural log-likelihood evaluated at the joint MLEs is not the behavioural MLE. If BIC is being computed for the behavioural data only, this breaks the assumptions underlying BIC. The only valid BIC here would be one defined for the joint model using the joint likelihood.

      We thank the reviewer for raising this important methodological point. We agree that, strictly speaking, the behavioral log-likelihood evaluated at the joint maximum likelihood estimates is not guaranteed to be equal to the behavioral maximum likelihood estimate. Therefore, if one were to interpret BIC<sub>Behavior</sub> for joint (NDT-informed) models as a conventional BIC derived from a purely behavioral maximum likelihood fit, this would indeed violate the standard assumptions underlying BIC. Our intention, however, was not to claim that BIC<sub>Behavior</sub> represents a formally valid BIC in the strict information-theoretic sense. Rather, we used it as a diagnostic measure to assess how well the jointly estimated parameters account for the behavioral data relative to the uninformed models. Importantly, in our datasets, the behavioral log-likelihood evaluated at the joint estimates is higher than that obtained from the uninformed models. This suggests that the additional constraint introduced by the non-decision time information helps guide the optimization procedure toward parameter regions that provide a better account of the behavioral data. In other words, uninformed models appear to converge to suboptimal parameter estimates, but the joint modeling framework helps regularize the estimation process and yields better estimates of the optimal parameters. We fully acknowledge that this use of BIC<sub>Behavior</sub> for the joint models departs from standard practice in the cognitive modeling literature. For this reason, we rely primarily on the joint likelihood–based model comparison as the formally valid criterion. The behavioral BIC is reported only to provide additional intuition regarding the goodness of fit on behavioral data and is used exclusively to compare NDT-informed models with their uninformed counterparts under identical evaluation criteria.

      Also, to warn readers about this limitation, we included the following statement in the section where we defined BIC measures:

      “It is worth noting that comparing joint and behavioral models using BIC<sub>Behavior</sub> is uncommon and constitutes a limitation of the model comparison study.”

      (4) Table 1 sets up the Study 1 comparisons, but there’s no row for the FT model. Similarly, Figures 10 and 13 would be more informative if they included FT predictions. This matters because, in Study 1, the FT model appears to fit aggregate accuracy better than the BIC-preferred collapsing model, currently shown only in Appendix 5. Some discussion of why would strengthen the argument.

      In the revised manuscript, we have included the results for NDT-informed FT-DDM in the main text. However, we kept the FT-DDM with drift variability in the appendix, since none of the models in the main text include drift rate variability.

      (5) In Figure 7, the degree of decay underestimation is obscured by using a density plot rather than a scatterplot, consistent with the other panels of the same figure. Presenting it the same way would make the mis-recovery more transparent. The accompanying text may also need clarification: when data are generated from an FT model with across-trial drift variability, the NDT-informed model seems to infer FT boundaries essentially. If that’s correct, the model must be misfitting the simulated data. This is actually a useful result as it suggests across-trial drift variability in FT models is discriminable from collapsing-threshold models. It would be good to make this explicit.

      Regarding the visualization of decay rate estimation in Figure 7, we would like to clarify that, in the data-generating process for the FT model with across-trial drift variability, the decay rate is fixed at zero. That is, unlike the other parameters, there is only a single true value on the x-axis. If we were to present the decay rate recovery using a standard scatter plot (true vs. estimated values), all points would lie vertically above the single true value (zero), resulting in a vertical strip of points. While such a plot would technically be consistent with the other panels, it would not clearly convey the distributional properties of the estimated decay rates, specifically, which values are more likely under model misidentification. For this reason, we chose a density plot to more transparently illustrate the distribution of inferred decay rates when the true generating process is an FT model. We believe this representation more effectively communicates the extent and structure of mis-recovery. To avoid confusion, we have added a clarifying footnote in the revised manuscript explaining why a density plot was used in this specific panel.

      Regarding distinguishing fixed-threshold models from collapsing-threshold models, as suggested in your comment 2, we conducted an additional simulation, and the results showed that when using the NDT-informed diffusion model, we can distinguish between both models with high accuracy. Specifically, we showed that incorporating the estimated decay rate value provided by the NDT-informed modeling approach can significantly improve the model recovery accuracy.

      (6) Given the large recovery advantage of the exponential NDT-informed function over the hyperbolic one, the authors may want to consider whether the results favour adopting the former more generally. Given these findings, I would consider recommending the exponential NDT-informed model for future use.

      Consistent with the reviewer’s argument, we included the following text in the general discussion:

      “Importantly, the exponential collapsing threshold exhibited substantially better parameter recovery and superior model recovery performance, suggesting that this specification may be preferable in future cognitive modeling applications.”

      (7) In Study 2 (Figure 13), all models qualitatively miss an interesting empirical pattern: under speed emphasis, errors are faster than corrects, while under accuracy emphasis, errors become slower. The error RT distribution in the speed condition is especially poorly captured. It would be helpful for the authors to comment, as it suggests that something theoretically relevant is missing from all models tested.

      Thank you for mentioning this point. Because the models considered in the main text do not include across-trial variability in the starting point, the model cannot predict the fast error pattern observed in the speed condition of the second study. In the revised manuscript, we included a note on this point:

      “The NDT-informed models’ predictions are depicted in Figure 13. This figure shows that the FT-DDM overestimates the last RT quantiles for both correct and incorrect responses in both speed and accuracy conditions. However, the qualitative predictions of CT-DDMs align more closely with the empirical data. It is also worth noting that all the considered computational models misfit the incorrect responses in the speed condition. This misfit is to be linked to the presence of fast errors, specifically in the speed condition. Including the starting-point variability parameter in the model enables the model to predict fast errors. However, as the aim here was not merely to fit the data with the best possible model, but to test the NDT-informed modeling framework, we did not include starting-point variability in the model.”

      (8) The threshold visualisations extend to 3 seconds, yet both datasets show decisions mostly finishing by 1.5 seconds. Shortening the x-axis would better reflect the empirical RT distributions and avoid unintentionally overstating the timescale of the empirical decision processes.

      In the new version of the manuscript, we shortened the x-axis in these plots. See Figures 9 and 12 in the new version of the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript should explicitly state how critical the log-normal assumption for NDT is, and whether there are caveats if it doesn’t hold.

      We agree with this suggestion. Therefore, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and replicated the main results reported in simulation study 1 (see Appendix 8 in the new version of the manuscript). These simulation results reveal that independent of distributional assumptions on non-decision time measurements, constraining non-decision time can improve the estimation of the collapsing threshold.

      (2) On page 14, one paragraph refers to five models and another to six. It is unclear which is correct.

      We are sorry for the confusion. In the new version of the manuscript, we addressed this issue.

      (3) In the Discussion: Hawkins & Heathcote (2021) found that NDT estimates in the TRDM recover well, but the timer-offset parameter does not (and hence is set to a fixed value). It would be interesting to test whether the NDT-informed approach could extend to that parameter. The TRDM can also generate error RT distributions that are faster than correct RT distributions, which is the qualitative pattern that none of the present models capture in the speed condition of Study 2.

      We appreciate the reviewer’s suggestion, as it would be another important demonstration of the framework we built in the paper. Nevertheless, given the number of models already considered in the manuscript and the focus on collapsing threshold diffusion models, we believe that this addition would blur the focus of the present manuscript. Therefore, we have decided not to include TRDM in the manuscript. Nevertheless, we have addressed the reviewer’s concern regarding the non-decision time estimation in the TRDM in the revised manuscript:

      “Notably, this model can also predict faster error responses than the correct response, the pattern that is observed in the speed condition of Study 2. However, the authors reported poor parameter recovery for the onset of the timing process (Hawkins and Heathcote, 2021). Thus, the reliability issue here specifically concerns the estimation of the shift parameter of the timing accumulator. Informing the model with external estimates of non-decision time might, therefore, improve parameter recovery in this model as well.”

      Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model on estimates of single-trial non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and, in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      We appreciate your feedback and comments. Below, we provided a response for each comment.

      Weaknesses:

      (1) This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model? Unfortunately, it is not clear whether and how the model structure itself differs between the NDTinformed and non-NDT-informed cases. Ideally, they are the same actual model, but with one getting extra guidance on where to place the tau and/or sigma parameters from external measurements. The absence of sigma (non-decision time variance) estimates for the non-NDT-informed model, however, suggests it is different in structure, not just in its lack of constraints. If they were the same model, whether they do or do not possess non-decision time variability (which is not currently clear), the only possible reason that the NDT-informed model could achieve better BIC is because the non-NDT-informed model gets lost in the fitting procedure and fails to find the global optimum. If they are in fact different models - for example, if the NDT-informed model is endowed with NDT variability, while the non-NDT-informed model is not - then the fit superiority doesn’t necessarily say anything about an NDT-informed reliability boost, but rather just that a model with NDT variability fits better than one without.

      To respond to this comment, we would like to note that the structural difference between NDT-informed and uninformed models is the assumption about non-decision time. In principle, the behavioral parts (i.e., parameters related to the diffusion part) of both NDT-informed and uninformed models are identical. However, during the estimation, the non-decision time in the NDT-informed model is subject to an additional constraint imposed by the neural data. In other words, in the uninformed model, non-decision time is estimated using the behavioral data by maximizing the likelihood of a CT-DDM. However, the NDT-informed model incorporates an additional data type, resulting in a different likelihood function (see Equation (3)). Specifically, in the NDT-informed model, we make an additional assumption regarding the non-decision time: it is set to the mean of the trial-level non-decision time measurements distribution derived from neural data (Equation (3) specifies the joint model structure). This additional assumption constrains the search space of the non-decision time parameter in the NDT-informed model. As in the main text, we assumed that non-decision time measurements obtained from neural data are log-normally distributed, where sigma is the shape parameter of the log-normal distribution. Therefore, sigma can represent the variability of the non-decision time measurements, and it is not the trial-to-trial non-decision time variability parameter. Indeed, the non-decision time parameter is fixed across trials in both models. Also, it should be clear from Equation (3) that sigma only appears in the second term of the joint likelihood, which corresponds to non-decision time measurements and not the CT-DDM term. Therefore, the behavioral parts of the NDT-informed models and Uninformed models are identical, and the NDT-informed models include one additional parameter (i.e., sigma) corresponding to the variability in non-decision time measurements.

      Also, to explain how constraining non-decision time improves the BIC, we would like to clarify that constraining non-decision time using an additional data source constrains the search space for non-decision time, thereby leading to better parameter identification and, consequently, a better fit to behavioral data. We included the following text in the discussion section to make this explicit.

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) One reason this is unclear is that Footnote 4 says that this study did not allow trial-to-trial variability in nondecision time, but the entire premise of using variable external single-trial estimates of nondecision times (illustrated in Figure 2) assumes there is nondecision time variability and that we have access to its distribution.

      We are sorry for the confusion. To respond to this comment, we would like to highlight that this modelling approach does not include any mechanism for across-trial variability in the behavioral part, as the likelihood of the choice behavior (see Equation (3)) does not include any variability parameter. As mentioned before, sigma represents the variability in non-decision measurements. However, the model does not propagate the across-trial variability of the neural data on the behavior side (unlike the models in Ghaderi-Kangavari et al. 2023). In other words, in this approach, none of the diffusion model’s parameters include across-trial variability, and we have only considered and estimated neural data variability in the NDT-informed model. Also, to improve the manuscript’s coherence, we removed this footnote.

      (3) It is good that there is an Intro section to explain how the tradeoff between NDT and collapsing bound parameters renders them difficult to simultaneously identify, but I think it needs more work to make it clear. First of all, it is not impossible to identify both, in the same way as, say, pre- and postdecisional nondecision time components cannot be resolved from behaviour alone - the intro had already talked about how collapsing bounds impact RT distribution shapes in specific ways, and obviously mean (or invariant) NDT can’t do that - it can only translate the whole distribution earlier/later on the time axis. This is at odds with the phrasing “one CANNOT estimate these three parameters simultaneously.” So it should be first clarified that this tradeoff is not absolute. Second, many readers will wonder if it is simply a matter of characterising the bound collapse time course as beginning at accumulation onset, instead of stimulus offset - does that not sidestep the issue? Third, assuming the above can be explained, and there is a reason to keep the collapse function aligned to stimulus onset, could the tradeoff be illustrated by picking two distinct sets of parameter values for non-decision time, starting threshold, and decay rate, which produce almost identical bound dynamics as a function of RT? It is not going to work for most readers to simply give the formula on line 211 and say ”There is a tradeoff.” Most readers will need more hand-holding.

      We are grateful for this comment. In response to this comment, which was also partly mentioned by the first reviewer, we first highlight that, in the presence of a nonlinear collapsing threshold, the effect of non-decision time is no longer linear, as it forces the threshold to take a specific value at the final stopping point. To illustrate the tradeoff between imprecise non-decision time estimation and collapsing threshold estimation, we conducted a simulation study. The results for the simulation study are presented in Appendix 1. In this study, we contaminated the non-decision time with noise at six levels (i.e., 0%,2%,4%,6%,8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters.

      Regarding the second point, we would like to clarify that in this paper, consistent with other works on collapsing threshold, we assumed that the collapsing starts with evidence accumulation and not with stimulus onset, as the decision makers need a short amount of time for perceiving and encoding the stimulus (i.e., perceptual encoding time). Fixing the collapsing onset to the stimulus onset introduces an ad hoc assumption into the model (that the encoding time is zero), which is cognitively implausible. Therefore, although this assumption can mathematically resolve the issue, it is not cognitively plausible. However, one important point we did not consider in the previous version is the two-stage accumulation process models in which the collapsing onset occurs later than the evidence-accumulation onset. We have discussed the estimation of such models as a limitation in the general discussion:

      “The collapsing threshold dynamics considered in this work (e.g., exponential and hyperbolic) impose a monotonically decreasing threshold over time. Although these dynamics are theoretically well motivated (Fudenberg et al., 2018; Frazier and Yu, 2007) and have been employed in several previous studies (e.g., Olschewski et al., 2025; Milosavljevic et al., 2010; Voskuilen et al., 2016), some research has proposed delayed-collapsing threshold models (e.g., Diederich and Oswald, 2016), in which the onset of threshold collapse does not coincide with the onset of evidence accumulation. Estimating such delayed-collapsing models may require more than simply constraining non-decision time, as the onset of threshold collapse must also be identified. A promising approach for addressing this challenge is the HMP method, which may provide additional temporal information about distinct cognitive processing stages. In particular, HMP may allow the onset of threshold collapse to be estimated as a separate cognitive stage. Future research should therefore investigate the estimation of multi-stage evidence accumulation models (e.g., Diederich and Oswald, 2016; Diederich and Colonius, 2021) within the HMP framework.”

      (4) A lognormal distribution is used as line 231 says it “must” produce a right-skew. Why? It is unusual for non-decision time distribution to be asymmetric in diffusion modeling, so this “must” statement must be fully explained and justified. Would I be right in saying that if either fixed or symmetrically distributed nondecision times were assumed, as in the majority of diffusion models, then the non-identifiability problem goes away? If the issue is one faced only by a special class of DDMs with lognormal NDT, this should be stated upfront.

      We would like to clarify that this assumption is about the non-decision time measurements and not about the across-trial variability parameter of non-decision time. Although for computational convenience, non-decision time is often assumed to follow a normal distribution in diffusion modelling literature, some empirical studies have shown that the perceptual encoding time and total approximated non-decision time follow a right-skewed distribution. Moreover, as HMP estimations of non-decision time are usually right-skewed, we formalized the model with a log-normal distribution. However, we would like to note that the identifiability issue in collapsing-threshold diffusion models is not related to the distributional assumption over the non-decision time measurements, as poor parameter recovery of the collapsing threshold was also reported in Evans et al. (2020). To show that the distributional assumption about the non-decision time measurements does not affect the results, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and showed that constraining non-decision time improves parameter recovery for collapsing threshold parameters. The results for this simulation are presented in Appendix 8 in the new version of the manuscript. We also revised line 231 as follows (see line 236 in the new version):

      “Empirical studies on non-decision time measurement usually have reported a right-skewed distribution for their measurements (e.g., Weindel, 2021; Weindel et al., 2025). For instance, the measured perceptual encoding time and motor execution time reported by Weindel et al. (2025) are right-skewed. Therefore, we assume that the observed non-decision time measurements Z<sub>n</sub> follow an approximate log-normal distribution, which is right-skewed (we will discuss how this distributional assumption can affect the results later). Thus, we model the non-decision time measurements Z<sub>n</sub>, using a log-normal distribution with parameters µ and σ<sub>z</sub>. The available measurements (i.e., observed data) at trial n can be represented as

      follows:”

      (5) In the simulation study methods, is the only difference between NDT-informed and non-informed models that the non-NDT-informed must also estimate tau and sigma, whereas the NDT-informed model “knows” these two parameters and so only has the other three to estimate? And is it the exact same data that the two models are fit to, in each of the simulation runs? Why is sigma missing from the uninformed part of Figure 4? If it is nondecision time variability, shouldn’t the model at least be aware of the existence of sigma and try to estimate it, in order for this to be a meaningful comparison?

      As mentioned in the response to your first comment, the difference between NDT-informed and uninformed models is the access to an additional source of data related to non-decision time, and sigma represents the shape parameter of the non-decision time measurements distribution. Therefore, sigma belongs only to the NDT-informed model, and, as in the uninformed model, there is no additional data source, so the model does not include sigma.

      (6) I am curious to know whether a linear bound collapse suffers from the same identifiability issues with NDT, or was it not considered here because it is so suboptimal next to the hyperbolic/exponential?

      Thank you very much for this comment. The main reason we did not include linear collapsing threshold models was the assessment by Evans et al. (2020), which indicated that, with sufficient trials, these models can be estimated reasonably well. However, to investigate whether constraining non-decision time can also improve the estimation of linear collapsing threshold models, we conducted an additional simulation study, which is reported in Appendix 3 in the new version of the manuscript. The simulation results confirm those reported by Evans et al. (2020) and show that the parameters of the uninformed linear models can be identified using more than 500 trials. Constraining the non-decision time using the NDT-informed diffusion modeling framework still improves parameter estimation in the linear collapsing threshold model and reduces the required number of trials for reliable estimation to 250. Especially, the estimation of the starting threshold improves significantly.

      (7) The approach using HMP rests on the assumption that accumulation onset is marked by the peak of a certain neural event, but even if it is highly predictive of accumulation onset, depending on what it reflects, it could come systematically earlier or later than the actual accumulation onset. Could the authors comment on what implications this might have for the approach?

      Thank you for mentioning this point. We first would like to point out that Weindel et al. (2024) showed that HMP can predict the underlying generative distribution of cognitive states with high precision and without systematic bias. Second, it is worth highlighting the results reported in Appendix 5 (i.e., “Bias in non-decision time measurements”). In this appendix, we discussed how bias in non-decision time measurements (i.e., systematic underestimation or overestimation) affects the estimation of the collapsing threshold. Particularly, see Figure 3 in Appendix 5. To make these results clearer in the main text, we included the following paragraph at the end of the results section in simulation study 1:

      “Additionally, we examined the effect of systematic bias in non-decision time measurement on parameter recovery. Appendix 5 presents the parameter recovery simulation results in the presence of biased non-decision time measurement (i.e., systematically underestimated or overestimated). The results revealed that, even in the presence of biased non-decision time measurement, the actual generating parameters show a high correlation with the estimated parameters. Underestimation in non-decision time leads to overestimation in the starting threshold and decay rate. Conversely, overestimation in non-decision time leads to underestimation in both the starting threshold and non-decision time.”

      (8) Figure 7: for this simulation, it would be helpful to know the degree to which you can get away with not equipping the model to capture drift rate variability, when the degree of that d.r. variability actually produces appreciable slow error rates. The approach here is to sample uniformly from ranges of the parameters, but how many of these produce data that can be reasonably recognised as similar to human behaviour on typical perceptual decision tasks? The authors point out that only 5% of fits estimate an appreciable bound collapse but if there are only 10% of the parameter vectors that produce data in a typical RT range with typical error rates etc, and half of these produce an appreciable downturn in accuracy for slower RT, and all of the latter represent that 5%, then that’s quite a different story. An easy fix would be to plot estimated decay as a scatter plot against the rate of decline of accuracy from the median RT to the slowest RT, to visualise the degree to which slow errors can be absorbed by the no-dr-var model without falsely estimating steep bound collapse. In general, I’m not so sure of the value of this section, since, in principle, there is no getting around the fact that if what is in truth a drift-variability source of slow errors is fit with a model that can only capture it with a collapsing bound, it will estimate a collapsing bound, or just fail to capture those slow errors.

      Thank you for this comment. We would like to first note that the aim of this section is to illustrate that NDT-informed modeling enables us to distinguish between CT-DDM and FT-DDM with drift variability (i.e., the two competing models that are relatively hard to distinguish). Therefore, we changed the name of this section to “Simulation study 2: Model recovery”. In the new version of the manuscript, we included a cross-model fitting simulation and a model recovery simulation in this section. The estimated decay rates in the cross-model fitting study indicate that, when the underlying generating model is FT-DDM with drift variability, parameter estimation using the NDT-informed approach yields precise inference. This is due to the estimated decay rate, which is very close to zero. The model recovery results also confirmed that incorporating the decay rate value into the model inference improves precision.

      Moreover, to address the concern regarding the slow-error pattern in the simulated data, we examined the relationship between the slow–fast accuracy difference and the estimated decay rate. Specifically, we computed the difference between the accuracy of responses with response times below the median (ACC1) and those above the median (ACC2), and plotted this difference against the estimated decay rate. Author response image 1 presents the resulting scatter plot, with color indicating the drift rate. As shown in Author response image 1, incorrect inferences about the decay rate primarily occur at high drift rates. This pattern emerges because high drift rates produce very fast responses, leaving little time for the threshold to meaningfully collapse. Consequently, the behavioral signatures of collapsing-threshold and fixed-threshold models become increasingly similar under high drift conditions.

      Author response image 1.

      Illustration of the difference between the accuracies of responses with response time below the median and those with response time above the median against the estimated decay rate. Colour shows the drift rate value.

      Reviewer #2 (Recommendations for the authors):

      (1) Abstract: improves fit to behaviour in what way? Reliability or absolute quant fit to behavior? I.e., is it just helping constrain it so it finds the global opt?

      (2) Line 72 - Are these “neuroimaging” studies? Perhaps use the broader term “neuroscience.”

      (3) Line 95 - It’s important because it may confuse readers how something dynamic like a collapsing bound could be resolved with, say, fMRI.

      (4) Line 85 – “greater variability”... in what?

      (5) Line 87 -Revise to avoid misconstruing a time on task effect - e.g., “higher error rate for trials with longer RT” is more explicit.

      (6) Line 89-94 - It’s not clear what findings are being referenced here.

      (7) Line 103 - External biases such as priors and relative value?

      (8) Line 175 - Clarify this applies to any ddm, not just ctddm.

      (9) Line 268 - Explain what portion N200 accounts for.

      (10) Line 327 - Is this equation supposed to be for Delta-X, as opposed to X(t+delta-t)? If you want X(t+delta-t) on the LHS, then X(t) must be added to the RHS.

      We are very thankful for such a precise evaluation and constructive comments. We addressed all the comments raised by the reviewer in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The current paper addresses an important issue in evidence accumulation models: many modelers implement flat decision boundaries because the collapsing alternatives are hard to reliably estimate. Here, using simulations, the authors demonstrate that parameter recovery can be drastically improved by providing the model with additional data (specifically, an EEG-informed estimate of nondecision time). Moreover, in two empirical datasets, it is shown that those EEG-informed models provide a better fit to the data. The method seems sound and promising and might inform future work on the debate regarding flat vs collapsing choice boundaries. As an evidence-accumulation enthusiast, I am quite excited about this work, although for a broader audience, the immediate applicability of this approach seems limited because it does require EEG data (i.e., limiting widespread use of the method or e.g., answering questions about individual differences that require a very large N).

      We are very grateful for your positive evaluation and your comments.

      Reviewer #3 (Recommendations for the authors):

      This is a very decent study, very well written and properly executed. Most of my comments below are suggestions for the authors as to how to make the story more compelling.

      (1) I think the authors can do more to explain why the NDT-informed models fit better to empirical data. If the NDT-estimates are equal to the ground truth, then isn’t it possible that such models win simply because they have one parameter less? However, this is not what the authors are claiming, though, on l.567 it says that the fit is better because parameters are better estimated. I think it should be possible to dissociate these two accounts.

      Thank you for this important comment. For clarification, both NDT-informed and uninformed models estimate the non-decision time parameter, and neither treats it as equal to the ground truth. However, the difference between these two models is that, in the NDT-informed model, the non-decision time is subject to an additional constraint; therefore, the number of behavioral parameters (i.e., parameters of the diffusion part) is identical for both NDT-informed and uninformed collapsing threshold models. Although the number of behavioral parameters is identical in the NDT-informed model, one parameter is constrained, thereby limiting the model’s complexity and flexibility compared to uninformed models. Consequently, the improvement in the fit can only be attributed to better parameter estimation in the NDT-informed models, resulting from constraining the non-decision time using neural measurements. To clarify that in the manuscript, we included the following text in the revised version:

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) Given that so much of the writing focuses on the conclusion that collapsing boundary models ¿ flat models, it is very odd that there is no comparison to a flat boundary model reported in the text. Why is the model in Supplement 5 not just included in the main text (and in Table 1)? This would make it so much easier for the reader.

      To address this comment and also the similar point raised by Reviewer 1, we included the NDT-informed FT-DDM in the results section and compared the FT-DDM with CT-DDMs with respect to BIC<sub>Joint</sub>. However, we retained the other FT-DDM in the appendix because this model includes across-trial variability in drift rate, whereas the models in the main text do not.

      (3) I would have appreciated a bit more background about the importance of the number of trials per participant. Given that the proposed method requires collecting EEG data (which is time and labour-intensive) I wonder to what extent you get similar improvements in parameter recovery by collecting more data per participant (which is usually cheap). Put differently, I would appreciate an additional matrix in Figure 5 for vanilla models.

      The new version of Figure 5 in the revised manuscript now contains the goodness of parameter recovery for the uninformed CT-DDMs for different numbers of trials. As the simulation results suggest, even with 1000 trials, the parameters of the uninformed CT-DDMs are still not reliably identifiable. Therefore, these results suggest that although increasing the number of trials can slightly improve the parameter recovery of the CT-DDMs, it cannot fully resolve the reliability issue in their parameter recovery.

      In addition, it is worth clarifying that, although the paper focuses on extracting non-decision time from the EEG signal using the HMP method, as discussed in the general discussion, the NDT-informed approach can also be employed with purely behavioral methods.

      (4) L116-117: minor detail: I don’t think it’s fair to write that it’s an open question whether or not thresholds collapse. I think it’s fair to say that this is hard to show, and that the conditions under which it appears are unclear; but saying that it’s still unclear whether this occurs at all seems unfair with regard to previous work.

      Thank you for mentioning this point. We revised the mentioned sentence as follows in the new version of the manuscript:

      “This issue is particularly critical given that the conditions under which individuals adjust their decision thresholds during a single trial remain an open question.”

      (5) Figure 5: Minor detail: It would be useful to mention in the figure or caption how many trials underlie these simulations, which is now somewhat buried in the text.

      As Figure 5 shows sensitivity to the number of trials, and the number of trials is explicitly mentioned in the figure, we assume that the reviewer intended Figure 6. In the revised manuscript, we explicitly specify the number of trials for the results reported in Figure 6. We simulated 500 trials for each noise level in this graph.

      (6) Figure 10 and similar figures: I tried to figure out how corrects and incorrects differ, but couldn’t see the difference. Can the authors use something more colorblind friendly (hope I didn’t give up on being anonymous here)?

      We are so sorry for the inconvenience. In the revised manuscript, we used different symbols to distinguish correct and incorrect data.

    1. eLife Assessment

      This important study investigates the peptide-binding principles by promiscuous chicken MHC molecules. The data from crystallography, mass-spec, and modeling are compelling. This paper will be of broad interest to immunologists and those interested in vaccine development.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

    3. Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates the peptide-binding principles of promiscuous chicken MHC molecules. The data from crystallography, mass spectrometry, and modeling are convincing. However, the presentation would benefit from streamlining and clear links between data and conclusions. This paper will be of broad interest to immunologists and those interested in vaccine development.

      Overall, we are delighted and grateful to the eLIFE editors and the two reviewers for the careful and thoughtful assessments and reviews of our paper. We are glad that the strengths of the paper were apparent and appreciated. And of course, every paper has weaknesses, especially for a story as complex as this one.

      We made only minor changes to accommodate the reviewer comments, along with additions for which we only became aware upon this submission of a revised manuscript. In particular, we shortened the title and abstract to fit what is usual for an eLIFE paper, added Key Resources table with accompanying references, changed the numbering of the figures throughout the manuscript to ensure that each page represented a figure (rather than panels of a figure), moved the figure legends from the embedded figures to a list near the end of the manuscript, and split the supplemental spreadsheet into two renamed Data Source files.

      Before answering the comments and questions directly, perhaps a few points would help clarify why the paper is as it is.

      First, the experiments cover over three decades of work, with the first gas phase sequencing results done in 1992. Unlike some of the chicken class I alleles which immediately gave completely clear stringent motifs (B4, B12 and B15 in Wallny et al 2006 PNAS, B19 in Han et al 2023 J Immunol), we harvested nothing but confusion from the B21 class I results (Fig. 1). Initially, we thought that the lack of a clear motif for B21 was due to multiple well-expressed class I molecules but only one dominantly-expressed class I molecule was found (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol) and, to our surprise, bacterially-expressed BF2*21:01 heavy chain and b<sub>2</sub>-microglobulin refolded with two synthetic peptides without sequence in common, and the crystal structures showed that this molecule remodeled the binding site to accommodate two such disparate peptides (Koch et al 2008 Immunity). This was the beginning of our understanding of the spectrum of class I alleles from promiscuous generalists to fastidious specialists, which we have explored in a series of further papers (in particular, Chappell et al 2015 eLIFE, Tresgaskes et al 2016 PNAS, Kaufman 2018 Trends Immunol, Tregaskes and Kaufman 2022 Mol Immunol).

      Second, over these many years, we continued to explore the binding properties of BF2*21:01 in ever more detail, resulting in the current manuscript. We learned only slowly how to probe this unexpected promiscuity, unprecedented in the MHC literature, so that the experiments proceeded with our best understanding at the time, including taking advantage of new approaches as they become available. Each experiment built on the previous set of experiments and each brought us closer to an understanding.

      Third, having amassed a collection of data, we chose eLIFE exactly because it allows us to present the entire story from beginning to end without compromise, not just the highlights with the major points illustrated by a few main figures and with the supporting data in many supplementary figures. We include all the data, because it is all part of the story, and so interested researchers to look at the data from their own perspective. Although mostly we provide bar graphs, we include spreadsheets for the raw data (or close to them) for the final experiments (illustrated by Figs. 10 and 14-22) in the two source data files, so these can be assessed easily by others in the field, perhaps using approaches that we may not feel competent to perform.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

      Weaknesses:

      The manuscript is sometimes difficult to follow because the authors present a large amount of peptide-library, structural and immunopeptidomics data. without always clearly explaining how these datasets support the proposed simplifying principles.

      We are delighted and grateful to the reviewer 1 for the careful and thoughtful comments and questions concerning our manuscript. We are glad that the strengths of the paper were apparent and appreciated, and acknowledge the weaknesses that come with such a complex story with experiments performed over decades.

      Major Issues - Points Requiring Clarification or Additional Support:

      (1) Line 282-301, 537-545)

      The immunopeptidomics conclusions are mainly based on one B21 cell line with one biological replicate and at least two technical replicates. Given the complexity of the BF2*21:01 peptide repertoire, this is a major limitation. The authors should either provide additional biological replicates or clearly state this limitation in the Abstract, Results and Discussion.

      This limitation is clearly stated in lines 537-545, as part of a paragraph covering the various ways in which the data presented in this manuscript could be improved. In fact, we have performed immunopeptidomics of several different B21 cell types, with many replicates and found similar data as presented, giving us confidence in our interpretations. However, these other experiments belong in different stories, so it is not appropriate that the data be reported in this manuscript.

      (2) (Lines 290-313)

      The B21 cell preparations contain both BF2 and the lowly expressed BF1 molecule. Some peptides, especially 8-mers or peptides with atypical motifs, may derive from BF1*21:01. The authors should clarify how BF2*21:01-bound peptides were distinguished from possible BF1-derived peptides, or interpret the immunopeptidomics motif more cautiously. The authors should also provide or cite evidence confirming the B21 haplotype identity of the cell line and chicken materials used for immunopeptidomics.

      The concern about the contribution of BF1*21:01 to the immunopeptidomics is clearly stated in the manuscript, both lines 290-313 and as part of the paragraph describing the limitations of the experiments (lines 542-543). In fact, the expression of BF1 molecules has long been known to be less than 10% of BF2 molecules at the RNA level, and much less at the protein level (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol). The proportion of 8mers identified by immunopeptidomics is also low (Fig. 14), and it is not impossible that most 8mers are due to BF1*21:01. We have used assembly assays with peptide libraries, immunopeptidomics and a crystal structure to determine the peptide motif for typical BF1 molecules, of which BF1*21:01 is one and found it may contribute to 8mer peptides but very seldom to longer peptides. This work is unpublished but gives us confidence that the characteristics of BF2*21:01 are not misrepresented by the data in this manuscript.

      The sources of the chicken samples and the cell lines are described in detail under Materials and Methods (lines 577-590), citing relevant publications.

      (3) (Lines 217-221, 243-253)

      The authors acknowledge that MALDI-TOF cannot reliably distinguish peptide combinations with identical or similar masses, nor determine residue positions in some cases. Therefore, MALDI-TOF results should not be overinterpreted as precise evidence for residue preference. The authors should clearly indicate which conclusions are supported by LC-MS/MS.

      As described, the experiments follow each other in temporal sequence, so that we started with single peptides, then peptide libraries that varied in one position, then peptide libraries that varied in two positions first analysed by MALDI-TOF and later by LC-MS/MS. The final experiment (Fig. 10, with the original data in the supplementary spreadsheet) directly compares MALDI-TOF and LC-MS/MS results for six peptide libraries, so that the strength of the evidence for residue preference is clear. Throughout the manuscript, we do our best to not to overstate conclusions based on the data of any particular experiment.

      (4) (Lines 297-301, 316-330)

      The authors suggest that longer peptides may bulge in the middle or extend out of the groove at the C-terminal end. The rationale for the C-terminal extension is not clearly explained. Why is the C-terminal extension considered rather than the N-terminal extension? If the binding register is uncertain, long peptides should be analyzed separately from canonical-length peptides.

      When the first sequence of a chicken class I cDNA was determined, an immediate mystery was why one of the so-called invariant residues that coordinate the N- and C-termini of the bound peptide is not conserved (Kaufman et al 1992 J Immunol). In fact, this residue Tyr at position 86 in HLA-A2 and the equivalent position in all mammalian classical class I molecules is an Arg in the classical class I molecules of all non-mammalian vertebrates and is common with class II molecules (Kaufman et al 1995 Semin Immunol). Similar to class II molecules, this Arg in chicken class I molecules allows the peptide to extend out of the C-terminus, as shown by a crystal structure (Xiao et al 2018 J Immunol). The concern that we might be misidentifying the C-terminal amino acid was the basis for the analysis in Figs. 23 and 24, but in the absence of crystal structures, we are not able to provide a final answer this question. Perhaps relevant is the fact that a chicken class II molecule can bind exactly the same peptide in two conformations, one with a canonical 9mer core and the other with an unexpected 10mer core (Goryanin et al 2026 J Virol).

      By contrast, N-terminal extensions are only found for some class I alleles and thus far depend on the substitution of small amino acid sidechains for W166 (Li et al 2011 J Virol for bovine, Ma et al 2020 J Immunol for Xenopus, Wei et al 2022 J Immunol for ovine). Thus far, no chicken BF2 sequences have this substitution, consonant with the many crystal structures, including those for BF2*21:01 (Koch et al 2008 Immunity, Chappell et al 2015 eLIFE, this manuscript). However, in unpublished data, we find that most BF1 sequences have sequence differences that could allow N-terminal extensions, although we have no crystal structures to support this possibility.

      (5) (Lines 406-439)

      In vitro assembly assays show that several hydrophobic residues can be tolerated at Pc, whereas immunopeptidomics shows a strong Leu preference at this position. The authors should clarify whether this Leu preference reflects intrinsic BF2*21:01 binding specificity, TAP-mediated peptide transport, antigen processing, peptide loading, or a cell-line-specific effect. Additional experimental support, such as TAP transport analysis, would strengthen this conclusion.

      The preference for Leu at the final position of the peptide by immunopeptidomics of the B21 cell line is strong but not absolute and is certainly affected at the least by the length of the peptide (Figs. 23 and 24). Unpublished immunopeptidomics results (mentioned above) show that this is not a cell line-specific result. The evidence from assembly assays of various peptides is that several hydrophobic amino acids are tolerated with sufficient stability of BF2*21:01 that they are detected in the assay (Figs. 3, 5, 9 and 10). Thermostability assays (Fig. 6) show that peptides with these same hydrophobic amino acids are stable to at least body temperature of chickens. These experiments show that such stability is peptide-dependent (that is, whether a particular amino acid is tolerated depends on the stability conferred by the rest of the peptide). Finally, peptide translocation assays using B21 cells have been done (Tregaskes et al 2016 PNAS) and show that peptides with several hydrophobic amino acids can be pumped into the lumen of the endoplasmic reticulum. However, the assays are with single synthetic peptides, so the data are not extensive enough to separate the effects of the final amino acid from the rest of the peptide. Certainly, peptides with amino acids other than Leu at the C-terminus can be translocated. So, it is not yet clear at which point the preference for Leu at the C-terminus of the peptide arises.

      (6) (Lines 172-178, 243-279, 442-457)

      The structural analysis explains some residue combinations, such as Arg at P2 with Glu at Pc-2 or Trp at Pc. However, the structural interpretation is not fully integrated with the large-scale peptide library and immunopeptidomics results. Representative high- and low-frequency combinations should be discussed structurally.

      Six crystal structures show that BF2*21:02 remodels the binding to accommodate a variety of anchor residues (Koch et al 2008 Immunity, Chappel et al 2015 eLIFE). These crystal structures are representative of sequences found by the immunopeptidomics from very frequent (H-E at roughly 15% 8-12mers) to moderately frequent (E-L at roughly 6% 8-12mers) to infrequent (N-F, A-D and E-D at roughly 1.5%, 1.6% and 0.7% 8-12mers) based on Fig. 18. All but one of the structures has Leu at the C-terminus, with the last one having Val which is found but not frequently by immunopeptidomics.

      Similar numbers are found by LC-MS/MS of double-substitution libraries of the two original peptide sequences in Fig. 10 with H-E found frequently (8.1% in P390, 3.8% in P498) and the others infrequently (0.1, 0.9, 1.0, 0.3% in P390, 0, 1.4, 1.0, 0.3% in P498), as calculated from the numbers in the Supplementary data spreadsheet. As discussed in the manuscript, for single-substitution peptide libraries of the two original peptides, Ile/Leu at the C-terminus was very frequent but at the same or slightly less level as Phe, with Met less frequent and Val even less so (Fig. 7).

      In addition, there are two more structures along with models explicitly testing some substitutions (Fig. 5). Attempting more current modelling approaches, we found AlphaFold 3 was unable to correctly predict most of the conformations that are found in the crystal structures of BF2*21:01, so we don’t feel confident in using them to predict unknown structures of this kind.

      (7) The inference of co-variation between P2 and Pc-2, as well as the modulatory effects of P3 and Pc-3, should be better explained. At present, some conclusions appear to be based mainly on residue-frequency patterns, and the logical connection between these observations and the proposed binding principles is not always clear. Statistical analyses, such as mutual information, chi-square tests or permutation tests, and representative structural explanations would strengthen this conclusion.

      We endeavored to do our best to explain the data, our interpretations and our reasoning, so we apologise if we have not managed to be as clear as might be desired. We have included as close to raw data as possible for the LC-MS/MS and MALDI-TOF (Fig. 10) and for the immunopeptidomics (Fig. 14 and 18) in the Supplementary Data spreadsheet, exactly so that competent practitioners can carry out further analyses (including the sophisticated statistical tests mentioned).

      Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

      Weaknesses:

      The data overall support the analysis and conclusion well. The only caveat is linked to Figure 4, which does not describe the stability of the peptide-MHC complex, but instead shows refold yield, and the two are not always linked.

      We are grateful for the clear understanding of the strengths of the work. With regards to Fig. 4, we agree with the reviewer that there are differences in refold yield but that measure may not be correlated with stability of the peptide-MHC complex. However, we were basing our interpretation of stability on the position and quality of the monomer peak, as illustrated by the trace in Fig. 2, in which a sharp peak at the monomer position represents a stable complex (as seen for the 10 and 11mer peptides) and later peaks represent unstable complexes falling apart during the chromatography (as seen for the 7, 8 and 9mer peptides).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Issues: Editorial and Data Presentation Modifications

      (1) Lines 53-62, 155-170, 303-314: The terms Pc, Pc-2 and Pc-3 should be clearly defined early in the manuscript and figure legends.

      The Abstract introduces the abbreviations as “peptide positions P<sub>2</sub> and P<sub>c-2</sub>” followed by P<sub>C</sub>, which are standard usage and seem clear. The first usage in the text is “anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub>” which again seems clear, particularly in the context of the text and the figure. However, a parenthetical description has been added to read “…anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub> (position 2 and the position two before the C-terminus) …”. Given these usages, it seems unlikely that the reader will fail to understand P<sub>C</sub> and P<sub>C-3</sub>.

      (2) Lines 255-279: The term "peptide backbone" should be defined as the fixed sequence context outside the randomized positions, if this is what the authors mean. Suggest clarifying the meaning of "peptide backbone".

      The phrase reads “double-substitution libraries based on four peptide backbones”, which in context of the figures seems clear. However, a parenthetical description has been added: “The examination of double-substitution libraries analysed by MALDI-TOF was expanded to other backbones (that is, other sequences in which two positions were randomised): …”.

      (3) Several figures are complex. The authors should add brief take-home messages to figure legends.

      Every figure legend starts with a (sometimes quite long) take-home message. It is not clear what more should be added.

      (4) A concise summary table of the proposed binding rules, including preferred peptide length, P2, Pc-2 and Pc preferences, and the effects of P3/Pc-3, would be useful for readers.

      A concise summary table of binding rules would be very helpful, but the rules are complex, both qualitative and quantitative. For example, immunopeptidomics shows that 10mers are preferred, but that fails to capture the quantitation. The co-variation of P<sub>2</sub> and P<sub>c-2</sub>, which is the strongest and best characterized correlation, nevertheless is complex since the occupancy of P<sub>2</sub> by different amino acids (presumably independently of P<sub>c-2</sub>) varies considerably. At our current level of understanding, it is hard to imagine a table that is both concise and precise. With time and more data, perhaps code for quantitative prediction might be constructed (dare one suggest machine learning…), which is the hope of presenting all the available data in one place.

      (5) The Results section contains a large amount of peptide-library, structural and immunopeptidomics data. The authors should improve the logical flow and add clearer transition sentences to explain how each dataset supports the proposed simplifying principles.

      The reviewer is of course correct that any written text can be improved (although each critic may have a different idea about which part should be fixed), but we have done our best with the material and time available. We could respond productively to a more detailed critique.

      (6) Some statements in the Abstract and Discussion should be softened, especially those related to in vivo peptide preferences, BF2*21:01 promiscuity and peptide prediction, given the limited biological replication and uncertainty in peptide assignment.

      The statements in the both the Abstract and Discussion are very general, reflecting what we believe to be careful interpretations based on the data. We could respond productively to concerns about specific claims.

      Reviewer #2 (Recommendations for the authors):

      Overall, the data presented in this study are interesting; however, it is complex, and some results could be merged and simplified, as well as the figures. The data provide an in-depth analysis, using mass spectrometry, of the interplay between the different positions of the peptide and the residues favourable to bind within the antigen-binding cleft.

      (1) From the abstract, the concept of "promiscuous generalists and fastidious specialists" is not explored after or defined within the results.

      The Abstract introduces the concept of promiscuous generalists and fastidious specialists to provide the basis for exploring the peptide-binding specificity of the most promiscuous class I known, BF2*21:01. This overall concept is described in enormous detail in several publications cited in the current manuscript, but it is not particularly germane to the analyses.

      (2) From the abstract "These simplifying principles may eventually allow predictions of pathogen peptides", I'm not sure how "simplifying" the principles are with the data, if anything, it does show a rather complex interplay between the different residues of the peptide that enable the MHC to bind a large and diverse number of peptides.

      Compared to any combination of anchor residues being permitted at equal frequency, there are clear preferences which the experiments identify. Instead of an enormous range of possible peptide lengths, roughly 50% of peptides are 10mers. Of course, the structural reasons behind these results are certainly complex and likely must be understood in detail in order to attempt peptide predictions. 

      (3) Line 48. "less well-expressed". Less than what? Do you mean the level of expression was lower? And if yes, are there values of fold change for comparison?

      The sentence in the Abstract reads “Chicken BF2 alleles … are less well-expressed on the cell surface … while certain human HLA-B alleles … are well-expressed …”. Read as a complete sentence, the meaning is clear. This difference for chicken BF2 alleles has been quantified as reported in several publications cited in the current manuscript, ranging from 3-5 fold for peripheral blood lymphocytes to ten-fold for erythrocytes (Kaufman et al 1995 Immunol Rev, Chappel et al 2015 eLIFE), with similar numbers for a few HLA-B alleles on human peripheral blood lymphocytes and monocytes (Chappell et al 2015 eLIFE).

      (4) Line 149. "with individual peptides" which peptides are we referring to here?

      This introductory sentence to a paragraph outlines the general method used for the experiments in this section of the Results, “refolding in vitro … with individual peptides.” Which “individual peptides” are described in the following paragraphs, with the next section of the Results using “refolding in vitro … with peptide libraries”.

      (5) Line 170. If 9 mers and below are not stable, which is not really quantified or shown with the data on Figure 4, why is refolding material observed for peptides with different lengths of 9 aa and below on Figure 4? A lower yield of refolded material can have a different origin, and there is no association between stability and refold yield. The notion of stability here probably needs to be changed, as it does not apply to the data.

      As described in our response to a similar concern above, we are not basing our interpretation of stability on the quantity of refolded material, but on the position and quality of the monomer peak, as described clearly in the legend to Fig. 4: “The original 10mer (REVDEQLLSV) and 11mer (GHAEEYGAETL) peptides refold with BF2*21:01 to give stable monomers as do 11mer and 10mer derivative peptides, but 9mer, 8mer or 7mer peptides give heavy chain only peaks” and “The peptides 11mer GHAEEYAETL (top panel), 10mer REVDEQLLSV (middle panel), 11mer GHAEAAAAETL and 10mer GHAEAAAETL (bottom panel) gave sharp monomer peaks, while the 9mer GHAEAAETL, 8mer GHAEAETL and 7mer GHAEETL gave a delayed broad peak indicative of unstable binding or heavy chain.” This concept is illustrated by the trace in Fig. 2 (discussed in the text at the beginning of this section of the Results), in which a sharp peak at the monomer position represents a stable complex, and later peaks represent unstable complexes falling apart during the chromatography or free heavy chains.

      (6) Line 175. As the 3BEV structure had a P2-His, is a comparable structure expected?

      This sentence reads “Structures with amino acid substitutions in the 11mer peptide GHAEEYGAETL bind with Asp at P<sub>c-2</sub> and either His or Arg at P<sub>2</sub> (5AD0 and 5ADZ), comparable to the original structure (3BEV) (Fig. 5A).” Minor changes in positions and orientations of individual amino acid sidechains are expected and are clear from the crystal structures presented in Fig. 5A, but are overall comparable to the original peptide in 3BEV, which has a His at P<sub>2</sub> and a Glu at P<sub>c-2</sub>.

      (7) Line 175 "modelling the substituted". How was the modelling done?

      The sentence reads “Modelling the substituted peptide with Arg at P<sub>2</sub> and Glu at P<sub>c-2</sub> shows a steric clash that can explain why this peptide did not refold with BF2*21:01 (Fig. 5A).” The legend to Fig. 5A states that “modelling done as detailed in Materials and Methods”, but apparently that section was omitted. A section has been added now to the Materials and Methods which states “Beginning with known crystal structures, modelling was carried out using PyMol with the protein mutagenesis Wizard, the rotamer toggle, show bumps and show surfaces.”

      (8) Line 176 "shows a steric clash" with what? Figure 5 is too small to see, and there is no label on the residue to follow where the steric clash is coming from.

      As stated in the legend to Fig. 5, “Structures were determined for GHAEEYGAETL (3BEV), GHAEEYGADTL (5AD0) and GRAEEYGADTL (5ACZ), which all refolded successful to give stable monomers, while GRAEEYGAETL did not (see Fig. 3), all models of which showed steric clashes (one depicted, red arrow).” In the model shown, the clash is between R9 of the BF2*21:01 with the Glu at peptide position 9. Parenthetically, this depiction of the key MHC residues for the co-variation has been used repeatedly, starting with Chappell et al 2015 eLIFE. The picture can be zoomed to make it large enough to see.

      (9) Line 247 "many combinations". It is not clear here if combinations are referring to a set of double-substitutions or different peptides?

      Each peptide in the library has a different double-substitution, so the meaning is the same either way. The point of Fig. 10 is to compare the identification of peptides from LC-MS/MS (in which each peptide is identified exactly) with the identification of sets of peptides from MALDI-TOF (in which the order of the double substitution is not clear, as well as the exact amino acid in the cases of I/L and Q/K). As the figure shows, the numbers are generally very similar, but this sentence describes the percentage of cases for which only one method or the other identified a peptide.

      (10) Lines 252-253. The conclusion of the MS data that the LC-MS/NS is more accurate and sensitive than MALDI-TOF is not very surprising. What was the rationale for using both?

      Our examination of the binding properties of BF2*21:01 for peptides and peptide libraries developed over a long time-span, so in the beginning we looked at single peptides with size exclusion chromatography peaks as the measure, then single- and later double-substitution libraries, first by MALDI-TOF and later by LC-MS/MS. Each set of experiments built on the previous work, so that together they tell the story. For example, after we optimized the use of double-substitution libraries, we were worried about the effects of temperature, so we tested that by MALDI-TOF. While repeating the temperature experiment once we optimized the LC-MS/MS approach might have yielded some additional data, there were other questions to answer.

      (11) Line 272-273 "support the idea that P3 is an important position within the peptide (Figures 8-10) despite not contacting the MHC molecule" The structure of 3BEV clearly shows interaction between the P3-Ala and the Tyr156 of the MHC. Residues that are fully or partially buried in the MHC cleft almost all contact the MHC molecule. Maybe I've missed something, but I think this statement is inaccurate.

      We agree with the reviewer that nearly every peptide residue contacts the MHC molecule (but of course some much more than others). The statement is now changed in the text to read “support the idea that P<sub>3</sub> is an important position within the peptide (Figures 8-10) despite not being an anchor residue."

      (12) Line 294. Figure 14 clearly shows that the number of 8 and 9-mer peptides eluted is at the same level as the 11mer and above, so how does the data fit with the statement that 9mer and shorter peptides are not stable with the MHC?

      Fig. 4 shows that 7, 8 and 9mer derivatives of the original 11mer failed to refold to give a single peak of stable monomers, while Fig. 6 shows that the 10mer derivative of the original 11mer was more thermostable than both the 9mer derivative and the original 11mer. The reason why the 9mer yielded so little monomer in Fig. 4 while giving enough to test by thermostability in Fig. 6 is no longer remembered. A key point is that these results are peptide sequence-specific, so it is not impossible to imagine stable binding of an appropriate 8mer (or perhaps even a 7mer). Another uncertainty, mentioned in the Results and Discussion, is that BF1*21:01 molecules bind primarily 8mers, and the contribution of peptides from BF1*21:01 is not known with certainty.

      (13) Line 316. What was the rationale for choosing peptides > 12aa to see if there is some overhang or bulge? As even 11-12mer can exhibit such features.

      We were just looking for any obvious patterns, but we didn’t find any.

      (14) The section starting at line 366 would have benefited from some structure prediction or modelling to illustrate the findings.

      We would have been delighted to model peptides, but we have used AlphaFold3 to compare the models to our crystal structures for seven chicken class I alleles (including BF2*21:01) with one or more peptides. The models sometimes fit the experimental data but they often didn’t, often by a wide margin, and with BF2*21:01 the worst (presumably because the system is not trained on MHC molecules which utilise charge transfer). Therefore, we do not feel confident in using any such modelling approaches except in the most defined situations (such as illustrated in Fig. 5).

      (15) Typo - Alleles should be italic, and in vitro as well.

      Alleles of genes are in italics, but alleles of proteins are not. To write in vitro in italics is customary, which we have corrected.

      (16) Figures

      (a) Some of the figures could be merged together.

      Of course, any presentation can be improved, but which figures we should merge (some already being three pages in length) is not clear. We could productively respond to more detailed suggestions.

      (b) Figure 1. I can't see the different colours mentioned in the figure legend

      Our apologies if the colours are not clear enough, but they are present only as an aid (as elsewhere in the manuscript), with red D and E, blue H, K and R, green N, Q, S and T, and all other amino acids black.

      (c) Figure 12. Is this figure only with 11mer peptides?

      Figure 12 shows the percentage of peptides with particular amino acids at P<sub>2</sub> and P<sub>c-2</sub> for three 11-mer double-substitution peptide libraries, the sequences of which are written on the graphs and described in the figure legends.

      (d) Table 1. The name of the protein should be added to the table.

      OK.

    1. eLife Assessment

      This study makes a valuable contribution by broadening the range of eukaryotic model systems and establishing Blastocystis, the most prevalent microeukaryote in the human gut, tractable for reverse-genetics investigations. The presented imaging data are convincing and informative. The work should interest readers studying host-microbe interactions in the human gut, as well as those developing new systems for eukaryotic research.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the reviewers' suggestions.]

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in Blastocystsis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did - so this data set is extremely useful indeed.

      Comments on previous version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystsis research.

    3. Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

    4. Author response:  

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystis research.

      We are grateful to Reviewer 1 for their careful and constructive assessment of our revised manuscript. We are pleased that the revisions have satisfactorily addressed their previous concerns, and we sincerely appreciate their recognition of the value of our work.

      Reviewer #3 (Public review):

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

      We are grateful to Reviewer 3 for their careful reading of the revised manuscript and for the additional constructive suggestions. We appreciate that these recommendations are aimed at improving clarity, accessibility, and presentation, and we have addressed them as detailed below.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      (1) Table 1

      (1a) The description of the column "Promoter Lengths Tested" is a bit confusing because it actually shows acronyms of the used promoters. Suggestion for clarity: either keep the name and mention just the lengths, or rename to "Promoter abbreviation" or "Promoter acronym".

      We thank the reviewer for pointing this out. To avoid confusion, we have renamed this column “Promoters Tested”, which better captures the information presented. The column contains both promoter identifiers and the corresponding promoter lengths, which are further clarified in the table notes.

      (1b) It is not entirely clear based on the formatting of the table, which columns are relevant for "Blastocystis ST4-WR1" and which for "Blastocystis ST7-B".

      We thank the reviewer for this helpful suggestion. We have revised the formatting of Table 1 to more clearly distinguish the data corresponding to Blastocystis ST7-B and Blastocystis ST4-WR1. Specifically, we have added a vertical divider between the relevant column groups to improve readability and make the species-specific information easier to follow.

      (2) Figure 2:

      If the authors wish to preserve the decorative colouring in the parts B and C, it would be better to at least keep the datapoint and boxplot hues consistent for individual conditions (or plot the datapoints in the same, neutral hue throughout, e.g., black or grey). Currently, this is only the case in the part C, but not in the part B.

      We thank the reviewer for this useful suggestion. We have revised Figure 2B to improve visual consistency between the data points and boxplots while preserving the overall colour scheme of the figure. This adjustment improves readability without changing the data or its interpretation.

      (3) Figure 3: With the modifications everything is clear. However, two issues are now apparent.

      (3a) After the part A was modified, the arrow colours changed (had been yellow and red, now are cyan and magenta), but the description in the legend remained (yellow and red), so the text should be corrected. [By the way, the original colours worked well.]

      We apologise for overlooking this inconsistency. The Figure 3 legend has now been corrected to match the revised arrow colours in panel A.

      (3b) The care taken by the authors to make the images more accessible is really greatly appreciated, but being a person on the colourblind spectrum, I can attest that the issue is not always just about red and green discrimination: the greyscale and blueish green used in the part A photo in particular are discernible only at huge magnification (to some people). If the signal in the part A were of the same hue as in the part B (yellowish green instead of blueish green), it would readily pop out. For that matter, what is the reason for the use of a different colour profile of the UnaG fluorescence in these two images?

      We sincerely thank the reviewer for this important accessibility-related comment. The difference in colour profiles between panels A and B was intentional because the images were acquired using different imaging modalities, as indicated in the figure legend. We wanted to avoid implying that the two panels were generated under identical imaging conditions. However, we appreciate the reviewer’s point that this distinction should not come at the expense of readability. We have therefore adjusted the display of the UnaG signal in panel A to improve contrast and visibility. 

      (4) Conclusion:

      The expression "bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements" feels a bit clunky. What the part "endogenous regulatory part discovery" alludes to is now clearer, but consider reformulating it to "discovery of endogenous regulatory elements". This would make it clear without the need to add the explanation "namely the identification of native promoter and terminator elements". [It is now clear that the accumulation of noun adjectives and the significance of the word "part" was what originally blurred the overall meaning.

      We thank the reviewer for this valuable suggestion. We agree that the original phrasing was unnecessarily clunky and have revised the sentence for clarity and flow. Lines 682–684 now read:

      “By integrating endogenous promoter and terminator discovery, DNA delivery, selection, clonal recovery, and reporter validation into a single pipeline, we provide a flexible foundation for routine transgene-based studies.”

    1. eLife Assessment

      This is a fundamental study on the sensory roles of cerebrospinal-fluid-contacting neurons (CSF-cNs) in mammals, revealing how the apical extension is used as an amplifier of chemical changes in content of the CSF. Specifically, the authors show compelling evidence that PKD2L1 is predominantly a pH-sensing channel in CSF-cNs and link its apical localization to dual phasic and sustained responses underlying CSF chemosensation.

    2. Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Comments on revised version:

      I thank the authors for their extensive revisions and detailed responses to the reviewers' comments. The manuscript has been substantially improved, and most of the major concerns raised in the initial review have been adequately addressed. In particular, the additional analyses of PKD2L1 channel activity, the incorporation of physiologically relevant pH conditions, the clarification of ASIC involvement, and the expanded Discussion have significantly strengthened the study.

      Major scientific concerns largely addressed:

      Quantification of PKD2L1 channel activity<br /> The authors appropriately addressed my previous concerns regarding the use of Po as the sole measure of channel activity. The inclusion of additional parameters such as apparent Po, open time, nmax, holding current, and membrane charge provides a more robust assessment of PKD2L1 activity and substantially strengthens the conclusions.

      Physiological relevance of pH modulation<br /> The inclusion of experiments at pH 6.5 and the additional analyses of holding current and resting membrane potential are valuable additions. These experiments considerably improve the physiological relevance of the study.

      ASIC contribution<br /> The additional pharmacological experiments using ASIC blockers are helpful and support the conclusion that the photolysis-evoked response in the apical process is predominantly mediated by PKD2L1 channels.

      Functional implications<br /> The expanded Discussion regarding Ca2+-dependent signaling, neurosecretion, and the potential physiological roles of CSFcNs considerably improves the manuscript.

      Remaining concerns:<br /> Continued overstatement regarding "exclusive" localization and function:

      Although the authors softened some statements in the revised manuscript, the term "exclusive" remains in several key locations, including the title.

      For example:<br /> "PKD2L1 channels segregated to the apical compartment are the exclusive dual-mode pH sensor..."

      The data clearly demonstrate strong enrichment of functional PKD2L1 channels in the apical process. However, the available evidence does not fully justify the term "exclusive," particularly because:

      - PKD2L1 immunoreactivity is still detectable outside the apical process.<br /> - ASIC-mediated responses are present in CSFcNs.<br /> - The authors themselves use more appropriate terminology such as "predominantly located" in the Discussion.

      Therefore, I recommend replacing "exclusive" with more conservative terminology such as:

      - predominant<br /> - predominantly localized<br /> - enriched<br /> - functionally segregated

      throughout the manuscript, including the title, Abstract, Introduction, Results, and Discussion.

      Use of the term "tonic current"

      The manuscript continues to use the term "PKD2L1 tonic current."

      While the dibucaine-sensitive holding current is clearly present, the precise mechanism generating this current remains uncertain. Indeed, the authors themselves acknowledge in the Discussion that:

      - an alternative conducting state may exist, or<br /> - unresolved brief channel openings may account for the current.

      Therefore, the data support the existence of a sustained PKD2L1-associated current, but do not yet definitively establish a distinct tonic gating mode of the channel.

      I therefore recommend replacing:

      "tonic current"

      with a more neutral expression such as:

      - sustained current<br /> - PKD2L1-associated holding current<br /> - sustained PKD2L1-mediated current throughout the manuscript.

      Continued use of "off-current" and "off-response":<br /> The revised manuscript has improved considerably in this regard. However, the terms "off-current" and "off-response" still remain in portions of the text and figure legends.

      Because the manuscript itself demonstrates that the response reflects recovery from transient acidification rather than a separate OFF signaling mechanism, these terms remain potentially misleading.

      I recommend replacing them with terminology such as:<br /> - photolysis-evoked PKD2L1 current<br /> - recovery current<br /> - proton-removal-induced current

      throughout the manuscript, including figure legends.

      Minor editorial corrections<br /> Figure 1Bd Please change: "po" to "Po" for consistency with standard channel physiology nomenclature.<br /> Figure 1Ca Please add units (mV) to the voltage labels shown on the left side of the traces.<br /> Figure 3E Please change: "Norm po" to "Norm Po".<br /> Figure 4Fb Please replace: "sec" with "s" to conform with SI unit conventions.

      The authors have addressed the majority of my previous concerns and the manuscript has been substantially improved. The remaining issues are primarily related to terminology and overinterpretation rather than experimental deficiencies.

    3. Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

      Weaknesses:

      Not weaknesses, there are only minor requests to complete the beautiful study.

      (1) The apical extension's response to removal of acidification is nicely illustrated in Figure 4C,G. There's something puzzling there: while the response to Glutamate is immediate, the channel responses to H+ is extremely delayed by 100ms - 2s, and even sometimes came in bursts separated by few hundreds of ms. H+ diffuse even faster than glutamate. Why is that?

      I don't quite understand how the response is so delayed & how to explain the recurring bursts of channel opening in the figure panel ?

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      - Could the authors use a fluorescent pH sensor to monitor pH in the extracellular space and in the cell ?

      - Could the authors investigate whether in the apical extension, PKD2L1 channels are mainly at the outer membrane in the apical extension OR whether many channels are located in inner membranes ?

      (2) Suppl Fig 4 is very cool and should be moved to main figure. The coupling of Soma and AP is very tight, yet there is a clear difference in targeting of channels that respond to cues in the CSF. In the context of an intact spinal cord, we can wonder how and when the contribution from ASIC in the some would be relevant to physiology. Can the authors think of experiments with an intact central canal to test the sensitivity and condition of recruitment of pH sensing in the soma (ASIC) versus the apical extension (PKD2L1)?

      (3) The Reissner fiber is missing after slicing the spinal cord. From our observations in fish, the fiber being under tension triggers lots of activity in CSF-cNs (Bellegarda et al Elife 2023) that also relies on PKD2L1 (Bohm et al NC 2016; Sternberg et al NC 2019). Could the authors discuss the contribution of the Reissner fiber to the PKD2L1 mediated modulation of CSFcN excitability ? Could the authors conceive a way to slice along the anteroposterior axis (sagitally) the spinal cord to keep the Reissner fiber in the central canal when recording CSF-cN apical extension ?

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Several aspects of data interpretation, however, require clarification or reanalysis-most notably the single-channel analyses (event counts, Po metrics, and mixed parameters), the statistical treatment, and the interpretation of purported "OFF currents." Additional issues include PKD2L1-TRPP3 nomenclature consistency, kinetic comparison with ASICs, and the physiological relevance of the extreme acidification paradigm. Addressing these points will substantially improve reproducibility and mechanistic depth.

      Overall, this is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing, provided that the analytical and interpretative weaknesses are satisfactorily corrected.

      (1) The authors should re-analyze electrophysiological data, focusing on macroscopic currents rather than statistically unreliable Po calculations. Remove or revise the Po analysis, which currently conflates current amplitude and open probability.

      We agree with the reviewer that the Po analysis has strong limitations, particularly in experiments where the recording times are short, like when extracellular pH is changed either by photolysis (Figure 4D) or puff applications (Figure 3Aa). In order to circumvent that problem and not to rely only on Po estimations, we used alternative methods as well, including the analysis of the current membrane charge that we have used extensively during the manuscript (Figures 3A and 4D, for example) or the analysis of the event latencies (Figure 4G). Nevertheless, single-channel recordings clearly contain information that is not included in the macroscopic current analysis. We intend to stress in the revised version that the elementary current amplitude is conserved by manipulations such as pH changes, leaving the total number of channels (N) and the channel open probability (Po) as possible culprits for the current changes. Since these changes are rapid and reversible, it is likely that N stays constant and that Po changes. In order to address the reviewer’s concern, we propose the following changes/reanalysis: i) to report in each condition the minimum N (maximum observed openings; for example, in Figure 3Aa the minimum N goes from 4 in control conditions to 1 during the puff of the pH 6.4 solution). This method (to estimate N by counting the maximum number of open channels), while imperfect, provides a tentative estimate of Po; ii) following the previous point, we propose to reword the text (and images) and use the expression “apparent Po” instead of “Po”; iii) to report the fraction of time that the channels remain open. Also, we acknowledge that some traces are confusing (Figure 3Aa, top) as they seem to show macroscopic currents. We will modify those figures by plotting the amplitude histograms (as in Figure 1Bb) in order to show unambiguously that current recordings from CSFcNs only show single-channel activities.

      (2) PKD2L1-TRPP3 nomenclature should be clarified and all figure labels, legends, and text should use consistent terminology throughout.

      We agree with the reviewer that the nomenclature concerning polycystin members is confusing. In this manuscript we have followed the nomenclature that has been proposed in a recent, comprehensive review on polycystin channels by Palomero, Larmore and DeCaen (Palomero et al. 2023), where the authors refer to the channels by their gene name. In this review, the authors indicate that PKD2L1 channels correspond to TRPP2 (formerly TRPP3, their table 1). In another, recent review on TRP channels, however, the authors refer to the PKD2L1 channel as TRPP3 (Zhang et al. 2023). In order to avoid any confusion we will remove from the text any reference to the TRPP nomenclature and stick to the PKD2L1 name.

      (3) The authors should reinterpret the so-called OFF currents as pH-dependent recovery or relaxation phenomena, not as distinct current species. Remove the term "OFF response" from the manuscript.

      We concur with the reviewer that the term “OFF response” is not very helpful from the biophysical perspective and conveys the idea that it is another current. We will remove the term “OFF response” or “OFF current” in the revised manuscript and replace it by the term “photolysis-evoked PKD2L1 current”. Also, we will condense two sections (“The proton-induced current is an off-current” and “The off-current is mediated by the activation of PKD2L1 channels”) into a single new section (“The photolysis-induced current is mediated by PKD2L1 channels”), as separating the description of the photolysis-evoked PKD2L1 current compromises its description. Finally, we will rewrite the discussion to better describe this current.

      (4) Evidence for physiological relevance should be provided, including data from milder acidification (pH 6.5-6.8) and, where appropriate, comparisons with ASIC-mediated currents to place PKD2L1 activity in context.

      This is partly addressed in Figure 3. The data there indicates that PKD2L1 channels are very sensitive to pH variations around physiological pH. In order to make this conclusion stronger, we will add to the figure the EC50 values drawn from the fittings. In terms of the ASIC-mediated currents, one of our main conclusions is that ASIC channels are not present in the ApPr, as the effects of proton photolysis in the ApPr and are not blocked by ASIC channels blockers. Our results indicate that PKD2L1 channels are the exclusive pH sensitive channels in the ApPr, while ASIC channels are probably the acid-sensitive channels in the soma, although we have not studied the latter in detail. Following this and the editor’s comments, the subsection “the involvement of ASICs” in the Discussion has been modified.

      (5) Terminology and data presentation should be unified, adopting consistent use of "predominant" (instead of "exclusive") and "sustained" (instead of "tonic"), and all statistical formats and units should be standardized.

      Following the suggestions of the reviewer, an exhaustive work will be performed to unify terminology, data presentation and correct the text following the reviewer’s suggestions.

      (6) The Discussion should be expanded to address potential Ca<sup>2+</sup> -dependent signaling mechanisms downstream of PKD2L1 activation and their possible roles in CSF flow regulation and central chemoreception.

      This is indeed a very interesting and currently unresolved point in the physiology of CSFcNs. Published data indicate that calcium flowing into the cell through PKD2L1 channels is a key regulator of apical process physiology: on the one hand, PKD2L1 channels are calcium permeable and at the same time, they are inhibited by intracellular calcium (DeCaen et al. 2016). Also, ultrastructural data indicate that the ApPr is rich in mitochondria and tubulo-vesicular structures resembling the Golgi apparatus (Bjugn et al. 1988; Bruni et Reddy 1987), which are intracellular organelles that contribute to calcium homeostasis. Altogether, this evidence suggests that intraApPr calcium concentration needs to be finely regulated, both in space and in time, in order for the ApPr to fulfil its physiological roles. Based on what has been published in the literature, we can speculate that these calcium signals can be decoded by several systems: i) calcium can act as the second messenger linking the activation of the multimodal PKD2L1 channels to changes in CSFcNs excitability, which in turn regulate spinal neuronal networks controlling locomotor activity; ii) calcium could initiate neurosecretion of different molecules from the ApPr to the central canal (as has been proposed by the Wyart group in the zebrafish in the context of bacterial infections(Prendergast et al. 2023)); iii) calcium could activate the Hedgehog signaling pathways (as has been shown by (Delling et al. 2013)); iv) calcium could modulate CSF flow directly (by modulating ciliary activity) or indirectly (by modulating ependymal cells ciliary activity through paracrine interactions). Resolving these downstream pathways is essential to fully define the role of CSFcNs as integrators of CSF homeostasis. We will expand this issue in the Discussion of the revised ms.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments relate to their multimodal CSF-sensing function remains unclear.

      The authors confirm that in GATA3:GFP mice, where these cells are labeled, that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show that functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding of how these sensory cells can detect the changes of pH in the CSF so finely.

      Weaknesses:

      The following do not constitute weaknesses; rather, they are minor requests that this reviewer considers would complete this beautiful study.

      (1) It would be nice to quantify further the relation in spontaneous as well as in acidic or basic pH between the effects observed on channel opening and holding current: do they always vary together and in a linear way?

      Following the reviewer’s suggestion, we have performed a Spearman’s rank correlation test, which shows a significant correlation between the changes in the apparent open probability and holding current (paired experiments; ctrl vs pH 6.4 pressure applications; p < 0.05, Spearman r = 0.72 and critical value = 0.67). The Pearson correlation coefficient calculated on the same data set = 0.63 and the critical value is 0.632, which indicates that the correlation is not linear. We will add this analysis to the manuscript.

      (2) Since CSF-cNs also respond to changes in osmolarity (Orts Dell Immagine 2013) & mechanosensory stimulations in a PKD2L1 dependent manner (Sternberg NC 2018), it would be nice to test the same results whether the same results hold true on the role of PKD2L1 in AP for pressure application of changes in osmolarity.

      This is a very important point. As the reviewer mentions, previously published experimental evidence indicates that CSFcNs are also sensitive to osmolarity changes and mechanical stimulation in a PKD2L1-dependent manner. It is therefore reasonable to assume that, as for the pH sensitivity, osmotic and mechanical sensitivity depends on channels segregated to the ApPr. For the mechanosensitivity, the spatial segregation could be tested by “touching” either the ApPr or the soma with a piezo-controlled blunted pipette (see, for example, Hao et al. 2013). However, the sensitivity to osmotic changes is much more difficult to assess, as pressure application does not have enough spatial resolution to discriminate among compartments in such a compact cell such as the CSFcNs. In theory, a highly spatially localized osmotic jump could be reached with photolysis, but a caged compound releasing many osmotic particles simultaneously should be used. In typical photolysis experiments, a localized osmotic jump is produced, but it is very low (in the order of 1 to 2 mOsm).

      In mice, like in fish (Sternberg et al, NC 2018), we can observe throughout the figures that a large fraction of the channel activity occurs with partial and very fast openings of the PKD2L1 channel. I recommend the authors analyse the points below:

      (a) To what extent do these partial openings of the channel contribute to the changes in holding current and resting potential?

      As the reviewer indicates, these partial and very fast openings are a characteristic of PKD2L1 single-channel activity that seems to be present in different species. However, estimating what is the exact contribution of these events to the sustained current would require a detailed model of the channel that it is still lacking. Indeed, the exact mechanism by which CSFcNs show this prominent sustained current is unknown and should definitely being addressed in future works.

      (b) In the trace from the outside out AP, it looks like the partial transient openings are gone. Can the authors verify whether these partial openings are only present in somatic recordings?

      The outside-out recordings from the ApPr also show some partial openings (please look at the upper trace in Figure 4Db). We will specifically mention this important point in the revised version of the ms.

      (3) Previous studies have observed expression of metabotropic Glutamate receptors in CSF-cNs (transcriptome from Prendergast et al CB 2023). The authors only used blockers for ionotropic glutamate receptors in their recordings: could it be that these metabotropic receptors influence the response to uncaging of MNI-Glu when glutamate is co-released with a proton?

      We thank the reviewer for pointing out the presence of metabotropic glutamate receptors in CSFcNs. However, our evidence indicates that there is no contribution of metabotropic receptors when uncaging MNIglutamate because: i) the response obtained when uncaging MNI-gLGG (where there is no glutamate release; Figure 5Ab) and ii) the response obtained when uncaging protons from DPNIGABA (a GABA cage that has similar photochemistry than MNI cages which also release a proton when photolysed; data not shown), are the same. Indeed, in both experiments (MNI-gLGG or DPNI-GABA uncaging) a clear photolysis-evoked PKD2L1 current can be observed.

      (4) In the outside out patch of the AP, PKD2L1 unitary currents appear rare. Could it be that the disruption in the cilium or underlying actin/myosin cytoskeleton drastically alter the open probability of the channel?

      Although we have not quantified it, the reviewer is right that the opening frequency of PKD2L1 channels in the outside-out patches is lower than in the whole-ApPr recordings. We interpreted this difference as a difference in channel number. However, another plausible interpretation is that, as the reviewer suggests, the biophysics of the channels are affected because the protein is taken out from its normal ionic environment and/or loses important interactions with regulatory proteins.

      (5) Could the authors use drugs against ASIC to specify which ASIC channels contribute to the pH response in the soma?

      As described in the manuscript, we did perform experiments with ASIC channel blockers, although we did not attempt to characterize the specific ASIC channel involved in the somatic response. Based on what has been published in the literature, we used both psalmotoxin-1 (which blocks ASIC1 channels) and APETx2 (which blocks ASIC3 channels). The presence of ASIC1 channels in mice CSFcNs has been shown by (Orts-Del’Immagine et al. 2012; Orts-Del’Immagine et al. 2016), while the presence of ASIC3 in the lamprey CSFcNs has been shown by (Jalalvand et al. 2016). When we puff an acidic solution aiming at the soma, we can record an inward current that is blocked by psalmotoxin-1, although there is always a small component remaining (as originally shown by Orts-Del’Immagine in the aforementioned articles); however, we have not attempted to block this small component that remains after psalmotoxin-1 bath application.

      (6) This is out of the scope of this study, but we did observe in fish a very rarely-opening channel in the PKD2L1KO mutant. I wonder if the authors have similar observations in the conditions where PKD2L1 is mainly in the closed state.

      We have never seen such kind of openings in our recordings (when the channel is closed or in the presence of dibucaine).

      Bjugn, R, H K Haugland, et P R Flood. 1988. “Ultrastructure of the mouse spinal cord ependyma.” Journal of Anatomy 160 (octobre): 117‑25.

      Bruni, J. E., et K. Reddy. 1987. “Ependyma of the Central Canal of the Rat Spinal Cord: A Light and Transmission Electron Microscopic Study”. Journal of Anatomy 152 (juin): 55‑70.

      DeCaen, Paul G., Xiaowen Liu, Sunday Abiria, et David E. Clapham. 2016. “Atypical Calcium Regulation of the PKD2-L1 Polycystin Ion Channel”. eLife 5 (juin): e13413. https://doi.org/10.7554/eLife.13413.

      Delling, Markus, Paul G. DeCaen, Julia F. Doerner, Sebastien Febvay, et David E. Clapham. 2013. “Primary cilia are specialized calcium signalling organelles”. Nature 504 (7479): 311‑14. https://doi.org/10.1038/nature12833.

      Hao, Jizhe, Jérôme Ruel, Bertrand Coste, Yann Roudaut, Marcel Crest, et Patrick Delmas. 2013. “Piezo-Electrically Driven Mechanical Stimulation of Sensory Neurons”. In Ion Channels, édité par Nikita Gamper, vol. 998. Methods in Molecular Biology. Humana Press. https://doi.org/10.1007/978-1-62703-351-0_12.

      Jalalvand, Elham, Brita Robertson, Peter Wallén, et Sten Grillner. 2016. “Ciliated Neurons Lining the Central Canal Sense Both Fluid Movement and pH through ASIC3”. Nature Communications 7 (janvier): 10002. https://doi.org/10.1038/ncomms10002.

      Orts-Del’Immagine, Adeline, Riad Seddik, Fabien Tell, et al. 2016. “A Single Polycystic Kidney Disease 2-like 1 Channel Opening Acts as a Spike Generator in Cerebrospinal Fluid Contacting Neurons of Adult Mouse Brainstem”. Neuropharmacology 101 (février): 549‑65. https://doi.org/10.1016/j.neuropharm.2015.07.030.

      Orts-Del’immagine, Adeline, Nicolas Wanaverbecq, Catherine Tardivel, Vanessa Tillement, Michel Dallaporta, et Jérôme Trouslard. 2012. “Properties of Subependymal Cerebrospinal Fluid Contacting Neurones in the Dorsal Vagal Complex of the Mouse Brainstem”. The Journal of Physiology 590 (16): 3719‑41. https://doi.org/10.1113/jphysiol.2012.227959.

      Palomero, Orhi Esarte, Megan Larmore, et Paul G. DeCaen. 2023. “Polycystin Channel Complexes”. Annual Review of Physiology 85 (Volume 85, 2023): 425‑48. https://doi.org/10.1146/annurev-physiol-031522-084334.

      Prendergast, Andrew E., Kin Ki Jim, Hugo Marnas, et al. 2023. “CSF-Contacting Neurons Respond to Streptococcus Pneumoniae and Promote Host Survival during Central Nervous System Infection”. Current Biology 33 (5): 940-956.e10. https://doi.org/10.1016/j.cub.2023.01.039.

      Zhang, Miao, Yueming Ma, Xianglu Ye, Ning Zhang, Lei Pan, et Bing Wang. 2023. “TRP (Transient Receptor Potential) Ion Channel Family: Structures, Biological Functions and Therapeutic Interventions for Diseases”. Signal Transduction and Targeted Therapy 8 (1): 261. https://doi.org/10.1038/s41392-023-01464-x.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers were very impressed with your work and definitely feel it is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing. They, however, request some re-analysis of the data and discussion with minimum new experiments, if any. I think, if feasible, this will improve the quality of the study and would look forward to receiving a revised version.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1 - Molecular identity and localization of PKD2L1

      Major

      Nomenclature clarity.

      Please clarify the distinction between PKD2L1 and TRPP3. Several parts of the text and figure labels appear to conflate these names. PKD2L1 corresponds to the TRPP3 subfamily member and should not be interchanged with PKD2/TRPP2. Please confirm and update the nomenclature consistently throughout the manuscript (text, figure labels, and captions).

      We have addressed this issue in response to reviewer 1. There seems to be some confusion in the literature concerning the nomenclature of PKD2L1 channels as in some recent publications the PKD2L1 channels are still named as TRPP3. However, the nomenclature of PKD2L1 channels or TRPP2, was updated in 2016 (Wu, Sweet and Clapham, Pharmacological Reviews, 2010). As indicated in response to reviewer 1, we have removed from the text any reference to the TRPP nomenclature and stuck to the PKD2L1 name.

      Physiological meaning of apical restriction.

      Expand the discussion of why apical-restricted localization matters. Specifically, address how segregation to the ApPr could support directional sensing of CSF flow and/or detection of localized pH gradients.

      Following the reviewers and editor’s comments, we have revised the discussion in order to take this and other comments into account.

      Open-state annotation (O3).

      Please include the O3 state in panels Ba and Bb; the figure clearly shows an additional open level consistent with O3.

      The editor is right in that there is another state that presumably corresponds to O3. Following the editor’s recommendation, we now indicate this 3rd level and add a short sentence explaining this in figure 1 legend.

      Minor

      Indicate the ROI definition and background-subtraction method used for fluorescence quantification (ApPr vs soma).

      Not applicable for this figure.

      Ensure the same intensity scale (lookup table and range) is used across panels to enable direct comparison.

      Done.

      (2) Figure 2 - Electrophysiological characterization of ApPr and somatic recordings Major

      Definition of "PKD2L1-dependent current."

      Define this term precisely at its first appearance. Specify whether it denotes currents inhibited by dibucaine, abolished in PKD2L1-knockout preparations, or both.

      Done

      Statistical power of single-channel analysis.

      The number of observed openings (< 1000 events) is too low to estimate open probability (Po) reliably. Please re-analyze the data using macroscopic current traces rather than Po-based kinetics.

      Confounded Po analysis.

      The current Po analysis mixes current amplitude and Po in the same calculation, conflating independent variables. Re-evaluate or remove this analysis.

      Unknown channel count.

      Because the number of channels in each patch is unknown, Po and "closed probability" values cannot be interpreted meaningfully. Focus instead on the averaged macroscopic current density.

      General analytical validity.

      The single-channel analyses in Figure 2 are not interpretable under these experimental conditions. Closed-time distributions and Po-based metrics (e.g., "Po1," "P2") depend critically on channel number and event sampling. Moreover, the manuscript applies essentially the same Po methodology across conditions (Po1 vs P2), which adds no mechanistic resolution and risks circular interpretation.

      Actionable recommendation:

      Remove Po- and closed-time-based analyses from Figure 2 and from the manuscript as a whole. Reanalyze the data using metrics that remain valid when the channel number is unknown:

      Macroscopic current analysis (leak-subtracted current density, I-V relationships, activation time constants).

      Single-channel conductance only (amplitude histograms and unitary slope conductance), without attempting Po or dwell-time inference.

      Report filtering bandwidth and sampling rate, and restrict statistical treatment to these robust parameters.

      Following the reviewers (see above) and editors’ recommendations, we have reanalyzed the data in order to avoid the analysis based on Po and Pc. We have instead calculated from the recordings other 2 parameters, n<sub>max</sub> (the maximum number of channels that open simultaneously during a 500 ms time window) and the total open time of a single channel during the same 500 ms time window. The main text, figures and corresponding figure legends, and the Materials and Methods section have been changed accordingly. Notably, the Po and Pc analysis were removed from Figures 3, 4 and Supplementary Figure 3, and replaced by the above-mentioned parameters. Also, the fact that the recordings are not long enough to calculate Po is now specifically mentioned in the Materials and Methods section, lines 785 to 790. In addition, the analysis in Figure 3Ce has been redone so that the activity of the channel as a function of pH is now plotted as the normalized apparent Po (relative to the apparent Po value at pH 7.4).

      Minor

      State whether input-resistance values (1.8-4.4 GΩ) were leak-subtracted and series-resistance-compensated.

      As already mentioned in the methodology section (line 693), series resistance was not compensated for during the experiments. We have now added a sentence in the methodology section indicating that in the voltage range that was chosen for the analysis of the input resistance, no voltage-dependent conductance was activated (lines 809 to 812).

      Ensure unit consistency: use Po or normalized Po rather than frequency (Hz) throughout.

      (3) Figure 3 - pH-evoked currents and kinetics

      Major

      Invalid Po analysis (Fig. 3Ca-Ce).

      The Po- and Pc-based single-channel analyses in panels 3Ca-3Ce should be deleted. As noted earlier, the event count is insufficient, and the number of active channels in each patch is unknown. Under these conditions, Po and Pc values have no quantitative meaning and could mislead readers. These panels do not contribute additional mechanistic insight beyond the macroscopic current data and therefore, should be removed. If retained for illustrative purposes, they must be explicitly labeled as representative traces without any statistical quantification.

      As we mentioned above, we removed the Po and Pc analysis from the manuscript.

      Minor

      Present regression equations and r<sup>2</sup> values for the linear fits shown in Fig. 3D.

      Done (lines 947951).

      Confirm that all axes include units and identical scaling between conditions for direct comparison.

      Done.

      (4) Figure 4 - Laser photolysis and local stimulation experiments

      Major

      Laser timing annotation.

      Clearly mark laser-pulse timing (e.g., arrow or shaded region) on all current traces to facilitate interpretation.

      We thank the editor for pointing out the inconsistencies in terms of the laser pulse timing. To indicate the laser pulses, we have now added an arrowhead in cases where a single sweep is shown (for example, Figure 3D), and an arrowhead and a dotted magenta vertical line in cases where multiple sweeps are shown (for example, Figure 3C).

      pH calibration within the laser spot.

      Provide quantitative calibration of pH changes induced by laser photolysis, including information on spot size, local diffusion, and estimated pH recovery kinetics.

      This is an important point and we thank the editor for mentioning it. We have now completed the subsection untitled “photolysis” where we provide information on the lateral and axial dimensions of the photolysis laser spot used in this work (lines 734 to 737). We have also rewritten part of Figure 5A legend to highlight the fact that the experiments presented there (photolysis on top of the ApPr and next to it) are compatible with a high spatial resolution of proton release (lines 1023 to 1026).

      On the other hand, we have attempted to perform pH calibrations in the setup using the pH-sensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2, and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, when trying to perform a calibration under the conditions of a real experiment,

      where photolysis occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml), we encountered the following problem, which made it impossible to obtain any useful data: the 405 nm uncaging pulse bleaches the dye, and any useful information is lost. Also, our imaging system is not fast enough to follow the pH change. As it is discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 814), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond time range.

      Repeated stimulation effects.

      Discuss whether repeated photolysis induces adaptation or desensitization of PKD2L1 currents, and indicate whether current amplitude decreases across successive trials.

      This issue is now specifically mentioned in the Materials and Methods section, lines 739 to 741.

      Invalid interpretation of the "OFF response."

      The interpretation of the so-called "OFF response" in Figure 4C is not supported by the presented data. There is no evidence for a bona fide OFF current, and the literature cited does not demonstrate such a phenomenon for PKD2L1 alone. Rather, previous studies implicate PKD1L3-dependent mechanisms in similar biphasic responses. Please reconsider the cited references and remove claims of an OFF current attributed to PKD2L1.

      Done.

      Actionable recommendations:

      Do not use the term "OFF response" throughout the manuscript. Recast these transients as pH dependent recovery or relaxation of current following cessation of acidification.

      Done. We have performed extensive rewriting and reorganization of the Results and Discussion in order to take into account both the reviewer’s and editor’s comments. Please also take a look at comment #3 of Reviewer 1 and point 11 below.

      Include continuous-illumination controls (sustained local acidification) to test whether a steady state current is maintained. This will clarify whether the post-stimulus transient reflects recovery kinetics rather than a distinct current species.

      We thank the editor for suggesting this experiment. However, continuous laser illumination is a difficult manipulation and does not necessarily lead to an acidification of the illuminated volume. Indeed, with continuous illumination the cage is lost from the illumination spot and needs to be replaced by diffusion from the non-illuminated volume, leading to non-homogeneous concentrations. Also, the chances of inducing photo damage are higher. We thus designed a similar experiment where instead of performing continuous illumination we photolysed with short and high frequency trains in order to produce a long-lasting acidification. The results of these experiments have been added to the manuscript as part of the results section and in Figure 5H. Similarly to what is seen with single illuminations, the photolysis trains induce a current that appears almost exclusively at the end of the train, implying that the current is indeed a PKD2L1-dependent recovery current.

      Align the current time course with measured or estimated local pH (or calibrated proxy) to demonstrate causal coupling and avoid implying a separate conductance.

      We have added the calculated pH change to the inset of Figure 4C as an example.

      Revise the schematic/model figure and textual description accordingly, restricting the framework to phasic vs sustained activation modes without invoking a separate OFF current for PKD2L1.

      Done.

      Minor

      Include scale bars, sample numbers (n), and laser parameters (duration, power) in all panels.

      In order not to make the figure and the panels very heavy in the original version, we tried to limit the number of scale bars. We have now performed some modifications, added the missing scale bars, and changed the figure legend in order to take into account the editor’s comments. We have also corrected a few values that were wrongly reported.

      Standardize p-value formatting (e.g., p = 6 × 10 ⁶) throughout the figure and legend.

      Done.

      (5) Figure 5 - Single-channel recordings

      Major

      Mixed parameters (current amplitude and Po).

      The current analysis improperly mixes single-channel current amplitude and Po within the same figure, conflating distinct parameters. These quantities must be analyzed and presented separately, or the Po data should be removed entirely if not independently supported.

      Insufficient event count.

      Given the very limited number of observed openings, Po-based statistics are not meaningful. Please report only representative single-channel traces and corresponding amplitude histograms without attempting quantitative Po estimation.

      Minor

      Convert frequency (Hz) values to Po for consistency with earlier analyses, or remove frequency metrics altogether if Po analysis is omitted.

      Figure 5 does not include Po or event frequency analysis, so we think there must be a misquotation of the figure. However, the Po issue has already been addressed before and alternative analysis have been proposed.

      (6) Introduction

      The introductory paragraph mentions the "five senses" as a framing concept. However, this statement lacks scientific grounding in the context of CSF-contacting neurons and chemosensory physiology. The traditional "five senses" classification is not an evidence-based neurophysiological framework and may be misleading to readers. I recommend removing or rephrasing this part, focusing instead on molecular and cellular mechanisms of sensory transduction (e.g., chemical, mechanical, and pH sensing) rather than on classical sensory categories.

      Following the editor’s recommendation, we have removed this part.

      The manuscript refers to PKD2L1 using the term TRPP2 in some parts of the introduction. This is incorrect, as PKD2L1 corresponds to TRPP3, not TRPP2. Please correct this nomenclature and ensure consistent use of "PKD2L1 (TRPP3)" throughout the entire manuscript to avoid confusion with the distinct PKD2/TRPP2 protein, which belongs to a different subfamily with separate physiological roles.

      We thank the editor for pointing this out. As we mentioned in the responses to the “public reviews”, the literature is confusing, so we decided to remove from the manuscript any mention to TRPP channels.

      (7) Discussion

      The current Discussion reads largely as a descriptive summary of results and lacks conceptual depth. It does not effectively integrate the biophysical properties of PKD2L1 with its physiological role as a neuronal pH sensor, nor does it develop a broader interpretation relevant to CSF homeostasis or chemoreception.

      Following the reviewers and editor’s recommendations, we have now added a new section in the Discussion untitled “PKD2L1 downstream signaling mechanisms”.

      (8) Insufficient biophysical analysis

      The discussion of channel gating and pH dependence is superficial and does not explore the energetic or structural mechanisms underlying proton sensitivity. The authors should analyze their data in the context of known PKD/TRPP family biophysics-for example, protonation sites, subunit composition, or gating kinetics-and explain how these confer bidirectional (acidic vs alkaline) sensitivity within physiological ranges.

      In this work, we studied the pH sensitivity of PKD2L1 channels in the context of CSFcN sensory physiology. From a pure biophysical perspective, the pH sensitivity of PKD2L1 channels has been studied by multiple groups; however, it is still unknown how the gating of the channel responds to pH changes, although it can be proposed that some polar residues in the protein regulate the state of the pore. Likewise, the mechanism of the “off-response” is also unknown. To the best of our knowledge, there is only one article in which the authors have attempted to relate pH, PKD2L1 channel structure, and function. In this work (Su et al., Nature Communications 2018), the authors compare PKD2L1 channels with another pH-sensitive member of the TRP family, TRPML3, whose structures at pH 7.4 and 4.8 are known (Zhou et al., Nature Structural and Molecular Biology, 2017). We have rewritten some sentences of the Discussion in order to be more specific about the pH dependence of PKD2L1 channels and its proposed mechanisms.

      (9) Weak physiological context

      The manuscript does not adequately address how PKD2L1 functions as a true physiological pH sensor. The discussion should connect channel activity to realistic CSF pH fluctuations (6.8-7.6) and to relevant physiological or pathophysiological conditions (e.g., respiratory acidosis, neurogenic regulation of CSF composition). Without this, the relevance of large, artificial acidification (pH 3-3.5) remains unclear.

      We have added a new section in the Discussion where we speculate on how PKD2L1 channels may be activated in physiological and pathophysiological conditions. However, we would like to insist here that the main goal of the photolysis experiments (which induce short and large acidifications) was to assess the spatial segregation of PKD2L1 channels. We now mention this point specifically and also speculate on the conditions that could eventually give rise to the “recovery” current.

      (10) Over-interpretation of unsupported points

      The paragraph describing voltage propagation from the ApPr to the soma/axon is speculative and unsupported by any data in the manuscript. Please delete this section entirely, including the citation to Orts-Del'immagine et al., unless new electrophysiological evidence is added.

      We think this point (the propagation of signals originating from the ApPr to the soma) is important in the context of our work, so we have decided to make new experiments in order measure directly the degree of coupling between the 2 compartments. To do that we made simultaneous, current-clamp and voltage-clamp recordings from the ApPr and the soma. In these conditions we were able to measure experimentally and for the first time both the coupling coefficient and coupling conductance, which confirm that the propagation of voltage signals from the ApPr to the soma is extremely efficient. These new results are now described in a new subsection and in a new Figure 6.

      (11) Clarify the role of "OFF currents."

      The Discussion repeatedly refers to an "OFF response," but this phenomenon is not experimentally demonstrated for PKD2L1 alone. It likely represents pH-dependent recovery rather than an independent current. All discussion of "OFF currents" should be removed or reformulated accordingly.

      Following the editor and reviewer’s comments, we have deleted the term “off response” and “off currents” from the ms and have replaced them with the term “recovery current”. We have also changed the discussion accordingly.

      (12) Integration with ASICs and compartmental sensing

      While the manuscript briefly mentions ASIC involvement, it does not articulate how PKD2L1- and ASIC-mediated signals might complement each other in different compartments (ApPr vs soma). The authors should discuss the potential division of labor between these sensors and how such compartmentalization enhances pH detection in CSFcNs.

      Following the editor’s comments, we have rewritten the part of the subsection ‘the involvement of ASICs’ in the Discussion.

      (12) Broadened physiological perspective

      The Discussion should close by considering Ca<sup>2+</sup> -dependent downstream pathways activated by ⁺ PKD2L1 and their implications for CSF flow regulation, neurosecretion, and central chemoreception. These translational aspects would substantially improve the impact and readability of the manuscript.

      Done

      Overall, the Discussion must evolve from a descriptive narrative to a mechanistically and physiologically integrative synthesis, highlighting why PKD2L1 is not merely present in the ApPr but is a key molecular transducer linking ionic microenvironment to neuronal excitability.

      As it has been detailed above, we have performed several changes in the Discussion that follow the reviewer’s and editor’s recommendations.

    1. eLife Assessment

      This valuable work introduces HSSM, a flexible and modular Python toolbox for hierarchical Bayesian estimation of sequential sampling models. By combining model simulation, neural likelihood approximation, established inference backends, formula-based regression, and support for trial-level covariates, HSSM could become a widely used resource in cognitive and computational neuroscience. Evidence for the framework's scope and core functionality is convincing, but support for its usability, computational performance, and reliability is incomplete. The contribution would be substantially strengthened by an executable end-to-end analysis, clearer exposition for non-specialists, parameter and model recovery demonstrations, guidance on diagnostics and model comparison, quantitative benchmarks, and a clearer delineation of HSSM's scope relative to broader claims about neurocognitive modelling.

    2. Reviewer #1 (Public review):

      Summary:

      This article describes a new software package, HSSM, for simulating and fitting sequential sampling models. The package consists of three modules - one for simulating the models, one to train neural networks on mappings from behavior to parameter values, and one for combining these tools to fit particular models to users' data.

      Strengths:

      This is a very detailed description of a new package that is building on an already highly successful package. It promises to become a go-to software package for cognitive modelers, experimentalists, and practitioners.

      Weaknesses:

      I have only a few critiques of the article, and some are a matter of taste:

      (1) I think it would be helpful for the authors to take the reader through one complete example at the end of the article, including loading in a dataset, fitting it with a regression model, checking model convergence, reading out parameter values, and doing posterior predictive checks (etc). It would be helpful to see it all in one place to get a sense of how much code is required to go through the whole process. One or two actual examples would help readers who are not already familiar with HDDM.

      (2) The article assumes a certain level of computing and modeling expertise, which somewhat limits its reach. There are many abbreviations and references to other software tools that a reader might not be familiar with. Such readers might feel like HSSM is beyond their reach. Below is a non-exhaustive list of such undefined or unexplained terms:<br /> DDM, API, LBA, fMRI, EEG, JAX, PyTorch, ONNX, MCMC, VI, MAP, PyMC, PyTensor, NUTS, ArviZ, WAIC, LOO, KDE, CLI, YAML, GUI, LBA, RDM, QP.

      (3) I found some of the figures/listings to be unpolished, unhelpful, and/or unnecessary. For instance, Figure 3 and Listing 3 seem to just be zoomed-in pieces of Figure 2. In Figure 4, what is the meaning of the little globe traveling between the within-trial and across-trial rows? Why is Figure 5 a figure and not a listing? Does HSSM not produce these plots directly? I have the same question for Figure 7. Figures 8 and 9 seem unnecessary to me, but perhaps they serve a function that I'm missing. Perhaps they could be turned into supplements for Figure 2.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript introduces HSSM, a Python-based toolbox for fitting cognitive models, with a specific focus on sequential sampling models. The toolbox brings together several components: a model construction interface, surrogate likelihoods, sampling tools, an inference backend, formula-based regressions, and tools for validation and visualization.

      One of the key advantages is that HSSM relies on well-established open-source packages. This ensures both a robust foundation and also opens a potential for future (community-driven) development. While HSSM is not the first publicly available toolbox for fitting sequential sampling models, it introduces several novel features that will be very valuable to researchers in the field.

      Strengths:

      The biggest strength of HSSM is its flexibility and ease of use for hierarchical modeling. Many existing toolboxes work as closed systems that are hard to modify. In contrast, HSSM's modular design allows it to be used either as a stand-alone tool or to pick out specific components to integrate into existing pipelines. In addition, the toolbox combines simulation-based inference, surrogate likelihoods, and formula-based regression. This opens up a lot of new modeling possibilities and makes it easy to incorporate trial-by-trial neural or physiological covariates alongside standard RT and choice data.

      Weaknesses:

      The paper provides a high-level overview of the toolbox, rather than a didactic walk-through that shows how to use it in practice. Additionally, despite being framed as a broad toolbox for "neurocognitive modeling", HSSM currently focuses on sequential sampling models. While these models are widely used, they represent only a small slice of neurocognitive modeling as a whole. Additionally, the toolbox is currently in beta phase, and lots of planned extensions are not implemented yet. Finally, there is currently no information on general performance benchmarks.

    1. eLife Assessment

      This study provides a useful single-cell atlas of the Clytia hemisphaerica planula, complemented by an updated medusa dataset, ultrastructural analyses and in situ validation of expression patterns. The cross-stage comparison and cluster-similarity framework offer a promising basis for investigating cellular diversification across the life cycle. The evidence is currently incomplete due to documentation deficits and insufficient cross-referencing with prior work. It should be relatively easy to address these issues, which should make this a valued resource for cnidarian researchers and for colleagues studying cell-type evolution.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript further explores the single-cell atlas of Clytia hemisphaerica by incorporating the planula larva. It compares the cell clusters with the previously established atlas of the medusa. It identifies similarities and differences between the two life stages.

      Strengths:

      The manuscript provides an important set of single-cell data that have not been assessed previously: the Clytia planula. The data is further supplemented with high-quality EM-based histology and an extensive in situ hybridisation of selected genes.

      Weaknesses:

      The detailed analysis does not go deep into the comparison between stages, nor does it provide an analysis of genes within the clusters; it could be described as remaining overall rather superficial.

    3. Reviewer #2 (Public review):

      Summary:

      The generation of alternate stages in the life cycle of a single species requires vast remodeling of the cellular complement of the individual during metamorphosis from one stage to another. In this paper, the authors provide a detailed description of single-cell RNA-seq data derived from the planula stage of the hydrozoan model Clytia hemispherica and compare this to an expanded dataset from the medusa stage to assess changes in transcriptomic identity of cell types between these two phases of the life cycle. The paper further includes valuable TEM data illustrating fine anatomy of cell types present at both stages investigated, and documentation of the retention of epithelial polarity from the planula through to the polyp stage, using a reporter line.

      Strengths:

      The study provides a solid and convincing transcriptomic characterization of planula cell types (including in situ validations and a planula-to-polyp mapping of epithelial polarity), and introduces a potentially valuable method for evaluating cluster similarity.

      Weaknesses:

      The work suffers from insufficient documentation of methodological approaches and missing code, lack of clarity regarding clustering resolution and nomenclature (thereby hindering cross-referencing with prior papers), and unclear plans for public, fully annotated data release.

      Full Review:

      The single-cell transcriptomic data analyzed include both previously published and newly generated data: two additional medusa libraries and two additional planula libraries were generated and integrated with the data from https://doi.org/10.1126/sciadv.abh1683, and https://doi.org/10.1126/sciadv.adv1159. The original release of the planula dataset in their 2025 Science Advances paper did not include analyses of all cell types. Here the authors provide this analysis for the planula stage. However, as both the number of clusters and the nomenclature of the clusters changed, this leads to some confusion and inability to cross-reference the two papers. There is no explanation given for the re-processing of the planula dataset in the current paper, and the fact that only some of the data is new is buried in the supplement, which is not referenced in the main document, while the text within the main article suggests that the entire dataset is new. The fact that the dataset in the current analyses contains fewer cells than presented in their Science Advances paper further adds to this confusion. The current paper would benefit from greater transparency in the origin of the data analyzed.

      The authors do try to apply the same nomenclature for the updated medusa dataset that is present in their 2021 Science paper. For example, the previously identified 'bioluminescent cells' are identified as 'gas-m8'. A look-up table that has all of the cluster id's cross-referenced would be useful (i.e. new: 8 = gas-m8 = previous: 28 = BC = "Tentacle GFP cells"). The inability to easily cross-compare with the published data is a major weakness of the current work and would benefit greatly from consistency between the three papers. Indeed, the clustering resolution is quite different across all three papers, and the current work does not adequately address how the clustering resolution was selected here. As an updated atlas, one would expect the entire transcriptomic diversity to be included here, so that the previous work can be transferred to the updated genomic mapping resource used in the current work. Nonetheless, presenting a unified nomenclature for moving forward would benefit the community as a whole and would increase the impact of the current work substantially.

      The paper also includes new TEM data of the planula cell types. The authors attempt to correlate transcriptomic profiles with these anatomical data through in situ hybridizations that provide spatial distribution of the profiles. While the TEM data are valuable to catalog the presence of cells with different morphologies within the planula, the association with the transcriptomic profiles is somewhat speculative. These valuable anatomical data should be provided at a high enough resolution to zoom in and see the details, and further description could be provided. For example, the paper states that vacuolated cells are characteristic of the basal gastrodermal cells adjacent to the mesoglea; please identify the vacuoles in Figure 4f/h for the reader.

      A novel method for reconstructing cluster similarity relationships is applied to grouping clusters into cell categories within the same life cycle stage, and also for matching cell types between stages. This is a valuable contribution to the field that is worthy of further evaluation. This is, however, difficult, as the methods for which DESeq2 was applied ("see code for details") are not present in the provided code, nor is it adequately described how the "binary matrix of marker gene presence/absence" was constructed. Similarly, there are additional details of other parts of the data analysis that are missing from the provided code, and the provided supplementary material is not referenced in the main document. More rigorous documentation of the methods is warranted.

      The description of the transcriptomic profiles present in the planula is solid, and the attempt to associate these profiles with anatomic locations and putative morphology provides a foundation onto which further studies can be developed. Mapping of the planula ectoderm through to the polyp stage is also an important step forward in characterizing the life cycle, and the evidence for the retention of the oral/aboral ectodermal axis is convincing. The paper falls short in describing the updated medusa dataset and could benefit from a minor restructuring of the paper. Introducing the new medusa data only after the planula dataset is fully described would mediate the shallower treatment of the updated medusa dataset, where only 22 of the original 36 transcriptomic states are recovered. In this way, the focus will shift onto the cross-life cycle stage comparisons, and it could be argued that the lower resolution of the medusa dataset is justified in order to simplify the comparisons.

      It will be essential that the datasets that are presented in this work be made available for public exploration in a fully annotated format. It is currently unclear how the authors intend to do this; however, there are many repositories available for this. The UCSC Cell Browser hosted at cells.ucsc.edu is one very good option if the authors do not wish to develop an interactive tool themselves. It is imperative that the gene annotations which correspond to the dataset, and the cluster annotations that are presented in this paper, are available and easily connected to the released dataset.

    1. eLife Assessment

      This important study presents an ultrastructural atlas of extracellular vesicles and non-vesicular particles in Drosophila olfactory sensilla, using cryofixation-based serial block-face electron microscopy to catalogue roughly 7,800 particles across four antennal regions. The evidence is compelling: the cryofixation minimizes fixation artifacts, and the rigorous, well-validated methodology and morphometric analyses provide strong support for the structural classifications and spatial distributions, though some conclusions require moderation or additional evidence. The work provides a potentially foundational resource for researchers studying extracellular-particle signaling in sensory organs.

    2. Reviewer #1 (Public review):

      Summary:

      Using cryofixation and serial block-face electron microscopy (SBEM), P. Vijayakumar and K. Cauwenberghs characterize extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) within native Drosophila olfactory sensilla. The study provides a unique and valuable dataset comprising approximately 7,800 extracellular particles, systematically describing their morphology, size, density, and distribution. The ultrastructural analysis across different sensillum classes offers insights into the potential biogenesis and functions of these extracellular particles.

      Strengths:

      Cryofixation preserves EVs and NVEPs within native tissue conditions. The detailed quantification of a very large dataset provides a unique source of information on extracellular particle number, categories, and distribution. The expertise of the group in the method and the tissue explored, as well as their detailed quantification, provide confidence in the dataset and observations.

      Weaknesses:

      Major comments

      (1) As the authors state, the rare observation of EV budding or MVB release events suggests that these are transient processes, whereas EVs and NVEPs are retained for relatively long periods within the sensillum lumen. The current analyses may overinterpret steady-state vesicle abundance as differences in vesicle production.

      (2) Given the above conclusion, differences between sensillum classes may be somewhat overstated.<br /> a) The absolute number of EVs and NVEPs per sensillum is highly variable, even within the same sensillum class (Figure 3D). For example, a substantial proportion of coeloconic sensilla have an empty lumen (Figure 3C). Consequently, expressing the data as ratios or proportions (Figures 3B, 4C, and 4E) may exaggerate differences between sensillum classes and should therefore be interpreted with caution.<br /> b) The rate of EV/NVEP production is unknown. For a similar rate of production across sensillum classes, Figure 3E suggests that the differences in lumen morphology and size may largely explain variation in EV density and distribution.

      That said, I agree that ab1 sensilla display a striking enrichment of large cargo-filled EVs compared with the other sensillum classes, while coeloconic sensilla display enrichment in small dense filled EVs (Figure 4C). Together, large and cargo-filled EV observation provides strong support for differences in EV biogenesis between ab1 sensilla and the other sensillum classes. In that context, I also think the EV size distribution shown in Figure 4 - Figure Supplement 1 should be moved into the main figure, as it demonstrates that the majority of EVs in the ab1 lumen are relatively large and are therefore likely to represent microvesicles. Could you clarify why ab1 sensilla are only included in Figure 4 and not analysed in Figure 3?

      (3) Approximately 10% of ORNs appear to be degenerating in 6-8-day-old flies, which seems unexpectedly high. This contrasts with the relatively infrequent occurrence of auxiliary cell apoptosis or complete sensillum degeneration. In these "degenerating ORNs", the authors state that the hallmarks of ORN apoptosis are restricted to the dendrites. As hallmarks, they state dendrite truncation, fragmentation and blebbing. Rather than apoptosis, I wonder whether these observations might instead represent ciliary truncation and ectosome shedding, followed by degradation of the shed ciliary membrane into EVs. Ciliary truncation and ectosome shedding, followed by ciliary regrowth, are dynamic processes that have been described across multiple species. This interpretation could explain large EVs that remain in the lumen long after the cilium has regenerated. It would reconcile this article with the general agreement that cilia are a prime site for the budding of EVs across species. Additional evidence supporting apoptosis of the ORNs would help distinguish between these possibilities. Otherwise, I believe the author should reconsider their interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      This paper presents a large structural survey of extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) in the olfactory sensilla of Drosophila melanogaster. Using high-pressure freezing and serial block-face SEM, the authors avoid many of the artifacts associated with conventional fixation and analyze more than 7,800 particles across 352 sensilla. The manuscript maps the distribution of these particles, describes their morphological heterogeneity, and examines their likely origins across different sensillum classes in both normal and degenerating tissue.

      Strengths:

      The strongest aspect of the paper is the imaging. Preservation is painstakingly controlled. The cryofixation appears to preserve the sensillum lymph in a more convincing native state than standard preparation methods, giving this work gravitas. Further, the authors characterized thousands of particles, further making this data strong.

      The figures are strong. They are clear, easy to read, and generally well designed; I think people will use this paper as a model for how to present complex data in a concise and straightforward manner. The manuscript is careful in how it presents the dataset and does not overinterpret the descriptive observations. As an ultrastructural resource, this paper will be useful to the field. The identification of auxiliary support cells as major secretory sites, together with the striking accumulation of EVs in degenerating tissue, will provide a useful starting point for future work.

      Weaknesses:

      The main point that could use more clarification is the vesicle categorization. In particular, the distinction between "dense," "cargo-filled," and "double EVs" is not always easy to follow from a biological perspective. Some additional discussion of how the authors think these categories relate to one another, and whether they are intended as purely morphological groupings or as distinct biological classes, would strengthen the manuscript.

    4. Reviewer #3 (Public review):

      Using cryofixation-based serial block face electron microscopy of several subregions of the Drosophila antenna, the authors segment and assemble a high-resolution atlas of the anatomical structure, density, and spatial distribution of extracellular vesicles (EVs) and non-vesicular extracellular particles (NVEPs) in different Drosophila olfactory sensilla types. This systematic and thorough description is an important prerequisite to understanding the function of extracellular particles in intercellular signaling in the nervous system. Additionally, they describe examples of putative biogenesis events (budding/fusion), as well as neuronal and axonal cell degeneration events and measure the changes in particle accumulation in these altered microenvironments.

      Overall, this is a significant and comprehensive analysis and represents an invaluable resource to this burgeoning field. The authors assemble an important dataset and their claims match the level of evidence provided.

      Strengths:

      (1) The authors use segmentations from four different patches of the antenna to provide a systematic ultrastructural survey of extracellular particles in native insect sensilla. The dataset captures the diversity of sensilla types and reconstructs ~7800 particles.

      (2) We commend the authors for making the EM volumes available in the public Cell Image Library with accession numbers. It would be helpful to the community to also make the segmentations for this great resource easily accessible.

      (3) The sample preparation technique appears to minimize typical artifacts associated with chemical fixation, as evidenced by the high reported sphericity of EVs.

      Specific points:

      (1) In Figure 3C, the authors should include a continuous measure of particle distribution in the sensilla. Currently, the authors define three categories of particle localization. In the five examples shown in Figure 3A, the spatial distribution of these particles appears quite distinct across classes. For example, EVs/NVEPs in large and small basoconic sensilla are largely restricted to the area proximal to the base, with a limited number located more distally. In contrast, intermediate sensilla show a marked concentration of particles more distally.

      (2) The conclusion of different EV ratios across sensillum classes stems from a Kruskal-Wallis of p = 0.0476, with none surviving pairwise comparisons. This is not a strongly supported conclusion and is probably better characterized as a trend.

      (3) The statement "selective enrichment of large, cargo-filled vesicles within the ab1 lumen suggests specialized EV-mediated communication adapted to the coordination demands of this neuronal population" seems speculative for a Results section without supporting functional evidence. It would seem better suited for the Discussion.

      (4) Figure 4: Criteria for defining the classes of EVs.<br /> a) The authors should explain the rationale for classifying EVs using relative density rather than absolute density? We would expect EVs with similar contents to have similar electron density (similar darkness in the images). Would classifying them relative to the background, which itself might vary across sensilla or regions, create a possible confound, especially when comparing across sensillum classes?<br /> b) The two example images (in Figure 4A) of the cargo-filled EVs appear to have different densities themselves. Do the cargo-filled ones also display systematic differences in density and, if so, why is this another class instead of being a subcategory within the dense and lucent classes (i.e. dense with/without cargo, lucent with/without cargo)? The dense and lucent classes are defined by their density, whereas this is a more structural property.<br /> c) Regarding "Double" and "Ball-and-Socket" EVs, does the density vary between the two particles involved (e.g., does the inner structure consistently differ in density from the outer)?

      (5) What was the rationale for the 200μm and 1000μm size cutoffs? A continuous distribution of maximum particle sizes would provide a clearer understanding of the data.

    1. eLife Assessment

      The authors investigate the activation mechanism of a gasdermin pore-forming effector of Callorhinchus milii, which is a shark that is one of the most basal members of the cartilaginous fishes. The valuable findings show that GSDMA/B can be cleaved by a lipopolysaccharide-activated caspase here termed CmiCASP1, analogous to the cytosolic LPS detection pathway already established in mammals. The strength of evidence supporting these conclusions is solid; however, the data indicating the functional consequences of these proteins remain incomplete.

    2. Reviewer #1 (Public review):

      Summary

      The authors present a valuable study of the gasdermins and caspases encoded by Callorhinchus milii, a shark that is one of the most basal members of the cartilaginous fishes. C. milii encodes GSDME and PJVK as well as another gene here called GSDMA/B (which has also been termed GSDMEc in other work). This latter gene is the ancestral gene for bird/reptile/amphibians GSDMA that, in turn, is the ancestral gene to mammal GSDMA, B, C, and D. Prior work had shown that more ancient animals have only GSDME and PJVK, and in these animals caspase-3 and caspase-1 can both independently cleave GSDME. Prior work had also shown that in birds/reptiles/amphibians, GSDMA is cleaved by caspase-1, and GSDME is only cleaved by caspase-3. Here, the authors demonstrate that the more ancient C. milii gene is similar to the bird/reptile/amphibian GSDMA in that it is cleaved by caspase-1. They further demonstrate that C. milii does not encode inflammasomes that would activate CmiCASP1, and instead this caspase is an LPS sensor through its CARD domain analogous to mammal caspase-4/5/11. The data supporting these conclusions are convincing, and could be strengthened by primary cell studies from C. milii in future studies. They further provide evidence that this gasdermin can kill bacteria directly, but the data supporting this conclusion are incomplete.

      Strengths:

      The data demonstrating that CmiCASP1 is an LPS sensor via its CARD domain is thorough and convincing.

      The data demonstrating that CmiCASP1 cleaves and activates GSDMA/B and that this causes pyroptosis is also thorough and convincing.

      Weaknesses:

      I think that the gene/protein referred to in this paper as GSDMA/B was in prior publications called GSDMEc (doi 10.3389/fcell.2022.952015). Is this correct? If not, the relationship or lack thereof to GSDMEc needs to be described. If the authors wish to rename the gene, this needs to be justified and discussed clearly. Also, a gene name with a slash is not typical and was initially confusing to me as it made me think the authors were referring to two different genes.

      The authors do not have data from primary cells from C. milii to demonstrate that the LPS sensing by CmiCASP1 and the pyroptosis induction by GSDMA/B is relevant in the native cell types. This is a common limitation in publications that seek to study diverse animals where tools may not be available. This issue can be studied in future publications.

      The ability of gasdermins to kill bacteria is controversial.

      This bactericidal effect was first shown by the cited article Liu et al. 2016 from Judy Lieberman's lab. I reviewed that manuscript at Nature, and I implored the authors to remove that data from the paper because the experimental design had a high risk of not being physiologically relevant. Indeed, my lab had previously published that bacteria survive the process of pyroptosis and they must be killed by secondary efferocytic phagocytes attracted to the pyroptotic corpse (doi 10.1084/jem.20151613). I have continued to consider whether gasdermins could kill bacteria over the decade since that 2016 paper, and wrote a detailed argument describing how this is unlikely to be physiologically relevant in a recent review article (see Box 3 in doi 10.1038/s41564-026-02272-z).

      In the author's current manuscript, the experiments performed show a very mild effect in the linear range of a reduction of perhaps 20% of the control bacteria in Figure 5A. This is a minimal effect compared to antimicrobial peptides, which will reduce colony-forming units by 99.9%. Take a look at the magnitude of effect in Figure 1 of an example paper looking at polymyxin or colistin killing of Acinetobacter (doi: 10.1128/AAC.00756-12), where CFUs are reduced by about 3 logs (1000-fold) in 30 minutes. The magnitude of effect in Figure 5A is not even 2-fold. Further, it would be very challenging to determine whether the concentration of gasdermin protein used in the assay is equivalent to the concentration that exists in cells. The methods section needs to be clearer to explain how many effective cell lysates of 293T cells were exposed to how many bacteria, because these concentrated lysates are of unspecified concentration.

      Regarding cardiolipin binding, this is a lipid that has a small head group attached to 4 lipid chains, resulting in a cone-like shape with the polar groups at the cone tip and the lipids forming the wide cone base. As such, it creates a larger lipid area than polar area, thus naturally creating a curved membrane shape such that cardiolipin is in the leaflet of the interior of a curvature. Thus, in the mitochondria, it exists in the inner membrane in the mitochondria and allows for the bends that form the cristae, where it faces the surface that is concave (nicely diagrammed in Figure 1 of doi 10.3390/biom16010071). Similarly, cardiolipin enriches in the inner leaflets of membranes at the poles of rod-shaped bacteria to allow for the curvature of the membrane at the poles. Therefore, cardiolipin is not exposed in the outer leaflet of the outer membrane of Gram-negative bacteria; instead, the primary lipid in the outer leaflet is LPS.

      A competing mechanism that could explain the results is that the opening of gasdermin pores in eukaryotic cell plasma membranes occurs concomitant with the generation of ROS from mitochondria. This could occur by gasdermins inserting into mitochondria, as supported by DOI: 10.1038/s41419-025-07760-4. The resulting ROS production due to mitochondrial dysfunction could cause the bactericidal toxicity seen in the cell extracts.

    3. Reviewer #2 (Public review):

      The authors investigate the mechanism by which a gasdermin pore-forming effector of the cartilaginous fish Callorhinchus milii, GSDMA/B (CmiGSDMA/B), is activated. This potentially provides information on the ancestral function of gasdermins, a class of proteins broadly important in human health and disease. By reconstituting components of this system in vitro using transfection models, they show that GSDMA/B is activated by cleavage by the caspase-1 homolog CmiCASP1, which directly senses lipopolysaccharide. This mechanism is broadly similar to the non-canonical pathway in mammals, wherein caspase-4/5/11 cleaves GSDMD upon cytosolic LPS sensing. The conclusions of these interactions are mostly well supported by data, but some aspects need clarification, and based on the experimental approaches, some of the broader interpretations have limitations that should be considered and further discussed.

      A more detailed analysis and discussion on the differences between caspases with regard to their LPS-binding capacity would be valuable for comparison. The analysis of Figure 4A and 4B effectively shows that there are similarities between CmiCASP1 and some of the studied mammalian caspases. However, part of this analysis is to make the point that some caspases do not bind LPS, and it would benefit from the inclusion of additional relevant LPS-insensitive caspases to show the connection between the chondrichthyan caspase residues highlighted and LPS-binding dependence. Modeling the LPS binding site (such as in Figure 1F) would further help clarify whether these are appropriately positioned for coordination, or for non-conserved residues, if there are alternate binding modes thought to have biological relevance.

      The authors note that two different cleavage products are formed, with variable function, which is of interest. The results of Figure 2b suggest that the 241A mutation (blocking the 30 kDa product) increases processing to the larger 35 kDa product, while the 288A mutation decreases processing of the 30 kDa product (also blocking the 35 kDa form). Paired with the lysis data (Figures 2C-2E), its not clear that the 35 kDa product is anything but inactive, but this is quite different from the observations in the experiments with each form (Figure 3N-3Q). A more detailed kinetic and stoichiometric analysis between full-length, N241, and N288 would be important for clarifying the potentially interesting observation of N288 inhibition of N241.

      The mechanism of bacteriocidal activity proposed in the final model and by the experiments of Figure 5 would benefit from further development to support the claim. The experiments do not adequately address whether, during pyroptosis, there is release of N241-like fragments that can kill bacteria. Figure 3G would indicate that it stays in the cell, either in the membrane or mitochondria, and it's not clear there would be circumstances where it could be extracted from it to then target bacteria. Figure 5B might require additional explanation and analysis, but the appearance of similar colonies between conditions would appear to support that there is not measurable antibacterial activity. More rigorous support would come from differences in bacterial killing by knockout Callorhinchus cells, but a minimal step to demonstrating the relevance would be MIC assays, and connecting the effective concentration with one that could naturally occur in Callorhinchus.

      Broadly, the methods of reconstitution of components of this system demonstrate the sufficiency of LPS for activating Casp1, and Casp1 for activating GSMDA/B. However, in more established models, it is clear that there are inhibitors, feedback mechanisms, alternative pathways, and regulation that could render these interactions irrelevant in Callorhinchus. For example, it's not clear where Casp1 and GSMDA/B are ever expressed in the same cell, at quantities sufficient for this mechanism, or that Casp1 doesn't induce more rapid death by acting on something other than GSDMA/B, or that Casp1 is irrelevant because GSDMA/B can be activated more readily by another mechanism. Therefore, while the insights into the evolution of the individual factors of GSDMA/B and Casp1 are interesting and of potential value to the field, reconstituting choice components by transfection of human HeLa and HEK293 cells introduces limitations to how far these experiments can be interpreted as a system. The abstract, for example, states this is a "pyroptosis pathway in cartilaginous fish". However, for all the interest of these data in the evolution of these proteins, the evidence falls short of this. It establishes a biological potential, but it's not clear this is an active pathway in fish.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors focused on Callorhinchus milii GSDMA/B (CmiGSDMA/B) and its upstream inflammatory caspase, CmiCASP1, and revealed that LPS directly engages the CARD domain of CmiCASP1, triggering its activation, which subsequently promotes the proteolytic cleavage of CmiGSDMA/B, yielding two N-terminal fragments with opposite functions. Moreover, consistent with GSDMD, the functional N241 of CmiGSDMA/B can mediate pyroptosis and exhibit bactericidal activity against Gram-negative bacteria in vitro. Based on these observations, the authors clarified that they uncovered an ancestral LPS-sensing CASP1-GSDMA/B axis in cartilaginous fish; however, several issues should be addressed.

      (1) The evidence for direct and functional LPS sensing by CmiCASP1 remains insufficient. Although the authors propose that LPS directly binds the CARD domain of CmiCASP1 to trigger a non-canonical inflammasome-like pathway, the current support mainly comes from pull-down, competition, and in vitro cleavage/activity assays. These results are suggestive but do not yet establish a direct, specific, and physiologically relevant interaction. Additional quantitative binding and specificity analyses are needed to exclude indirect association, aggregation, or other assay artifacts. Therefore, the claim that CmiCASP1 functions as a bona fide direct LPS sensor appears overstated at this stage.

      (2) The proposed antagonistic role of N288 is not yet convincingly supported. While the "dual-fragment antagonistic regulation" model is interesting, it currently relies mainly on overexpression/co-expression, co-IP, and localization analyses, which do not clearly distinguish a physiological inhibitory mechanism from a non-specific dosage or sequestration effect. Stronger support would require evidence for the relative generation and timing of N241 and N288, as well as quantitative data showing that N288 interferes with N241 membrane targeting, oligomerization, or pore formation. Testing the effect of selectively blocking D288 cleavage in the full-length protein would also strengthen this conclusion. At present, the antagonistic model remains premature.

      (3) The physiological and evolutionary claims are stronger than the available evidence. Although the study shows that CmiCASP1 can cleave CmiGSDMA/B and that this module can be reconstituted in heterologous mammalian systems, these data do not demonstrate that such a pathway operates in elephant shark cells or tissues under physiological conditions. A similar concern applies to the antibacterial assays, which use HEK293T lysates rather than purified N241, making it difficult to exclude contributions from host-derived factors. The authors should either provide more direct evidence in a relevant chondrichthyan context or substantially tone down the evolutionary and physiological interpretations.

      (4) The inhibitor data do not convincingly demonstrate suppression of CmiCASP1 activation. Although the authors state that Z-VAD-FMK blocks CmiGSDMA/B cleavage and pyroptotic phenotypes, Figure 1L and Figure 2H do not clearly show that CmiCASP1 activation or processing itself is inhibited. If CmiCASP1 remains processed in the presence of the inhibitor, it becomes unclear whether Z-VAD-FMK blocks CmiCASP1 activation, catalytic activity, or only downstream substrate cleavage. This point should be clarified with more direct biochemical evidence.

      (5) The dosage control for GSDM-derived proteins in the antibacterial assays is unclear. In Figure 5, antibacterial activity is tested using HEK293T lysates or concentrated supernatants containing full-length CmiGSDMA/B, N241, or N288, but it is not clear how protein amounts were normalized across conditions. Differences in expression, stability, or recovery could substantially affect the apparent antibacterial activity. The authors should clarify how input was controlled and ideally provide quantitative normalization or matched-concentration assays to support the comparison.

    1. eLife Assessment

      This important study presents a new approach to individualised functional brain mapping using multi-task batteries, and introduces a toolbox to support its implementation. Convincing evidence from simulations and empirical analyses across multiple datasets demonstrates clear advantages over traditional single-task approaches and provides practical guidance on task selection. The extent to which these improvements translate into more accurate recovery of individual-specific functional boundaries beyond well-characterised systems remains to be established.

    2. Reviewer #1 (Public review):

      Summary:

      In this well-written and well-presented manuscript, Arafat and colleagues describe the proper use and advantages of multi-task batteries to understand the organization of the human brain. The authors present both simulation and empirical results suggesting a substantial advantage in using many short tasks vs a single localizer in identifying specific task-engaged regions. The natural question arises as to which tasks should be used within the battery, and how they should be organized. The authors address this question by demonstrating a data-driven strategy for task selection that outperforms random selection, and they further demonstrate the advantages of highly interspersed tasks over a more typical one-task-per-run strategy.

      Strengths:

      In general, I find this work highly compelling. The topic itself should be of high interest to the majority of researchers conducting human functional neuroimaging studies. The manuscript itself is comprehensive and sound. The analyses and data are truly excellent, with only a few, relatively minor issues that can be improved. The authors do an exceptional job of laying out the motivation and logic for almost every analysis and conclusion in the manuscript.

      I want to point specifically to the potential impact of this work. While the claims made here are very appropriately constrained to the conclusions that can be drawn from the actual analyses, their impact is potentially far-reaching. By the end of this manuscript, we are left with a set of ideas that in effect overturns 2-3 decades of received knowledge about how functional neuroimaging tasks should be designed to optimally understand the organization of the human brain.

      Thanks to this paper, I personally will be rethinking how I design all of my fMRI studies in the future after reading this work. The authors are to be commended for this excellent contribution to the literature.

      Weaknesses:

      I struggled to understand the motivation and logic of the "connectivity modeling" section of the analyses.

    3. Reviewer #2 (Public review):

      Summary:

      This paper presents theoretical and empirical insights into the use of multi-task batteries for precision functional brain mapping and offers practical guidelines for optimal task design. Specifically, the authors evaluate differences between single-contrast and multi-task localizers, explore data-driven strategies for battery selection, such as minimizing collinearity, and compare grouped and interspersed stimulus-presentation designs. Through a combination of simulations and analyses of empirical fMRI data, the study provides a systematic set of recommendations for improving the reliability and specificity of individualized functional mapping.

      Strengths:

      Traditional functional mapping has long relied on single-contrast localizers or resting-state fMRI. However, there is growing recognition that diverse batteries of general tasks can yield more detailed functional maps with higher signal-to-noise ratios (SNRs). This manuscript systematically evaluates these advantages using both simulations and empirical data. The contribution is timely and provides the community with not only a theoretical justification for multi-task designs but also practical tools, in the form of the MultiTaskBattery toolbox, for implementing them.

      Weaknesses:

      Although the results are robust, they are largely consistent with existing expectations in the field, and the conceptual novelty or "surprise" factor is therefore somewhat limited. Nevertheless, synthesizing these findings into a coherent set of design recommendations provides significant value to researchers.

      Additionally, there appears to be a slight mismatch between the content of the manuscript and its designated article type. Although the manuscript was submitted as a "Tools and Resources" article, its extensive empirical analyses and theoretical evaluation make it read more like a "Research Article." I defer this categorization to the Editor's judgment.

      Finally, the authors use inter-subject overlap as a primary metric for validating the accuracy of functional mapping (Figure 3). However, given that genuine inter-individual variability in brain organization is a central premise of precision mapping, greater overlap across subjects may not necessarily indicate more accurate individual-level localization. A more detailed analysis or discussion of how to distinguish measurement noise from genuine individual differences would make the paper more comprehensive and strengthen its overall contribution.

    4. Reviewer #3 (Public review):

      Summary:

      This study introduces a principled framework for optimizing multi-task batteries for individualized functional brain mapping. Through simulations and empirical validation, the authors show that selecting tasks to maximize differences in regional response profiles can substantially improve the identification of functional brain regions. The work represents a valuable methodological advance for precision functional mapping, although some assumptions underlying the broader applicability of the framework would benefit from further discussion.

      Strengths:

      The manuscript addresses an important methodological challenge in precision functional mapping using a rigorous combination of theoretical analyses, simulations, and empirical validation. The framework is practical and well supported by open-source software and a publicly available task library, making it readily accessible for adoption and further development by the research community. The manuscript is well written, logically structured, and clearly presents both the methodological framework and its practical implementation.

      Weaknesses:

      (1) The abstract and introduction emphasize the application of the framework to individualized brain parcellation. While the presented analyses convincingly demonstrate improved prediction of held-out task responses using atlas-guided parcel assignments, they do not directly validate whether the optimized task batteries improve the estimation of an individual's true functional boundaries. The empirical validation relies on atlas-defined parcel identities as the reference standard, yet substantial inter-individual variability in the location and extent of functional regions - particularly within association cortex - has been well documented. Consequently, improved recovery of atlas-defined parcel labels does not necessarily imply more accurate recovery of an individual's functional organization. It would therefore be valuable to clarify this distinction in the abstract and discussion and to discuss how inter-individual variability may influence the interpretation and generalizability of the parcellation analyses.

      (2) The framework assumes that informative task batteries can be designed to distinguish neighboring functional regions. While this is compelling for well-characterized systems with distinct functional response profiles, it is less clear how the approach generalizes to finer-scale subdivisions within association cortex (e.g., subnetworks), where neighboring regions may exhibit highly similar task-response profiles and their functional roles remain incompletely understood. In these settings, the relevant functional dimensions may not yet be known, making it difficult to design optimized task batteries a priori. It would therefore be valuable for the authors to discuss how the framework could be extended to such cases.

      (3) More generally, the framework assumes that the sampled task space adequately captures the functional dimensions that differentiate cortical regions. However, particularly within the association cortex, neighboring regions may exhibit similar task-response profiles while differing in the information they represent, their interactions with other regions, the computations they perform, or their cortical layer-specific response patterns. In such cases, the dimensions that best distinguish cortical organization may not be fully reflected in task-response profiles alone, but instead become apparent through complementary approaches such as representational analyses, task-evoked or resting-state connectivity, computational modelling, or laminar response profiles. It would therefore be valuable to discuss how the proposed framework relates to these complementary perspectives.

      (4) Many task batteries inherently contain tasks that vary substantially in cognitive demand. Given that task difficulty is itself a major organizational axis in association cortex, it would be helpful for the authors to discuss how the optimization framework accounts for this. Specifically, could differences in task difficulty drive regional differentiation, even when tasks probe similar underlying cognitive processes? If so, how does the framework distinguish between organizational differences arising from a common demand axis and those reflecting more specific functional specializations?

      (5) The Discussion places the proposed framework in the broader context of precision functional mapping and refers readers to a companion paper demonstrating advantages over resting-state ("inside-out") approaches. Given that these comparisons motivate several of the broader recommendations made in the Discussion, it would be helpful to provide a brief summary of the main findings of the companion paper here. This would allow readers to better understand the basis for these conclusions without relying on a separate manuscript.

    1. eLife Assessment

      This useful study builds on the previously identified terminal Schwann cell (tSC) marker Col20a1 to generate a novel Col20a1-CreERT2 knock-in mouse line that enables the selective genetic labelling and ablation of tSCs at the neuromuscular junction. This model represents a new tool for distinguishing tSCs from myelinating Schwann cells. The authors conclude that loss of tSCs does not cause major alterations in the overall structure of the neuromuscular junction, but that tSC ablation intrinsically compromises the presynaptic reserve pool. However, the study remains incomplete because the evidence provided does not adequately support these conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript's major strength is the identification of Col20a1 as a novel marker for terminal Schwann cells (tSCs) and the generation of the Col20a1-CreERT2 knock-in mouse line, which represents a valuable new genetic tool for studying tSC biology at the neuromuscular junction. However, the central conclusion that terminal Schwann cells regulate presynaptic vesicle homeostasis is not sufficiently supported because the electrophysiological and ultrastructural analyses are based on limited sample sizes. Increasing the number of animals and NMJs analyzed would substantially strengthen the conclusions. In addition, further validation of Col20a1 expression using RNAscope or immunostaining, together with comparisons to established tSC markers such as Kir4.1 and NG2, would help establish its specificity. Finally, the developmental appearance of Col20a1-positive axonal Schwann cells is intriguing and warrants further investigation to determine whether these cells migrate and differentiate into terminal Schwann cells during postnatal development.

      Strengths:

      The authors identified Col20a1 as a specific marker of terminal Schwann cells (tSCs) at the neuromuscular junction (NMJ) in mice and generated a Col20a1-CreERT2 knock-in mouse line for in vivo labeling of tSCs. This represents a novel and significant technical advance for the NMJ field, providing a valuable genetic tool for studying the development, maintenance, and function of terminal Schwann cells in vivo.

      Weaknesses:

      The major weakness of this manuscript is that the sample sizes are too small to support the authors' conclusion that "Terminal Schwann Cells Regulate Presynaptic Vesicle Homeostasis but Not Neuromuscular Junction Integrity in Mice."

      Although the authors state that "Quantification of the tdTomato-positive NMJ ratio showed an ablation efficiency of approximately 80%, with only ~20% of NMJs retaining escaper tSCs" (page 9), they do not provide the sample size or sufficient quantitative information for the Col20a1-tdTomato/DTA ablation experiment shown in Figure 4. This information is essential for evaluating the robustness and reproducibility of the ablation strategy.

      The electrophysiological analyses are also based on very limited sample sizes. According to Figure 6, mEPP recordings were obtained from 10 NMJs from 3 control mice and 14 NMJs from 5 tSC-ablated mice. The EPP recordings were based on similarly small numbers (control, n = 10 NMJs from 3 mice; tSC-ablated, n = 14 NMJs from 5 mice). Thus, only approximately 2-3 NMJs were analyzed per tSC-ablated mouse on average. Given the inherent variability among individual NMJs and animals, these sample sizes are insufficient to support broad conclusions regarding the effects of terminal Schwann cell ablation on synaptic transmission.

      In addition, the authors describe the phenotype as "leaky" presynaptic spontaneous release, but this terminology is not defined and lacks mechanistic explanation. It is therefore unclear what specific physiological alteration the authors intend to describe.

      The sample sizes for the electron microscopy analyses (Figure 7) also appear to be limited, making it difficult to determine whether the reported changes in synaptic vesicle distribution are representative or statistically robust. Because the central conclusion relies heavily on these electrophysiological and ultrastructural data, the evidence presented is not sufficient to support the claim that terminal Schwann cells regulate presynaptic vesicle homeostasis. At present, this conclusion is overly broad and not adequately supported by the available data.

      Minor comments:

      (1) While several figures contain high-quality NMJ images (e.g., Figures 1B and 4), the image quality in other figures should be improved. For example, the S100B immunostaining appears overexposed in some panels, making it difficult to distinguish individual Schwann cells or visualize the boundaries between adjacent cells.

      (2) Some neuromuscular junctions shown in Figure 2C appear to be partially denervated. The authors should clarify whether these represent normal variability, effects of the experimental manipulation, or imaging artifacts.

      (3) The authors should specify the muscle preparation used in Figure 3, as this information is necessary for interpreting the results and comparing them with previous studies.

    3. Reviewer #2 (Public review):

      Summary:

      Two types of Schwann cells (SCs) ensheath motor axons - myelinating SCs along the axonal length and terminal SCs (tSCs) that cover nerve terminals at the neuromuscular junction (NMJ). Therefore, the NMJ is, like other synapses, tripartite, with specialized presynaptic, postsynaptic, and glial cells. Many studies have shown that tSCs play roles in the development and function of the NMJ, but for some of these, interpretation is difficult because it is hard to manipulate tSCs without also manipulating myelinating SCs. To circumvent this problem, Kong et al. make use of a gene selectively expressed in tSCs, Col20a1 (Figure 1), to generate a knock-in mouse line, Col20a1-CreER, that gives them genetic access to tSCs. They cross this to Cre-dependent lines that mark tSCs with a red fluorescent protein (Figures 2 and 3) or ablate them by expression of diphtheria toxin along with the fluorescent protein (Figure 4). They show that ablation at postnatal day (P) 10 does not affect the overall structure or function of the NMJ (Figures 4 and 5). It does, however, affect some aspects of neuromuscular transmission over the following few weeks (Figures 6 and 7). Long-term effects cannot be studied by this method, however, because terminal SCs are replaced, presumably from the preterminal population (Figure 8).

      Strengths:

      The work is done to a high technical standard, including detailed characterization of the knock-in model. Results are presented clearly and illustrated beautifully. The finding that some early reports of synaptic alterations may result from concurrent loss of axonal SCs is important in rethinking the role of tSCs.

      Weaknesses:

      (1) The authors claim that tSCs are dispensable for some aspects of NMJ maturation, including synapse elimination (called pruning here), formation of "pretzel-like" postsynaptic topology, and generation of junctional folds in the postsynaptic membrane (lines 223 and 363). However, this conclusion is based on injection of tamoxifen to initiate tSC ablation at P10, which is necessary because Col20a1 is expressed in some preterminal SCs at earlier times. It presumably takes a few days for CreER to translocate to the nucleus and activate the toxin transgene, and some more time for the toxin to be generated and act. This is problematic because synapse elimination and other aspects of maturation mentioned occur during the first two postnatal weeks and are largely complete by P14. Therefore, one cannot conclude that these aspects "proceeded normally despite the loss of tSCs....".

      (2) Effects on synaptic transmission are modest at best, being significant at a level of p<0.05 but not p<0.01 (Figure 6E, G, H and most of L). Effects on vesicle density are more robust (Figure 7).

      (3) The authors use red fluorescent protein from the Col20a1 to label tSCs, and antibodies to S100b to label all SCs. This is appropriate in normal muscle and soon after tSC ablation. At later times, however, the NMJ is repopulated by S100+ Col20a1- SCs (Figure 8B). It is therefore important to show when this repopulation begins, because a modest recovery of SC coverage could have a big effect. For example, Figure 4C quantifies loss of NMJs with residual RFP+ cells but not S110+ cells; both should be quantified at this and slightly later stages.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript reports a novel genetic model, Col20a1-CreERT2 knock-in mouse, to target terminal Schwann cells (tSCs) at mouse neuromuscular junctions in a cell-type-specific and temporal manner. The authors analyzed multiple publicly available single-cell transcriptome databases to identify Col20a1 as the tSC marker. The authors crossed Col20a1-CreERT2 and Rosa26-LSL-tdTomato to label tSCs successfully. In addition, the authors generated Col20a1-CreERT2; Rosa-tdT/ diphtheria toxin subunit A (DTA) to specifically ablate tSCs and analyze the role of tSCs in motor behavior, neuromuscular junction electrophysiological function, and the histology and ultrastructure of neuromuscular junctions.

      Strengths:

      The Col20a1-CreERT2 x Rosa26-LSL-tdTomato mice successfully labeled the tSCs at NMJs and reported the developmental distribution of the Col20a1-positive cell population. The Col20a1-CreERT2; Rosa-tdT/DTA mice successfully ablated tSCs, which did not cause changes in gross neuromuscular junction architecture, neuromuscular synapse physiology, or motor behavior.

      Weaknesses:

      The conclusion of this manuscript will be strengthened by additional analysis showing time-course data of tSC ablation and replacement by non-recombined Schwann Cells. Currently, it is not clear when and how long the tSCs are ablated, which makes it difficult to interpret the data and phenotype. Detailed review comments are provided to the authors in the "recommendations for the authors" section.

    1. eLife assessment

      The findings documented here are interesting and valuable with respect to revealing genetic players that may underlie AD. However, there are concerns related to the analysis pipeline and the basic assumptions underlying the modelling. There is an overarching lack of controls in terms of comparator diseases and/or positive and negative control findings; hence, the main claims are only partially supported, and the study is incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      The authors build a reusable, disease-agnostic pipeline that retrieves GWAS risk genes from the GWAS Catalog by ontology terms, filters them, and maps them onto Human Protein Atlas (HPA v24) co-expression modules at three biological scales (tissue/organ, brain region, cell type). Enrichment is assessed by a consensus of Fisher's exact test with Benjamini-Hochberg correction and 10⁶-iteration Monte Carlo simulation. Applied to AD, DLB/PD, and FTD/ALS, the analysis reports convergent neuronal-module enrichment across all three diseases, AD-specific enrichment in liver- and immune-associated modules, and DLB-specific enrichment in ciliary modules, followed by a DrugBank-based survey of compounds targeting module gene products.

      Strengths:

      (1) The core premise is sound and well-motivated: risk-gene lists are hard to interpret because most variants are low-penetrance and broadly expressed, and projecting them onto a multiscale expression atlas is a reasonable route from statistical association toward tissue/cell context.

      (2) The pipeline is delivered as reusable, open code (GitHub) built entirely on public inputs (HPA, GWAS Catalog, DrugBank), which is a significant contribution to the field and provides opportunities for replication and expansion.

      (3) The dual-enrichment design (fold enrichment with FDR correction plus a 10⁶-iteration Monte Carlo empirical null) is more defensible than any single test, and the "consensus" logic is sound.

      (4) The multiscale framing (organ → brain region → cell type), with UMAP module projections, is genuinely helpful and makes the mapping legible to broad scientific backgrounds.

      (5) The authors are commendably restrained on one key point: they explicitly report that most risk genes are broadly expressed and not brain-selective, rather than overstating neuronal specificity.

      (6) The AD liver/immune convergence is nicely discussed and integrated, and the NAFLD-AD discussion (Kupffer-cell/hepatocyte Aβ clearance, locus coeruleus noradrenergic parallels, "type 3 diabetes") is thorough and well-referenced, even where it remains speculative.

      Weaknesses:

      (1) Gene-to-variant mapping via the author-reported gene field is unclear. The Methods assign genes using the GWAS Catalog author-reported gene field, which predominantly reflects the nearest gene to the lead SNP and may not be the effector gene; it is also inconsistent across studies and different time periods of publication (i.e., changing methodologies in genomics and GWAS procedures). Because every downstream result depends on the gene set, this choice likely introduces noise and bias into all enrichment, specificity, and drug claims. More modern practices link variants to genes via fine-mapping plus eQTL/pQTL colocalization, or integrative scores. At minimum, the sensitivity of the main signatures to nearest-gene versus colocalization-based assignment should be demonstrated.

      (2) The suggestive threshold (p < 1×10⁻⁵) trades specificity for coverage in the analysis that requires specificity. Relaxing from 5×10⁻⁸ substantially raises the false-positive fraction of the gene set. This is defensible for exploratory coverage in under-powered DLB/FTD, but the headline claims concern disease specificity (limited gene-set overlap; distinct signatures). Non-overlap among partially false-positive lists can lead to biological specificity conclusions that may not actually exist. A genome-wide-threshold sensitivity analysis is needed to show the signatures persist. Also, see concerns below about the DLB designation.

      (3) "Disease specificity" is confounded by GWAS power. AD GWAS (e.g., Bellenguez; Kunkle; Sherva ~205,500 cases) vastly outpower DLB and FTD discovery. The gene counts (453/278/219) and the minimal three-way overlap (only two genes) track sample size and locus density as much as biology. The claim of "disease-specific genetic architectures" should be tempered and ideally power-matched (e.g., subsampling AD, or restricting to comparable effective N) before specificity is asserted. Also, see concerns below about the DLB designation.

      (4) Linkage disequilibrium structure at gene-dense loci is not addressed and may inflate the lipid/liver signal. The overlap genes named as driving the liver/lipid theme include APOC1, APOC2, and APOE, all of which reside within the same chromosome-19 linkage disequilibrium (along with TOMM40). Counting co-regulated, physically clustered genes from one association signal as independent risk genes risks inflation of enrichment for lipid/lipoprotein modules. The five "liver-enhanced" genes flagged in the text again lead with APOE. Evaluation of one gene per independent signal is essential before the liver/lipid signature can be interpreted as multi-gene convergence rather than a single strong gene driving the effect.

      (5) The enrichment background is not clear. Risk genes are filtered to HPA brain-detected transcripts, but the Methods do not state whether the Fisher/Monte Carlo background (N_Total) is likewise restricted to brain-expressed/HPA-detected genes or reflects all protein-coding genes, which may introduce bias. Specific information on the Fisher/Monte Carlo should be provided to ensure that the enrichment background is the same. If they are not, additional analyses should be performed to ensure robustness of the findings when restricted to brain-expressed only or the broader background.

      (6) Cross-scale "consensus" is not independent confirmation. The same genes reappear across tissue, brain, and cell modules, so agreement across scales is partly built-in rather than corroborating. The neuronal-signature counts (82 AD / 62 DLB / 53 FTD; 186 "unique" genes) should be accompanied by a clear statement of how much cross-scale evidence is non-redundant. Also, see concerns below about DLB designation.

      (7) Temporal claims are overstated. Using control HPA tissue avoids end-stage confounds but, by construction, cannot capture disease-state programs central to AD. More importantly, framing these modules as "baseline vulnerability hotspots that precede clinical neurodegeneration" is not tested, as nothing here is longitudinal. This assumption is presented as a finding and should be reworded as a hypothesis.

      (8) DLB and PD should not be merged, and the cilia-associated risk genes cannot be termed causal based on the study design. The supporting literature is almost entirely PD (Schmidt iPSC-NPCs from sporadic PD; LRRK2 PD striatum), yet DLB and PD are merged, and the signature is branded "DLB." These should not be merged, particularly as it relates to sporadic PD. Although there are similar genetic risk factors (i.e., GBA, SCNA), of which GBA is unfortunately not mentioned in the manuscript, there are major genetic differences in sporadic PD and even PD with dementia (PDD) and DLB. It is not clear why PDD was not incorporated.

      Cross-sectional expression overlap with GWAS genes cannot establish that cilia-associated risk genes are "causative vulnerabilities rather than secondary effects." Please soften to association and rename to reflect the synucleinopathy grouping. Further, additional discussion of important differential genes in this category (i.e., APOE and GBA) is needed, and the pathological description of DLB is incomplete in the introduction (i.e., focuses only on synuclein).

      (9) The drug/repurposing analysis rests on a weak targeting rationale. Most of the 1,777 drugs target co-expression-module neighbors of risk genes, not risk genes themselves, so "substantial repurposing reservoir" likely overstates the impact. The anesthetics example (sevoflurane/halothane/desflurane hitting GABA_A subunits and ATP2B2) is a near circular argument, as anesthetics necessarily engage neuronal ion channels. This finding does not independently confirm the neuronal signature.

      Statin-dementia and hydroxychloroquine/amodiaquine links are pre-existing, and the Tirzepatide→cilia→DLB inference is highly speculative. This section should be labeled hypothesis-generating, with direct-risk-gene targets separated from module-neighbor targets.

      (10) Interpretation leans heavily on nominal (p < 0.05) modules. Several of the most novel claims (parts of the liver and immune signatures) rest on nominally significant modules that do not survive FDR (itself set leniently at p_adj < 0.1). The text should make consistently explicit which claims are FDR/Monte-Carlo-supported versus nominal-only, and de-emphasize conclusions resting solely on the latter.

    3. Reviewer #2 (Public review):

      Summary:

      Genes associated with risk for a specific disease commonly have widespread expression and functions across the body. Surveying patterns in these effects may reveal novel mechanisms, organs, and systems implicated in a disease, amongst other associations that are truly independent. In this work, Husen and coauthors use the Human Protein Atlas to explore such associations in Alzheimer's disease (AD), Lewy body dementia (DLB), and Frontotemporal dementia (FTD). Focusing on human non-disease tissue expression may avoid the effects of disease progression obscuring initial vulnerabilities. However, associations in non-diseased tissues do not necessarily reflect mechanisms causally related to the diseases themselves.

      The work describes patterns of enrichment of genes across tissue types, brain regions, and cell types. A relatively small set of classes of each show enrichment for disease. While neural signatures are unsurprisingly prevalent, these classes are largely distinct across the three diseases. Alzheimer's disease is linked to liver and central and peripheral immune cells, while DLB shows interesting enrichments associated with cilia, which are linked to an existing literature. Results from the drug repurposing approach are then presented, with 1777 drugs linked to protein products of any of the modules enriched by the risk genes using DrugBank, categorised according to key signatures.

      The authors developed R code (the HPA GeneSet Explorer) to automate the production of multi-system summaries of the organs, brain regions, cells and gene modules associated with traits and diseases and associated gene sets, within the HPA. Risk gene sets for the three dementia types were derived from the GWAS Catalog.

      Strengths:

      While many studies of how risk genes contribute to disease take a narrow approach focusing on organs, cell types, and processes already associated with a disease, it is a sensible approach to start with a system-agnostic approach that assesses tissues that are not ostensibly affected by disease. Here, this approach reveals a range of associations for 3 neurodegenerative diseases, identifying disease-associated modules and drug candidates that might be prioritized for subsequent confirmatory inference across biological scales. Results highlight key organs and cell types, most of which have established associations with the disease. Perhaps the most intriguing results are the links of DLB to cilia-related processes, which can be linked to some prior reports of DLB/PD but are not a core element of current theories of pathogenesis.

      Weaknesses:

      A difficulty with broad, multi-dataset surveys of disease associations is the need to distinguish novel and robust patterns - even if they lack causal evidence - from those that are unsurprising or do not stand out statistically. The work is exploratory in nature, but it is often hard to know how strong the evidence is for particular observations reported.

      The work combines nominal, FDR<0.1, and Monte Carlo-based inference (with no apparent multiple assessment control across all tested modules) p-values throughout the paper, with patterns of effects of nominal significance. In some places, modules appear to be retained if they meet any of these criteria, muddying inference. This makes it difficult to weigh the different reported associations. Results report numbers of risk genes showing nominal p<0.05 enrichment across gene modules and biological scales - it is difficult for the reader to determine null expectations for false positives here. Similarly, it is unsurprising that thousands of drugs can be linked to the risk genes and their signatures using nominal significance.

      The results have limited mechanistic specificity. The modules identified often reflect biological processes implicated in the diseases. This provides some validation of the approach, but the modules are often broadly defined, providing little mechanistic insight. For example, many aspects of ciliary biology may overlap with DLB, but can the HPA provide more specific insight? More generally, it is difficult to determine how much relevance that enrichment in non-disease tissue has for disease processes. Similarly, it is hard to determine whether overlap of drug targets from DrugBank with these modules realistically increases their prioritization.

      Methodologically, there could be more detail. The paper - in particular the methods - is partially presented as a tool/pipeline paper, but thorough descriptions of the HPA models that are employed and modules reported for the analyses should still be presented in detail. The drug repurposing approach is described in a couple of sentences without a precise reference to the tool or statistical methods.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript presents the development of a software tool and a computational workflow for the comparison and biological interpretation of GWAS results among three neurodegenerative diseases - AD, PD and FTD - using the data from the Human Protein Atlas. The multi-scale content analysis produces a representation of GWAS results highlighting enrichment at the level of brain regions, organ systems/tissues and cell types as well as in terms of molecular pathways. The manuscript argues that DLB, AD and FTD have differential 'modules' revealed by the procedure.

      Strengths:

      The system is leveraging vast knowledge sources including both the GWAS catalog and the HPA, and it is integrating these data. This synthesis of information is useful. It is making use of large investment in data generation and data warehouses in order to yield interpretation of epidemiologic results from GWAS in biological terms that may yield insight into disease processes and differences between diseases. The multiscale analysis recognizes the various biological lenses at which the implications of GWAS can be evaluated.

      Weaknesses:

      There is a lack of controls and/or disease comparators in the study. It is hard to assess that the workflow is performing 'as expected' without a set of positive or negative controls - or at least comparators - to gauge the performance of the tool. The statistical methods are simplistic and rely on Fisher's exact tests, UMAP analyses and clustering with little justification. There is a lack of power analysis and specification of the number of genes required for the procedure to 'work'. The mapping of GWAS hits to genes is simplistic and may in many cases be erroneous, and the implications of errors in these mappings are not considered. Many of the hits are not brain-specific - highlighting the complexity of gene function and the pleiotropic nature of gene activity. Moreover, the mapping between 'tissue enrichment' and 'tissue that is the functional driver of the GWAS signal' may be a logical flaw in the reasoning of the authors. Just because a tissue - such as liver enrichment in AD - is enriched in the GWAS gene-mapping analysis does not mean that that tissue found to be enriched is in fact the functional tissue that gave rise to the GWAS signal. Genes have different isoforms, functions, and regulatory mechanisms in different parts of the organism in different parts of development. In a phrase, the enrichment observed could be correlative and not causative, and in fact the enrichment could be driven by some hidden variable not considered. The etiologic tissue for neurodegenerative disease is the brain. The findings are not necessarily surprising or novel in the distinction between PD, AD, and FTD. Finally, the figures are perhaps not the best way to present results. There are many small pie charts that lack interesting results; the figures are in general hard to read and could use refinement in terms of fonts.

    1. eLife Assessment

      This important study shows how hunger alters avoidance of harmful heat in C. elegans by reconfiguring the activity of key sensory neurons. The evidence is convincing, with well-designed behavioural, genetic, and imaging experiments that support the main conclusions. The work will be of interest to neuroscientists studying how internal states shape sensory processing and behaviour across species.

    2. Reviewer #1 (Public review):

      This study by Thapliyal and Glauser investigates the neural mechanisms that contribute to the progressive suppression of thermonociceptive behavior that is induced under conditions of starvation. Several previous studies have demonstrated that when starved, C. elegans alters its preferences for a variety of sensory cues, including CO2, temperature, and odors, in order to prioritize food seeking over other behavioral drives. The varied mechanisms that underlie the ability of internal states to alter behavioral responses are not fully understood, however there is growing evidence for a role by neuropeptidergic signaling as well as capacity for functionally distinct microcircuits, formed by distinct internal states, to trigger similar behavior outcomes.

      Within the physiological range of C. elegans (~15-25C), starvation triggers a profound reduction in temperature-driven thermotaxis behaviors. This reduction involves the recruitment of the amphid sensory neuron pair AWC. The AWC neurons primarily act to sense appetitive chemosensory cues, however under starvation conditions begin to display temperature responses that previous studies have linked to the reduction in thermotaxis navigation. Here, Thapliyal and Glauser investigate the impact of starvation on thermonociceptive responses, innate escape behaviors that are triggered by exposure to noxious temperatures above 26C or rapid thermal stimuli below 26C. They compare the strength of thermonociceptive behaviors, specifically heat-triggered reversals, in worms experiencing either early food deprivation (1 hour off food) or prolonged starvation (6 hours off food). Their experiments demonstrate a progressive loss of heat-triggered reversals that is mediated by AWC and ASI neurons, as well as both glutamateric and neuropeptidergic signaling.

      At the level of neural activity, this study reports that the transition from early food deprivation to prolonged starvation reconfigures the temperature-driven activity of AWC neurons from mostly excitatory to a heterogenous mix combining excitatory and inhibitory responses. This finding is interesting in light of previous work that reported the opposite transition in temperature-driven AWC responses when comparing well-fed worms to those kept from food for 3 hours. Specifically, these differences highlight the differences between temperature responses within the C. elegans physiological temperature range (previous studies) and their noxious temperature response (this study). This study also identifies neural and genetic mechanisms that contribute to differences in thermonociceptive responses at +1 versus +6 hours starvation; interestingly, these mechanisms are also partially distinct from those that contribute to differences in negative thermotaxis behaviors in well-fed and +3 hours starvation worms. A limitation of this manuscript is that these differences are not particularly acknowledged or addressed, other than the hypothesis that independent mechanisms underlie negative thermotaxis versus thermonociceptive stimuli.

      In this revised article, the authors commanding knowledge of the distinction between thermotaxis navigation (especially negative thermotaxis) and thermonociceptive behaviors is communicated with an admirable depth and clarity to the readers; this study's new findings are helpfully contextualized within the previous literature.

      This study represents an important addition to the growing evidence that C. elegans sensory behaviors are strongly impacted by internal states, and that neuropeptigergic signaling plays a key role in mediating behavioral plasticity. To that end, the authors have provided compelling evidence of their claims.

    3. Reviewer #3 (Public review):

      Thapliyal, Gopinath, and Glauser show that starvation alters how C. elegans respond to noxious thermal stimuli. Using targeted neural ablation, mutant analysis, and live-cell functional imaging the authors demonstrate that hunger changes the properties of AWC sensory neurons, which sense noxious heat. The authors further show that effects of hunger on nociception require ASI neurons, which are known to respond to hunger and mediate effects of food deprivation on behavior. Finally, the study uses mutant analysis to implicate glutamate and specific neuropeptides in thermal nociception and in modulation of nociceptors by hunger-responsive neurons.

      The study clearly shows a strong effect of hunger on nociception and documents a striking effect of hunger on the intrinsic properties of AWC sensory neurons, which respond to noxious heat. The study also clearly and compellingly demonstrates that ablation of hunger-responsive ASI neurons blocks effects of hunger on nociceptive AWCs. These data, which constitute the kernel of the manuscript, are striking and exciting. This revised manuscript analyzes effects of starvation on AWC physiology and clearly shows that starvation alters the way AWCs respond to thermal stimuli by decreases the probability that AWCs will be activated and increases the probability that they will be inhibited. New data also identify ASI-derived neuropeptides that are required for modulation of AWCs by starvation. This study reveals a mechanistic link between an animal's metabolic state and sensory processing and establishes modulation of AWC function as a powerful model to study the molecular basis of this link.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We believe that the manuscript has been substantially strengthened through the revision process. The main changes are summarized below:

      We substantially revised the Introduction and Discussion sections to better position our work relative to previous studies on starvation-dependent thermotaxis plasticity, neuropeptidergic modulation, and AWC function.

      We clarified throughout the manuscript the distinction between negative thermotaxis in innocuous thermal ranges and thermonociceptive responses to noxious heat. We now discuss more explicitly that these behaviors involve at least partly distinct molecular, cellular, and circuit-level mechanisms.

      We performed new experiments in ins-1 mutants. Unlike what was previously reported for thermotaxis plasticity, ins-1 does not appear required for starvation-dependent thermonociceptive plasticity in our paradigm (new Figure 6—figure supplement 1).

      We revised the analysis and terminology used for AWC calcium imaging data. We no longer use the “deterministic/stochastic” terminology and instead describe a starvation-induced shift from predominantly excitatory responses to a mixed distribution of excitatory and inhibitory responses. We also added new quantitative analyses and histogram representations of response distributions, as directly suggested by reviewers, to better illustrate this point.

      We performed new genetic interaction experiments using eat-4; flp-6 double mutants. These analyses revealed that glutamatergic and FLP-6 signaling act largely in parallel to mediate heat-evoked reversals after early food deprivation, while prolonged starvation reveals a hierarchical interaction between these pathways.

      We revised and clarified the mechanistic model figures accordingly, particularly regarding the proposed ASI → AWC signaling pathway and the role of ASI-derived neuropeptides.

      We improved the presentation and statistical rigor throughout the manuscript, including:

      - Replacement of heating power values by corresponding temperature increases,

      - Clarification of the rationale for using the 1-hour off-food condition as reference,

      - Expanded statistical reporting and multiple-comparison procedures,

      - Additional methodological details for calcium imaging, rescue validation, and cell ablation approaches,

      - Clarification of replotted datasets in figure legends.

      We also simplified the manuscript by removing experiments whose interpretation remained ambiguous (notably the nsy-1 and nsy-7 analyses).

      Below, we provide a detailed point-by-point response to all reviewer comments.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Thapliyal and Glauser investigates the neural mechanisms that contribute to the progressive suppression of thermonociceptive behavior that is induced under conditions of starvation. Several previous studies have demonstrated that when starved, C. elegans alters its preferences for a variety of sensory cues, including CO2, temperature, and odors, in order to prioritize food seeking over other behavioral drives. The varied mechanisms that underlie the ability of internal states to alter behavioral responses are not fully understood; however, there is growing evidence for a role of neuropeptidergic signaling as well as the capacity for functionally distinct microcircuits, formed by distinct internal states, to trigger similar behavior outcomes.

      Within the physiological range of C. elegans (~15-25{degree sign}C), starvation triggers a profound reduction in temperature-driven thermotaxis behaviors. This reduction involves the recruitment of the amphid sensory neuron pair AWC. The AWC neurons primarily act to sense appetitive chemosensory cues; however, under starvation conditions begin to display temperature responses that previous studies have linked to the reduction in thermotaxis navigation. Here, Thapliyal and Glauser investigate the impact of starvation on thermonociceptive responses, innate escape behaviors that are triggered by exposure to noxious temperatures above 26{degree sign}C or rapid thermal stimuli below 26{degree sign}C. They compare the strength of thermonociceptive behaviors, specifically heat-triggered reversals, in worms experiencing either early food deprivation (1 hour off food) or prolonged starvation (6 hours off food). Their experiments demonstrate a progressive loss of heattriggered reversals that is mediated by AWC and ASI neurons, as well as both glutamatergic and neuropeptidergic signaling.

      At the level of neural activity, this study reports that the transition from early food deprivation to prolonged starvation reconfigures the temperature-driven activity of AWC neurons from largely deterministic to stochastic. This finding is interesting in light of previous work that reported the opposite transition (from stochastic to deterministic) in temperature-driven AWC responses when comparing well-fed worms to those kept from food for 3 hours. This study also identifies neural and genetic mechanisms that contribute to differences in thermonociceptive responses at +1 versus +6 hours of starvation; confusingly, these mechanisms are partially distinct from those that contribute to differences in negative thermotaxis behaviors in well-fed and +3 hours of starvation worms (Takeishi et al, 2020). A limitation of this manuscript is that these differences are not particularly acknowledged or addressed, other than the hypothesis that independent mechanisms underlie negative thermotaxis versus thermonociceptive stimuli. However, this suggestion is not experimentally verified.

      We thank this reviewer for pointing to the interest of our work. The difference between previous work focusing on negative thermotaxis in the range of innocuous temperatures and our work focusing on thermo-nociceptive response is important and indeed deserves further clarification and a deeper discussion in the manuscript.

      Two major empirical evidence for a distinction between negative thermotaxis (as assessed in previous studies) and thermonociceptive plasticity (as assessed in our paradigm) were already included in the initial article version. First, we reported a decrease in average response in AWC neurons due to a shift in the distribution of response polarities from mostly up-response to a mix of ‘up-response’ and ‘downresponse’ after starvation, while previous results showed an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to starvation-evoked thermotaxis adaptation, ASI neurons are required to orchestrate starvation-evoked plasticity in thermonociception. These observations already indicate differences at the circuit and cellular level. For the revision, we conducted further experiments to address the molecular level. We tested ins-1 mutants (see also specific point 3 by reviewer 2, below) and deepened this aspect in the discussion section of the revised manuscript. Previous study found that INS-1 signaling from the intestine is a major mediator of negative thermotaxis plasticity. In contrast, our new data show that INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (see new Figure 6-supplement 1). Taken together these three lines of empirical evidence support the notion that negative thermotaxis and thermonociceptive starvation-evoked plasticity involves at least partially distinct mechanisms and it seems therefore inappropriate to qualify this notion as purely hypothetical.

      Modification in the revised manuscript include extended introduction about the known thermotaxis regulation mechanisms (Introduction section), new Figure 6-supplement 1 about ins-1 and accompanying text in the result section as well as extended discussion about these differences (Discussion section)

      Multiple additional aspects of this study make the results difficult to synthesize with existing knowledge, including

      (1) Differences in - and insufficient discussion of - the magnitude and kinetics of thermal stimuli;

      We have included a better description of the stimuli characteristics in the revised methods section. The discussion section was deepened to better emphasize that different types of thermal stimuli have been used in different studies.

      (2) This study's use of "heating power" rather than temperature values when presenting behavioral results;

      Thanks for noting this point, which was indeed an unnecessary complication in the result display of the initial manuscript. We have changed ‘power values’ to corresponding ‘temperature increase’ in the revised figures.

      (3) The use of +1 hours starvation as a baseline instead of well-fed worms. Indeed, this last point reflects a noticeable experimental result that differs from previous studies, namely that at room temperature, the basal movements of well-fed and starved worms are not different. Such a surprising result warrants further quantification of worm mobility in general and could have prompted a set of experiments directly testing previously published thermal conditions to demonstrate that the new effects reported arise specifically from the use of thermonociceptive stimuli, as hypothesized.

      The consideration of on-food and off-food behavioral state is an important point indeed. We found that, at room temp, worms shift from dwelling on food (a state with high spontaneous reversal rate) to global search off-food (a state with low spontaneous reversal, after 1hr starvation) (see Figure 1B). Therefore, unlike the reviewer’s statement, the reported data highlighted key differences in the basal locomotion of worms in fed and 1hr starved conditions. Furthermore, these behavioral states have been characterized very deeply using high-content worm behavioural tracking in our recent publication: Thapliyal et al. 2023 (PMID: 37236963). Our choice of using 1hr as a baseline is primarily driven by the fact that an elevated baseline of spontaneous reversals on food decreased the dynamic range to monitor changes in heat-evoked reversals. 1-hour early food deprivation reduced spontaneous reversals and led to a mild attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. Additionally, we observed a clear progressive decrease in heat-evoked reversals with increased duration of food-deprivation, which we further used to dissect the mechanism underlying this plasticity. The choice of the 1hr food deprivation timepoint as a reference is further justified below.

      Finally, a previous report (Yeon et al, 2021) demonstrated differences in the impact of chronic versus acute neural silencing on starvation-dependent plasticity in the context of negative thermotaxis. We therefore wonder whether similar developmental compensation impacts the neural circuits that contribute to starvation-dependent plasticity in the thermonociceptive responses.

      Indeed, this is an interesting question. Our conclusions are so far based on ablation (with chronic effects). In order to gain insight on this question, future studies could address the impact of chronic vs acute silencing approach in starvation-dependent thermonociceptive responses. We have added an opening on this question in the discussion, as follows:

      “Another open question is whether ASI action takes place during development (prior to starvation), or more acutely with active signaling after starvation.”

      A weakness of this manuscript is that the introduction is insufficiently scholarly in terms of citations and the description of current knowledge surrounding the impact of internal state on sensory behavior, particularly given previous work on the impact of feeding state on thermosensory behavioral plasticity (Takeshi et al 2020, Yeon et al 2021) and chemosensory valence (Banerjee et al 2023, Rengarajan et al 2019, etc).

      To address this weakness, we have revised the introduction section of the manuscript and cited previous relevant research on the impact of internal states on animal behavior, including the papers suggested by the reviewer. We note that 2 out of 4 suggested citations were already present in the initial manuscript (though in the discussion section).

      Similarly, the authors' commanding knowledge of the distinction between thermotaxis navigation (especially negative thermotaxis) and thermonociceptive behaviors could be communicated in more depth and clarity to the readers, in order to contextualize this study's new findings within the previous literature.

      As mentioned above, we have deepened this aspect in the discussion section of the manuscript (with a dedicated paragraph). It is quite clear that starvationinduced plasticity in negative thermotaxis and thermonociceptive behaviors engage distinct mechanisms (at least in part). These differences include distinct alterations in AWC calcium activity, role of ASI neurons and INS-1 neuropeptide.

      Nevertheless, this study represents a solid addition to the growing evidence that C. elegans sensory behaviors are strongly impacted by internal states, and that neuropeptidergic signaling plays a key role in mediating behavioral plasticity. To that end, the authors have provided solid evidence of their claims.

      We thank this reviewer for the efforts in evaluating our manuscript, for the positive assessment of our work, and for highlighting some weaknesses which, we believe, have been addressed through the revision.

      Reviewer #2 (Public review):

      In this work, Thapliyal and Glauser tried to provide a mechanistic understanding by which animals modulate their neural circuit responses to control nociceptive behavior on the basis of the dynamic internal feeding state. It is an important study that adds to the growing body of evidence coming from multiple model systems. They have used elegant genetics, behavioral, and Ca-imaging experiments to demonstrate how the auxiliary thermosensory neuron pair, AWC, and one of the internal state-sensing interneuron pairs, ASI, respond to dynamic internal starvation state to modulate behavioral response to noxious heat. Interestingly, these neuron pairs use distinct molecular mechanisms along with some other unidentified neurons to suppress heat-induced reversal response under short-term and prolonged starvation. The experiments are well performed, supporting most of the claims and providing an important framework for future studies.

      I have some queries that, if answered, will certainly enhance the study.

      (1) The results suggest that ASI is one of the primary drivers for the starvation-evoked behavioral plasticity, which regulates AWC activity under prolonged starvation. It raises many important questions, including: (a) how starvation modulates ASI response to heat?, and (b) under prolonged starvation, whether ASI also promotes other, non-AWC, glutamatergic inhibitory neurons to suppress heat-induced reversal, and how?

      We agree with this reviewer that the mechanisms by which ASI detects and mediates starvation-evoked changes in our model is a very interesting (unsolved) question. However, addressing these questions empirically represents a substantial body of work that would go beyond the scope of the present report. E.g., is temperature-dependent activity in ASI even relevant? At present, we envision that ASI could either work acutely (during heat stimuli) or be modulated over much longer time frames (hours of starvation) as an internal state sensor. Therefore, there will be quite some exploration needed before we figure out the ASI-level regulation more fully (including the critical temporal aspect regarding cell activity, as well as quantitative and qualitative transmission aspects). It will be very interesting in future work to address these questions.

      (2) How does ASI regulate AWC activity? In the proposed model (Figure 8) authors suggested an independent, unknown signal, other than INS-32 and NLP-18, from ASI to regulate AWC activity. However, from the results, the existence of another signal is not very clear.

      Thanks for raising this point, which reveals a weakness in our graphical representation (in Fig. 8) that was not properly conveying our point. Our current work shows INS-32 and NLP-18 to be important in modulating heat-evoked reversals upon starvation. However, at the moment, we don't know if INS-32, NLP-18, both, and/or other neuropeptides from ASI modulate AWC activity patterns. The calcium imaging experiments in single, double and potentially triple mutants would answer these questions but are not within our current reach, given the time needed to carry out these experiments. However, we acknowledge this point and have changed the figure and its legend to state that the arrow connecting ASI to AWC activity pattern could potentially reflect the action of these neuropeptides.

      (3) Previously, Takeishi et. al. showed that ins-1 dynamically modulates AWC-AIAmediated thermotaxis behavior based on the feeding state of the animal. It raises questions whether ins-1 also contributes to noxious heat-induced reversal behavior.

      We thank the reviewer for this question. We have now quantified the phenotype of ins-1 mutant in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity, unlike plasticity in thermotaxis behavior (See Figure 6- Supplement 1). Together with the differential activity patterns in AWC and the differential need for ASI neurons, these new data further consolidate the notion that starvation-evoked thermotaxis adaptation and noxious-heat avoidance engage separable molecular, cellular and circuit-level modulatory mechanisms. A specific discussion paragraph was added too.

      (4) Experiments with AWC fate conversion mutants (nsy-1 and nsy-7) were very good ideas; however, the results obtained were confusing. flp-6 mutant data suggest AWCoff would be essential for heat-induced reversal, especially at the low intensity stimulus level. However, the nsy-1 mutant-forming two AWCon neurons showed complete rescue at the low heat level, which is quite opposite. Similarly, although less prominent, eat-4 rescue experiments suggested both nsy-1 and nsy-7 should behave normally at high heat conditions, which was not the result observed.

      We appreciate this comment and the legit attempt to infer what we should expect from a worm with two AWCon or two AWCoff, respectively. From previous studies so far, it's not quite clear if cellular properties of newly formed AWCs in nsy-1 and nsy-7 mutants, including response to sensory cues, formed synapses and their partners, expression of neuromodulator and gap junctions, synaptic output are similar or different. We think further studies are required to first establish if FLP-6 and glutamate signaling (expression, release and action) from altered AWCs in nsy-1 and nsy-7 mutants are the same or different. Therefore, direct comparison between cell fate conversion mutants with flp-6 and glutamate would rely on too many assumptions at this stage. Considering this comment, the limited additional value of the data with nsy1 and nsy-7 mutants (in the absence of additional analyses) and the confusion it could trigger, we have decided to remove these non-essential data of the manuscript.

      Reviewer #3 (Public review):

      Summary:

      Thapliyal and Glauser show that hunger alters how C. elegans responds to noxious thermal stimuli. Using targeted neural ablation, mutant analysis, and live-cell functional imaging, the authors demonstrate that hunger changes the properties of AWC sensory neurons, which sense noxious heat. The authors further show that the effects of hunger on nociception require ASI neurons, which are known to respond to hunger and mediate the effects of food deprivation on behavior. Finally, the study uses mutant analysis to implicate glutamate and specific neuropeptides in thermal nociception and in the modulation of nociceptors by hungerresponsive neurons.

      Strengths:

      The study clearly shows a strong effect of hunger on nociception and documents a striking effect of hunger on the intrinsic properties of AWC sensory neurons, which respond to noxious heat. The study also clearly and compellingly demonstrates that ablation of hunger-responsive ASI neurons blocks the effects of hunger on nociceptive AWCs. These data, which constitute the kernel of the manuscript, are striking and exciting.

      Weaknesses:

      The study has some weaknesses that the authors should address.

      (1) Ablation of AWC neurons alters the basal sensitivity to noxious heat stimuli. This should be clearly noted in the description of the result and warrants some discussion.

      We thank this reviewer for raising this legitimate point. We have clarified this aspect in the results section of the revised manuscript, reading as follows:

      “Removal of AWC nearly abolished heat-evoked reversal behavior across all stimulus intensities and timepoints (Figure 2B and E). While one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines, this observation suggests that AWC plays an essential role in mediating the thermonociceptive response under both early food deprivation and prolonged starvation. Notably, in AWC-ablated animals, the residual response level was unaffected by starvation, suggesting that AWC might also be required for the expression of starvation-dependent plasticity.”

      The contrast with known function of the best-characterized sensory neurons mediating thermal nociception (AFD and FLP) is discussed as follows:

      “...Therefore, noxious heat-evoked activity in AWC varies widely according to context, which is in line with previous literature [18, 20, 41]. Interestingly, the role of AWC is distinct from that of AFD and FLP neurons, which are canonically linked to thermosensation and nociception [5, 14, 42, 43], but contribute only modestly to heat-evoked behavior in our assay conditions with between 1 and 6 hrs of food deprivation.”

      (2) Throughout the study, it seems that data are replotted in multiple figure panels. The authors should clearly indicate in the figure legends when this occurs. Also, the authors should ensure that statistical tests requiring multiple comparisons are correctly implemented and reflect the number of times experimental data are compared to a single set of control data.

      Thanks for raising this important point. We have clarified this aspect in the revised figure legends of the manuscript, and in the method section. In some instances, we reconducted some analyses to be perfectly rigorous in multiple comparison accounting. This did not lead to significantly different conclusions. The one exception was that the small effect of eat-4 mutation on spontaneous reversal went below significance threshold. We therefore removed this aspect of the result reporting and of the corresponding interpretation scheme, which became slightly simpler (Figure 3). Globally, this makes the story more focused.

      (3) How ASIs modulate AWCs remains unclear. The authors find that loss of INS-6, an insulin-like peptide provided by ASIs, partially recapitulates the effect of ASI ablation. This observation is not further developed, and instead, the authors characterize other secreted factors that seem to mediate sensitization of animals to noxious heat stimuli. While it is interesting that there are multiple opposing inputs into the nociceptor circuit, the essential connection between ASIs and AWCs that underlies the foundational observations in Figures 1 and 2 is not sufficiently characterized.

      Whereas we agree that how ASI modulates AWCs is only partially solved by our study, we should emphasize that our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal upregulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI ablated line, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint.

      Following this reviewer’s comment, we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data (with a negative result) remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS-6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required go deeper in this question, which could be performed in future work.

      “...Additional neuropeptides (such as INS-6) may also be involved, but in the absence of direct evidence for their origin from ASI, they were not included in this scheme.”

      (4) The assertion that 'starvation reshapes AWC responses from deterministic to stochastic' is not clearly supported by the data. AWC neurons seem capable of showing different responses to thermal stimuli, and the probabilities associated with these responses change after fasting. The different kinds of responses are seen under basal and fasted conditions.

      We thank this reviewer for the comment. There is an activity response shift that is quite solidly described, including with new quantitative analyses of distributions (histograms in new Fig. 4CD and new Fig. 5C-D, accompanied by Kruskal-Wallis tests). Yet, we totally agree that the wording choice was inappropriate. We have furthermore changed our terminology to avoid using the terms “stochastic” or “deterministic” that were indeed a cause of confusion. We now use the terms “stimulus-locked responses” and describe the shift as “shift from mostly excitatory responses to a mix of both excitatory and inhibitory responses”. We have also included detailed methodology for characterization of traces and statistical analysis in the revised method section of the manuscript, together with the new analyses on peak polarity distribution.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers agree that the study is clearly presented and makes good use of behavioral, genetic, and imaging approaches to link starvation state with changes in AWC and ASI function. To strengthen the manuscript and ensure clarity for readers, we ask you to address the following points in revision:

      (1) Positioning and citations.

      Clarify how your findings relate to Takeishi 2020, where the opposite trend in AWC activity was reported, and make a clear distinction between thermonociception and thermotaxis. The introduction should also include additional citations in two specific areas: prior work on AWC and noxious thermal stimuli, and studies demonstrating starvation-dependent behavioral changes via altered neuropeptide release (e.g., Banerjee 2023; Rengarajan 2019).

      We have clarified this aspect with extension of the work cited in the introduction and extensive rewriting of the discussion sections.

      Our data shows that mechanisms underlying starvation dependent changes in thermonociception and thermotaxis show differences at the molecular, cellular and circuit levels. First, we see a decrease in average response in AWC neurons due to shift from mostly excitatory to a mix of excitatory and inhibitory responses in response to noxious heat after starvation, while previous study found an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to thermotaxis behavior ASI neurons are required to orchestrate starvation evoked plasticity in thermonociception. And, finally, previous study found that INS-1 signaling from the intestine regulates thermotaxis behavioral plasticity while INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (new data in Figure 6 supplement 1).

      We have revised the introduction section of the manuscript and cited previous relevant research on AWC and noxious thermal stimuli and studies demonstrating starvationdependent behavioral changes via altered neuropeptide release including the papers suggested by reviewers. The extended discussion section reads as follows:

      “Starvation regulates thermonociceptive and negative thermotaxis plasticity via at least partly different mechanisms

      Previous studies showed that AWC plays an important role in starvationdependent plasticity in the negative thermotaxis behavior in an innocuous thermal range between 15 and 25°C [26, 33]. Negative thermotaxis involves the detection of thermal changes created by animal movement in spatial thermogradient (0.5°C/cm), the magnitude of the expected thermal changes approximating 0.01°C/s [26]. The starvation impact on negative thermotaxis was shown to (i) involve an up-regulation of AWC cell activity, (ii) rely on INS1 neuropeptide produced in the intestine and (iii) to occur independently of ASI neurons. In contrast, our study used thermo-nociceptive stimuli, with faster raising thermal slopes (~0.5-2°C/s, hence 50-200 times faster than those occurring for thermotaxis) and covering noxious temperatures (up to 28°C). Our results indicate that the regulation of thermo-nociceptive response by starvation (i) is linked to a shift in the distribution of AWC activity response polarities from mostly excitatory to a mix of excitatory and inhibitory response, (ii) relies on ASI and specific neuropeptide produced in ASI, and (iii) works independently of INS-1 neuropeptide. Therefore, our study complements our understanding of the modulation of temperature-dependent behavior in C. elegans with previously undocumented mechanisms at the circuit, cellular and molecular levels.”

      (2) ASI → AWC mechanism.

      Because ASI is central to your conclusions, please expand on how ASI is thought to act on AWC and/or other neurons. If an additional ASI signal is proposed beyond INS-32/NLP-18, mark this as speculative unless further rationale can be provided, and adjust the model figure accordingly.

      We have revised Figure 8 and its legend to clarify what is still hypothetical in the way ASI could affect AWC activity patterns and reversals. Note that the figure was also modified to integrate the conclusions made from epistasis analysis of eat-4 and flp-6.

      (3) AWC response description.

      The data support a shift in response distributions rather than a categorical switch from "deterministic to stochastic." Please adjust the language accordingly and provide a clear description of how traces were classified as "up, variable, or down," ideally with a quantification of the distributional shift.

      We agree that the term “stochastic” can convey different things, and because it was used for something different for AWC in the past, we should have avoided it. We have revised the nomenclature. What we observe can indeed be better described as a shift in the response polarity distribution. The article was revised accordingly. We also included the quantitative analysis and histogram representation, suggested in one of the specific comments, and added detailed methodology on the categorization of traces.

      When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (4) Presentation and statistics.

      In figure legends, indicate where datasets are replotted across panels and confirm that multiple-comparison corrections take account of repeated comparisons to the same controls.

      We have included these details in the revised figure legends, and a statement in the method section.

      (5) Methods clarity.

      Provide justification for using 1-h off-food as the baseline, with quantification of baseline mobility/reversal rates. Expand the calcium-imaging methods to describe the processing pipeline (ΔR calculation, baseline period, drift correction), and add a brief rationale if the approach deviates from common normalization procedures. Clarify how cell ablations were performed and verified for specificity, and how cell-specific rescues were confirmed. Please also acknowledge the potential for developmental compensation with chronic ablation.

      The justification of using 1hr off-food as baseline was made more prominent in the revised manuscript.

      Revised result section:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with much-reduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvation-dependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      In addition, the method section was modified as follows:

      - Calcium imaging details were added regarding ΔR calculation, baseline period, drift correction.

      - We now explicitly refer to the original respective articles describing the neuroablation lines.

      - We clarify that cell-specific transgene expression for rescue was confirmed using SL2::mCherry co-marker

      In the result section, we now explicitly address potential developmental compensation in genetic ablation backgrounds in the result section as follows: “…one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines”.

      Reviewer #1 (Recommendations for the authors):

      (1) The data availability statement is missing from the reviewed manuscript and should be included.

      Thanks, we have included the data availability statement in the revised manuscript.

      (2) We request additional information on how n's were determined for individual experiments, as well as the inclusion of post-hoc power measurements for all quantification.

      n were determined in agreement with previous studies using similar measures. No a priori power analyses were performed. A posteriori power analyses are not informative beyond the reported effect sizes and p-values (now reported in File S2). We clarified this in the statistical subsection of the method section.

      (3) In many cases, the specific statistical tests used are not clear or justified; more details should be provided, including the non-post-hoc test used. Are all tests one-way ANOVAs? For comparisons across genotype and starvation duration, two-way ANOVAs would likely be more appropriate. Also, the authors switch between Bonferroni post-hoc tests and Holm-Bonferroni post-hoc tests. What determined the use of one versus another?

      We have now clarified the statistical analyses used and provided full details in File S2. We have now more systematically applied two-way ANOVAs across all relevant analyses (with detailed parameters reported in File S2). When particularly relevant (e.g epistasis analysis between eat-4 and flp-6 mutations) the results of the two-way ANOVAs, is also explicitly stated in the result section.

      We also note that all multiple-comparison corrections were performed using the Bonferroni method. Previous mentions of Holm-Bonferroni correction were inaccuracies, and we apologize for this confusion; these mentions have now been corrected throughout the manuscript.

      (4) The use of heating power instead of the temperature experienced by the worms is an unwelcome abstraction. We strongly recommend revisiting that choice.

      We do agree. We have revised the figures to label the axis with temperature increase.

      (5) For calcium imaging, how are the traces categorized into "calcium up", "calcium down", or "no change"? Were those determined blindly - i.e., by individuals unaware of the experimental condition? Did the response direction need to be consistent across different temperatures? Did the change from baseline need to hit a specific threshold, consistent with previous studies in the field (i.e., +/- 3xSD for a minimum amount of time)? We encourage the authors to include these details in their methods section.

      We have complemented the method section to clarify the criteria for the qualitative classification of traces. More importantly, new quantitative peak polarities comparisons were added (see specific points below and above, about histograms).

      (6) For the experiments showing that exposure to food odor does not prevent response reduction, we suggest that feeding worms heat-killed bacteria would be a helpful control for the importance of bacterial nutritional status. In addition, showing that the impact of starvation was reversible with re-feeding would have been a useful experiment in line with standard experimental design in the starvation field.

      Thanks, indeed with our current work we cannot pinpoint the role of additional sensory cues (except food odor) to be mediating starvation-evoked plasticity. Together with refeeding, these are all extremely interesting questions that we aim to answer and potentially link with ASI and AWC activity in our future work.

      (7) For the various AWC rescue experiments, we found it curious that there wasn't an AWCon+off rescue, only each neuron individually.

      Previous studies have identified similar or opposite responses of both AWCs for distinct sensory cues. Though our calcium imaging experiments point to both AWC on and off having similar response patterns to heat, we cannot rule out the possibility that their output (ability of control reversals) is distinct possibly due to recruited neuromodulators. Therefore, in the present work, we examined where these cell types act via the same or distinct combinations of neuromodulators to control reversals.

      Reviewer #2 (Recommendations for the authors):

      Experiments suggested:

      (1) The authors should look into the Ca-dynamics in ASI.

      How does the spontaneous and heat-evoked activity of ASI differ in fed, early food-deprivation and prolong starvation and its link to releases of neuromodulators, modified AWC activity to alter output of thermal nociception are very interesting questions. However, these questions are extremely exploratory (see argumentation above in response to the public review) and addressing them goes beyond the scope of our current manuscript.

      (2) The authors should check AWC activity in ins-32 and nlp-18 mutant animals.

      In this study, we focused on the roles of INS-32 and NLP-18 released from ASI in modulating heat-evoked reversals, as these mutants exhibit relatively strong behavioral phenotypes. However, these effects remain less pronounced than those observed following ASI ablation. In addition, we cannot exclude the contribution of additional signaling molecules, including INS-6 and other neuropeptides.

      A comprehensive analysis of AWC activity in this context would require calcium imaging across multiple genetic backgrounds, including single, double, and potentially higher order peptide and receptor mutants, combined with cell-specific rescue experiments. While we appreciate the suggestion, such an approach would represent a substantial extension of the present work and will be important to pursue in future studies to further elucidate the underlying mechanisms.

      (3) Short-term food deprivation completely eliminated heat heat-induced reversal response to 100W stimulus, while the response to 400W stimulus remained unaffected. This suggests fed, short-term starvation, and prolonged starvation are three distinct states, and authors should also test the response of AWC and ASI ablated animals in the fed conditions.

      We agree that analyzing thermal nociception in fed states, in addition to short-term and prolonged food deprivation states is an important and interesting question, as these 3 conditions likely represent 3 distinct internal states that may recruit different neural pathways.

      Several reasons led us to set the fed condition aside for this study, and we realize we insufficiently explain them in the initial manuscript. There are 2 main reasons.

      (1) It is of paramount importance to consider the ‘baseline’ reversal rate (spontaneous reversals not triggered by heat, but visible in our dataset as the first point in the ‘dose-response’ curve). In Fed animals spontaneous reversal rate is very high (~38%) compared to the 1hr and 6hr food-deprivation conditions (>10%). This has two consequences: first a decreased dynamic range for quantify heat-evoked reversal, and, second, the difficulty in judging quantitative differences in heat-evoked reversals with such major differences in baseline reversals.

      (2) Experimentally, assessing calcium responses in truly fed animals presents technical challenges. With our current setup, animals must be removed from food for at least ~5 minutes prior to recording (followed by ~5 minutes of imaging), which effectively corresponds to a “freshly starved” condition rather than a fully fed state. While previous studies have used serotonin to mimic aspects of the fed state, such manipulations can be difficult to interpret in this context.

      A systematic comparison including fully fed animals, as well as AWC- and ASI-ablated conditions across these states, would be a valuable direction for future work, in particular once the methodological barriers associated with point 2, have been overcome.

      We have clarified these choices in the result section as follows:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with muchreduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvationdependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      (4) The authors should test the effect of ins-1 in noxious heat-mediated dynamic reversal behavior.

      We thank the reviewer for this valuable suggestion. We have now tested the phenotype of ins-1 mutants in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity unlike plasticity in thermotaxis behavior (New Figure 6 Sup1). This molecular aspect adds to our initially presented evidence at the cell activity and circuit levels, that noxiousevoked reversal and thermotaxis behaviors are regulated in a clearly separable manner.

      (5) Whether Glutamate and flp-6 work in parallel or in the same pathway to regulate reversals?

      We thank the reviewer for this question. We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat-evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4; flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup></sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig.

      8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food-deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (6) The authors should perform flp-6 and eat-4 mutant/rescue experiments in the nsy-1 and nsy-7 background to clarify the results.

      Our results indicate that glutamate release via EAT-4 from both AWC<sup>ON</sup> and AWC<sup>OFF</sup>, as well as FLP-6 from AWC<sup>OFF</sup>, contributes to heat-evoked reversals regulation. The experiments suggested by the reviewer would, in principle, provide further insight into the interaction between AWC subtype identity and the respective roles of glutamatergic and peptidergic signaling.

      However, as discussed in more details above in the public review, the extent and nature of AWC<sup>ON/OFF</sup> remodeling in nsy-1 and nsy-7 mutant backgrounds remain incompletely understood. This introduces significant uncertainty in interpreting results obtained from combining these mutations with eat-4 and flp-6 manipulations. As a result, such experiments would be difficult to interpret in a definitive manner at this stage.

      We therefore consider this an important direction for future work, once the roles of nsy1 and nsy-7 in AWC subtype specification and function are more clearly established. As our preliminary results with nsy-1 and nsy-7 mutants added more confusion than clarity, we have chosen to set them aside (former Fig. 3-figure supplement 2 has been removed).

      Minor comments:

      (1) What is food odor? The experiment should be clearly mentioned.

      Thanks for spotting this unintended omission. Food odor experiments were performed by adding OP50 bacteria on the inward side of the petri dish lid instead of the NGM surface. We have added this description in the method section of the revised manuscript.

      (2) Panel 3C is coming before 3B. This should be rearranged.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 3, with the addition of eat-4 x flp-6 genetic interaction analysis.

      (3) In Figure 6, if the panels are arranged horizontally, it would be easier to follow.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 6, with the addition of eat-4 flp-6 genetic interaction analysis.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors should consider moving measurements of AFD-ablated animals into the main Figure 1. AFD is a well-known thermosensor, and it is worth showing that responses to noxious thermal stimuli persist in animals lacking AFD.

      Thank you for this suggestion. We have moved measurements of AFDablated animals to the main figure (revised Fig. 2).

      (2) AFD ablation does affect responses to noxious heat. The authors could consider ablating/silencing AFDs and AWCs simultaneously to determine whether these two neurontypes account for the behavior.

      We agree that investigating the combinatorial contributions of thermosensory neurons, including AFD and AWC, to thermal nociception is an important and interesting question. In principle, simultaneous ablation or silencing of these neuron types could indeed reveal unexpected interactions.

      In our experimental paradigm, however, we observe only a minimal contribution of AFD neurons to heat-evoked reversals, whereas ablation of AWC nearly abolishes the response. Based on these observations, we chose to focus the present study on AWC, which appears to play a more prominent role in this behavior.

      A more detailed dissection of the potential interactions between AFD and AWC, including combinatorial manipulations, would be a valuable direction for future work. In particular, the possibility that AFD exerts a modulatory influence remains an interesting hypothesis to explore.

      (3) Given that EAT-4/VGLUT and FLP-6 neuropeptides each contribute to nociception, the authors should consider testing an eat-4; flp-6 double mutant to determine whether this combination of neurochemical signals accounts for AWC signaling to downstream circuits.

      We thank the reviewer for this suggestion, which is similar to point 5 of Reviewer 2 (above).

      We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4;flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup>2</sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig. 8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (4) The authors should consider representing AWC responses to thermal stimuli as histograms to illustrate how fasting increases the probability of some responses and decreases the probability of others.

      We thank this reviewer for the suggestion. The proposed histograms nicely convey the concept of “shift in response polarity distribution” that we observed (using the new terminology we now use, instead of using the term “stochastic”). To create such histograms, we computed the magnitude of the peaks on a trial-by-trial basis. When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (5) It seems important to better understand the ins-6 mutant phenotype and determine whether ASI-to-AWC signaling involves this insulin-like peptide (ILP). The authors should consider using some of the tools available for disrupting ILP signaling to more clearly demonstrate that a specific neurochemical signal mediates modulation of AWCs by ASIs.

      Our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal up-regulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint (as follow up data indicated at least for nlp-18).

      Following this reviewer’s comment (and the similar one in the public review), we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required to go deeper in this question, which could be performed in future work.

    1. eLife Assessment

      This important study shows that an odorant that is typically thought of as a repellant actually activates both attractant and repellant olfactory neurons in C. elegans. Convincing evidence is provided that nematode worms can integrate signals in different sensory pathways to drive different behavioral responses to the same cue. These findings will be of interest to scientists interested in combinatorial coding in sensory systems.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      The authors investigated the response of worms to the odorant 1-octanol (1-oct) using a combination of microfluidics-based behavioral analysis and whole-network calcium imaging. They hypothesized that 1-oct may be encoded through two simultaneous, opposing afferent pathways: a repulsive pathway driven by ASH, and an attractive pathway driven by AWC. And the ultimate chemotactic outcome is likely determined by the balance between these two pathways.

      It is not surprising that 1-octanol is encoded as attractive at low concentrations and repulsive at higher concentrations. However, the novel aspect of this study is the discovery of the combinatorial coding of 1-oct in the periphery, where it serves as both an attractant and a repellent. Furthermore, the study uses this dual encoding as a model to explore the neural basis of sensory-driven behaviors at a whole-network scale in this organism. The basic conclusions of this study are well supported by the behavioral and imaging experiments, though there are certain aspects of the manuscript that would benefit from further clarification.

      A key issue is that several previous studies have demonstrated a combinatorial and concentration-dependent coding of odorant sensing in the nematode peripheral nervous system. Specifically, ASH and AWC are the primary receptors for repellent and attractive responses, respectively. However, other neurons such as AWB, AWA, and ADL are also involved in the coding process. These neurons likely communicate with different interneurons to contribute to 1-oct-induced outputs. The authors' conclusion that loss of tax-4 reduces attractive responses and that osm-9 mutants reduce repulsive responses is not entirely convincing. TAX-4 is required for both AWC (an attractive neuron) and AWB (a repulsive neuron), and osm-9 is essential for ASH, ADL, and AWA (attraction-associated). Therefore, the observed effects on the attractive and repulsive responses could be more complex. Additionally, the interpretation of results involving the use of IAA to reduce the contribution of AWC at lower concentrations lacks clarity.

      The authors did not observe any increased correlation between motor command interneurons and sensory neurons, which is consistent with the absence of a consistent relationship between state transitions and 1-oct application. Furthermore, they did not observe significant entrainment of AIB activity with the 2.2 mM 1-oct application. This might be due to the animals being anesthetized with 1 mM tetramisole hydrochloride, which could affect neural activity and/or feedback from locomotion.

    3. Reviewer #2 (Public review):

      Summary:

      The authors used whole-network imaging to identify sensory neurons that responded to the repellant 1-octanol. While several olfactory neurons responded to the initial onset of odor pulses, two neurons consistently responded to all the pulses, ASH and AWC. ASH typically activates in response to repellants, and AWC typically activates in response to the removal of attractants. However, in this case, AWC activated in response to the removal of 1-octanol, which was unexpected because 1-octanol is a harmful repellant to the worm. The authors further investigated this phenomenon by testing different concentrations of 1-octanol in a chemotaxis assay and found that at lower (less harmful) concentrations the odor is actually an attractant, but becomes repulsive at higher concentrations. The amplitude of the ASH response appeared to be modulated by concentration, but this was not true for AWC. The authors propose a model where the behavioral response of the worm is the result of integrating these two opposing drives, where repulsion is a result of the increased ASH activity over-riding the positive drive from AWC. The authors further tested this theory by testing mutants that ablated the AWC response (tax-4 or AWC::HisCl) or ASH response (osm-9 or ASH::HisCl). The chemo-silencing (HisCl) and tax-4 experiments were consistent with their hypothesis, while the osm-9 mutation had a limited impact on chemotaxis behavior, highlighting the potential role of osm-9-independent signaling in ASH in response to 1-octanol. While the interneuron(s) that integrate these signals to influence behavior were not identified, the authors did find that increasing concentrations of 1-octanol did increase the likelihood of AVA activity, a neuron which drives reversals (and hence, behavioral repulsion).

      Strengths:

      This was simple and elegant work that identified specific neurons of interest which generated a hypothesis, which was further tested with mutants that altered neuronal activity. The authors performed both neuronal imaging and behavioral experiments to verify their claims.

      Weaknesses:

      The authors note that other sensory neurons likely contribute to 1-octanol chemotaxis. Given the NeuroPAL data, it would have been nice to identify these other neurons as well. However, the reviewer is aware that this is tangential to the primary focus of this study.

    4. Reviewer #3 (Public review):

      Summary:

      This work describes how two chemosensory neurons in C. elegans drive opposite behaviors in response to a volatile cue. Because they have different concentration dependencies, this leads to different behavioral responses (attraction at low concentration and repulsion at high concentration). It has been known that many odorants that are attractive at low concentrations are aversive at high concentrations, and the implicated neurons (at least AWC for attraction and ASH for repulsion) have been well established. Nonetheless, by studying behavior and neural responses in a common context (odor pulses, as opposed to gradients) this provides a clear picture of how these sensory neurons may guide the dose dependent response by separately modulating odor entry and odor exit behaviors.

      Strengths:

      (1) This work provides good evidence that worms are attracted to low concentrations and repelled by high concentrations of 1-oct. Calcium imaging also makes it clear that dose-dependence of this response is stronger for ASH than AWC.

      (2) This work presents calcium imaging and behavior with the same stimulus (sudden pulses in volatile odor concentration), while previous studies often focus on using neuronal responses to pulses to understand navigation of gentle gradients.

      Weaknesses:

      (1) As a whole it is not clear precisely how important AWC is (compared to other cells) for the attractive response (as the authors correctly acknowledge).

      (2) The evidence that AIB minus AVA contains relevant information is weak. It appears the entrainment index in Fig. 6H for AIB-AVA could easily be explained by the negative entrainment between AVA and the stimulus (along with no effect or role for AIB). This is suggested by the similar p-values and similar distribution of random EIs (stretched and mirrored) between the first and last rows of this figure.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The authors investigated the response of worms to the odorant 1-octanol (1-oct) using a combination of microfluidics-based behavioral analysis and whole-network calcium imaging. They hypothesized that 1-oct may be encoded through two simultaneous, opposing afferent pathways: a repulsive pathway driven by ASH, and an attractive pathway driven by AWC. And the ultimate chemotactic outcome is likely determined by the balance between these two pathways.

      It is not surprising that 1-octanol is encoded as attractive at low concentrations and repulsive at higher concentrations. However, the novel aspect of this study is the discovery of the combinatorial coding of 1-oct in the periphery, where it serves as both an attractant and a repellent. Furthermore, the study uses this dual encoding as a model to explore the neural basis of sensory-driven behaviors at a whole-network scale in this organism. The basic conclusions of this study are well supported by the behavioral and imaging experiments, though there are certain aspects of the manuscript that would benefit from further clarification.

      A key issue is that several previous studies have demonstrated a combinatorial and concentration-dependent coding of odorant sensing in the nematode peripheral nervous system. Specifically, ASH and AWC are the primary receptors for repellent and attractive responses, respectively. However, other neurons such as AWB, AWA, and ADL are also involved in the coding process. These neurons likely communicate with different interneurons to contribute to 1-oct-induced outputs. The authors' conclusion that loss of tax-4 reduces attractive responses and that osm-9 mutants reduce repulsive responses is not entirely convincing. TAX-4 is required for both AWC (an attractive neuron) and AWB (a repulsive neuron), and osm-9 is essential for ASH, ADL, and AWA (attraction-associated). Therefore, the observed effects on the attractive and repulsive responses could be more complex. Additionally, the interpretation of results involving the use of IAA to reduce the contribution of AWC at lower concentrations lacks clarity.

      The authors did not observe any increased correlation between motor command interneurons and sensory neurons, which is consistent with the absence of a consistent relationship between state transitions and 1-oct application. Furthermore, they did not observe significant entrainment of AIB activity with the 2.2 mM 1-oct application. This might be due to the animals being anesthetized with 1 mM tetramisole hydrochloride, which could affect neural activity and/or feedback from locomotion.

      Comments on revisions:

      The authors have addressed all my previously raised concerns.

      Reviewer #2 (Public review):

      Summary:

      The authors used whole-network imaging to identify sensory neurons that responded to the repellant 1-octanol. While several olfactory neurons responded to the initial onset of odor pulses, two neurons consistently responded to all the pulses, ASH and AWC. ASH typically activates in response to repellants, and AWC typically activates in response to the removal of attractants. However in this case, AWC activated in response to the removal of 1-octanol, which was unexpected because 1-octanol is a harmful repellant to the worm. The authors further investigated this phenomenon by testing different concentrations of 1-octanol in a chemotaxis assay, and found that at lower (less harmful) concentrations the odor is actually an attractant, but becomes repulsive at higher concentrations. The amplitude of the ASH response appeared to be modulated by concentration, but this was not true for AWC. The authors propose a model where the behavioral response of the worm is the result of integrating these two opposing drives, where repulsion is a result of the increased ASH activity over-riding the positive drive from AWC. The authors further tested this theory by testing mutants that ablated the AWC response (tax-4 or AWC::HisCl) or ASH response (osm-9 or ASH::HisCl). The chemo-silencing (HisCl) and tax-4 experiments were consistent with their hypothesis, while the osm-9 mutation had a limited impact on chemotaxis behavior, highlighting the potential role of osm-9-independent signaling in ASH in response to 1-octanol. While the interneuron(s) that integrate these signals to influence behavior were not identified, the authors did find that increasing concentrations of 1-octanol did increase the likelihood of AVA activity, a neuron which drives reversals (and hence, behavioral repulsion).

      Strengths:

      This was simple and elegant work that identified specific neurons of interest which generated a hypothesis, which was further tested with mutants that altered neuronal activity. The authors performed both neuronal imaging and behavioral experiments to verify their claims.

      Weaknesses:

      The authors note that other sensory neurons likely contribute to 1-octanol chemotaxis. Given the NeuroPAL data, it would have been nice to identify these other neurons as well. However, the reviewer is aware that this is tangential to the primary focus of this study.

      Reviewer #3 (Public review):

      Summary:

      This work describes how two chemosensory neurons in C. elegans drive opposite behaviors in response to a volatile cue. Because they have different concentration dependencies, this leads to different behavioral responses (attraction at low concentration and repulsion at high concentration). It has been known that many odorants that are attractive at low concentrations are aversive at high concentrations, and the implicated neurons (at least AWC for attraction and ASH for repulsion) have been well established. None the less, by studying behavior and neural responses in a common context (odor pulses, as opposed to gradients) this provides a clear picture of how these sensory neurons may guide the dose dependent response by separately modulating odor entry and odor exit behaviors.

      Strengths:

      (1) This work provides good evidence that worms are attracted to low concentrations and repelled by high concentrations of 1-oct. Calcium imaging also makes it clear that dose-dependence of this response is stronger for ASH than AWC.

      (2) This work presents calcium imaging and behavior with the same stimulus (sudden pulses in volatile odor concentration), while previous studies often focus on using neuronal responses to pulses to understand navigation of gentle gradients.

      Weaknesses:

      (1) As a whole it is not clear precisely how important AWC is (compared to other cells) for the attractive response (as the authors correctly acknowledge).

      (2) The evidence that AIB minus AVA contains relevant information is weak. It appears the entrainment index in Fig. 6H for AIB-AVA could easily be explained by the negative entrainment between AVA and the stimulus (along with no effect or role for AIB). This is suggested by the similar p-values and similar distribution of random EIs (stretched and mirrored) between the first and last rows of this figure.

      (3) The model in Figure 7 would be strengthened if it was demonstrated that IAA is attractive when worms are saturated in a 1/10^4 concentration. Panel 7G (and ref. 39) indicate that 10^-4 IAA activates ASH, which would suggest a different explanation for the change from attraction to repulsion in 7C.

      It was previously published that 1x10^-4 IAA is attractive in a similar microfluidics context (Albrecht and Bargmann), and we have confirmed its attractiveness in our experimental setup (not shown). Specifically, worms accumulate in zones of the arena where 1x10^-4 IAA is present: they are attracted upon first encounter and maintain attraction after prolonged exposure. A sentence to this effect has been added to Results (line 402).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      This looks great! My only small recommendation is to include the 1-octanol concentration used (2.2 mM) in Fig 4 F,G so that the reader can easily compare these data to Fig 4 D,E. You can do this in the figure or legend, though in the figure would be preferable (as you did for Fig 5 and Fig 6).

      [1-oct ] added to Fig. 4D, E

      Reviewer #3 (Recommendations for the authors):

      (1) The AWC traces in Fig. 3C do not match, despite them being from the same animal. There appear to be too many ASH traces. NeuroPAL data/traces from other animals were not available on Zenodo when this review was prepared. This should be fixed.

      Single-animal traces added to figure so reader can see the relationship in the single animal shown and across the full dataset. The averaged ASH/AWC traces are retained in the updated figure because this observation is central to the main hypothesis, so it is important to demonstrate it is a consistent phenomenon. Legend modified (lines 1016-17), ‘consistent’ added to Results (line 147). Data will be uploaded to Zenodo upon finalization of the Version of Record

      (2) The normalization of traces in Fig. 4 is not clear (panel A is not min-max, as panel B appears to be). This normalization may be critical to the interpretation that ASH responds dose-dependently. The code linked on Zenodo was not available when this review was prepared. This should be fixed.

      The reviewer correctly points out a slight error in normalization of the individual worm traces for the ascending concentration series (4A, left panel; the maximal value shown was slightly under 1.0). This is now corrected, interpretation unchanged.

      (3) Chemogenetic data (4F,G and 5D) should specify the promoters used, as they are not ASH and AWC-specific.

      Figures 4F, G and 5D altered to specify the promoters used, as requested. As before, the full range of neurons in which these promoters are active is stated in the Discussion

    1. eLife Assessment

      Feeding, the circadian rhythm, and the gut microbiota are all intimately linked, motivating new approaches to identify causal relationships while minimizing confounding factors. The authors employ an innovative combination of the osmotic laxative lactulose and a defined 3-member gut microbiota to acutely induce gut bacterial metabolism in mice during the daytime, resulting in changes in the ileal expression of clock genes and altered feeding behavior. Together, this study utilizes convincing methods to provide important new insights into the role of gut microbiota in the circadian rhythm, setting the stage for follow-on studies aimed at better understanding the mechanisms responsible.

    2. Reviewer #2 (Public review):

      Summary

      The authors aimed to investigate how microbial metabolites, including hydrogen and short-chain fatty acids, influence feeding behavior and clock-gene expression in mice. Specifically, they examined these effects across different microbial environments, including a reduced-community model, germ-free mice, and specific-pathogen-free mice. The study addresses an important and poorly understood question concerning how microbial metabolism may influence host daily rhythms and feeding patterns.

      Strengths

      The manuscript presents a thoughtful and innovative investigation into the relationship between microbial metabolism, feeding behavior, and host clock-gene expression. A major strength is the use of a reduced microbial community together with real-time measurements of hydrogen production, which provides a useful experimental framework for separating microbial metabolic activity from direct host nutrient intake. The inclusion of germ-free and specific-pathogen-free mice also allows the authors to examine how these effects depend on microbial complexity.

      The revised manuscript has addressed several concerns raised in the original review. The authors have clarified aspects of stool collection, provided additional methodological information, refined their use of the terms "circadian" and "diurnal," and added experiments examining the osmotic effects of lactulose across different microbial environments. These revisions improve the clarity and interpretation of the work. The finding that lactulose-induced microbial activity is associated with altered feeding behavior and clock-gene expression in the reduced-community model, but not consistently in germ-free or specific-pathogen-free mice, remains intriguing and highlights the complexity of microbial-host interactions.

      Weaknesses

      Despite these improvements, several important concerns remain unresolved. Most notably, the response regarding excluded food-intake measurements is insufficient. The authors report that the probability of excluding measurements differed significantly between treatment and control groups in the key feeding experiment shown in Figure 4A/S5A. However, they do not provide the underlying number or percentage of excluded observations, the direction of the imbalance, how many animals were affected, whether the exclusions occurred before or after treatment, or which exclusion criteria accounted for the removed values. Because food-intake rate was calculated from repeated measurements of cumulative intake, differential exclusion of observations could influence both the estimated slope and the reported treatment effect. The authors should provide these details and demonstrate that the principal feeding result is robust to the exclusion procedure.

      There is also ambiguity surrounding the analysis and presentation of the quantitative polymerase chain reaction data. Figure 3B-C and Supplementary Figure 4A-C use conflicting mathematical labels, the raw cycle-threshold and replicate-level delta cycle-threshold values are not provided, and the rebuttal appears to conflate logarithmic transformation with exponentiation. I suspect that this may primarily reflect terminology or figure-labeling errors rather than an incorrect underlying analysis. The negative values shown in Figure 3 suggest that the authors may have plotted negative delta-delta cycle-threshold values, corresponding to log2 fold change. Nevertheless, the current description makes the workflow difficult to verify.

      Several mechanistic and interpretive issues were acknowledged but only partially addressed. The authors appropriately softened their interpretation of the clock-gene findings and clarified that hormone concentrations were measured at only one time point. However, they did not provide baseline evidence that the positive and negative limbs of the clock-gene network are normally in counterphase, did not discuss the possible involvement of AMP-activated protein kinase, and did not address the limitation of measuring total rather than active glucagon-like peptide 1. These issues should be explicitly discussed as limitations of the current study and as priorities for future work.

      Overall assessment

      The authors have mostly achieved their aims by providing novel evidence that acute changes in microbial metabolism may influence feeding behavior and clock-gene expression in a simplified microbial environment. The experimental approach is creative, and the study has the potential to contribute meaningfully to the fields of microbiome research and circadian biology. However, unresolved concerns regarding differential data exclusion and the transparency of the gene-expression analysis currently limit confidence in the strength of the evidence supporting the principal conclusions.

      Addressing these reporting and analytical issues would substantially strengthen the manuscript and improve its value to researchers studying microbial metabolism, feeding behavior, and host biological rhythms.

    3. Reviewer #3 (Public review):

      Summary:

      In the manuscript by Greter, et al., entitled "Targeted induction of gut-microbial metabolism acutely affects feeding patterns and clock gene expression in the host" the authors investigate whether acute exposure to a non-nutritive disaccharide (lactulose) promotes microbial metabolism that feeds back onto the host to impact circadian networks. The premise of the study is interesting, and the experiments are thoughtfully designed to dissect these relationships. The evidence presented generally supports the authors' conclusions regarding the impact of lactulose administration during the fasting period, which is intended to mimic a feeding-associated perturbation of the gut microbiota, and its comparison with lactulose administration during the fed state. The studies employ complementary model systems, including germ-free mice, mice colonized with a simplified three-member microbial community, and conventionally colonized animals. These approaches support the authors' conclusions regarding the relationship between diurnal rhythms of microbial fermentation and host circadian clock gene networks. Overall, the work provides a useful experimental framework for developing a deeper mechanistic understanding of how microbial fermentation products contribute to diurnal host-microbe interactions.

      Strengths:

      Attempting to disentangle nutrient acquisition from microbial fermentation and its impact on diurnal dynamics of gut microbes on host circadian rhythms is an important step for providing insights into these host-microbe interactions.

      The authors utilize a novel approach in leveraging lactulose coupled with germ-free animals and metabolic cages fitted with detectors that can measure microbial byproducts of fermentation, particularly hydrogen, in real time.

      The authors consider several interesting aspects of lactulose delivery, including how it shifts osmotic balance as well as providing calculations that attempt to explain the caloric contribution of fermentation to the animal in the context of reduced food intake. This provides interesting fundamental insights into the role of microbial outputs on host metabolism.

      The authors employ complementary systems, including a simplified three-member microbial community, providing insight into the minimal set of functionally distinct community members necessary to promote the rhythmic production of fermentation products that can affect host physiology.

      Residual limitations:<br /> Hypothesis and study framing: The manuscript still does not clearly articulate a specific, testable hypothesis. While the Introduction provides motivation and objectives (e.g., line 53 onward), it remains unclear what precise hypothesis was being evaluated. A more explicit statement would strengthen the conceptual framework of the study.

      Interpretation of circadian gene expression changes: The authors have not fully reconciled the differing effects of lactulose treatment on circadian gene expression in the 3MM and SPF settings. In particular, it remains unclear how the increased expression of certain circadian genes observed in lactulose-treated 3MM mice, particularly Cry1, relates to the decreased expression seen in SPF mice, and how the reduction in Arntl expression observed in lactulose-treated SPF mice fits within the proposed model. The authors acknowledge that resolving these mechanistic differences is beyond the scope of the current study, but the limitation should be discussed more explicitly.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We thank the three reviewers for their encouraging and constructive comments. We have addressed them by increasing clarity of the writing and adding details to the methods section that were previously missing or not stated clearly enough. We have included additional experimental data. Specifically, we tested how the osmotic effect of lactulose depends on lactulose dosage and colonization state, added data on cecum sizes of mice with different microbiomes, and tested how lactulose treatment in the active phase affects feeding behavior.

      Public Reviews:

      Reviewer #1 (Public Review):

      Greter et al. provide an interesting and creative use of lactulose as a "microbial metabolism" inducer, combined with tracking of H2 and other fermentation end products. The topic is timely and will likely be of broad interest to researchers studying nutrition, circadian rhythm, and gut microbiota. However, a couple of moderate to major concerns were noted that may impact the interpretation of the current data:

      (1) Much of the data relies on housing gnotobiotic mice in metabolic cages, but I couldn't find any details of methods to assess contamination during multiple days of housing outside of gnotobiotic isolators/cages. Given the complexity of the metabolic cage system used, sterility would likely be incredibly challenging to achieve. More details needed to be included about how potential contamination of the mice was assessed, ideally with 16S rRNA gene sequencing data of the endpoint samples and/or qPCR for total colonization levels relative to the more targeted data shown.

      We thank the reviewer for pointing out that we have not made the experimental setup clear in the text. One of the unique features of our metabolic cage setup is that the mice do not need to be housed outside gnotobiotic isolators, but that the whole system is placed inside an isolator. We have developed and published this system recently (Hoces et al, PLOS Biol 2022), including extensive testing for sterility/gnotobiosis. We have now adapted the main text to increase clarity on this issue (lines 76ff).

      Given that 16S sequencing of germ-free mice will typically produce false-positive reads, we used Blautia pseudococcoides as an indicator strain for contamination. This strain is present in our SPF mouse colony, forms spores that are highly resilient to decontamination measures, and has been the most likely contaminant in our gnotobiotic system. We have checked for presence of this strain in the cecum content of all our animals at the end of each experiment, and only included experiments which had a B. pseudococcoides signal below threshold level. We have now added this information to the methods section (lines 389ff).

      (2) The language could be softened to provide a more nuanced discussion of the results. While lactulose does seem to induce microbial metabolism it also could have direct effects on the host due to its osmotic activity or other off-target effects. Thus, it seems more precise to just refer to lactulose specifically in the figure titles and relevant text.

      We have adapted all figure legends to not contain interpretations, but rather state what was done in the experiments shown in the figures. We have also adapted the text in multiple places to soften the language and avoid overinterpretation of our experimental results.

      Additionally, the degree to which lactulose "disrupts the diurnal rhythm" isn't clear from the data shown, especially given that the markers of circadian rhythm rapidly recover from the perturbation. It is probably more precise to instead state that lactulose transiently induces fermentation during the light phase or something to that effect.

      We tried to make the argument that what we call disruption of the diurnal rhythm is acute, meaning that it is not disrupting the rhythm "chronically" (i.e., for longer), but that it recovers rapidly from this transient disruption. Given the confusion this wording is causing we are introducing this conceptually in the new version of the manuscript (lines 56ff).

      The discussion could also be expanded to address what methods are available or could be developed to build upon the concepts here; for example, the use of genetic inducers of metabolism which may avoid the more complex responses to lactulose.

      We also appreciate the mention of concepts from our study that can be built on in future studies, and we added a paragraph on potential further research. (lines 301ff).

      Despite these concerns, this was still an intriguing and valuable addition to the growing literature on the interface of the microbiome and circadian fields.

      We thank the reviewer for all their encouraging and constructive remarks!

      Reviewer #2 (Public Review):

      Summary:

      The authors aimed to investigate how microbial metabolites, such as hydrogen and short-chain fatty acids (SCFAs), influence feeding behavior and circadian gene expression in mice. Specifically, they sought to understand these effects in different microbial environments, including a reduced community model (EAM), germ-free mice, and SPF mice. The study was designed to explore the broader relationship between the gut microbiome and host circadian rhythms, an area that is not well understood. Through their experiments, the authors hoped to elucidate how microbial metabolism could impact circadian clock genes and feeding patterns, potentially revealing new mechanisms of gut microbiome-host interactions.

      Strengths: 

      The manuscript presents a well-executed investigation into the complex relationship between microbial metabolites and circadian rhythms, with a particular focus on feeding behavior and gene expression in different mouse models. One of the major strengths of the work lies in its innovative use of a reduced community model (EAM) to isolate and examine the effects of specific microbial metabolites, which provides valuable insights into how these metabolites might influence host behavior and circadian regulation. The study also contributes to the broader understanding of the gut microbiome's role in circadian biology, an area that remains poorly understood. The experiments are thoughtfully designed, with a clear rationale that ties together the gut microbiome, metabolic products, and host physiological responses. The authors successfully highlight an intriguing paradox: the significant influence of microbial metabolites in the EAM model versus the lack of effect in germ-free and SPF mice, which adds depth to the ongoing exploration of microbial-host interactions. Despite some methodological concerns, the manuscript offers compelling data and opens up new avenues for research in the field of microbiome and circadian biology.

      We thank the reviewer for their encouraging remarks, specifically on the surprising findings that microbial metabolism seems to affect circadian clock gene expression and behavior differently in EAM and SPF mice.

      Weaknesses:

      The manuscript, while providing valuable insights, has several methodological weaknesses that impact the overall strength of the findings. First, the process for stool collection lacks clarity, raising concerns about potential biases, such as the risk of coprophagia, which could affect the dry-to-wet weight ratio analysis and compromise the validity of these measurements.

      We thank the reviewer for pointing out that our description of the specific methods used for collecting feces were presented in a somewhat confusing manner. In short, dry and wet faecal weights were determined based on fecal pellets that were freshly produced and directly collected from restrained mice. To determine total fecal output over time, we collected all fecal pellets produced in a 5-hour window in a cage, determined their dry weight, and then used the water content determined for fresh faeces to calculate wet weight. Using this method, we cannot account for potential differences in coprophagia between the groups. However, this is not likely to affect the dry-to-wet ratio of faecal output in our results. We have now adapted the section in the methods to increase clarity (lines 440ff), and changed the quantity shown in Figure S2C and E to "water content", which is a more intuitive measure for the same thing.

      Additionally, the use of the term "circadian" in some contexts appears inaccurate, as "diurnal" might be more appropriate, especially given the uncertainty regarding whether the observed microbiome fluctuations are truly circadian.

      Similarly to our answer to reviewer 1 above, we appreciate this remark about imprecise language and have addressed this issue in the text and the figure legends. Indeed, we do not think the fluctuations in microbiota activity are truly circadian, but likely a result of the entrainment through the host's food intake.

      Another significant issue is the unexpected absence of an osmotic effect of lactulose in EAM mice, which contradicts the known properties of lactulose as an osmotic laxative. This finding requires further verification, including the use of a positive control, to ensure it is not artifactual.

      This is a good point. We have used this lactulose dosage specifically to induce microbial metabolism without causing osmotic diarrhoea and went to some lengths do demonstrate this (FigureS2C-E). In response to this comment (and one by reviewer 3 below about transit time), we have now performed additional experiments using higher lactulose dosage (new FigureS3). Our results indicate that the effect of lactulose on water content and transit time depends on microbiota complexity, with no change in fecal water content in 3MM mice even when treated with higher lactulose doses, and a stronger change in SPF mice. We now address this in the main text (lines 127ff).

      The presentation of qRT-PCR data as log2-fold changes, with a mean denominator, could introduce bias by artificially reducing variability, potentially leading to spurious findings or increased risk of Type I error. This approach may explain the unexpected activation of both the positive and negative limbs of the circadian clock.

      While we agree that our description of the qPCR method used for measuring circadian clock gene expression was lacking detail, we do not see how our analysis would lead to an increased risk of Type 1 error.

      Briefly, we use the standard ΔΔCt method to analyze our RT-PCR results. We first normalize gene expression values for each gene of interest to an internal housekeeping gene. Then, we use these normalized values to compare gene expression values in treatment vs control groups (or to the control group at time point 0 in the case of Figure 3C). We then use a log2 transformation to convert the logarithmic RT-PCR values to fold changes. We apologize for the confusing labeling in the figures, where we called the values "log2(fold changes)", which we have now changed to "log2(ΔΔCt) of expression" in Figures 3B, C and S4.

      The simultaneous activation of both limbs of the circadian clock is indeed a surprising result and somewhat complicates interpreting the effect of lactulose treatment on clock gene expression. We take it as evidence that it generally interferes with clock gene expression, while a clearer understanding of the effects would require further research.

      Moreover, the lack of detailed information on the primers and housekeeping genes used in the experiments is concerning, particularly given the importance of using non-circadian housekeeping genes for accurate normalization.

      It seems like the resource table describing these important experimental details was omitted in the original submission. We have now included it in the revised version (Table S1).

      The methods for measuring metabolic hormones, such as GLP-1 and GIP, are also not adequately described. If DPP-IV/protease inhibitor tubes were not used, the data could be unreliable due to the rapid degradation of these hormones by circulating proteases.

      We thank the reviewer for pointing out this omission. We have now added details of how we measured the metabolic hormones to the methods section, including the fact that we have added a DPP-IV inhibitor to the tubes at sampling (lines 469ff).

      Finally, the manuscript does not address the collection of hormone levels during both fasting and fed phases, a critical aspect for interpreting the metabolic impact of microbial metabolites.

      While we agree that it would be interesting to measure hormone levels also in the fed phase, a more thorough examination of hormone levels over the diurnal cycle, as suggested by reviewer 3, would be relevant for a full-scale follow-up. Given our data, we of course cannot exclude that there may be time-point-specific differences and therefore have softened the language around this conclusion to state that hormone levels are not acutely changed after a lactulose intervention at the time-points examined. (lines 318ff).

      These methodological concerns collectively weaken the robustness of the study's results and warrant careful reconsideration and clarification by the authors.

      Because of these weaknesses, the authors have partially achieved their aims by providing novel insights into the relationship between microbial metabolites and host circadian rhythms. The data do suggest that microbial metabolites can significantly influence feeding behavior and circadian gene expression in specific contexts. However, the unexpected absence of an osmotic effect of lactulose, the potential biases introduced by the log2-fold change normalization in qRT-PCR data, and the lack of clarity in critical methodological details weaken the overall conclusions. While the study provides valuable contributions to understanding the gut microbiome's role in circadian biology, the methodological weaknesses prevent a full endorsement of the authors' conclusions. Addressing these issues would be necessary to strengthen the support for their findings and fully achieve the study's aims.

      We thank the reviewer again for their careful and critical reading of our work, and for their constructive input. In the revised version of our manuscript, we address the reviewer's concerns by providing more methodological detail and additional experimental data.

      Despite the methodological concerns raised, this work has the potential to make a significant impact on the field of circadian biology and microbiome research. The study's exploration of the interaction between microbial metabolites and host circadian rhythms in different microbial environments opens new avenues for understanding the complex interplay between the gut microbiome and host physiology. This research contributes to the growing body of evidence that microbial metabolites play a crucial role in regulating host behaviors and physiological processes, including feeding and circadian gene expression.

      We thank the reviewer for their encouraging remarks!

      Reviewer #3 (Public Review):

      Summary:

      In the manuscript by Greter, et al., entitled "Acute targeted induction of gut-microbial metabolism affects host clock genes and nocturnal feeding" the authors are attempting to demonstrate that an acute exposure to a non-nutritive disaccharide (lactulose) promotes microbial metabolism that feeds back onto the host to impact circadian networks. The premise of the study is interesting and the authors have performed several thoughtful experiments to dissect these relationships, providing valuable insights for the field. However, the work presented does not necessarily support some of the conclusions that are drawn. For instance, lactulose is administered during the fasting period to mimic the impact of a feeding bout on the gut microbiota, but it would be important to perform this treatment during the fed state as well to show that the effects on food intake, etc. do not occur.

      We thank the reviewer for this important point. In the revised version, we include an experiment where we administer lactulose during the fed state and do not observe a significant change in food intake. We describe this in the text (lines 189ff) and in the new Figure S5C and D.

      To truly draw the conclusion that the current outcomes are directly connected to and mediated via an impact on the host circadian clock, it would be ideal to perform these studies in a circadian gene knock-out animal (i.e., Cry1 or Cry2 KO mice, or perhaps Bmal-VilCre tissue-specific KO mice). If the effects are lost in these animals, this would more concretely connect the current findings to the circadian clock gene network.

      We agree that these would be interesting experiments to follow up on the question how the observed effects are actuated by host functions. However, they would require a large amount of preparatory work (including rederiving the KO mice to get them germ-free in our gnotobiotic facility), we argue that they are beyond the scope of this study.

      Despite these reservations, the work is promising.

      We thank the reviewer for their encouraging assessment.

      Strengths:

      Attempting to disentangle nutrient acquisition from microbial fermentation and its impact on diurnal dynamics of gut microbes on host circadian rhythms is an important step for providing insights into these host-microbe interactions.

      The authors utilize a novel approach in leveraging lactulose coupled with germ-free animals and metabolic cages fitted with detectors that can measure microbial byproducts of fermentation, particularly hydrogen, in real-time.

      The authors consider several interesting aspects of lactulose delivery, including how it shifts osmotic balance as well as provides calculations that attempt to explain the caloric contribution of fermentation to the animal in the context of reduced food intake. This provides interesting fundamental insights into the role of microbial outputs on host metabolism.

      Thank you!

      Weaknesses:

      While the authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery, the authors have not accounted for the enlarged cecum and increased cecal surface area in germ-free mice. The authors could consider an additional control of cecectomy in germ-free mice.

      We thank the reviewer for pointing out the potential effect of the anatomical differences of germ-free and conventionally colonized mice. We agree that when comparing germ-free mice to SPF mice, the enlarged cecum area in germ-free animals could lead to differences in water release or uptake. However, this difference is smaller in gnotobiotic mice colonized with our minimal microbiota, even though their ceca are still slightly smaller than those of germ-free mice (new Figure S2F). While we agree that including control of cecectomy in germ-free mice could be interesting, we do not have the option of doing surgery on germ-free mice given our current experimental setup. We have now added information on cecum weight, a good proxy for cecum size, in the new Figure S2F done.

      The authors have examined GI hormones as one possible mechanism for how food intake is altered by microbial fermentation of lactulose. However, the authors measure PYY and GLP-1 only at a single time point, stating that there are no differences between groups. Given the goal of the studies is to tie these findings back into circadian rhythms, it would be important to show if the diurnal patterns of these GI hormones are altered.

      We fully agree that a deeper investigation of the diurnal fluctuations of hormone levels would be an interesting next step in studying whether perturbations in food intake can disturb these rhythms. Doing this for the whole rhythm would really require a full second study.

      In response to the reviewer's comments, we have changed the statements made around these data to point out just that hormone level fluctuations could not be detected during specific time points after lactulose treatment and therefore do not seem to explain the imminent behavioral changes (lines 318ff).

      Considerations of other factors, such as conjugated vs. deconjugated bile acids, microbial bile salt hydrolase activity, and bile acid resorption, might be an important consideration for how lactulose elicits more influence on ileal circadian clock genes relative to cecum and colon.

      We absolutely agree that investigation of microbial bile acid modification and their metabolism by the host would be an interesting topic for a follow-up study.

      Measurements of GI transit time (both whole gut and regional) would be an important for consideration for how lactulose might be impacting the ileum vs. cecum vs. colon.

      This is also an interesting point. While we did not add an experiment in which we specifically measure transit time to the revised version, we now measure total faecal output in a 5 h time period after PBS or lactulose treatment (Figure S3C). Faecal output is known to be a good proxy for transit time, and we see no significant difference between lactulose treatment (even with a two-fold higher dose than used before, new Figure S3) and PBS treatment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) Line 126 - see the point in the public review, this data argues against disrupting the rhythm.

      See our response above

      (2) Line 156 - the metric used for water content is confusing. Why not just subtract dry weight from wet weight to get water content? The ratio is much harder to think about for me. Perhaps more importantly, this data is very confusing given that colonization seems to impact the activity of lactulose, which complicates the interpretation. Could be an interesting area for future study that you might highlight more in the discussion.

      We thank the reviewer for pointing out that our presentation of water content could be clearer. We have changed the dry/wet ratio we have used in the previous version to the "water fraction" (new Figure S2C, E; new Figure S3A,B), i.e., the per cent of weight of the wet sample that is made up by water. We would argue that this is a measurement that is easier to interpret than the difference suggested by the reviewer, because it is independent of the absolute sample weight.

      We also agree that it is surprising that colonization state changes the osmotic state of the gut environment and have now added text discussing that (lines 127ff).

      (3) Line 188 - the lack of expression changes in the distal gut (cecum/colon) potentially conflicts with the model, warranting additional discussion/qualifications. Is lactulose getting metabolized in the small intestine? Alternatively, does lactulose have a direct effect on the host? The current text seems to imply that lactulose is fermented in the colon, the fermentation products are absorbed, and then they only impact the ileum through circulation, which doesn't seem physiologically possible.

      It is possible that lactulose affects small intestinal tissue directly, but we show that the effect depends on microbial activity. Microbial activity is much larger in the large intestine than in the small intestine, which is why we hypothesize that it acts through systemic signals rather than locally. These systemic signals might be fermentation products impacting the ileum through circulation. We would argue that this is not implausible, given that there are well-documented systemic effects of fermentation products in circulation (e.g., den Besten et al, 2013). It is, however, also possible that, e.g., metabolism of fermentation products in the liver triggers a secondary signal that acts systemically.

      (4) Line 241 - need to weaken this sub-heading. The experimental design shows that fermentation products impact feeding behavior, but this does not necessarily imply that fermentation products are responsible for the lactulose effect.

      Done.

      (5) Line 256 - The lack of an effect in SPF mice is surprising and potentially conflicts with the model proposed. Given this and other confusing results (gene expression site specificity, osmolarity effects, etc) it seems prudent to be more cautious as to the potential mechanisms through which lactulose supplementation impacts host gene expression and feeding behavior, which would likely require a lot more experiments to provide a definitive answer.

      While the lack of an effect in SPF mice is surprising, we would argue that this is rather points towards the need for a better understanding of host-microbiota-diet interactions than a conflict with our interpretation. We do agree with the reviewer that more work is necessary for providing definitive proof of the underlying mechanisms of the observed effects and have adapted the language throughout the text.

      (6) Line 260 - A lot of text is devoted to the potential caloric effects of the fermentation products and the lactulose itself. Was this in response to a prior reviewer? Either way, it seems too speculative to me and I would recommend trimming it down and moving it to the discussion.

      We have rewritten this section to make it more concise and increase readability (lines 224ff).

      (7) Line 305 - Not fasting, just lower caloric intake, as shown by Figure 1c.

      We have changed the text to reflect that.

      (8) Line 371 - Too definitive given the current data. Need to qualify the interpretation here.

      We have adapted the text and qualified the interpretation.

      (9) Figure 2 - Need to revise the title, no data showing that the rhythm is disrupted.

      Done.

      (10) Figure 3b - Clarify what timepoint is shown in the legend.

      Done.

      (11) Figure 3c - label when the treatment started. Consider changing to 2-way ANOVA which is probably more appropriate than t-tests.

      We have added information on treatment start to the figure legend and have changed the statistical analysis to a two-way ANOVA (time, treatment).

      (12) Figure 4c - move to supplement as this negative data isn't sufficient to rule out an impact on the hypothalamus or liver. Even when only considering transcript levels it's possible that the timepoint is just not ideal.

      We agree that this data only represents a snapshot of gene expression at one timepoint after treatment and does not rule out involvement of hypothalamus or liver in this process. We have now moved the previous Figure 4C to the supplementary information (new Figure S6A).

      (13) Figure 5 - defined the "fermentation products" and their concentrations in the legend. Consider moving panels d, and e to the supplement - negative results with unclear interpretation. Clarify in the legend how many calories/g were assumed for the fermentation products and provide a scientific rationale for this decision. Modify the title to be more cautious - as discussed above.

      We thank the reviewer for pointing out that this was not clear. We have now added the formulation of the fermentation products to the figure legend and moved panels D and E to the supplement. We have also combined the former Figures 4AB and 5ABC into the new Figure 4.

      (14) Figure S1 - The patterns in panel a are really intriguing and could be discussed more. For example, what do you think is driving the rapid oscillations in E. rectale?

      We agree that the patterns are potentially interesting, but the fact that the patterns we observe do not replicate well between the two light-dark cycles we monitor suggest a large contribution of experimental noise. We therefore refrain from interpreting too much into this dataset.

      (15) Figure S2 - Could changing the metric for water content be more easily interpretable? Modify the title to better match the data shown.

      We thank the reviewer for pointing that out. We have now changed the metric in Figure S2 (and the new Figure S3) to % water in feces/cecum content, which is easier to interpret.

      (16) Figure S3 - need to specify the multiple testing correction used.

      Done.

      (17) Figure S4 title - replace "inducing microbial metabolism" with "lactulose".

      Done.

      (18) Figure S5 title - modify to weaken the causal claim.

      We have adapted all figure titles to conform to this comment.

      Reviewer #2 (Recommendations For The Authors):

      Greter and colleagues present an insightful manuscript investigating the effects of microbial metabolites, such as hydrogen and SCFAs, on feeding behavior and circadian gene expression at a single time point. Notably, they observed that these metabolites exert a significant influence on behavior in a reduced community model (EAM). However, this effect was not evident in germ-free or SPF mice, highlighting an intriguing paradox. The manuscript is well-written, with experiments that are thoughtfully designed and clearly rationalized. The study contributes valuable data to the poorly understood relationship between the gut microbiome and host circadian rhythms. While I find the manuscript compelling, I believe that providing additional experimental and analytical details would enhance clarity and rigor.

      Major Comments:

      (1) Additional clarity is needed regarding the stool collection process. Were the samples collected as fresh specimens, or was there a possibility of coprophagia before collection? Clarification on this point is important, as it could impact the results, particularly the dry-to-wet weight ratio analysis. Ensuring the collection process did not introduce this bias is crucial for the validity of these measurements.

      Thank you for pointing out that this was not stated clearly. Wherever we assessed water content in faeces, we used fresh samples that were directly collected from a live animal and immediately frozen in a closed container or analyzed. When we assessed total faecal output, we collected the total bedding from a cage, sorted out the faecal pellets, and only measured dry weight. Total wet weight output was then assessed by using the water content of fresh faeces of animals with the same microbiota and the same treatment as a correction factor. We have now made this clear in the methods section (lines 435ff).

      In cases where we performed total output measurements by collecting bedding, there was the possibility for coprophagia. We did not control for this but assumed that coprophagia will have a small effect that is likely similar between groups and should thus not affect the comparison.

      (2) Caution is advised in the use of the term 'circadian,' which is sometimes used when 'diurnal' might be more appropriate. For example, the title of the first results section could be revised to 'Host Feeding [or Diurnal] Rhythms Influence Microbial Metabolic Fluctuations.' Additionally, line 357 should likely use 'diurnal' instead of 'circadian.' It's important to note that 'circadian' implies that cyclical fluctuations would persist without environmental cues (e.g., feeding). Since it's not clear whether most microbiome compositional or functional fluctuations are truly circadian, 'diurnal' is likely the more accurate term.

      We thank the reviewer for pointing out our imprecise use of the term circadian. We have now adapted this throughout the manuscript.

      (3) The lack of an osmotic effect of lactulose in EAM mice is quite surprising, given that lactulose is known to be an osmotic laxative. Was this finding specific to EAM mice, or was a similar lack of osmotic effect observed in SPF mice? A positive control is necessary to verify that this unexpected result is not artifactual. If there is no osmotic laxative effect in SPF mice, an explanation is needed as to why this medication is not functioning as expected in these mice.

      This is a good point. While we initially thought that the lack of an osmotic effect was due to the specific lactulose dose we were using, we also did not observe an osmotic effect (measured by the water content of fresh faecal pellets produced after treatment) when we used twice the amount of lactulose in EAM mice (new Figure S3). However, in SPF mice, we did see an increase in faecal water content after treatment. This intriguing result suggests that the effect of lactulose as an osmotic laxative depends on the presence of a complex microbiota. We now show this data in the new Figure S3A and discuss it in the text (lines 127ff).

      (4) The authors state that 'To account for faulty measurements due to disruptive events and for measurement noise, some datapoints were excluded from the raw datasets.' It would be important for the authors to confirm that this data exclusion was unbiased, meaning it did not disproportionately affect one group over another, and that any exclusions affected groups randomly.

      This is a good point, and we analyzed this for the experiments we show in Figs 4A/S5A, 4B/S5E, 4C, and S5F, and discuss this in the methods part (lines 493ff). The resulting statistics is not fully conclusive: using a Chi-square test to check whether the probability of excluding values differs between experiments, we get significant differences (p=3.9 x 10<sup>-18</sup>). We would, however, argue that this is not surprising, as different experiments sometimes different in the number of times we needed to do maintenance work on the isolators, which could lead to actuation of the scales measuring feed values, and thus faulty measurements.

      We face the same problem within experiments: we found a significant difference in the probability to exclude values between treatment and control in the experiment shown in Figs 4A/S5A (p=1.7 x 10<sup>-6</sup>), but no significant differences in the experiments in Figs 4B/S5E and S5F (p=0.71, and p=0.10, respectively). While it is hard to strictly exclude an influence of our data exclusion strategy, these findings speak against a systematic effect of treatment.

      (5) The authors should provide details on the primers and housekeeping genes used in their experiments. It's crucial that the housekeeping gene is not circadian and is stable at all time points (PMID 17878933).

      In the previous submission, the main resources table was omitted by accident. We have now added it (Table S1), including details on the primers and reagents used for all experimental work.

      (6) The presentation of qRT-PCR data as log2-fold change is confusing. It's unclear what the numerator and denominator represent for this ratio or why such normalization was deemed necessary. Ideally, transcripts should be normalized to a housekeeping gene (as noted in a previous comment), not to a baseline measure of other same genes acquired from other mice. Log2-fold change is typically appropriate when comparing two measures from the same mice; however, in this study, the mice were euthanized, and the denominator is a mean of genes measured from other samples. This approach could introduce bias and might explain why both the positive and negative limbs appear to be activated by the microbial metabolites. It would be more rigorous to present these values as absolute gene expression levels.

      We thank the reviewer for pointing out that this was not described clearly in the previous version of our manuscript. We have now adapted the methods part to explain that all data showing RT-PCR data is normalized to a housekeeping gene. Only after that, we compare the gene expression levels of the treatment group to the control group (or to gene expression of the control group at timepoint 0 in the case of Figure 3C).

      (7) The use of log-ratios, with a mean as the denominator, could artificially reduce the variability in the data, potentially leading to spurious findings or an increased risk of Type I error.

      As pointed out above, we have used internal normalization to a housekeeping gene before comparing the resulting values of the treatment and control groups. This method (commonly known as ΔΔC<sub>t</sub> method) is a standard way of comparing gene expression values obtained by RT-PCR. The use of a log2 transformation is commonly used to convert the logarithmic data resulting from the RT-PCR measurement to a linear fold-change measurement. We do not see that this data analysis strategy should lead to an increased risk for producing false positives. We now explain this better in the methods section of the manuscript (lines 541ff), and have adapted the labels in Figures 3B, C and S4 to avoid confusion.

      (8) It is unusual that both the positive and negative limbs of the circadian clock are overexpressed following lactulose administration. The authors should provide data confirming that these genes are in counter phase to each other at baseline. This clarification would help readers better understand the effects of the experimental interventions. As it stands, this critical part of the results is quite confusing.

      We agree that this result does not allow a clear interpretation of the effect of lactulose on the diurnal rhythm of the host. While we agree that this would be interesting to understand in detail, we are merely taking this as a first indication that actuation of microbial metabolism during the inactive phase of the diurnal rhythm can lead to changes in clock genes. This claim is supported by our data.

      (9) Could the effects of the microbial metabolites be mediated by AMPK, a known nutrient sensor that can influence the post-translational modification of Cry proteins? It would be beneficial for the authors to explore whether these metabolites have a more direct, previously unknown mechanism of affecting the circadian clock, or if their effects are mediated through known signaling pathways such as AMPK.

      We agree that this would be a valuable path to continue investigating the effect of microbial metabolism on clock gene activity.

      (10) The methods section does not specify how the metabolic hormones (e.g., GLP-1, GIP, leptin, ghrelin) were measured in the experiments. It is important for the authors to confirm that DPP-IV/protease inhibitor tubes were used for hormone measurement, as these proteins can be rapidly degraded by circulating proteases. Without the use of appropriate tubes, this data cannot be reliably interpreted. Additionally, it would have been ideal to collect these hormone levels during both the fasting and fed phases, but it appears this was not done. This represents a significant limitation of the study and should be addressed in the discussion.

      We thank the reviewer for pointing out this omission, we have now added a description of our protocol to measure metabolic hormones to the methods section (lines 469ff).

      (11) If the samples were appropriately collected in DPP-IV/protease inhibitor tubes, the authors should consider measuring active GLP-1, as this would likely provide a more accurate assessment of GLP-1 activity.

      We agree that this would be a valuable next step, in addition to testing the effect of changes in microbial metabolism on the time traces of hormone levels.

      Minor Comments:

      (12) Line 350 appears to have an incomplete sentence, as it seems part of the first sentence in the paragraph has been inadvertently deleted. This should be reviewed and corrected for clarity.

      Done.

      Reviewer #3 (Recommendations For The Authors):

      Major comments:

      (1) Could the authors provide a deeper description about what they are referring to in the following statement? "...higher order interactions and microbial metabolism are variable..." it is difficult to interpret as written. Do the authors mean cross-feeding interactions?

      We have changed this sentence to clarify the meaning.

      (2) Could the authors explicitly state their hypothesis in the introduction and provide a brief, but deeper explanation of the intervention prior to the results section?

      This is a good point, we have adapted the text accordingly (lines 53ff).

      (3) Could the authors include a bit more information regarding the diet provided to the mice? If grain-based chow, please provide insights into the fiber source, etc.

      While we agree that it would be interesting to know what part of the mouse diet is available to the microbes, this is hard for the standard mouse chow that we (and most others doing experiments with mice) feed the experimental animals. We have now added more detailed information on the specific type of chow the mice were fed (lines 347f). We would argue that, because the control and treatment groups were always fed the same chow, the effect of the fiber source and other specifics are controlled for, even if we do not know them.

      (4) In figure 1A - cells/g does not seem to be the correct unit - # of copies/g feces perhaps?

      Cells/g is the appropriate unit, but it seems like we have not explained the way we arrive at this unit in sufficient detail. In short, we use a qPCR run on a known standard curve of bacterial counts (known cells/g values) to estimate these numbers from qPCR results from faeces. We have now explained this better in the methods section (lines 385f).

      (5) Figure 1B/C and Figure S1B/C are confusing - the legend states these measurements were taken over two days, however, the plot shows a single 12:12 LD period. Was the data averaged? It might be best to show each day separately (i.e., over a 48-hour period) rather than in one 24-hour plot. Then, the authors could also show the averages in the light period vs. the dark period in a separate, complementary graph.

      We thank the reviewer for pointing this out. The previous figures were indeed averaged over the two days of measurement and projected onto one 24h period for plotting. We have now changed Figures 1B,C and S1B,C to show the full 48h time windows.

      (6) Line 109 - The reviewer concurs that lactulose is a non-nutritive, synthetic disaccharide, however, in theory, lactulose may have a high heat increment, which could cause the animal to undergo metabolic responses to defend core body temperature (which also exhibits diurnal rhythmicity). Have the authors considered core body temperature rhythms, their connection to microbial metabolism, and the core circadian clock gene network in their model?

      This is an interesting thought. We have not measured body temperature in our experiments. As the heat increment from food is typically associated with metabolic activity, and lactulose is not metabolized, we do not expect its heat increment to be high, at least in GF mice. In mice with a microbiota, we agree that metabolic heat will be produced upon lactulose metabolism by the microbes, which could be a contributor the observed effect.

      (7) Figure 2A and corresponding text in lines 121 - 123 - indicate at what time the bar and whisker plots were taken (assuming at ZT8, but please be explicit in the figure and corresponding text). Could the authors also include statistics for these waveforms?

      Done.

      (8) In Figure 2B, the authors state that microbial metabolism had been restored to normal levels by ZT15, however, did this persist into the next light phase? It would be ideal if the authors could present these data in a similar manner to that shown in Figure S2A for SPF mice. Further, what are the statistical considerations here to describe changes in phase, amplitude, periodicity, etc.

      This is a good point. Unfortunately, we do not have H2 measurements for EAM and GF mice over comparable time periods as shown for SPF mice in Figure S2A. However, the food intake data shown in Figure S5BCD is a indicates that the food intake normalized in the second dark phase after treatment.

      (9) In lines 147 - 149 and in Figure S2B figure legend - assuming these measurements are from individual bacteria? Could this be stated clearly in the text or legend? Also, what are the statistical considerations? Were there significant differences in SCFA production between bacteria?

      We thank the reviewer for pointing out that this was unclear. We have now adapted the legend to explain the way these data were collected and added a statistical analysis.

      (10) The authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery - however, have the authors accounted for the enlarged cecum and increased cecal surface area in germ-free mice? Would an additional control be cecectomy in germ-free mice to be more in-line w/ SPF animals? Further, could the authors tie in these findings more explicitly and state how they pertain to the overall goal of the study? Is this simply to draw the conclusion that microbial biomass is increased w/ lactulose?

      We wanted to make sure that the effect we see with lactulose is due to microbial metabolism and not due to the induction of osmotic diarrhoea or other host-dependent effects, and we have added a statement to that effect (lines 127ff). We have not corrected for the change in cecum size between GF and EAM mice, as EAM size (and many other gnotobiotic mouse models) also have enlarged ceca relative to conventionally colonized mice, but have added a dataset showing how cecum size of EAM and GF mice compare and discuss this in the text (new Figure S2F, lines 135ff).

      (11) Is Figure 3A necessary?

      It might not be strictly necessary, but it can help with understanding the relation of the genes tested in B and C, and we would therefore like to keep it in.

      (12) Line 155 - 157 - the authors make the statement that dry/wet feces weight ratio is decreased in GF mice, but this does not appear to be statistically significant. Please adjust to state numerical differences were observed or provide statistics.

      We have changed the statement in the text.

      (13) Could Figure S3A be moved to the main Figure 3 as this may provide a more logical flow? qRT-PCR data is expressed as - log2 (fold expression), but relative to what? Could the authors provide further info about the control?

      The former Figure S3A (new S4A) and Figure 3B show the same data in slightly different ways. We therefore opted to keep only one of those illustrations in a main figure. The fold changes are always relative to the average of the PBS control group, which we now state explicitly in the figure legend.

      (14) The authors state that the qRT-PCR data shows that microbial metabolism of lactulose impacts peripheral circadian gene expression, but this conclusion seems simplified. Lactulose treatment only impacted the ileum circadian gene expression. Additional peripheral tissues (liver, adipose tissue, etc.) could be moved to this figure, i.e., move Figure 4C data. Why do the authors think the ileum was most impacted beyond GI hormones as discussed later in the manuscript? Could changes in bile acid deconjugation (i.e., BSH activity?) and/or bile acid resorption by the host in the distal ileum due to lactulose delivery be involved? Or is it simply due to differences in GI transit time (which was not measured in the current study)? Further, lactulose had minimal impact in SPF ileum, and in fact, shifted Cry1 in the opposite direction relative to EAM mice. Could the authors provide more insight into these disparate observations (line 196 - 200)?

      We agree that the statement "lactulose impacts peripheral gene expression" is oversimplified, and we have now adapted the text to avoid the impression that this is our conclusion. No tested tissues other than the ileum showed significant differences in gene expression at the time point tested, which does not rule out that other tissues would react to the treatment at that time point, or the tested tissues would do so at the tested time point. As we don't have a good enough understanding of what mechanism causes the gene expression changes in the ileum, we refrain from speculating in the text, even though we agree that this is an intriguing question.

      (15) Could the authors provide more insight into the statistical approaches used to assess amplitude, peak, nadir, etc. in Figure 3C? Was the co-sinor waveform tested?

      This is a good question. Even though this was a highly work-intensive experiment using many animals, we would argue that the noise level is too high and the coverage of the time analyzed too sparse to infer meaningful statistics on the fluctuations of gene expression over time. In the new version, we have changed our statistical analysis of this dataset to a two-way ANOVA (treatment, time) to better analyze this dataset.

      (16) The food intake decrease and interpretation following treatment (Figure 4A and S4A) is curious - all animals were gavaged and in EAM mice, many animals, regardless of PBS or lactulose are trending down in food intake rate/total intake. It seems to be more of an impact of gavage and not of treatment, which the authors somewhat acknowledge in Lines 228 - 230.

      We agree that gavage is a possible factor in future food intake of experimental animals, which is precisely why we used the PBS gavage as a control. Even though the difference is not large, we see a significant change in food intake when lactulose is given, but not when PBS is given (Figure 4A, S4A).

      (17) Could the authors provide a deeper rationale for their line of thinking for lines 234 - 240? What is the evidence that systemic effects are likely to occur 3 hours after lactulose delivery? Further, as stated in comment 13, could brain and liver data be moved to Figure 3/Figure S3 as an additional example of peripheral tissue clocks?

      We have added an explanation for the rationale we use to justify the 8h time point (it is 3h after the peak of H2 production, as shown in Fig2AB, which happens 5h after lactulose delivery). While we agree that the brain and liver data would also fit into the Figure 3, we have now moved all negative data to the supplementary information (in response to a comment by reviewer 1, and in a general effort to clean up the data in the manuscript).

      (18) The authors measure PYY and GLP-1 at a single time point and state there are no differences, yet, the goal of the studies is to tie this back to circadian networks. Would it be possible to measure these GI hormones over a 24-hour period to show that the diurnal patterns are altered?

      We fully agree that measuring the metabolic hormones over time would be very interesting. It is possible but would represent a major effort using many animals and a large amount of work. We would therefore argue that it is beyond the scope of this revision, but a good starting point for a follow-up study.

      (19) The authors state that the administration of fermentation products acutely altered circadian food intake, but the studies do not support that this change is connected to the circadian network. Suggest softening the interpretation of the findings.

      We have changed the language there to soften the interpretation.

      Minor comments:

      (1) The authors should consider when it is appropriate to refer to rhythms as diurnal vs. circadian, as each has a distinct meaning. Diurnal follows entrainment cues while circadian is endogenously driven (i.e., line 39, line 58).

      We thank the reviewer for pointing out this important difference, we have adapted this in the whole text accordingly.

      (2) Circadian rhythm should be plural throughout the manuscript (circadian rhythms).

      Thank you, we have changed that where we refer to host circadian rhythms generally.

      (3) Lines 54 - 63. Fermentation should be capitalized when used at the beginning of a sentence.

      Done.

      (4) Line 289 - This should be Figure 5D and 5E.

      Done.

      (5) Line 290 - heart should be cardiac.

      Done.

    1. eLife Assessment

      This paper introduces a valuable optical method for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photostimulation in different axial planes. The evidence for effective optical separation of excitation and imaging is convincing, although the in vivo data in the olfactory bulb suggest potential confounding factors arising if stimulation not only affects cell bodies but also neuronal processes. The work will be of broad interest to neurobiologists working in circuit and systems neuroscience, as well as to specialists in optical microscopy.

    2. Reviewer #1 (Public review):

      In this methods paper, the authors introduce a novel and innovative imaging approach for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photo-stimulation in different axial planes. This is a highly exciting technique that enables the axial decoupling of optical imaging of deep neural circuits from surface photo-stimulation of spatially precise (tens of micrometres) brain spots. This method builds on previous developments from the same laboratory, combining DMD-based patterned photo-stimulation with in vivo electrophysiological recordings. To my knowledge, this is the first instance in which patterned photo-stimulation has been combined and axially decoupled from two-photon (2P) imaging.

      Beginning with a thorough characterisation of the optical resolution of the photo-stimulation system, the authors applied this method to the olfactory bulb (OB) network, in which sensory inputs are topographically organised at the surface of the OB and thus ideally suited to demonstrate the relevance of this approach. They first showed that this technique can be used to rapidly reveal connectivity patterns of OB output neurons and to identify sister mitral cells. In addition, they manipulated a specific glomerular inhibitory population and demonstrated that these neurons provide spatially heterogeneous long-range inhibition of OB output neurons, with differential effects on mitral and tufted cells (a result previously observed in a paper from the same lab: Banerjee et al., 2015, Neuron). Altogether, the data demonstrate that this technique is well-suited for high-throughput functional mapping of neural circuit properties. The results are compelling and illustrate both the significant advance represented by this method and its feasibility.

      Despite my initial enthusiasm, there are several concerns in the present study that must be addressed in order to rule out confounding observations and to resolve remaining uncertainties regarding photo-stimulation resolution. These include the following:

      (1) Spatial resolution: Although the authors provide convincing data on spatial resolution in vitro, several observations throughout the paper suggest that the effective photo-stimulation precision may be lower than initially reported. For instance, in Figure 2, the authors observe repeated responses in neighbouring glomeruli (e.g., glomeruli #3 & #5, #4 & #6). To what extent could light scattering along the X/Y/Z-axis above the targeted glomerulus recruit en passage axons, resulting in the inadvertent activation of multiple glomeruli?

      A further observation concerns the presence of "inhibited" sister mitral cells (Figure 3). The authors claim this is reminiscent of the differential spike-timing reported between sister cells (Dwawale et al., 2010, Nat Neuro). However, observing both excitatory and inhibitory responses following stimulation of glutamatergic inputs is an altogether different matter, particularly given that sister mitral cells are reciprocally connected via gap junctions. This observation requires further clarification and raises serious questions about the effective resolution of the stimulation. Could the inhibited cell simply correspond to a non-sister mitral cell receiving disynaptic feed-forward inhibition? To verify sister cell identity, the authors could confirm that the predicted sister cells share a similar odour receptive field compared to randomly selected mitral cell pairs. In their previous study employing analogous DMD-based photo-stimulation (Dhawale et al., 2010, Nat. Neurosci.), sister mitral cells did not exhibit such opposite response profiles (firing rate correlation of ∼0.7 between sister cells). Could the authors verify that a comparable activity correlation is also observed among the sister cells identified using ADePT in the present study? In Figure S6, the authors show recordings and stimulation of the same neurons co-expressing GCaMP and ChR2. Applying this experimental design to the mitral/tufted cell population (using a Tbet-Cre mouse transduced in the OB with both GCaMP and Chrimson virus) would constitute a valuable control to clarify the nature of these "inhibited" sister cells.

      An additional concern relates to the 21 out of 162 mitral cells that were activated by two distinct glomeruli - a finding that is incompatible with the established OB wiring diagram and that further challenges the claimed stimulation resolution.

      A critical control experiment is also absent: in a Thy1-GCaMP6 mouse lacking any light-sensitive opsin, do the authors observe any unintended side effects of photo-stimulation?

      Regarding sister cells (Figure 3), tufted cells are not analysed alongside mitral cells in this dataset, whereas this is elegantly performed in Figure 5 using the DAT+ model. Could the authors also demonstrate how the technique can reveal the complete family portrait of sister mitral and tufted cells?

      (2) The authors have explored only a limited set of photo-stimulation parameters, primarily varying light intensity. They should present additional tests, such as varying the spot size (which appears to be arbitrarily fixed at 30-50 µm) and the z plane of stimulation. The level of activation can vary considerably: for example, in Figure 3a(iii), identical stimulations elicit responses of markedly different amplitudes (see glom#3 and #4). In Figure 2, 5 out of 15 glomeruli failed to respond - could the choice of z-plane account for this variability? The stimulation duration (50-150 ms) also appears somewhat arbitrary: can the authors demonstrate that the technique is compatible with finer temporal patterns (e.g., 10 Hz stimulation for 500 ms using 20 ms light pulses)? What are the spatiotemporal and axial scanning limits of this approach, and can two or three glomeruli be targeted simultaneously with temporally patterned stimulation?

      (3) One particularly relevant application of this method would be to guide photo-stimulation based on prior functional measurements - for instance, by generating a photo-stimulation mask specifically targeting odour-responsive glomeruli. In the DAT-Cre × Thy1-GCaMP6 experiment shown in Figure 5e, which glomeruli are activated by a given odour, and how does this odor responsiveness influence the efficiency of DAT+ cell-mediated inhibition?

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Koh and colleagues describe ADePT (Axially Decoupled Photo-stimulation and Two-photon Readout), a modular approach for combining patterned one-photon optogenetic stimulation with two-photon calcium imaging in independently controlled axial planes. The method relies on a digital micromirror device together with a motorized holographic diffuser to generate spatially confined stimulation patterns while imaging deeper neuronal populations. As proof-of-principle applications, the authors use the system to map excitatory and inhibitory functional connectivity in the mouse olfactory bulb by stimulating superficial glomerular circuits and recording responses from mitral and tufted cells in deeper layers.

      This is a well-executed Tools and Resources manuscript. The technical implementation is described in considerable detail, the optical performance is systematically characterized, and the biological experiments provide convincing demonstrations of the types of circuit questions that can be addressed using the method.

      Strengths:

      The greatest strength of the manuscript is the comprehensive technical characterization of the optical system. The authors carefully benchmark the spatial resolution, axial confinement, registration accuracy, calibration procedure, and practical operating limits of the setup. I found the extensive optical benchmarking particularly helpful, as it gives readers a realistic sense of the operating regime and practical limitations of the approach.

      Another strength is the high level of methodological transparency. The optical design, calibration procedures, stimulation strategies, and analysis pipeline are described in sufficient detail that an experienced laboratory could realistically evaluate whether the system is suitable for its own applications. This level of documentation is particularly appropriate for a Tools and Resources article.

      A further strength is the clear positioning of ADePT relative to existing approaches. The authors are transparent about the trade-off between spatial resolution and implementation complexity: ADePT does not provide single-cell photostimulation, but offers flexible axial separation, a large stimulation field, and cellular-resolution two-photon readout in deeper planes without requiring a full holographic stimulation system. This defines a credible and potentially useful experimental niche.

      The biological applications convincingly demonstrate the utility of ADePT. The experiments identifying sister mitral/tufted cells through selective glomerular stimulation and the mapping of heterogeneous inhibitory influences from DAT-positive interneurons illustrate the types of functional connectivity questions that become experimentally accessible with this approach. Importantly, the authors generally avoid overstating these biological findings and appropriately present them as proof-of-principle demonstrations of the technology.

      Weaknesses:

      The primary limitation is inherent to the method itself rather than the execution of the study. Because ADePT relies on one-photon patterned illumination, photo-stimulation remains restricted to relatively superficial structures and does not achieve single-cell spatial resolution. The authors appropriately acknowledge these constraints and clearly position the method within this operating regime. Consequently, ADePT occupies a useful niche for interrogating spatially organized functional units such as olfactory glomeruli or cortical barrels, rather than applications requiring single-cell precision or deeper tissue penetration.

      Although the manuscript describes the approach as relatively simple and cost-effective, implementation still requires careful optical alignment, registration, calibration, and optimization. This does not diminish the value of the approach, but terms such as modular or accessible may better reflect the practical implementation than simple. Likewise, a brief bill of materials, approximate add-on cost, and indication of which components are essential versus substitutable would help prospective users assess the accessibility of the system.

      Finally, the manuscript provides an impressive level of technical characterization, but much of the practical guidance for adopting the system is distributed across the Results and Discussion. Bringing together the principal limitations, recommended operating regime, expected calibration workflow, evidence for long-term alignment stability, and the circumstances in which ADePT is preferable to alternative approaches would further strengthen the manuscript as a community resource.

    4. Author response:

      eLife Assessment

      This paper introduces a valuable optical method for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photostimulation in different axial planes. The evidence for effective optical separation of excitation and imaging is convincing, although the in vivo data in the olfactory bulb suggest potential confounding factors arising if stimulation not only affects cell bodies but also neuronal processes. The work will be of broad interest to neurobiologists working in circuit and systems neuroscience, as well as to specialists in optical microscopy.

      We thank the reviewers for their constructive feedback. We are planning to address their concerns as detailed below. Specifically, in the revised manuscript, we will streamline and consolidate the text to include:

      (1) An in-depth discussion of the awake recording results in the main text.

      (2) Further discussion of light scattering, photo-stimulation specificity of targeting individual glomeruli and resolution.

      (3) A summary (including also a table) of the operating regime and comparisons with alternative techniques for patterned photo-stimulation and imaging of the ensuing responses. We will highlight the advantages and constraints of the current implementation of ADePT. In particular, here we explored a small set of spatiotemporal parameters to understand the limits of our technique and provide a proof of principle of the strategy. These parameters can be varied further depending on the exact research question.

      (4) Implementation considerations and technical guidelines for calibration and long-term stability of the rig for ADePT (i.e. ‘a how-to guide’).

      (5) A bill of materials and estimated hardware costs.

      Furthermore, we will provide additional controls and rephrase some of the statements in the text as suggested (e.g. replace ‘accessible’ with ‘simple’, etc.). We will correct the unfortunate grammatical errors, improve clarity of text, and update the references accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      In this methods paper, the authors introduce a novel and innovative imaging approach for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photo-stimulation in different axial planes. This is a highly exciting technique that enables the axial decoupling of optical imaging of deep neural circuits from surface photo-stimulation of spatially precise (tens of micrometres) brain spots. This method builds on previous developments from the same laboratory, combining DMD-based patterned photo-stimulation with in vivo electrophysiological recordings. To my knowledge, this is the first instance in which patterned photo-stimulation has been combined and axially decoupled from two-photon (2P) imaging.

      Beginning with a thorough characterisation of the optical resolution of the photo-stimulation system, the authors applied this method to the olfactory bulb (OB) network, in which sensory inputs are topographically organised at the surface of the OB and thus ideally suited to demonstrate the relevance of this approach. They first showed that this technique can be used to rapidly reveal connectivity patterns of OB output neurons and to identify sister mitral cells. In addition, they manipulated a specific glomerular inhibitory population and demonstrated that these neurons provide spatially heterogeneous long-range inhibition of OB output neurons, with differential effects on mitral and tufted cells (a result previously observed in a paper from the same lab: Banerjee et al., 2015, Neuron). Altogether, the data demonstrate that this technique is well-suited for high-throughput functional mapping of neural circuit properties. The results are compelling and illustrate both the significant advance represented by this method and its feasibility.

      We thank the Reviewer for their constructive input.

      Despite my initial enthusiasm, there are several concerns in the present study that must be addressed in order to rule out confounding observations and to resolve remaining uncertainties regarding photo-stimulation resolution. These include the following:

      (1) Spatial resolution: Although the authors provide convincing data on spatial resolution in vitro, several observations throughout the paper suggest that the effective photo-stimulation precision may be lower than initially reported. For instance, in Figure 2, the authors observe repeated responses in neighbouring glomeruli (e.g., glomeruli #3 & #5, #4 & #6). To what extent could light scattering along the X/Y/Z-axis above the targeted glomerulus recruit en passage axons, resulting in the inadvertent activation of multiple glomeruli?

      Indeed, we cannot rule out this possibility. Fibers of passage are a potential concern. This is why we systematically sample different light intensities and assess their impact on specificity of dendritic mitral and tufted cell responses within the glomerular layer (same axial-plane optical stimulation and imaging experiments, Fig. 2). We identify a range of intensities that on average result mostly in activation of the targeted glomeruli. Within the range of intensities used for identifying sister cells, >90% of responses were on the diagonal (targeted glomeruli) and ~5% pixels that cleared the signal significance criterion used were in off-target glomeruli, as stated in the text and quantified in Fig. 2f. In the revised manuscript, we will further clarify and expand on these points.

      A further observation concerns the presence of "inhibited" sister mitral cells (Figure 3). The authors claim this is reminiscent of the differential spike-timing reported between sister cells (Dwawale et al., 2010, Nat Neuro). However, observing both excitatory and inhibitory responses following stimulation of glutamatergic inputs is an altogether different matter, particularly given that sister mitral cells are reciprocally connected via gap junctions. This observation requires further clarification and raises serious questions about the effective resolution of the stimulation. Could the inhibited cell simply correspond to a non-sister mitral cell receiving disynaptic feed-forward inhibition?

      This is indeed what we think it is happening (i.e. disynaptic feed-forward inhibition as the Reviewer points out). We observe inhibition in some of the mitral cells in the field of imaging when we stimulate not their parent glomerulus, but other glomeruli in the neighbourhood. As the Reviewer points out, sister cells are connected via gap junctions, but they also receive inhibitory chemical synaptic inputs via their secondary (and primary dendrites) from other (not-their-parent) glomeruli mediated by numerous types of interneurons including the DAT+/GABAergic (a.k.a. superficial short axon cells) and granule cells. Our data is consistent with differential inhibitory input from other glomeruli on sister cells getting input from the same parent glomerulus. As it appears that we failed to present this point clearly in the initial submission, we will further expand along these lines in the revised manuscript. Briefly:

      First, we identify a photo-stimulation regime that results mostly in the activation of a given targeted glomerulus (and not of other glomeruli in the field of stimulation). To this end, we photo-stimulate and image ensuing neuronal responses in the same axial optical plane. We strobe (alternate) between monitoring dendritic mitral and tufted cell (enhanced) GCaMP responses within the targeted glomerulus and other glomeruli in the field of imaging, while varying systematically the light intensity (Figs. 2,3a; Suppl. Fig. 4b, Suppl. Fig, 5b-d;h-j). We use as criterion for specificity a condition when >95% of significantly responding pixels (above a statistically defined signal response threshold) lie within the anatomical boundaries of the targeted glomerulus. For each glomerulus (or pixel within a glomerulus) we compared the average light response across trials with the baseline reference distribution in the absence of light stimulation. If this value crossed the 99th percentile of the baseline distribution, the glomerulus/pixel within glomerulus was classified as responsive to the photo-stimulation.

      Second, using the minimal light intensity regime experimentally identified as ‘specific’ for targeting individual glomeruli in the field of photo-stimulation (< 5% significant activation of off-target pixels), we decouple photo-stimulation in the glomerular layer from monitoring responses of mitral and tufted cells in the deeper layers of the olfactory bulb (100-250 µm axial displacement). This approach enables us to map cohorts of sister (daughter) cells associated with any specific target glomerulus in the field of view (1,2,3…n) by monitoring excitatory (enhanced) responses of mitral and tufted cell bodies. A cohort of sister cells associated with glomerulus x<sub>i</sub> (daughters of glomerulus x<sub>i</sub>) is defined by those cells which show statistically significant excitatory responses (4 SD - standard deviations - above their baseline fluctuations) specifically in response to photo-stimulation of glomerulus x<sub>i</sub>.

      Third, in the process, as we photo-stimulate different glomeruli in the field of stimulation, we also observe at times suppressed (inhibitory) responses in a subset of the mitral and tufted cells (exceeding 3 SD in the negative direction their baseline fluctuations). These suppressed responses occur in response to photo-stimulating not the parent glomerulus of a given cell, but other glomeruli in the field. These experiments revealed that within a cohort of sister cells (daughters of glomerulus x<sub>i</sub>), only a subset of cells are suppressed by activation of glomerulus x<sub>j</sub>, and, in a few example cases, different cells are suppressed by activation of different glomeruli (e.g. x<sub>j</sub> vs. x<sub>k</sub>), presumably through disynaptic feed-forward inhibition (Fig. 3b iii; 3d; Suppl. Figs. 5f,g; l,m). These preliminary observations suggest that sister cells receive differential inhibitory inputs from glomeruli in the neighborhood. In the revised manuscript, we will expand to further clarify these points.

      To verify sister cell identity, the authors could confirm that the predicted sister cells share a similar odour receptive field compared to randomly selected mitral cell pairs. In their previous study employing analogous DMD-based photo-stimulation (Dhawale et al., 2010, Nat. Neurosci.), sister mitral cells did not exhibit such opposite response profiles (firing rate correlation of ∼0.7 between sister cells). Could the authors verify that a comparable activity correlation is also observed among the sister cells identified using ADePT in the present study? In Figure S6, the authors show recordings and stimulation of the same neurons co-expressing GCaMP and ChR2. Applying this experimental design to the mitral/tufted cell population (using a Tbet-Cre mouse transduced in the OB with both GCaMP and Chrimson virus) would constitute a valuable control to clarify the nature of these "inhibited" sister cells.

      We thank the Reviewer for the suggestion. We consider that the experiments shown here are proof-of-principle in nature, highlighting the potential of ADePT for mapping functional neural circuit connectivity. In our opinion, further investigating the logic of similarities and differences in the odor responses of sister mitral cells and the nature of inhibitory glomerular interactions forms the focus of future studies. We also note that firing rate correlations can be notoriously difficult to compare and interpret across experimental regimes (spikes vs. calcium imaging).

      An additional concern relates to the 21 out of 162 mitral cells that were activated by two distinct glomeruli - a finding that is incompatible with the established OB wiring diagram and that further challenges the claimed stimulation resolution.

      Indeed, this reflects some degree of non-specific activation of the targeted glomeruli as discussed above. In the revised manuscript, we will further highlight this issue.

      A critical control experiment is also absent: in a Thy1-GCaMP6 mouse lacking any light-sensitive opsin, do the authors observe any unintended side effects of photo-stimulation?

      In the revised manuscript, we will include an additional control as suggested by the reviewer. Within the range of intensities used, we did not observe significant modulation of GCaMP6s activity in mitral and tufted cells in mice lacking light-sensitive opsins.

      Regarding sister cells (Figure 3), tufted cells are not analysed alongside mitral cells in this dataset, whereas this is elegantly performed in Figure 5 using the DAT+ model. Could the authors also demonstrate how the technique can reveal the complete family portrait of sister mitral and tufted cells?

      We thank the Reviewer for the suggestion. We think that the differences between mitral and tufted cells are indeed very interesting to investigate, but in our opinion form the subject of future studies.

      (2) The authors have explored only a limited set of photo-stimulation parameters, primarily varying light intensity. They should present additional tests, such as varying the spot size (which appears to be arbitrarily fixed at 30-50 µm) and the z plane of stimulation. The level of activation can vary considerably: for example, in Figure 3a(iii), identical stimulations elicit responses of markedly different amplitudes (see glom#3 and #4). In Figure 2, 5 out of 15 glomeruli failed to respond - could the choice of z-plane account for this variability? The stimulation duration (50-150 ms) also appears somewhat arbitrary: can the authors demonstrate that the technique is compatible with finer temporal patterns (e.g., 10 Hz stimulation for 500 ms using 20 ms light pulses)? What are the spatiotemporal and axial scanning limits of this approach, and can two or three glomeruli be targeted simultaneously with temporally patterned stimulation?

      Indeed, here we explored a limited set of spatiotemporal parameters to understand the limits of our technique and provide proof of principle. These can be varied depending on the exact question. In the revised manuscript, we will clearly state what the constraints of the current implementation are and provide context for further optimizations. Briefly, in the current version, individual as well as multiple glomeruli can be photo-stimulated together and 20 ms per pulse regime in trains of pulses is feasible.

      (3) One particularly relevant application of this method would be to guide photo-stimulation based on prior functional measurements - for instance, by generating a photo-stimulation mask specifically targeting odour-responsive glomeruli. In the DAT-Cre × Thy1-GCaMP6 experiment shown in Figure 5e, which glomeruli are activated by a given odour, and how does this odor responsiveness influence the efficiency of DAT+ cell-mediated inhibition?

      We thank the Reviewer for the suggestion. We are thinking along exactly the same lines. In particular, we would like to investigate the relationship between the degree of overlap in odor responses of individual glomeruli and the strength and specificity of their inhibitory interactions mediated by DAT+ interneurons. In the revised manuscript, we will further expand on discussing this venue of study. We feel however that this investigation is beyond the scope of this technical report.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Koh and colleagues describe ADePT (Axially Decoupled Photo-stimulation and Two-photon Readout), a modular approach for combining patterned one-photon optogenetic stimulation with two-photon calcium imaging in independently controlled axial planes. The method relies on a digital micromirror device together with a motorized holographic diffuser to generate spatially confined stimulation patterns while imaging deeper neuronal populations. As proof-of-principle applications, the authors use the system to map excitatory and inhibitory functional connectivity in the mouse olfactory bulb by stimulating superficial glomerular circuits and recording responses from mitral and tufted cells in deeper layers.

      This is a well-executed Tools and Resources manuscript. The technical implementation is described in considerable detail, the optical performance is systematically characterized, and the biological experiments provide convincing demonstrations of the types of circuit questions that can be addressed using the method.

      Strengths:

      The greatest strength of the manuscript is the comprehensive technical characterization of the optical system. The authors carefully benchmark the spatial resolution, axial confinement, registration accuracy, calibration procedure, and practical operating limits of the setup. I found the extensive optical benchmarking particularly helpful, as it gives readers a realistic sense of the operating regime and practical limitations of the approach.

      Another strength is the high level of methodological transparency. The optical design, calibration procedures, stimulation strategies, and analysis pipeline are described in sufficient detail that an experienced laboratory could realistically evaluate whether the system is suitable for its own applications. This level of documentation is particularly appropriate for a Tools and Resources article.

      A further strength is the clear positioning of ADePT relative to existing approaches. The authors are transparent about the trade-off between spatial resolution and implementation complexity: ADePT does not provide single-cell photostimulation, but offers flexible axial separation, a large stimulation field, and cellular-resolution two-photon readout in deeper planes without requiring a full holographic stimulation system. This defines a credible and potentially useful experimental niche.

      The biological applications convincingly demonstrate the utility of ADePT. The experiments identifying sister mitral/tufted cells through selective glomerular stimulation and the mapping of heterogeneous inhibitory influences from DAT-positive interneurons illustrate the types of functional connectivity questions that become experimentally accessible with this approach. Importantly, the authors generally avoid overstating these biological findings and appropriately present them as proof-of-principle demonstrations of the technology.

      We thank the Reviewer for their constructive input.

      Weaknesses:

      The primary limitation is inherent to the method itself rather than the execution of the study. Because ADePT relies on one-photon patterned illumination, photo-stimulation remains restricted to relatively superficial structures and does not achieve single-cell spatial resolution. The authors appropriately acknowledge these constraints and clearly position the method within this operating regime. Consequently, ADePT occupies a useful niche for interrogating spatially organized functional units such as olfactory glomeruli or cortical barrels, rather than applications requiring single-cell precision or deeper tissue penetration.

      We agree. In the revised manuscript, we will further expand on these points, highlighting the limitations and advantages of ADePT compared to other techniques in a table discussing various operating regimes.

      Although the manuscript describes the approach as relatively simple and cost-effective, implementation still requires careful optical alignment, registration, calibration, and optimization. This does not diminish the value of the approach, but terms such as modular or accessible may better reflect the practical implementation than simple. Likewise, a brief bill of materials, approximate add-on cost, and indication of which components are essential versus substitutable would help prospective users assess the accessibility of the system.

      We agree. We will change the text accordingly and provide the additional information as suggested.

      Finally, the manuscript provides an impressive level of technical characterization, but much of the practical guidance for adopting the system is distributed across the Results and Discussion. Bringing together the principal limitations, recommended operating regime, expected calibration workflow, evidence for long-term alignment stability, and the circumstances in which ADePT is preferable to alternative approaches would further strengthen the manuscript as a community resource.

      We agree. We will proceed accordingly in the revised manuscript.

    1. eLife Assessment

      This study presents a valuable finding on the spatio-temporal neural dynamics of language processing modulated by story context and presentation speed. The presentation of evidence in the version of the original submission is solid, as providing further clarifications on methodological details and addressing potential confounds would strengthen the study. The work will be of interest to cognitive neuroscientists working on reading and language comprehension.

    2. Reviewer #1 (Public review):

      Summary:

      This paper reports an important MEG study that is interesting from many different angles. By manipulating presentation rate (Fast vs. Slow) and "Structure" (word list vs. sentence list vs. story), the authors revealed the spatiotemporal dynamics of processing constituents at different speeds and semantic-conceptual scales. The methods are solid, and the results would be interesting to both researchers interested in the neural basis of structure building and researchers interested in the consequences of presentation rate, which, as the authors note, is understudied for visual sentence presentation. Overall, I enjoyed reading this paper and agree that it reports valuable findings for language processing, but there are a few points where clarity can be improved for readers to better evaluate and appreciate this work.

      Strengths:

      This paper studies the effects of presentation rate and linguistic structure building at different scales with MEG, which is perhaps one of the first MEG studies approaching these questions, and MEG is an appropriate technology to investigate spatiotemporal dynamics in the brain. The authors conducted a breadth of analysis, enabling us to fully understand what is shown by their data.

      Weaknesses:

      Here I would like to point out a few points where clarity can be improved (e.g., analytical details) for readers to better appreciate this paper.

      (1) The "mean proficiency" was reported on page 5, but I did not see the details of the proficiency task, which also need to be reported.

      (2) I was confused about the baseline correction part of the MEG analysis on page 10, and I would appreciate it if the authors laid out the rationale more clearly. From what I understand, the authors did not perform baseline correction, which is completely understandable, as word-lists, sentence-lists, and stories would have different baselines, and the baselines also keep building up as the trial evolves (which is a part of the study, instead of something to be parsed out). In this context, it becomes confusing why the authors singled out the Fast presentation condition to be unsuitable for baseline correction ("Since the prestimulus period cannot be analyzed during the Fast presentation..."). If the rationale is that the Fast condition did not have a long-enough blank screen between segments, this would not be a problem, as baselining is done in ERP RSVP studies anyway all the time. I am also curious what is meant by "The Fast and Slow presentation conditions were separated and demeaned independently". Would this make comparing the Fast and Slow conditions trickier? Overall, I think the analytical choices are potentially defensible, but more explanations of the rationale are needed.

      (3) For the Results section, the authors should check the text again to increase clarity for statistical reporting for readers. For example, on page 14, the p-values were not always reported, and the dfs for t-values were also not always reported. For "pairwise comparisons", one would expect a reporting like "ps < ..." instead of a singular "p < 0.001".

      (4) For behavioral results, was the analysis of RT directed at all trials or correct trials only? It would be worth checking whether the RT results hold true when analyzing only correct trials, or whether they were primarily driven by incorrect trials. Note that I don't find it a problem to include all trials for the MEG analyses (which the authors did), as we would believe that the participants were doing the task anyway despite it being challenging, and task difficulty was an inherent component of the research goal instead of something to be parsed out.

    3. Reviewer #2 (Public review):

      Summary:

      The main contribution of this study is that brain activations related to linguistic context varied as a result of presentation speed, with the main finding that increased activity for coherent stories relative to other conditions was reduced in fast presentation relative to slow. The results thus challenge certain assumptions about the nature of the brain dynamics of language processing, with certain effects even disappearing under faster presentations, which may be related to the processing mode of the participant. The results continue to establish the viability of a parallel presentation design, which generally produces results congruent with those of the literature.

      Strengths:

      The study contains a somewhat novel presentation method, illustrating its viability. The results are bolstered by a strong sample size (N=33) and robust analytic techniques. The conclusions are measured and appropriate to the results, and the manuscript is exceedingly clearly written and accessible to readers.

      Weaknesses:

      The spatial specificity of the effects is hampered by the use of MEG, particularly with minimal structural MRIs for participants. Thus, the conclusions of the study in the spatial domain are tentative and more general than might result from other studies.

      In addition, the general finding that faster presentation speed reduced activity overall (and eliminated it in the frontal cortex) appears to be somewhat contradictory to existing literature, which finds that sentences which are complex or difficult to process generally produce greater activation, particularly in the frontal cortex. These studies might be reviewed, and this (seeming) contradiction could be addressed.

    1. eLife Assessment

      Chen and colleagues use in vivo electrophysiology in rats to characterize distance-encoding neurons in the retrosplenial cortex during both random foraging and goal-oriented navigation. This valuable work identifies a subset of retrosplenial neurons that encode distance to a hidden goal, and it reports that head-direction coding is enhanced during goal-directed navigation relative to foraging. Although the analytical framework aims to isolate the unique contribution of goal distance to retrosplenial firing patterns, strongly covarying self-motion and spatial signals leave the main claims about distance coding incompletely supported.

    2. Reviewer #1 (Public review):

      Summary:

      Chen and colleagues utilize in vivo electrophysiology to characterize distance-encoding neurons in the retrosplenial cortex of rats during both random foraging and goal-oriented navigation. They observe a subset of RSC neurons that encode distance to a hidden goal location and that HD coding is enhanced during goal-directed navigation when compared to foraging. They also demonstrate that goal distance coding is preserved in the dark. Distance encoding has been shown in RSC in prior publications, but examining it with respect to a behaviorally relevant location will be of interest to the field. That said, the manuscript needs substantially more methodological detail before I am convinced that goal-distance-to-goal (DTG) coding is not an artifact of self-motion or of other spatial tuning already known to exist in the area. The authors also introduce several new machine-learning approaches that are hard to interpret without clearer justification or a demonstrated need. Addressing the points below would provide more convincing evidence for DTG coding and enhance readability.

      Strengths:

      The task is useful for determining whether the goal distance is encoded in neural populations.

      Retrosplenial cortex is an excellent candidate region for examining representations related to goal distance.

      The analytical framework is state-of-the-art and is useful for determining the contribution of goal distance to complex activation in retrosplenial cortex that possesses mixed selectivity.

      Weaknesses:

      The analyses/simulations intended to ascertain the relationship between other known spatial/self-motion codes in RSC and DTG coding are insufficient. I struggle with how DTG and speed can be convincingly disentangled given task structure. I suspect many DTG cells are in fact speed-modulated, and that if the task was flipped such that the animal had to run through the goal location rather than stop, DTG tuning curves would be mirrored. There are a couple of simple things that could help:

      (1) Show significantly more examples of DTG neurons (perhaps all of them) alongside their corresponding spatial (e.g., EBC and HD) and self-motion tuning curves (e.g., speed and angular speed) for multiple goal locations.

      (2) The relationship between the DTG tuning curve and the speed tuning curve should be presented.

      (3) Report what fraction of DTG cells are tuned to a random, non-goal location, what fraction qualify as DTG by chance, and how much better goal decoding is than decoding to a random location (Figure 2a).

      (4) We need visualizations of DTG reliability both within and across goal locations (see below).

      The GLM analyses are meant to address the unique contribution of distance to goal, but there are issues with this approach.

      (1) The five behavioral variables included in the analysis are not independent and will covary strongly. Forward selection is greedy, so once one member of a correlated set is admitted, the remaining members' unique contribution to held-out log-likelihood may fall below the 0.01 bits/spike threshold even if they are genuinely encoded. The pairwise correlation (or mutual information) structure among the five variables should be reported for all sessions, as well as the Δllh values of the variables rejected at each step, so readers can judge how close the near-misses were.

      (2) What is the L1 penalty and how was it selected? L1 shrinkage lowers each candidate's Δllh, so a stronger penalty yields smaller selected models. This is also important when considering that the basis sets for each variable have different dimensionality. Is a single L1 penalty shared across candidate models?

      Reliability of DTG responses.

      (1) The 1D DTG tuning curves lack error bars, which should be presented to convey reliability. These should be shown for individual DTG neurons for multiple goal locations.

      (2) The 2D DTG ratemaps in Figure 1 are not especially compelling; it would be helpful to see all examples in the supplement.

      Several observations suggest that some DTG neurons may actually be encoding boundaries.

      (1) The distribution of DTG peaks is strongly bimodal, with peaks either at the goal or at the boundary. Together with the mixed-selectivity results, this suggests that some DTG cells may show a boundary response or activity at the goal related to speed or acceleration. The manuscript at present does not effectively rule out this possibility. This could be addressed by showing that DTG coding is preserved across different goal locations.

      (2) An arena-expansion condition would clearly distinguish DTG responses from boundary responses.

      (3) It would be useful to see the peak sorted plot (1E) cross-validated within and across goal locations. I suspect that DTGs with intermediate-distance peaks shift with goal location while the others do not. If they remain fixed, it would solidify the presence of the phenomenon.

      More characterization of the neurons that are 'important' for goal distance decoding that are not DTG cells is needed (i.e., the IMP population). As I understand it, IMP and DTG cells were grouped for decoding analyses, but the two populations do not overlap. The IMP population is much larger than the DTG cells, and it is unclear what these neurons are doing. Moreover, it seems that the IMP sub-class could alone be used to decode distance to goal. This is difficult to reconcile with the framing of DTG cells as the substrate of goal-distance coding and deserves comment. Decoding should be conducted with the IMP cells alone to show what the DTG cells actually add.

    3. Reviewer #2 (Public review):

      Summary:

      Chen et al. consider the activity of retrosplenial cortex (RS) neurons during performance of an open-field navigation task in mice. Using a Ca-++ transient imaging approach to examine activity, the authors claim to find tuning to distance of the animal to a hidden reward location. The question of tuning to distance in RS is of much interest of late, with other works making claims. In this respect, the present work is interesting in that it utilizes an actual navigational task that does not explicitly demand encoding of distance and does consider an open-field environment. I do have reservations concerning the robustness of distance tuning.

      Strengths:

      Testing of distance coding in open fields during performance of an actual navigational task.

      Weaknesses:

      Lack of robust evidence for distance coding and head direction coding.

    1. eLife Assessment

      This valuable study is a rigorous and comprehensive anatomical description of populations of spinal-projecting neurons in the brain of the larval zebrafish. The manuscript consists of compelling data that reveal a far more complete and robust visualization of these populations than past work. The scholarship is also notable, featuring thoughtful contextualization of the findings with respect to both non-mammalian and mammalian literature.

    2. Reviewer #1 (Public review):

      Summary:

      The authors have achieved an excellent, thorough anatomical characterization of all spinal projecting neurons in the larval zebrafish. The comprehensive nature of their labeling approach and their quantification will make this work an instant reference benchmark for a wide number of zebrafish researchers. In addition, the scholarly approach in comparisons with other work in and outside of fish makes the manuscript valuable to researchers outside the field who would like to know how to translate between zebrafish and mouse terminologies.

      Strengths:

      The figures are clear and easy to follow. The literature review is impressive. The authors are careful to note the few limitations of their approach (eg the absence of Mauthner cell labeling and associated large neuron weak label). The manuscript does the whole field a major service.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    3. Reviewer #2 (Public review):

      Summary:

      The vertebrate spinal cord receives inputs from many supraspinal regions. The authors used optical backfilling to trace neurons in the zebrafish larval brain sending axons to the spinal cord. With two-photon microscopy, they managed to render a comprehensive 3D map of these neurons and drew homologs with mammalian brain structures.

      Strengths:

      The main strength lies in the precise 3D mapping. The fact that most of the previously reported neuron groups have been confirmed by their approach is a solid endorsement of their methodology.

      This study provides a comprehensive alternative anatomical reference framework for studying individual groups of supraspinal neurons with projections to the zebrafish spinal cord.

      Weaknesses:

      The whole approach could be enhanced by counter-staining their preparation to profile brain structures, including many nuclei more precisely.

      Also, the backfilling approach does not reveal the full trajectories of axons, which is already available to some degree by ZExplorer Atlas.

    4. Reviewer #3 (Public review):

      In this study, the authors aim to provide the most comprehensive and detailed topographic map to date of spinal projection neurons in the larval zebrafish brain. They achieve this by retrogradely photoactivating, in the rostral spinal cord, a photoconvertible GFP expressed pan-neuronally, and by constructing a template larval zebrafish brain atlas to regionalize the location of all labeled somata across the brain. The labeling strategy, together with the chosen animal model, provides strong support for the completeness of the dataset. The generation of a standardized anatomical atlas establishes a rigorous framework for analysis. Molecular and anatomical evidence suggesting evolutionary conservation of selected regions of interest appears solid.

      Overall, the authors successfully achieve their aim. By generating this atlas of spinal projection neurons, they provide not only an anatomical framework of the regions involved, but also an important reference for improving the orientation and regionalization of the zebrafish brain, which has historically been difficult to define. This work may serve as a valuable resource for future evolutionary, developmental, and comparative studies of spinally projecting neuronal populations implicated in diverse functions.

    1. eLife Assessment

      This important study introduces EvoDiff, an order-agnostic diffusion foundation model for protein sequence generation, and demonstrates its potential across unconditional generation, motif scaffolding, MSA-conditioned design, and IDR inpainting. The evidence supporting the principal claims is solid, combining computational evaluation with experimental validation, although the comparative advantages and generality of some applications would benefit from further clarification and calibration.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript presents a new foundation model, EvoDiff, for designing primary protein sequences. By leveraging an evolutionary-scale dataset, the model can be applied to evolution-guided sequence generation, sequence inpainting, and functional scaffolding.

      Strengths:

      The model provides an efficient approach for designing protein sequences and could be useful for developing protein therapeutics, engineering enzymes, designing biomaterials, and many other applications. The manuscript presents solid results showing that proteins designed by EvoDiff can achieve the same biological functions as their wild-type counterparts.

      Weaknesses:

      Compared with other sequence-generation models, EvoDiff does not substantially improve the success rate, suggesting that significant experimental effort is still required to screen and identify successful hits.

    3. Reviewer #2 (Public review):

      In this work, Alamdari et al. present EvoDiff, which provides the capability to generate protein sequences directly in sequence space, using a discrete diffusion model. There are several versions. EvoDiff-seq is trained on UniRef50 sequences (~42 million), and EvoDiff-MSDA operates instead by using sequence alignment methods to generate new members of protein families. The authors demonstrate many modes of sequence generation, including unconditional and conditional, inpainting of disordered regions, and also generating scaffolding of functional motifs. Their evaluation is also multifaceted, covering foldability, folding self-consistency, language embeddings, secondary structure distributions, and experiments for a set of different scenarios.

      The paper has many notable strengths. It is comprehensive in breadth, and the experimental component is distinctive, although I am not personally suited to review the rigor of that element.

      I would suggest that the paper's results certainly support the conclusion that order-agnostic sequence generation can yield useful candidates for multiple conditional design tasks. I am not totally convinced that it necessarily establishes diffusion as a generally superior approach to other competitors, like the conventional protein language models- EvoDiff is certainly competitive, and I think that the demonstration of diffusion is nice. I also am not sure that it is fair to say that sequence alone is sufficient for the broad design capabilities claimed (other than the "in principle" statement).

      I am overall quite supportive of the work and its demonstration, but I have a few comments for consideration in any revision.

      (1) I did not work through all dates of everything, but it appears to me that there are several recent conceptual and methodological competitors. These include DPLM and ProtBFN - both of these seem to be after the first preprint of EvoDiff, but given the time gap, there probably deserves to be some additional discussion or comparison. I would say, ideally, they should offer direct benchmarking. If the authors are disinclined, then I would think they should just temper their claims of contemporary SOTA performance or general superiority. Instead, the paper would still remain valuable as an early and experimentally demonstrated sequence diffusion framework. I don't think it needs to be more than that.

      (2) Related to the above, the manuscript should more carefully distinguish the demonstrated advantage of order-agnostic generation from the quality of unconditional generation. Regarding Figure 3, the authors argue that Evodiff's diffusion objective is necessary, but this does not seem to account for or address the LRAR baselines in Tables S1 and S3. Unless I am misunderstanding, the 640M LRAR model exhibits several better scores. The authors later suggest that EvoDiff's principal advantage is conditioning on arbitrary positions, which is valid, but that's a little different that what is being claimed. I suggest that the LRAR results should be shown or discussed alongside Figure 3, and then the authors revise to say that they have flexible conditional generation as the principal empirical benefit.

      (3) I really like the IDR experiment, but I'm not sure it demonstrates that Evodiff can design functional IDRS generally. Cox15 is a favorable target because its mature sequence strongly identifies a conserved mitochondrial protein, and EvoDiff-MSA is supplied directly with its orthologous family. The eight tested sequences were also selected from hundreds of candidates using both DR-BERT and MitoFates, making the experiment a test of the full generation-and-prediction pipeline rather than of EvoDiff alone. Moreover, mitochondrial targeting is tested, but the disordered character of the generated sequences is not experimentally established. In any case, I think it would certainly be more convincing if there were other examples, with unrelated proteins or IDR functions. I would appreciate that this is again a step beyond what the authors might be compelled to do, but their claim could be simply more calibrated.

      (4) The authors might benefit from explaining the advantages or complementarity of EvoDiff to other property-directed approaches for exploring sequence space. This has been done, for example, by using Bayesian optimization and genetic algorithms to tune properties of IDP condensates (DOI: 10.1021/acs.jpcb.8b03822). In my understanding, these are addressing a different problem from EvoDiff by optimizing sequences explicitly towards physical targets, while EvoDiff is a generative framework that can be used for sampling/inpainting/ etc. Is it clear how these strategies might be plausibly integrated? If so, that would be a relevant point of discussion and a potential advantage for EvoDiff.

    1. eLife Assessment

      This study examines how adolescent adversity influences maternal caregiving and social development across generations. The findings are potentially valuable, but several conclusions currently extend beyond what is directly supported by the data, rendering the strength of the evidence incomplete. Key mechanistic claims are limited by the communal reassuring design, and by several unresolved questions regarding causality and interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      In the manuscript by Francis-Oliveira et al., the authors investigated whether mild adolescent social isolation in female mice produces a latent vulnerability that emerges during the postpartum period as impaired maternal caregiving. They further tested the hypothesis that maternal deficits alter offspring social development. Based on their findings, the authors propose that adolescent psychosocial adversity disrupts maternal behavior and that resulting alterations in offspring social function are mediated through dysfunction of the midcingulate cortex (mCg) to prelimbic cortex (PrL) pathway. They further suggest that exposure to experienced parous females during the postpartum period can rescue maternal behavior and normalize offspring outcomes through restoration of activity in this circuit.

      To test these hypotheses, the authors exposed female mice to mild social isolation during late adolescence and subsequently bred those females. Maternal behaviors were assessed postpartum, and offspring were evaluated in adulthood using assays of sociability, social novelty recognition, social odor recognition, anxiety-like behavior, locomotion, and non-social memory. The authors also measured corticosterone levels in control and stress-reared offspring. To test circuit-specific effects, the authors employed chemogenetic activation and inhibition of the mCg→PrL pathway using DREADDs and performed electrophysiological recordings from identified projection neurons. Finally, stressed dams were co-housed with experienced parous females during the postpartum period to determine whether maternal and offspring phenotypes could be rescued, as well as the social deficits previously observed in offspring.

      The authors found that adolescent isolation selectively impaired pup-directed maternal behaviors, including nursing, licking, nest building, and pup retrieval, while leaving self-directed behaviors intact. Adult offspring of stressed dams exhibited deficits in sociability, social novelty recognition, and social odor discrimination, but showed no impairments in locomotor activity, anxiety-like behavior, or novel object recognition. Chemogenetic activation of the mCg→PrL pathway restored social behavior in stressed offspring, whereas inhibition of the pathway induced social impairments in controls. Co-housing stressed dams with experienced parous females restored maternal caregiving, normalized offspring social behavior, and rescued reduced firing of mCg→PrL neurons observed in offspring of stressed dams. Collectively, these findings support the authors' model that adolescent psychosocial adversity disrupts maternal caregiving and contributes to offspring social deficits through dysfunction of the mCg→PrL circuit.

      Strengths:

      The study includes multiple levels of assessment, including behavioral measures, behavioral intervention, the use of DREADDs for circuit manipulation to both test effects of activation versus inhibition on behavioral outcomes as well as physiology experiments. The multilevel approach is a strength.

      Weaknesses:

      (1) Interpretation of the parous co-housing experiment:

      The principal limitation of the study is that the communal housing paradigm does not distinguish rescue of maternal behavior in the stressed dam from direct caregiving provided by the experienced parous female. The authors interpret the rescue experiment as evidence that social support and/or social learning from experienced mothers restores maternal behavior in stressed dams, which in turn normalizes social behavior in offspring. However, pups were continuously housed with both the stressed dam and the parous female from P0-P7, and the parous female had unrestricted access to the pups throughout the intervention period. The parous female was removed only briefly during maternal behavior testing. This design raises an important alternative interpretation. The experienced parous female may have directly provided substantial maternal care to the pups, supplementing or compensating for deficits in the stressed dam. The rescue of offspring phenotypes may reflect care received from the parous female rather than improved caregiving by the stressed dam.

      Were caregiving behaviors of the stressed dam and parous female quantified separately during the co-housing period? What proportion of licking, nursing, retrieval, and nest maintenance was performed by each animal? Can the authors exclude the possibility that direct maternal care from the parous female, rather than social learning or social support, accounted for the rescue of offspring outcomes? Without such controls, the central claim that restoration of maternal behavior in the stressed dam mediates normalization of offspring phenotypes is not fully supported.

      (2) Specificity of the behavioral phenotype and rescue:

      The manuscript repeatedly frames the findings as restoration of offspring outcomes and intergenerational vulnerability. However, the behavioral phenotype appears highly selective and restricted primarily to social behaviors. Offspring exhibited impairments in sociability, social novelty recognition, and social odor discrimination, but showed normal locomotion, anxiety-related behavior, and non-social memory. Thus, the authors should more explicitly acknowledge that maternal adversity produced a domain-specific social phenotype rather than broad behavioral dysfunction. Interestingly, the rescue studies only evaluated a subset of the affected behaviors, making it unclear whether co-housing with parous females restored broader aspects of offspring neural or behavioral function or selectively improved specific social behaviors.

      Why was social olfactory recognition not included in the rescue experiments? Why were additional behavioral measures not reported following circuit activation or parous co-housing (e.g., anxiety-like behavior and novel object test) - were these improved in controls by enriched early parenting? Did manipulation of the mCg→PrL pathway or co-housing with parous females influence anxiety-like or depressive-like behaviors despite the absence of baseline group differences?

      (3) Strength of the DREADD-mediated causal claims:

      The DREADD experiments implicate the mCg→PrL pathway in regulating social behavior; however, several aspects limit the strength of the causal conclusions. Sample sizes were relatively small (n = 6/group). In addition, the variance observed in the DREADD cohorts appears substantially reduced relative to that observed in the non-surgical cohorts. For example, vehicle-treated groups appear more clearly separated than would be expected based on the original behavioral data presented in Figure 2. Can the authors comment on potential reasons for this discrepancy?

      Second, although activation and inhibition experiments support involvement of the mCg→PrL pathway in social behavior, the manipulations do not fully recapitulate the broader phenotype observed in stressed offspring. Thus, the data support a role for this pathway but may not justify the stronger conclusion that dysfunction of this circuit alone accounts for the entirety of the offspring phenotype.

      The authors argue that altered maternal behavior is causal for social deficits in offspring. However, in the absence of a cross-fostering experiment, the study cannot fully exclude alternative explanations, including direct influences of caregiving by the co-housed parous female, gestational effects, altered maternal physiology during pregnancy, or germline-mediated influences. A cross-fostering design would substantially strengthen the causal interpretation of the findings.

      Additionally, clarification of litter effects is important: For example, how many litters contributed to each experimental group? Was litter treated as a random effect in statistical analyses? Were multiple offspring from the same litter analyzed as independent observations? Because maternal behavior is manipulated at the litter level, litter rather than individual offspring may represent the appropriate experimental unit for many analyses.

      (4) Integration of corticosterone findings into the mechanistic model:

      The corticosterone findings appear somewhat disconnected from the central mechanistic narrative. The authors report elevated corticosterone levels in stressed offspring and suggest that HPA-axis dysregulation may contribute to the observed behavioral phenotype. However, the manuscript does not establish whether corticosterone plays a causal role in the social deficits or instead represents a parallel physiological consequence of altered maternal care.

      Were corticosterone levels normalized by co-housing with parous females? Did DREADD-mediated activation of the mCg→PrL pathway normalize corticosterone levels? Could corticosterone manipulation alone drive aspects of the behavioral phenotype independent of circuit manipulation? Do corticosterone levels correlate with the severity of social behavioral impairments? It is difficult to determine whether corticosterone is mechanistically relevant or simply serves as an associated physiological marker. The authors should either more directly integrate the endocrine findings into their mechanistic framework or temper discussion suggesting a causal role for HPA-axis dysfunction.

      In summary, this manuscript addresses an important and understudied question concerning how adolescent adversity influences maternal caregiving and offspring social development. The behavioral, circuit, and electrophysiological findings are generally coherent and support a role for the mCg→PrL pathway in mediating offspring social outcomes. However, the strongest mechanistic claim, that social support rescues offspring phenotypes by restoring maternal behavior in stressed dams, is weakened by the communal rearing design, which allows direct caregiving by parous females. Additional clarification regarding caregiver-specific behaviors, litter effects, the role of corticosterone, and the specificity of the DREADD-mediated phenocopy would strengthen the causal interpretation of the findings. Overall, the study is potentially impactful, but several conclusions currently extend beyond what is directly supported by the data.

    3. Reviewer #2 (Public review):

      Summary:

      Studies in rodents have demonstrated that early life adversity (ELA) impacts many aspects of the exposed offspring's brain and behavior. Work in this field has traditionally focused on how<br /> stress in very early life can impact cognitive and emotion-related behaviors in the ELA-exposed offspring. By contrast, this manuscript focuses on how stress in a slightly later adolescent period can produce latent and intergenerational effects by impacting maternal caregiving from female offspring, as well as social outcomes of the next generation of animals born to ELA-exposed females. Specifically, the manuscript describes that female mice exposed to ELA in the form of social isolation in late adolescence show reduced pup-directed maternal behaviors, while self-directed behaviors remain intact. Offspring reared by these dams in turn show deficits in social behavior, which are linked to reduced activity in an excitatory connection between the medial cingulate cortex (mCg) and prelimbic cortex (PrL). Further, the authors find that social behavior can be rescued by chemogenetic activation of the mCg-PrL pathway in the offspring of ELA-exposed/stressed mice or recapitulated in control mice by chemogenetic inhibition of this connection. Importantly, co-housing ELA-exposed/stressed dams with experienced parous females during the early postpartum period restores pup-directed maternal behaviors in these mice and normalizes offspring social outcomes as well as mCg-PrL activity.

      Strengths:

      Strengths of the manuscript include the focus on an important and novel question about intergenerational effects of adolescent ELA transmitted via subsequent maternal care, and the use of multiple techniques to link circuit function to behavior, including slice electrophysiology and chemogenetics. While the findings that maternal care can influence offspring behavior and that experienced females can instruct and improve maternal care of less experienced mice are not novel, they add support to this important area of literature.

      Weaknesses:

      Weaknesses of the paper include the lack of validation that the viral chemogenetic paradigm was appropriately targeted in the brain and impacted the excitability of mCg to PrL projections as anticipated, the use of inappropriate statistical tests that do not account for non-independence of pups from the same litter or cells measured from the same pup or categorical versus continuous data, and the lack of important descriptions of methods or experimental paradigms in several places that altogether make it difficult to judge the rigor of the findings in its current state.

      If these weaknesses are addressed, these findings will provide important information about circuit mechanisms underlying intergenerational effects of adolescent stress on social behavior in next-generation offspring.

    1. eLife Assessment

      This is a valuable study of the effects of selective or broadband amyloid deposition in medial septum (MS) cholinergic neurons on neuropathology, cognition, sleep, and hyperexcitability in aging mice carrying risk factors for Alzheimer's disease (AD). The investigation is still incomplete, with some weaknesses related to conceptualization and methodology that need to be addressed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

    4. Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ⁻ is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ⁻ and AAV‑Aβ⁺ is marked by asterisks (***) placed above AAV‑Aβ⁺. In addition, the single asterisk above KI-Aβ⁺ does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ⁻ vs WT‑Aβ⁻ or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ⁺ mice lack a preference for the novel object compared with AAV‑Aβ⁻ controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ⁻, and the notation over AAV‑Aβ⁺ is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ⁺, KI-Aβ⁺⁺, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Thank you for the positive assessment and for recognizing the novelty of our findings, the complementarity of our experimental approaches, and the breadth of our multi-modal readouts. We will address the weaknesses raised below point by point.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      This is an insightful comment. We fully agree that early functional changes preceding structural degeneration represent an important and exciting avenue, and we will explicitly acknowledge this in the revised manuscript, including the relevant literature on early cholinergic hyperactivity. Examining earlier time points in our model to capture such changes is a compelling perspective that we will discuss as a key direction for future work.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      We acknowledge this limitation and will increase the sample size for the tracing experiments in the revised manuscript. However, we wish to clarify that the tracing was performed in a separate cohort of young animals specifically to characterize baseline MS cholinergic projections independently of any amyloid-related disturbances. While we agree that replicating these experiments in APP mice would be of great interest, this falls outside the scope of the current study and will be highlighted as an important direction for future work.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      We agree with this pertinent suggestion. We will perform and report region-specific analyses of the hippocampal formation in the revised manuscript, in line with our discussion of specific amyloid accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      We respectfully maintain our analytical approach. As the test phase was terminated upon reaching a predefined cumulative exploration time of 20 seconds rather than using a fixed trial duration, computing a discrimination index is not appropriate in this context, as total exploration time is constrained by design. We instead followed the validated protocol described by Leger et al. (2013, Nature Protocols; PMID: 24263092), which controls for inter-individual differences in exploratory motivation by fixing cumulative exploration time, ensuring equivalent sampling conditions across animals. We will clarify this methodological choice in the revised manuscript.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

      We appreciate this constructive suggestion. We will provide additional methodological detail on interictal spike detection, include representative examples, and perform a vigilance state-specific analysis in the revised manuscript. We agree this will provide valuable mechanistic insight and will update the reference accordingly.

      Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Thank you for this positive assessment and for recognizing the rigor of our experimental design, the value of our multi-modal approach, and the mechanistic significance of our cholinergic lesion comparison. We will address the weaknesses below point by point.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

      Regarding amyloid broadcasting, we fully acknowledge that the precise mechanism remains to be elucidated; while this was not a primary objective of the study, it represents a fascinating and unexpected finding that we will discuss more carefully as an open question for future investigation. Regarding the causal relationship between MS<sup>ChAT</sup> cell loss and the observed phenotypes, we agree that the specific circuits involved remain to be fully resolved; however, we would like to emphasize that the convergent evidence from our three complementary models (and in particular the recapitulation of cognitive and REM sleep deficits by selective cholinergic ablation) provides strong causal support for MS<sup>ChAT</sup> neuronal loss as a key driver of these phenotypes, independent of amyloid deposition per se.

      Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Thank you for acknowledging the compelling conceptual premise of our study, the thoughtful and broad experimental framework, and the value of our longitudinal sleep analysis and lesion comparison. We will address all concerns raised below point by point.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      This is a valid point. We agree that the term “prodromal phase” requires more careful framing in the context of animal models, and we will revise the title and relevant statements accordingly. We would like to note, however, that the REM sleep disturbances we report seem to emerge prior to overt cognitive decline in our longitudinal analysis, which we consider to reflect a prodromal-like feature of the model. Nevertheless, we will adopt more precise terminology throughout the manuscript to avoid conflation with clinical staging criteria.

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      We appreciate this comment and will carefully nuance our opening statement to better reflect the current state of evidence. However, we respectfully note that several reports, including Pase et al. (Neurology, 2017; PMID: 28835407) and Ibrahim et al. (Sleep, 2024; PMID: 38001022), have demonstrated that REM sleep loss is associated with increased risk of incident neurodegenerative disorders, particularly Alzheimer's disease, supporting the broader validity of our framing. We will revise the statement to more accurately capture the complexity of this relationship while preserving its scientific relevance.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      We agree and will revise this sentence to better reflect the multi-system complexity of sleep/wake regulation, acknowledging the interplay between multiple neurotransmitters and neuromodulatory systems.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      We agree and will revise this sentence to clarify that ACh is a crucial component of a broader REM sleep-generating circuitry, rather than a sole requirement, while better contextualizing its specific contribution to REM sleep regulation.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      We will revise this statement to more accurately reflect the evidence. We would like to note, however, that while the Pase et al. (2017) cohort included a mixed dementia population, 75% of incident dementia cases (24 out of 32) were consistent with Alzheimer's disease, lending meaningful support to the relevance of REM sleep alterations specifically in the context of AD. We will ensure this nuance is clearly conveyed in the revised manuscript.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      We agree and will revise this section to acknowledge the well-established contribution of tau pathology to BF cholinergic neuronal vulnerability in human AD, including its early accumulation at Braak stages I-II. We wish to clarify, however, that the present study focuses exclusively on amyloid-driven mechanisms, and the introduction will be revised to ensure readers do not infer a purely amyloid-dependent mechanism in the broader human disease context.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      As noted in our response to weakness (1), we will systematically revise the manuscript to replace “prodromal phase” with more precise terminology that clearly distinguishes our animal model findings from clinical disease staging.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      Thank you for raising these important points. Regarding age matching, we acknowledge that chronological ages are not perfectly aligned across groups; however, we wish to emphasize that the duration of amyloid pathology is carefully matched across models. Indeed, AAV injection in MS<sup>ChAT</sup>-AppNL-G-F mice marks the onset of amyloid expression, directly paralleling the onset of pathology from birth in App<sup>NL-G-F/NL-G-F</sup> knock-in mice. We believe pathology duration represents the most biologically relevant variable for comparison in this context, and we will clarify this in the revised manuscript. Regarding the MS<sup>ΔChAT</sup> cohort, animals were culled upon reaching a comparable degree of REM sleep loss, providing a functionally meaningful matching criterion. Finally, regarding sham-operated controls, we acknowledge this limitation; however, based on our experience, surgical procedure alone has negligible effects on the cellular populations under study, and the inclusion of an additional sham group across all experimental cohorts would have required a prohibitive number of animals, raising significant ethical concerns under the 3R principles. We will address these points more explicitly in the revised manuscript.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      Thank you for this observation, we will correct the figure citation accordingly in the revised manuscript.

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      We appreciate this important observation. We will carefully re-examine our tracing data to determine whether MS<sup>ChAT</sup> axons specifically innervate the LHA, and whether amyloid deposition was detectable in this region in our MS<sup>ChAT</sup>-App<sup>NL-G-F</sup> model. We agree that clarifying the potential anatomical relationship between MS<sup>ChAT</sup> projections and LHA-MCH neurons is important to properly position our proposed circuit mechanism relative to this well-established REM sleep-promoting node, and we will address this in the revised manuscript.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      Thank you for catching this inconsistency. We confirm that this is an error in the manuscript: line 125 should read “15-16 month-old” referring to the chronological age of the animals at the time of perfusion, rather than “13-14 months,” which corresponds to the duration of AAV expression. We will correct this in the revised manuscript.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      Thank you for flagging this ambiguity. We will revise the sentence to clearly distinguish between AAV-injected and global knock-in cohorts in the revised manuscript.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      The supplementary information referenced in the main text will be provided in full in the revised manuscript, together with analyzed datasets and analysis scripts, in accordance with eLife's data sharing policy.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      We agree that a comprehensive description of our histological quantification pipeline is essential. We will provide full methodological details in the revised manuscript, including: the metric used for Aβ load quantification, ROI definitions, Fiji/ImageJ thresholding algorithms (accounting for batch effects and autofluorescence normalization), as well as the strategy for identifying and counting ChAT-positive neurons, section spacing, number of sections per animal, hemisphere coverage, normalization, and blinding procedures. All ImageJ scripts will be made available to ensure full transparency and reproducibility.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ<sup>-</sup> is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ<sup>-</sup> and AAV‑Aβ<sup>+</sup> is marked by asterisks (***) placed above AAV‑Aβ<sup>+</sup>. In addition, the single asterisk above KI-Aβ<sup>+</sup> does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ<sup>-</sup> vs WT‑Aβ<sup>-</sup> or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      We acknowledge that the current statistical annotations can be difficult to interpret when multiple experimental groups are displayed within a single plot. While the notation is consistent across figures, we agree that clarity can be improved, and we will revise all figure annotations to explicitly link significance indicators to their corresponding pairwise comparisons, using standardized brackets throughout, with all contrasts clearly specified in the figure legends.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ<sup>+</sup> mice lack a preference for the novel object compared with AAV‑Aβ<sup>-</sup> controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ<sup>-</sup>, and the notation over AAV‑Aβ<sup>+</sup> is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ<sup>+</sup>, KI-Aβ<sup>++</sup>, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      We wish to clarify that in Figure 5D, the absence of novel object preference manifests as equal exploration of both objects (approximately 10 seconds each), such that comparisons against the 10-second chance level reflect this lack of preference. We will revise the annotations and figure legend to make the statistical comparisons explicit and fully consistent with the results text.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      This is a constructive feedback, and we agree that the Discussion would benefit from pruning and recalibration. We will streamline it to focus on the most evidentially supported points, particularly those linking cholinergic loss to REM sleep and cognitive outcomes, while toning down or removing speculative statements regarding disease-stage equivalence, prion-like spreading mechanisms, and direct therapeutic implications.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      Thank you for raising this important point. We will expand the Discussion to address this apparent contradiction with the human literature. Indeed, while increased anxiety is reported in early AD stages, it tends to normalize or decrease at later stages (Botto et al., 2022; PMID: 35461471), which may partly reconcile our findings. In addition, anxiety-like phenotypes are highly inconsistent across AD mouse models, varying with model type, age, sex, and behavioral assay (Pentkowski et al., 2021; PMID: 33979573). Anxiety-related changes in human AD may reflect damage to brain regions beyond the MS cholinergic system, involving additional circuits and mechanisms not captured by our model.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

      We will nuance our wording at line 654 by replacing “comparable durations of amyloid pathology” with “comparable amyloid burden,” acknowledging that while both cohorts were aged for matched durations, the kinetics of amyloid progression may inherently differ between an AAV-driven focal model and a germline knock-in model.

    1. eLife Assessment

      This important study investigates how viral macrodomains in dual-host viruses functionally contribute to their infection in the vector, which was previously understudied. The findings that chikungunya virus (CHIKV) macrodomain mutants have unique impacts on virus replication in mammalian cells and in live mosquitoes are convincing and significant; however, the study can be strengthened by further investigating the role of ADP-ribose binding and potential other compensatory mutations. The work will be of interest to virologists and biologists who study macrodomains and ADP-ribosylation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

    3. Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      We thank the reviewer for raising this important point. The emergence of second-site mutations at residue D31 in Vero cells was actually observed in two independent experiments, each initiated by transfection of viral RNAs encoding either the N24A or N24D mutant. In the second experiment, additional timepoints were sampled for sequencing. Comparison of the two experiments revealed that, in the second replicate, only the D31N variant was recovered, whereas D31H was not detected. We will include the results of both independent experiments in the updated Figure 1 and will modify the text accordingly to improve clarity.

      In addition, we will present new experiments performed in other cell types that further confirm the reproducibility of D31 mutation emergence across different cellular contexts. Specifically, we will include results from new viral RNA transfections in BHK-21 mammalian cells and C6/36 mosquito cells, which will be included in the updated Supplementary Figure 1.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      We thank the reviewer for this correction. The sequencing was indeed focused on the region of the macrodomain surrounding the introduced mutations, covering the first 119 amino acids of nsP3. This region was selected to confirm the stability of the introduced mutations at position 24 and to monitor the emergence of potential second-site mutations in its immediate vicinity. We acknowledge that this approach does not rule out the emergence of compensatory mutations elsewhere in the macrodomain or in the full-length nsP3 protein. We will correct the text accordingly to accurately reflect the region that was analyzed.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      This is an excellent point. To address it, we introduced N24A and N24D mutations into an infectious clone of the Indian Ocean strain, a representative of the ECSA lineage, and assessed the emergence of D31 second-site mutations during viral stock production. As observed with the Caribbean strain, the D31N secondary mutation consistently emerged in both N24A and N24D Indian Ocean mutant viruses. These results will be included in the updated Supplementary Figure 1 and in the main text.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      We thank the reviewer for raising this point. We acknowledge that previous studies have reported limited infection of A549 cells by certain CHIKV strains. However, in our hands, with our viral stocks and under our experimental conditions, we observe an increase in viral titers following infection of A549 cells with WT Caribbean strain virus, indicating productive viral replication. We also confirmed that the Indian Ocean strain replicates in A549 cells under the same conditions, and we will include growth curve data for this strain in the updated version of the manuscript. Importantly, we did not observe detectable cytopathic effects in A549 cells infected with either WT or N24 mutant viruses, which is consistent with the notion that CHIKV replicates less efficiently in this cell line compared to other mammalian cell lines. We will include a sentence acknowledging this in the discussion of the revised manuscript.

      Regarding the suggestion to use human fibroblasts, we respectfully note that the primary focus of this study is the role of macrodomain catalytic activity in the mosquito host, and the experiments in A549 cells were performed to confirm the known importance of the macrodomain in interferon-competent mammalian cells. We therefore consider the current data in A549 cells sufficient to support this conclusion within the scope of the manuscript.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      We acknowledge that the full viral genome was not sequenced, and we cannot rule out the presence of additional mutations elsewhere in the genome that could contribute to the observed phenotype. However, we note that the enhanced dissemination phenotype was also observed in independent experiments performed with the Indian Ocean strain mutant viruses in Ae. albopictus, which were generated independently from the Caribbean strain stocks. The consistency of the phenotype across two independently generated sets of mutant viruses from different CHIKV lineages strongly supports the conclusion that the enhanced dissemination phenotype is linked to the macrodomain mutations. These new data will be included in the updated manuscript.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

      We thank the reviewer for pointing this out. The text will be modified accordingly to accurately reflect that we assessed transmission potential, based on viral dissemination to heads, rather than actual transmission.

      Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      We agree that the mutants may indeed have stronger binding for modified substrates than the WT protein; however, given that DSF is not a quantitative measure of binding affinity and that free ADPr is not the relevant ligand (in fact we do not know the relevant ADPr-modified molecule), we have refrained from speculating further than saying in the discussion:

      “The progressive selection of D31H over D31N in the mosquito host further suggests that subtle differences in ADP-ribose substrate recognition may influence viral fitness in the mosquito environment in ways that are not yet understood.” (Lines 368-370)

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

      Regarding the title, we agree with the reviewer that it could be misleading, as both mutants lack catalytic activity yet show distinct phenotypes in mosquitoes. We will therefore modify the title to: "Loss of macrodomain catalytic activity modulates Chikungunya virus dissemination and transmission potential in Aedes mosquitoes", which more accurately reflects that the observed phenotypes arise as a consequence of the loss of catalytic activity at position N24.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      We thank the reviewer for this important comment. Regarding the generation of viral stocks in A549 cells, transfection of N24A and N24D viral RNAs into A549 cells did not yield sufficient viral titers to produce usable stocks. However, as described in our response to Reviewer #1, we confirmed the reproducible emergence of D31 second-site mutations in C6/36 mosquito cells and BHK-21 mammalian cells, which will be included in the updated Supplementary Figure 1. These results demonstrate that D31 mutations spontaneously emerge in stocks generated directly in mosquito cells, addressing the reviewer's concern.

      Regarding the question about insect-specific alphaviruses, we examined the sequence at position 31 using a multiple sequence alignment of 14 alphaviruses with diverse host ranges, including dual-host, insect-specific, and aquatic alphaviruses. We observed that Yada Yada virus (GenBank: QGR15362.1) has an asparagine (N), Tai Forest alphavirus (GenBank: YP_009333615) has an aspartic acid (D), Mwinilunga alphavirus (GenBank: BBC45634.1) has an aspartic acid (D), Eilat virus (GenBank: QBG67155.1) has an aspartic acid (D), and Agua Salud alphavirus (GenBank: QEV83787.1) has a lysine (K) at this position. These results suggest that insect-specific alphaviruses do not share a conserved residue at position 31, and therefore no clear conclusion can be drawn regarding a direct correspondence with the compensatory mutations observed in our study. This sequence alignment with the corresponding text will be included as supplementary data in the revised manuscript.

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      We thank the reviewer for this observation. We acknowledge that the prevalence of WT virus infection shows variability between experiments and timepoints. Based on our experience with infectious blood meal experiments, this variability is sometimes observed between independent experiments even under identical experimental conditions and with the same virus. In this particular case, each set of experiments was performed with independently produced viral stocks. Specifically, for the experiments shown in Figure 3, WT, N24A-D31N, and N24D-D31H/N viral stocks were produced in parallel from transfection of viral RNAs, and the same stocks were used for sequencing, growth curves, and mosquito infections. Subsequently, when we generated single D31H and D31N mutant viruses, a new WT viral stock was produced in parallel with the D31 mutant stocks, and these independently produced stocks were used for the experiments shown in Supplementary Figure 2. Differences in absolute infection rates between experiments are therefore expected, as they reflect both the use of independently produced viral stocks and the inherent variability in the efficiency of midgut infection and dissemination between mosquito cohorts. Importantly, the comparisons between WT and mutant viruses are always made within the same experiment, using stocks produced in parallel.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    1. eLife Assessment

      This important study identifies and characterizes a set of amino acid states that can rescue protein function in the presence of substantially deleterious mutations. Some of these super-compensatory substitutions also confer substantial mutational robustness, with broader implications for understanding epistasis, protein evolution, and protein engineering. The evidence is convincing, supported by a creative reanalysis of a large deep-mutational-scanning dataset, statistically rigorous treatment of anticipated error rates, experimental validation, and analyses of additional proteins, although clearer presentation and deeper investigation of the evolutionary implications and structural mechanisms would further strengthen the study.

    2. Reviewer #1 (Public review):

      Summary:

      The study identifies and characterizes a set of amino acid states that make the protein robust to other mutations, to the point of being able to compensate mutations that render wildtype proteins entirely non-functional. The study uses a previously published dataset and uses it to find and study such super-compensators. It then analyzes the biophysics and fitness landscape structure of what may be behind the compensation, identifying stability as an important parameter that, nevertheless, is not sufficient to explain all of the compensatory effect. These findings have important implications for our understanding of protein evolution, with these super-compensators possibly acting in a role of "permissive mutations" and opening up evolutionary trajectories that may be closed without them. Perhaps the identification of such super-compensator substitutions can be incorporated into various protein design approaches.

      Strengths:

      The paper presents a compelling case with a rigorous analysis of the expected error rates of observation. While not unique, the current state-of-the-art in the field typically does include experimental error rate estimation like this work. The paper also does a good job in exploring the issue, including looking at plausible biophysical basis of super-compensators.

      Weaknesses:

      The paper lacks rigor in talking about evolutionary-related issues of the state of the fitness landscape. As an example, the paper mentions that these super-compensators flatten the landscape. While I understand where this is coming from, I think that the fitness landscape in this context is a static entity and cannot be flattened or otherwise altered. A much more accurate description is that a sequence with a super-compensator is located in a flatter-than-expected segment of the fitness landscape, or on a flat fitness ridge. These issues are more semantic in nature, and while the manuscript would benefit from it being shown to an expert in molecular evolution or fitness landscapes, this issue does not take away from the importance of the results.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents an interesting and conceptually valuable analysis of compensatory evolution using a large combinatorial deep-mutational-scanning dataset for yeast His3p.

      Strengths:

      I particularly like the identification of "super compensatory" substitutions that improve fitness across diverse genetic backgrounds and apparently reduce the sensitivity of the local fitness landscape to subsequent mutations. The work connects epistasis, protein stability, mutational robustness, and evolvability in a clear and potentially broadly relevant manner.<br /> The authors provide several complementary lines of evidence in support of this central conclusion. In particular, the new experimental validation of S189A is an important strength because it directly demonstrates that a predicted super compensator can buffer the effects of diverse deleterious substitutions, while analyses of additional DMS datasets from other proteins and assay systems suggest that the phenomenon is not restricted to the original His3p landscape.

      Weaknesses:

      The structural analysis currently relies primarily on correlations with RSA, weighted contact number, conservation, and Rosetta-predicted changes in folding or binding energy. For super compensators, the mechanistic evidence is largely limited to predicted stabilization and individual examples, such as the proposed salt bridge between 110D and R112. I believe that the newly developed structure-aware deep-learning approaches could provide useful information on the mechanism of super compensators. For example, an inverse-folding model such as ESM-IF1 could score complete multi-mutant sequences conditioned on the His3p backbone and test whether adding a super compensator restores sequence-structure compatibility across backgrounds. More recent multimodal mutation-effect or stability models could similarly be used to cross-check the Rosetta results, including models that explicitly support combinatorial mutations. I would not recommend simply comparing AlphaFold confidence scores between mutants, because current structure predictors are not necessarily sensitive to subtle mutation-induced energetic or conformational changes.

      The manuscript states that the pipeline was applied to 217 ProteinGym datasets and concludes that super compensators are broadly distributed across proteins and assays. However, this central generalization is described in only a few sentences and is largely relegated to Figure S7. The Methods do not explain which datasets contained sufficient combinatorial mutants to calculate compensatory ability or buffering, how many genotype pairs or quadruplets were available per substitution, or how differences in assay scale and library design were handled. This point requires clarification because supercompensation is inherently a background-dependent property and cannot be established from single-mutant measurements alone. ProteinGym is widely used as a substitution-effect benchmark, and many of its constituent assays primarily contain single substitutions; for example, an analysis of an earlier ProteinGym collection reported that 76 of 87 assays contained only single substitutions. It is therefore unclear how the same compensatory-interaction pipeline could be applied uniformly to all 217 datasets.

      The analysis of 335 His3p orthologs in Discussion is potentially very interesting, but co-occurrence between super compensators and putatively deleterious amino-acid states does not by itself demonstrate evolutionary compensation. Closely related species share substitutions through common ancestry, and both states could be associated with a particular lineage or ecological context. A tree-aware analysis would considerably strengthen this result. The authors could reconstruct ancestral states and ask whether acquisition of a super compensator tends to precede or accompany otherwise deleterious substitutions. Alternatively, they could use phylogenetically informed permutations that preserve substitution frequencies and shared ancestry.

    4. Reviewer #3 (Public review):

      The manuscript by Jiang and co-authors presents an analysis of experimental measurements (about 400k variants) from a deep mutational scan of the HIS3 enzyme. The authors assess the ability of a genotype to be "rescued" and show that this depends on mutation sites (in particular their solvent accessibility) and mutation effects (should be mild on folding stability or binding affinity). They further identify a set of super-compensatory mutations, and their results suggest that these mutations flatten the fitness landscape.

      This finding is interesting and likely of interest to a broad community. The analysis seems sound.

      However, I have a number of major concerns regarding the presentation and positioning of the work.

      (1) It would improve the manuscript to clarify the present contribution with respect to a previous study by the same authors, namely Pokusaeva et al. 2019. Did the authors apply the same protocol to generate a new library of mutants, or did they re-analyse an already published library? If the library is not new, ambiguous sentences like "Nevertheless, to our knowledge, the His3p library remains one of the largest and most comprehensive resources that contains multi-site mutants" should be reformulated.

      (2) Pokusaeva et al. 2019 is cited for the library and also for the deep neural network. It would be beneficial to briefly describe the architecture, the inputs and outputs, and the training procedure. Was the network trained on the current library? What is the purpose of this network? It looks more like an additive linear model (except for the global sigmoid) than a deep neural network. How does it relate to global epistasis models? The sigmoid function is designed to capture plateauing effects; doesn't that introduce some circularity issue in the reasoning?

      (3) Are the super-compensatory mutations observed (conserved) across evolution? Beyond the fact that they are accompanied by mildly deleterious mutations in natural sequences. Can we predict them with variant effect predictors?

      (4) The AAindex mention should be accompanied by a citation.

      (5) Equations should be numbered. WCN formula seems to contain misformatting issues.

      (6) A more explicit description of the structural data analysed (which PDB entry?) should be provided.

      (7) I believe the citation Van Cleve and Weissman 2015 for the ProteinGym benchmark is incorrect. Additionally, is the Rosetta citation adequate?

      (8) How is the definition of rescueability sensitive to the threshold choice?

    1. eLife Assessment

      This important study investigated whether the adoption of different explicit strategies influences implicit recalibration during visuomotor adaptation. Through a series of increasingly controlled experiments, the authors demonstrated that implicit recalibration is relatively insensitive to the specific type of explicit strategy employed but is influenced by the variability of strategic motor plans. However, the evidence supporting these conclusions is currently incomplete. The findings will be of interest to cognitive scientists studying motor control and its relationship to higher-level cognitive processes, provided that the key claims are supported by more rigorous analyses and experimental approaches.

    2. Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

    3. Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback).

      Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

    5. Author response:

      Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      We appreciate the reviewer’s thoughtful and comprehensive summary of our study.  One thing we would like to clarify is that the non-critical targets used in Experiment 1 received the same type of online continuous feedback as the critical target.Therefore the extent of implicit recalibration should be comparable between the critical and non-critical targets in Experiment 1. In contrast, for Experiment 2, we tightened the control for implicit recalibration at the non-critical targets by: 1) delivering delayed endpoint feedback for all non-critical targets while keeping the online continuous feedback for the critical target, which is known to suppress implicit recalibration, and 2) we widened the spatial gap between the critical and non-critical targets, to minimize spillover 2) As a result, the implicit recalibration was diminished at the non-critical targets, contributing to the shrinkage of the generalization curve around the critical target. 

      One other issue that we would like to make clear is that Experiment 1 was not a straight replication of a previous study, at least to our knowledge. We believe that the reviewer is referring to our previous study (McDougle and Taylor 2019), which found broader generalization for algorithmic strategies (Experiment 4). However, in that study cursor feedback was always delayed. As such, the observed broader generalization was most likely due to the strategy itself and not implicit recalibration. 

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      We are glad to know that our primary finding and conclusion appears reasonable. While we acknowledge that the finding doesn’t appear to be particularly exciting at face value, it does speak to larger questions regarding the independence of different learning systems and how just statistical or surface-level differences in training can result in relatively large differences in apparent behavior that could be easily misinterpreted as the result of system interactions.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      This is a fair concern, as RT is an indirect marker of strategy use and, by itself, cannot establish that participants used different strategies at the Critical target. Prior work, however, provided guidance for the design of our experimental manipulations. Algorithmic strategies, in which an aiming solution is computed online, are associated with longer RTs, whereas retrieval of a previously cached stimulus–response association produces substantially shorter RTs (McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Moreover, caching becomes increasingly difficult as the number of target-specific solutions increases, particularly beyond approximately four targets (Velazquez-Vargas and Taylor, 2024; Bejjanki and Taylor 2026). Our manipulation was designed around these findings: participants in the Algorithmic condition learned the 45° rotation across 10 targets, whereas participants in the Retrieval condition repeatedly encountered the 45° rotation only at the Critical target. As expected with this experimental design, RTs at the Critical target were significantly longer in the Algorithmic condition across all three experiments.

      We agree with the reviewer, however, that this RT difference could in principle reflect a more general carryover of cognitive demands from the Non-Critical targets rather than online computation at the Critical target itself. We can address this possibility more directly by asking whether RT at the Critical target exhibits the parametric signature expected of an algorithmic process. A defining feature of mental rotation is that RT scales with the magnitude of the computed aiming solution (Georgopoulos and Massey, 1987; Bhat and Sanes, 1998; McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Although rotation magnitude was fixed in the present experiments, participants’ actual reach angles varied naturally from trial to trial. Indeed, McDougle and Taylor (2019) originally demonstrated this relationship using actual reach angle rather than imposed rotation magnitude. We can therefore test whether trial-by-trial RT covaries with reach angle at the Critical target in the Algorithmic condition but not in the Retrieval condition. Such a relationship would be difficult to explain as a nonspecific carryover of cognitive load and would instead provide direct evidence that preparation time at the Critical target reflects the computation of the aiming solution.

      We also observe a second, independent difference at the Critical target: reach angles are consistently more variable in the Algorithmic condition than in the Retrieval condition across all three experiments (Figure S6). This pattern is consistent with repeated online computation producing variability in the selected aiming solution, whereas retrieval of a cached stimulus–response association produces a more stable response. Importantly, a general carryover account based solely on increased cognitive demands does not readily explain why movements to the Critical target should also be systematically more variable. Nor is this pattern easily explained by a speed–accuracy tradeoff: the Algorithmic group had more preparation time yet nevertheless exhibited greater variability. Consistent with this notion, Velázquez-Vargas and Taylor (2024) found that retrieving cached solutions produced less variable and more precise movements than movements that are not cached in the memory trace. 

      Third, we can conduct additional analysis to compare RT variability between algorithmic and retrieval groups at the critical target location. According to Logan instance theory (Logan, 1988), the retrieval of cached stimulus-response associations produces a stable RT profile, in contrast, trial-by-trial algorithmic computation can result in more variable trial-by-trial RT differences. 

      While we agree that RT differences alone should not be taken as definitive evidence of distinct strategies, our specific experimental design, the longer and more variable RTs at the Critical target, the greater trial-by-trial variability at that same target, and, if confirmed, a parametric relationship between RT and reach angle provide converging evidence that participants in the two conditions relied on different strategy implementations when preparing movements to the Critical target.

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

      We appreciate the reviewer’s thoughtful comment and agree that this is an important distinction. First, we would like to clarify the relationship among reaching variability, movement-plan variability, and the width of the implicit generalization function. Previous work has shown that when implicit recalibration occurs at a particular target location without an explicit aiming strategy, its generalization across the workspace can be described by a Gaussian-shaped function centered near the trained target location (Morehead et al., 2017). When an explicit aiming strategy is involved, however, implicit recalibration is centered closer to the planned aiming location rather than the visual target itself (McDougle et al., 2017). Thus, implicit recalibration is greatest near the direction in which the movement is planned, and a broader spatial distribution of movement plans can, in principle, produce a broader aggregate implicit generalization function. In the present study, we therefore use trial-to-trial variability in endpoint hand angle as a behavioral proxy for variability in planned movement direction, while recognizing that endpoint variability may also contain contributions from execution-related noise.

      This framework motivated the progression from Experiments 1 to 3. In Experiment 1, participants in the algorithmic condition exhibited substantially greater reaching variability, consistent with the idea that they sampled a wider range of movement plans across trials. Because error feedback associated with these different movement plans can induce implicit recalibration around each planned direction, greater variability in strategy use could contribute to the broader implicit generalization observed in the algorithmic group. In Experiment 2, we imposed stricter controls on spillover from noncritical targets, which reduced the overall breadth of generalization; nevertheless, model fits still suggested a modestly broader implicit generalization function in the algorithmic group, consistent with the remaining difference in reaching variability.

      Experiment 3 was designed to further reduce the direct influence of strategic variability on the induction of implicit recalibration by using a modified error-clamp paradigm. Error-clamp feedback has been shown to elicit implicit recalibration independently of task success and the participant’s explicit strategy (Morehead et al., 2017). We therefore used error-clamp feedback as an incidental signal to induce implicit recalibration while participants implemented either algorithmic or retrieval-based strategies. Importantly, however, Experiment 3 did not completely eliminate between-group differences in reaching variability: the algorithmic group continued to show greater variability around the critical 45-degree location than the retrieval group. We agree with the reviewer that, in principle, comparable generalization widths could arise if this broader distribution of movement plans were offset by a narrower local generalization of implicit recalibration in the algorithmic group. 

      However, not all variability is equivalent—or well described by a Gaussian distribution. Depending on the direction of the skew, variability in aiming can produce different effects on the implicit recalibration function, appearing as either broader generalization or greater amplitude. These effects are difficult to appreciate in Figures 3 and 5 for Experiments 2 and 3, respectively. We therefore sought to illustrate this more clearly in Figure 6, which shows how differences in the underlying aim distributions can shape the resulting implicit recalibration function in directionally complex ways. What complicates matters further is that implicit recalibration can asymptote (Morehead et al., 2017; Kim et al., 2018; Wilterson & Taylor, 2021). As a result, plan-based generalization can distort the implicit recalibration function in different ways depending on which side of the aim the error falls. The block-by-block analysis, suggested by reviewer 2, may shed light on this issue because we can get a sense if implicit recalibration has reached asymptote. 

      At a minimum, in the revised manuscript, we work to make clearer that subtle changes in the reach distribution may have a corresponding impact on the shape of implicit recalibration’s generalization function.  

      Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      We appreciate the reviewer raising this potential confound. We agree that the five pre-exposure trials in the Retrieval condition introduce a small difference in initial task familiarity. However, this account makes a straightforward prediction: if the shorter RTs in the Retrieval condition simply reflect five additional trials of general task experience, then the RT difference should disappear once the Algorithmic group has received a comparable amount of practice.

      We can test this directly by comparing the five pre-exposure trials in the Retrieval condition with the first five trials of the Algorithmic condition. We will also compare these pre-exposure trials with a later five-trial window from the Algorithmic condition to determine whether additional practice substantially reduces Algorithmic RTs. Assuming the observed pattern is as expected, RTs in the Algorithmic condition remain substantially longer even after considerably more than five trials of practice. Thus, the group difference cannot be explained simply by the Retrieval group having five additional trials of task familiarity. This persistent RT difference, together with our prior work showing characteristic RT differences between algorithmic computation and retrieval of cached aiming solutions, supports our interpretation that the groups relied on different strategy implementations.

      We are less certain what the reviewer means by the “early training advantage.” If this refers to angular error, we agree that the Retrieval group shows somewhat better performance very early in training, but this difference is not a central focus of the present study and largely disappears by the second or third training block. We will clarify the text so that we do not overinterpret this transient difference.

      If instead the reviewer is referring to RT, then the matched-trial analysis directly addresses the concern. Five additional familiarization trials could plausibly produce a brief initial advantage, but such an effect should dissipate within a small number of subsequent trials. In contrast, the RT difference between the Algorithmic and Retrieval conditions remains robust throughout training. We therefore do not think that general task familiarity provides a sufficient explanation for the observed RT differences.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      The reviewer raises an interesting possibility. In McDougle and Taylor (2019), however, the transition from algorithmic computation to retrieval occurred in a condition with only two targets in the task set. With repeated practice, participants needed to retain only two target-specific aiming solutions, making it feasible to replace online computation with retrieval of cached stimulus–response associations. By contrast, the Algorithmic condition in the present study contained 10 target locations. Our prior work suggests that caching becomes increasingly difficult once the number of target-specific solutions exceeds approximately four, at least over the timescale of several hundred trials (Velázquez-Vargas and Taylor, 2024; Bejjanki and Taylor, 2026). Thus, our task was designed to maintain pressure toward an algorithmic strategy throughout training.

      The reviewer nevertheless raises a more specific possibility that is not ruled out simply by the size of the target set: participants might selectively cache the aiming solution for the frequently sampled Critical target while continuing to use an algorithmic strategy at the remaining targets. We think the existing behavioral data argue against such a clear transition.

      First, reaction times at the Critical target in the Algorithmic condition remained substantially longer than those in the Retrieval condition throughout training. If participants had cached the aiming solution for the Critical target, we would expect preparation times at that location to be the same as the Retrieval condition. However, the RTs for the Algorithmic and Retrieval conditions are significantly different in the last block of training. 

      Second, within the Algorithmic condition, reaction times at the Critical target remained similar to those at the Non-Critical targets. Selective caching of the Critical target predicts a different pattern: preparation should become faster at the Critical target than at the surrounding locations, where participants would still need to compute the appropriate aiming solution. However, we do not observe a significant difference between RTs at Critical and Non-critical targets for the Algorithmic conditions at the end of the training block. Taken together, these two observations suggest that the Critical target continued to be treated similarly to the other members of the 10-target set rather than becoming a privileged, cached stimulus–response association.

      Third, participants would have to single out the Critical target as being distinct. All targets had the same visual appearance, the Critical target was never presented on consecutive trials, and participants were not informed that it had a special role in the experiment. However, we acknowledge that its higher sampling frequency, its somewhat greater separation from neighboring targets, and the location of the subsequent exclusion trials could nevertheless have made it more salient. Thus, we cannot rule out selective caching solely from the task structure.

      For this reason, we agree that the reviewer’s proposed analysis provides a useful additional test. If the Critical-target strategy progressively transitioned from algorithmic computation to retrieval, one prediction is that the generalization function in the Algorithmic condition should become narrower across successive exclusion blocks and increasingly resemble that of the Retrieval condition. We will therefore attempt to estimate the width of the generalization functions as a function of the training block between the Algorithmic and Retrieval conditions.

      There is, however, an important limitation to interpreting block-by-block generalization functions in this experiment. Implicit recalibration is both plan-based and temporally labile. Generalization is centered around the planned aiming direction (McDougle et al., 2017), and recent work indicates that implicit adaptation can decay over relatively short intervals (Zhou et al 2017; Hadjiosif et al 2023). Consequently, the first trial of an exclusion block provides the cleanest sample of the current state of implicit recalibration. Across later trials in the block, the measured response can be influenced both by temporal decay and by where the sampled target falls relative to the participant’s current aiming direction.

      To minimize systematic sampling bias, the starting exclusion target was randomized across participants. This means that these effects should average out at the group level, but individual exclusion blocks do not provide equally precise samples of the entire generalization function. A fully balanced estimate of every position within each exclusion block would require substantially more participants than were included in the present experiments. We will therefore present the blockwise analysis while interpreting changes in the estimated breadth cautiously.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback). Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

      The reviewer raises an important point. By design, the error-clamp manipulation in Experiment 3 decouples the participant’s planned movement from the visual consequence of that movement. While this gives us precise control over the error driving implicit recalibration, it could reduce agency over the cursor and thereby alter the interaction between explicit strategy and implicit learning in ways that are difficult to rule out completely. In particular, as the reviewer notes, reduced agency could potentially weaken either the influence of the explicit plan on implicit recalibration or the strategies themselves. Because Experiment 3 was intended to test for the absence of a strategy-dependent difference in implicit recalibration, we acknowledge that higher-order interactions of this kind represent an inherent limitation of our study if Experiment 3 is taken in isolation. 

      There are nevertheless several observations that make us think that reduced agency is unlikely to provide the primary explanation for the convergence between groups. First, the progression across Experiments 1–3 is informative. In Experiment 1, where participants retained normal control over the cursor, the broader generalization function in the Algorithmic condition closely mirrored the broader distribution of reach directions. This relationship suggests that the apparent difference in implicit generalization could arise from differences in where participants planned their movements rather than from a direct effect of strategy type on the implicit system. In Experiment 2, we sought to reduce the difference in the distribution of planned movements while preserving normal action–outcome contingencies and, importantly, the generalization functions became correspondingly more similar. Experiment 3 then controlled the error signal itself and again produced similar generalization across strategy conditions. Taken together, this progression favors the interpretation that strategy affects the measured generalization function indirectly, through differences in the distribution of movement plans, rather than directly altering the underlying implicit recalibration process.

      We nevertheless agree that these experiments cannot exclude all possible interactions between explicit and implicit learning systems. Indeed, whether these systems interact directly has been an important and persistent question in the sensorimotor adaptation literature. Several studies have reported evidence consistent with direct interactions (e.g., Albert et al., 2022; Maresch and Donchin 2021; t’Hart and Henriques 2024), whereas our own work has generally pointed toward indirect interactions mediated by factors such as movement planning and the current state of implicit adaptation (Taylor et al 2010; Taylor and Ivry 2011; McDougle et al 2017). Indeed, our recent study was designed specifically to distinguish these possibilities under tighter experimental control (Chen and Taylor, 2026), yet we found that the interaction between explicit and implicit processes is more complex than a simple independent-versus-interacting dichotomy. Going forward, we think it is more cautious to first rule out low-level statistical or distributional differences that could account for apparent effects before invoking higher-order interactions between learning systems.

      Finally, one motivation for the present study was that much of the literature on implicit generalization trains participants at a single target location before measuring generalization across the workspace. Under such conditions, participants have ample opportunity to retrieve a stable target-specific aiming solution. If algorithmic and retrieval strategies fundamentally alter implicit generalization, then many existing estimates of generalization may characterize implicit learning under retrieval-like conditions rather than providing a strategy-independent property of the implicit system. Across the present experiments, we find little evidence for such a fundamental difference once the distribution of movement plans and the experienced error are better controlled. We therefore think the most parsimonious interpretation of the current results is that algorithmic and retrieval strategies primarily influence implicit generalization indirectly through how movements are planned. 

      We plan to revise the manuscript to acknowledge that Experiment 3 cannot completely rule out higher-order effects associated with reduced agency under error-clamp feedback. We will also provide additional validation that participants were implementing distinct strategies, beyond the group-level RT differences, by testing whether RT in the Algorithmic condition scales with the instructed rotation magnitude, and whether RT variability shows group-level difference (Logan, 1988). If present, this relationship would provide stronger evidence that participants continued to engage the intended strategy under the clamp. We agree, however, that confirming distinct strategy use would not by itself rule out the possibility that reduced agency altered how those strategies interacted with implicit recalibration.

      Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      We thank the reviewer for this thoughtful and constructive assessment of the manuscript. We especially appreciate their recognition of the progressive experimental logic across the three experiments and of the broader theoretical motivation for distinguishing algorithmic and retrieval-based strategies. We are also grateful for the reviewer’s comments on the clarity and transparency of the work.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      We agree with the reviewer that our original interpretation of several non-significant effects was too strong, especially without providing some form of equivalence test. Our central hypothesis predicts little or no difference between algorithmic and retrieval-based strategies under conditions in which movement plans and error exposure are controlled, and we therefore face the inherent difficulty of drawing conclusions from an expected null effect. As the reviewer notes, a non-significant conventional hypothesis test does not by itself provide evidence that two conditions are equivalent.

      We therefore plan to supplement the existing analyses with quantitative assessments of the strength of evidence for the null/equivalence, using Bayesian factor analyses to confirm whether two conditions are equivalent. These analyses will allow us to distinguish between effects that are sufficiently small to support our theoretical interpretation and effects for which the data are simply inconclusive. We will also revise the manuscript throughout to avoid treating p > .05 as evidence of equivalence in the absence of such supporting analyses.

      We also now appreciate that the magnitude of implicit recalibration decreases progressively across experiments. This reduction could diminish our sensitivity to differences between the Algorithmic and Retrieval conditions and therefore represents an important qualification on the null result. Because the experiments used similar trial structures and were conducted with the same experimental equipment, the source of this reduction is not immediately clear. The block-by-block analysis suggested by Reviewer 2 may provide useful insight into how implicit recalibration evolves over the course of training and whether this attenuation emerges gradually within experiments.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Based on prior work characterizing implicit generalization in relative isolation from explicit strategy (Morehead et al., 2017; Poh and Taylor, 2019), we expected a relatively narrow generalization function, with a full width at half maximum of approximately 30°. We therefore expected probes spanning −45° to +45° around the Critical target to capture most of the function. At the same time, prior work on plan-based generalization predicts that the function should shift toward the participant’s aiming direction (Day et al., 2016; McDougle et al., 2017; Chen and Taylor, 2026), which complicates the choice of where to center the probes. Expanding the range and density of probe locations is also not cost-free, because additional exclusion trials increase temporal decay (Hajiosif et al., 2023) and begin to overlap with trained locations.

      We nevertheless agree with the reviewer that, when the fitted center approaches the edge of the sampled range, estimates of Gaussian width can become poorly constrained and may partly reflect parameter trade-offs or extrapolation beyond the observed data. We therefore plan to test whether the group differences persist when the Gaussian fits are constrained so that their centers fall within the sampled range. We will also examine complementary nonparametric measures of generalization breadth, such as the area between the group generalization curves across the sampled probe locations. Convergence across these approaches would provide stronger evidence that the reported differences reflect the observed shape of the generalization functions rather than instability in the Gaussian parameter estimates. If the results are not robust across approaches, we will revise the manuscript to qualify the interpretation of the fitted width estimates accordingly.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      This concern is closely related to that raised by Reviewer 2. We are fairly confident that participants in Experiment 3 were nevertheless engaging in the intended strategy manipulation. Participants in the Algorithmic condition showed substantially longer RTs than those in the Retrieval condition, and they were able to accurately generate the instructed angular reach directions across trials. The two conditions also differed in the variability of both RT and reach direction, consistent with online computation of an aiming solution in the Algorithmic condition and retrieval of a more stable cached response in the Retrieval condition. We can provide an additional validation by examining the relationship between RT and reach angle and comparing RT variability across two conditions. If RT scales parametrically with instructed reach angle in the Algorithmic condition but not in the Retrieval condition, this would provide stronger evidence that participants were engaging an online mental-rotation-like computation rather than simply following arbitrary spatial instructions. Moreover, RT yielded by the retrieval strategy would tend to be less variable than the algorithmic strategy.

      We agree, however, that this does not fully address the reviewer’s broader concern. Experiment 3 necessarily changed the context in which the strategy was implemented. In Experiments 1 and 2, mental rotation was used to counteract a visuomotor perturbation and was therefore directly tied to successful control of the cursor. In Experiment 3, the error clamp removed this instrumental relationship: participants still had to compute and execute different angular reach directions, but those computations no longer determined the visual cursor outcome. In that sense, the algorithmic process in Experiment 3 was less naturally embedded in the task and could reasonably be viewed as a somewhat different instantiation of the task.

      We therefore acknowledge that Experiment 3 cannot establish with certainty that the same higher-order cognitive process was engaged in exactly the same way as in Experiments 1 and 2, nor can it rule out the possibility that this change in task relevance altered how explicit strategy interacted with implicit recalibration. At the same time, when considered together with Experiments 1 and 2, we think the overall pattern remains informative. The apparent strategy-dependent difference in generalization was largest when movement plans and error exposure differed most, became smaller when these factors were better controlled while normal action–outcome contingencies were preserved, and was eliminated when the error signal itself was experimentally controlled. This progression is more consistent with an indirect influence of strategy through differences in movement planning and error exposure than with a robust direct effect of strategy type on implicit recalibration.

      Nonetheless, we agree that Experiment 3 should not be interpreted as a definitive test of whether algorithmic strategy, in its more relevant visuomotor adaptation context, can lead to different interactions with implicit recalibration compared to a retrieval strategy. We plan to revise the manuscript to make this limitation explicit.  

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      We agree that the RT difference in Experiment 3, by itself, cannot rule out the possibility that the Algorithmic condition imposed greater spatial precision demands because participants were reaching to invisible locations specified by angular instructions. We can, however, test for a more diagnostic signature of algorithmic computation by examining whether RT scales parametrically with the instructed reach angle. A general cost associated with reaching to invisible targets could increase overall RT, but it would not necessarily predict the characteristic increase in preparation time with the magnitude of the required angular transformation.

      We will therefore examine the relationship between RT and instructed reach angle in Experiment 3. If RT increases systematically with angular displacement in the Algorithmic condition, this would provide additional evidence that the longer RTs reflect online computation of the instructed aiming direction rather than simply the greater difficulty of reaching invisible targets.

      We can also compare this RT–angle relationship across experiments. If participants are engaging the same underlying algorithmic computation in Experiments 1–3, we would expect the slope relating RT to angular displacement to be similar across experiments, even if the overall intercept differs because of differences in task structure and spatial demands. A comparable slope would therefore provide converging evidence that the same computational process was engaged despite the altered task context in Experiment 3.

      We acknowledge, however, that similarity of the slopes would itself require yet another inference from a null difference and should therefore be interpreted cautiously. As with the generalization functions, we will use a Bayes factor analysis to quantify the strength of evidence for the null.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

      Based on the reviewers’ comments and the additional analyses they have suggested, we think we will be able to place our conclusions on a firmer empirical footing while also tightening and narrowing them. In the revised manuscript, we will strengthen the generalization analyses, use more appropriate inferential tools for interpreting null effects, and temper our broader claims about the insensitivity of implicit recalibration to strategy type. We will also more explicitly acknowledge the limitations of the present experiments, especially Experiment 3.

    1. eLife Assessment

      Shin et al present significant new observations highlighting a novel oscillatory window during REM sleep that impacts neuronal dynamics and cross-regional communication and could have implications for our understanding of sleep's role in memory consolidation. This important study identifies and characterizes high-frequency oscillations (HFOs) in the prefrontal cortex (PFC) during REM sleep, reporting their temporal dynamics and coordination with hippocampal area CA1, and an intriguing dissociation between activation of CA1 neurons during REM HFOs and non-REM ripples that may impact sleep-dependent firing rate decreases. The main claims are supported with convincing evidence including an impressive range of analyses.

    2. Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high frequency events during nonREM or in hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from down regulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will open become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique, I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this, comparison it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs? What percentage of theta waves contain HFOs and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods for often each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and grossly what calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of overall calculations being done for each panel.

      The authors mention in discussion that they see increased functional connectivity between mPFC and CA1, but most data suggesting that seems to be based on LFP rather than spiking. Functional connectivity is defined best by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions. 


      Comments on revised version.

      Previously raised concerns are addressed.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Comments on revised version.

      I do have one remaining concern, which is about their continued use of the term "reactivation" during REM sleep, whereas it still seems "activation" is more appropriate. The only place they show more Post vs. Pre activation is in Figure 6F/6G which includes NREM sleep where indeed reactivation is robust (but not the main focus of this paper). There is no evidence offered that the REM ensembles are not already "pre-configured" and active at similar levels (with similar activation patterns) during Pre sleep. Notably Louie and Wilson 2001 found greater "replay" during Pre than Post during REM. Also, the first half vs. second half comparisons (e.g. Fig 6C) could be more effectively performed in Figure 6A, showing that the same ordering persists across the periods. If this point were addressed, the significance of the findings could potentially increase.

    4. Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      Distinction from wake HFOs<br /> The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent more general theta-associated phenomenon.

      Link to memory consolidation<br /> The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      Comments on revised version.

      The authors have since addressed these weaknesses. In supplementary figure S11 the authors now show that while HFOs were detectable during wake, they were not associated with gamma/theta oscillations or theta modulation of unit activity. This suggests that HFOs during REM are a distinct feature of REM sleep and not comparable to HFOs during NREM or wake. It would be interesting for future work to identify the significance of wake PFC HFOs, whether there are differences between HFOs during running compared to stationary behaviour, and their relationship to hippocampal sharp-wave ripples and memory consolidation.

      Regarding the link between REM HFOs and memory consolidation, the authors have further acknowledged this as a limitation of the study and requirement for a more specific experimental design to test related hypotheses. Nevertheless, they do show a clear trajectory of learning in the rats and corresponding increase in reactivation of task-related activity which could be associated with REM sleep HFOs. This study paves the way for future experiments to more directly test this link.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high-frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high-frequency events during non-REM or in the hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from downregulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique: I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this comparison, it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs?

      Here, we provide additional analyses to demonstrate that the phasic, theta-modulated PFC activity that we observe during HFOs is specifically tied to the occurrence of HFOs and not a strong phenomenon during non-HFO-associated theta periods. In Figure S5 M (middle and right), we find that aligning PFC multiunit activity to theta periods in phasic REM but temporally distant from HFOs does not elicit the same degree of theta-modulated activity as aligning to HFOs (as in Figure 2A and Figure S5 M, left).

      Additionally, we provide analyses of the theta periods adjacent to HFOs (at different temporal thresholds) and demonstrate that this theta-modulated spiking activity is largely absent (Figure S5; compare to Figure 2A). Unlike Figure S5, this analysis was not restricted to putative phasic REM.

      We have now added Supplementary Figures S5L-N to the revised manuscript.

      What percentage of theta waves contain HFOs, and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      Although theta oscillations are continuously expressed during REM sleep, HFOs occur only intermittently, such that only a small subset of theta cycles contain HFOs. Across all animals and epochs included for analysis, we found that ~7.4% of theta cycles contain HFOs. We present an epoch-level quantification of this in Author response image 1, where the proportions were calculated across all, tonic, and phasic theta cycles. As expected, a higher proportion of putative phasic theta cycles contain HFOs.

      Regarding differential firing rate modulation, we refer reviewer to normalized MUA plots in the manuscript (Figures 2A and 4F). We would like the emphasize that what we show is PFC multiunit activity that is normalized by the mean population firing rate during REM sleep. Thus, these figures, specifically the chain HFO aligned figure, indicate that there are peaks in activity around baseline level in the background of an overall decrease in population activity relative to baseline (i.e. the troughs surrounding the peaks have lower activity compared to baseline, as in Figure 9C). While this suggests that there is an overall decrease in firing rates during HFOs as compared to baseline theta periods without HFOs, this simply provides a qualitative account of this difference. In Figure S5N, we present firing rate comparisons during HFO chains versus theta periods during putative phasic REM bouts at least 4 theta cycles away from HFOs. We find that HFO chain-associated PFC neuron firing rates are lower compared to non-HFO theta periods, supporting our finding of activity suppression during HFOs.

      Author response image 1.

      Proportion of theta cycles with HFOs. (A) Proportion of theta cycles with HFOs across all, tonic, and phasic cycles (***p = 4.90e-05, rank sum test).

      Lastly, we appreciate the reviewer's suggestion that REM HFO quantifications could be framed relative to phasic theta cycles without HFOs. We agree that such comparisons are informative and ensure that the results we present are specific to periods with detected HFOs. In response, we have added additional analyses of theta periods outside of HFOs in Figure S5. Furthermore, while we found that a larger proportion of chain HFOs occurred during bouts of putative phasic REM compared to isolated events (Figure S5 and Figure S3C), most of the analyses that we performed were on events pooled across putative tonic and phasic, since putative phasic REM is relatively scarce (<10%).

      We also refer the Reviewer to our response to Reviewer 2’s major comment #1 below where we reiterate several of our findings that demonstrate the HFO-specificity of the reported dynamics, as well as our extended response to Reviewer 3’s Public Review comment #1 where we show that the dynamics associated with REM HFOs are absent during HFOs detected during awake behavior (comparable theta state) on the W-Track (Figure S11). We hope that the additional control analyses we present as well as our expanded explanations adequately address the Reviewer’s concerns.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods, often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods, often for each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and what gross calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of the overall calculations being done for each panel.

      We have now split and updated the figures in a more logical progression of ideas:

      Figure 1: Prefrontal cortical HFOs in REM sleep using spectral analyses.

      Figure 2: Characteristic spiking modulation in PFC during REM HFOs.

      Figure 3: HFO and gamma distinction in theta cycles, and PFC-CA1 coherence (including chain/ isolated HFOs in phasic and tonic REM stages).

      Figure 4: Differential modulation of PFC spiking activity during REM PFC HFOs vs. NREM PFC ripples.

      Figure 5: Characteristic PFC population activity during REM HFO chains.

      Figure 6: Comparison of PFC reactivation during REM PFC HFOs vs. NREM PFC ripples.

      Figure 7: Differential engagement and excitability modulation of CA1 neurons by REM HFOs.

      Figure 8: REM theta phase shifting CA1 neurons preferentially engaged by REM HFOs.

      Figure 9: (Model) Network model with ACh reproduces spiking modulation during REM PFC HFOs vs NREM PFC ripples.

      Figure 10: (Model) Model reproduces restricted REM coactivity vs. widespread NREM coactivity.

      We have also rearranged the figures to parallel the main figures and results section.

      The authors mention in the discussion section that they see increased functional connectivity between mPFC and CA1, but most data suggesting this seems to be based on LFP rather than spiking. Functional connectivity is best defined by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions.

      We have updated the manuscript accordingly. Specifically, we have modified the text on Page 13, Lines 23-24:

      “These chains are associated with increased measures of oscillatory coupling between PFC and CA1…”

      Reviewer #1 (Recommendations for the authors):

      (1) Please ensure that analytical methods are presented in the same order in the methods section as they are in the results section - panel by panel. That said, the methods section is well-written and presented.

      We apologize for any confusion this may have caused. We have now reorganized the methods section to ensure they presented in the same order as the panels in the main figures.

      (2) Please specify whether the recordings/behaviors occur during the animal's light circadian phase.

      We have added a statement in the methods under the “Behavior” section on Page 34, Lines 9-12 indicating that the experiments took place during the light phase:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Please specify how many tetrodes are in mPFC and how many in CA1?

      We have added this clarification on Page 33, Lines 27-28:

      “Tetrodes were split equally between PFC and CA1 (15, 16, or 32 tetrodes in each region).”

      (4) All mentions of "coherence" should have frequency bands specified. "Theta coherence", for example.

      We have now added this information to all relevant mentions of “coherence”.

      (5) The intuitive logic of the phase slope index (2K) should be briefly explained for maybe half a sentence in the results section. The intuition should be explained better in the methods section devoted to it.

      We have added more clarification on the phase slope index method, both in the legend of Figure 3 on Page 21, Lines 22-23:

      “Phase slope index (PSI), which is a measure of phase lag consistency across different frequencies…” and in the methods section under “Phase slope index” on Page 39, Lines 21-26:

      “In practice, PSI is used to assess the consistency of phase lag relationships between two signals across different frequency bands and is a measure that is weighted by oscillatory coherence. We opted to use PSI to estimate the directional flow of information instead of other methods, such as Granger causality, since it has been demonstrated that PSI is less prone to false positives.”

      (6) "Cofiring" and "coactivity" should be defined clearly as measures - preferably in the results section if possible. They sound similar and are somewhat jargon-y without self-explanatory meaning (or difference from each other). How should readers understand and interpret them?

      We apologize for the confusion regarding these two terms, which are both used throughout the manuscript. Here, we specifically used to term “cofiring” to specify the explicit quantification of coincident activity between pairs of neurons (e.g. Figure 4E, left) as described in the Methods section under “Ripple and HFO co-firing”. There were instances where “coactivity” was used to refer to this quantification, and they have been changed to “cofiring”. We have now added a statement to the manuscript to clarify that “cofiring” is a quantification of coincident activity between neurons on Page 7, Lines 34-35:

      “Overall cortical co-firing, which is a measure of coincident activity between neuron pairs during discrete events”

      Additionally, the term “coactivity” in the manuscript is used when describing neuronal activity in the model or when there are mentions of coincident activity other than the explicit quantification described above.

      (7) The temporal threshold for "cofiring" should be stated in the results to enable interpretation of the results.

      We apologize for the lack of clarity regarding the cofiring metric. Here, the temporal threshold that we are imposing is determined by the ripple/HFO event times (see Methods section “LFP event detection”). If both neurons in a pair emit spikes within the defined window of an event, they are considered to have “cofired” (Cheng and Frank 2008, Singer and Frank 2009, Sosa, Joo et al. 2020).

      We have now clarified this in the results section on Page 7, Lines 34-35:

      “Overall cortical cofiring, which is a measure of coincident activity between neuron pairs during discreet events…”

      We have also added a clarifying statement in the Methods section under “Ripple and HFO cofiring” on Page 40, Lines 24-25:

      “Here, cofiring assesses coincident activity between neuron pairs within the start and end times of events, and thus, no explicit temporal threshold was implemented.”

      (8) Please clarify more systematically in which region the NREM ripples were detected. The natural assumption is the hippocampus, but at times it is mentioned that they are detected in the cortex. Are they cortical in all analyses? Readers could easily get confused about this and misinterpret "ripples". To clarify further, if these are always cortical events, I suggest renaming "ripples" in this text to "cortical NREM HFOs".

      We apologize for the confusion. For all mentions of “ripples” in the text, we are referring to PFC ripples specifically in NREM sleep. When hippocampal ripples are mentioned, we differentiate them by explicitly using “sharp-wave ripples” or “SWRs”. Our decision to use “ripples” for NREM sleep was based on previous studies that investigated NREM high-frequency events in cortex (Khodagholy, Gelinas et al. 2017, Helfrich, Lendner et al. 2019, Vaz, Inati et al. 2019, Aleman-Zapata, Morris et al. 2022, Ghosh, Yang et al. 2022, Shin and Jadhav 2024). Furthermore, we elected to use the “HFO” nomenclature for high-frequency events in REM sleep, since it has been used in previous studies to describe these events (Tort, Scheffer-Teixeira et al. 2013, Bueno-Junior, Ruckstuhl et al. 2023). Thus, to remain consistent with the literature, we decided to use these terms to describe these sleep-state-specific events throughout the manuscript:

      Hippocampal sharp-wave ripples (SWRs) in NREM sleep

      PFC ripples in NREM sleep

      PFC HFOs in REM sleep

      We have now added the following statement on Page 5, Lines 1-3:

      “However, to avoid ambiguity, and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout.”

      (9) The analysis performed for 4B is not explained clearly. It is somewhat better explained in the methods. I believe the reader should understand that each HFO is treated as an event, and spiking participation per unit was measured, and then the similarity of that pattern was assessed between HFOs. I also find the x-axis being quartiles makes understanding this graph particularly difficult.

      Why not label by raw lag and show a correlation plot rather than break down by quartiles? Alternatively, labeling the millisecond values of these quartiles on the x-axis labels may make comprehension much easier.

      We apologize for the lack of clarity regarding this method. We have added points to clarify the analytical procedure used for Figure 4B (Now Figure 5B) on Page 8, Lines 25-28:

      “We represented each HFO as a binary vector of PFC neurons active during the event, and computed the Pearson correlation between the vectors of every pair of consecutive HFOs” As well as in the Figure 5 legend on Page 24, Lines 7-10:

      “Here, the PFC spiking activity during each HFO was binarized across all neurons and the Pearson correlation coefficient was calculated between adjacent events as a measure of pattern similarity. Then the relationship between pattern similarity and IEI was reported.”

      In addition to the quartile labels, we have now added the average inter-event interval (IEI) of each quartile to the x-axis labels of Figure 5B to improve comprehension as well as a statement in the figure legend on Page 24, Line 11:

      “Below each quartile is the average IEI of that quartile in milliseconds.”

      (10) "Rank order analysis" should be again defined in the results in a manner that the reader can follow the point of the analysis and figure. For example, the goal is to assess the spike sequence across HFO events by looking at the regularity of spike timing rank for each neuron in each HFO event. Also, could this correlation be more simply calculated and presented as just a standard deviation around the mean of that cell's rank?

      We apologize for not including an explanation of this analysis in the results section. We have now added more detail about the rank order analysis to the Figure 5 legend on Page 24, Lines 18-21:

      “The mean rank-order correlation from the leave-one-out cross-validation procedure. Each event’s rank was correlated with the averaged rank across all other events. The average across all events compared to a distribution of means generated by jittering (n = 1000) spike times is shown.”

      We also provided a short description of the procedure in the results section on Page 8, Lines 32-36:

      “Furthermore, we examined whether the order in which individual PFC neurons fired during HFOs within a chain was preserved across chains. For each chain, we extracted each cell's first-spike rank order and compared it to a leave-one-out template constructed from the average normalized rank across all other chains.”

      Regarding the Reviewer’s second comment about the correlation, if we presented the result as the standard deviation around the mean of the cell’s rank, the result would be similar to the example rank order in Figure 5C, which we provided as a visualization of the rank-order template procedure we used. However, this would not necessarily demonstrate the consistency of population activity across HFO chains. To demonstrate consistency in activity across HFO chains, we used a leave-one-out cross-validated approach where each chain event was assessed separately. For each event, the neurons firing during that event was ranked and normalized 0-1. Then, the average rank of each neuron across all other events was calculated. Lastly, the correlation between the ranks of the left-out event and the average ranks of the template was taken to assess similarity in sequential activity. We opted to use this method since it has been demonstrated to be effective for evaluating similarities in sequential across events (Stark, Roux et al. 2015, Valero, Viney et al. 2021).

      (11) In Figure 5B and the related results section, I gather that "spatial" relates to place field location rather than anatomical/tetrode location of the neuron? This was not my original understanding and should be stated clearly.

      We apologize for the lack of clarity regarding this result. Yes, the term “spatial” refers to spatial rate map correlation between CA1 neuron pairs, specifically during the behavioral W-Track session prior to the sleep session where cofiring was assessed. We have updated the y-axis label of Figure 5B (Now Figure 7B) to include “rate map” and have updated the results section on Page 9, Lines 32-33:

      “…we observed a higher degree of spatial rate map correlation for high cofiring pairs…”

      (12) Figure 5E and its caption are almost totally unable to be explained since axes aren't explained well in either. It is only by inference from the results text that meaning can be assumed.

      We apologize for the lack of clarity regarding Figure 5E (Now Figure 7E). We have now updated the y-axis labels on both updated figures (Figures 7E and 8B) and have reworded and added more detail to the figure legend on Page 28, Lines 1-5:

      “High REM PFC HFO cofiring CA1 neurons exhibited a greater degree of suppression during NREM PFC ripples. For a description of modulation index, see Methods section Ripple/HFO aligned modulation. Here, since CA1 neurons exhibit a robust decrease in activity in response to NREM PFC ripples, we refer to the modulation as suppressive.”

      (13) Figure 5F is interesting, supporting the concept that HFOs "protect" neurons from downscaling. However, how can a "neuron" be cofiring? Would cofiring not be defined in a pairwise manner, and so each unit of the cofiring measure would be a pair of neurons? This question applies to other panels in this figure. Please clarify this.

      We apologize for the lack of clarity regarding the exact cofiring metric that we use. For Figures 7D-F and 8A, since we wanted to relate cofiring to changes in firing rate and modulation state during NREM PFC ripples, we calculated single cofiring values for each CA1 neurons by averaging across all pairings with PFC neurons. Thus, this metric gives us an estimate of the overall cofiring strength of each CA1 neuron. We have now added a new section in the Methods under “Calculation of a single cofiring metric and separation into populations of high and low cofiring neurons” on Page 44, Lines 34-43:

      “Since we wanted to relate the above cofiring metric to other measures, we needed to obtain a single-value cofiring metric for each neuron. To do this, we averaged the cofiring values across all pairings for a neuron (e.g. 1 CA1 neuron paired with all PFC neurons) and reported it as the cell’s cofiring. Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0 (High cofiring > 0; low cofiring < 0). Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions (High cofiring > mean; low cofiring < mean).”

      We have also clarified this in the Figure 7 legend on Page 27, Lines 25-28:

      “For comparisons between cofiring and other metrics (e.g. firing rate), a single cofiring value was calculated for each neuron by averaging the cofiring metric across all neuron pairings. Additionally, high and low cofiring CA1 neurons were split based on average cofiring values > 0 and < 0, respectively.”

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Weaknesses:

      The primary weakness of the study lies in the lack of a clear distinction between global brain states and the specific events being analyzed. Because the authors compare HFOs across different sleep stages (NREM, tonic REM, and phasic REM) without sufficient controls, it is difficult to determine if the observed differences are intrinsic to the HFOs themselves or simply a reflection of the different physiological states in which they occur.

      We would first like to note in case it was unclear – as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly spiking activity patterns observed in prefrontal cortical-hippocampal circuits during these event-types that occur in the two sleep stages. As noted in response to Reviewer 1’s Comment #8 and Reviewer 3’s Comment #1, to avoid ambiguity and to remain consistent with existing literature, we use the following terms to describe these sleep-state specific events throughout the manuscript: Hippocampal sharp-wave ripples (SWRs) in NREM sleep, PFC ripples in NREM sleep, and PFC HFOs in REM sleep.

      We refer the Reviewer to our Response to Reviewer 1’s first comment where we provide additional analyses of theta periods outside of HFO times (i.e. baseline REM periods), including Figure S5. We also further address these comments as a response to Reviewer 2’s major comment #1 below.

      Furthermore, the evidence for "structured reactivation" is not yet convincing. The temporal alignment of these reactivation events appears inconsistent, with peaks occurring well before the HFO itself, and the analysis does not sufficiently control for pre-existing cellular assembly strengths.

      We have now addressed this as a response to Reviewer 2’s major comments #3 and #8 below, including Figure 6.

      Additionally, some of the sleep architecture presented appears atypical, such as very short REM bouts and direct NREM-to-REM transitions that bypass standard progression, raising questions about the consistency of the sleep detection across animals.

      We have now addressed this as a response to Reviewer 2’s major comment #2 below, including Figure S1 and Figure S3. We expect that addition of these figures will mitigate concerns about our sleep staging procedures and will clarify certain points raised by the Reviewer.

      Finally, the study does not account for potential confounds like baseline firing rates when interpreting the behavior of "high-cofiring" neurons, which may simply be the most active cells in the population.

      We have now addressed this as a response to Reviewer 2’s major comment #6 below.

      Reviewer #2 (Recommendations for the authors):

      In this study, the authors detect activity periods during REM sleep that feature high-frequency oscillations in the prefrontal cortex. They term these REM HFOs and report that they occur in "chains" phase-locked to theta oscillations. They contrast these with HFOs that appear in isolation and with ripples observed during NREM sleep, either in the PFC or in the hippocampus. The data presented is very interesting in places. The authors show that these HFOs are observed primarily during "phasic REM" periods that have higher theta power. It appears that overall firing is substantially lower surrounding these events. Intriguingly, it appears that CA1 cells that co-fire with the PFC during HFOs increase their firing rates over the course of sleep, whereas other neurons show a firing decrease. This seems to be the most striking finding of this study.

      Major:

      (1) The findings are generally intriguing, but the study makes some choices that are hard to understand. They begin comparing HFOs that occur in chains to HFOs that occur in isolation, even though it appears that chained and isolated HFOs likely occur at different times, some during tonic REM, some during phasic REM, and others during NREM. The lack of control for sleep state makes it difficult to determine if the reported differences are intrinsic to the HFO patterns or merely reflections of the underlying global brain state. Overall, I just didn't quite understand the motivation for comparing isolated and chain HFOs, as it seems natural that chains would occur during periods of greater synchrony.

      We appreciate the Reviewer for raising these points and agree that the properties of ripples and HFOs cannot be interpreted independently of the underlying brain state. As we mentioned in the initial response, we do expect that the generation of these ripples/HFOs in NREM and REM sleep are inextricably linked to global brain state (ex., cholinergic tone, as shown in the model in Figures 9-10), which results in differing patterns of activity across sleep states. We noted in the response to the public review comment #1 above that as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly, spiking activity patterns observed in cortical-hippocampal circuits during these event types that occur in the two sleep stages.

      Sleep state: Our primary goal in comparing isolated and chain HFOs (chain vs. isolated HFOs are only compared in REM periods, not NREM periods) was not to suggest that the differences that we observe are state-independent, but rather to provide evidence that temporal clustering of HFOs underlies distinct PFC dynamics, as well as enhanced CA1 engagement. Rather than attempting to dissociate ripple/HFO occurrence and the differential physiology that underlie NREM and REM sleep states, we view them as complementary—both sleep states are permissive to generation of high-frequency oscillations. Similarly, putative tonic and phasic REM substates (low vs high theta power) are differentially permissive for isolated and chain HFOs (Figure S3C). We add additional detail on tonic vs. phasic REM substages in Supplementary Figure S3, showing that the rate of HFOs and HFO chains is significantly elevated in putative phasic REM. While we would have liked to analyze REM HFOs during tonic and phasic states separately to be able to make more concrete conclusions, the scarcity of phasic REM sleep made it difficult to make accurate comparisons between the two, especially for spiking modulation. We have therefore removed the qualifying term “phasic REM” from the abstract.

      Also, we would like to clarify that while we investigated chains of events (ripples) in NREM sleep as a comparison to REM HFOs, we do not refer to them as HFOs in NREM anywhere in the manuscript. In NREM, while there are PFC ripples that are clustered into chains based on our definition (separation of <200 ms), we do not observe a prominent peak in the IEI distribution that would suggest entrainment by other oscillations (e.g. theta or spindles). This analysis was only to show that associated results are specific to REM HFO chains, and not seen during comparable chains of NREM cortical ripple events.

      Chain vs isolated HFOs: Regarding the Reviewers comment about why we chose to compare isolated and chain HFOs, the motivation is clearly demonstrated by differences in these events in spectral properties (Figure 3H-I) and spiking modulation (Figure 4F-G). Our initial motivation to investigate these chains of events in REM sleep came from our observation that PFC population activity aligned to REM HFOs was theta-modulated (multipeaked, suggesting multiple high-frequency events over a short duration). This led us to hypothesize that there may be chaining of events in REM sleep, which in line with previous studies demonstrating the clustering of events, such as spindles in NREM sleep (Darevsky, Kim et al. 2024). In that study (Darevsky, Kim et al. 2024), the authors showed that reactivation of motor patterns was more persistent during trains of spindle events as compared to isolated spindles. Furthermore, trains/chains of multiple hippocampal SWRs have been shown to underlie the replay of extended experience (Davidson, Kloosterman et al. 2009). Thus, the separation of high-frequency events into isolated and chained events has precedence and may have functional significance. Furthermore, since we see that isolated and chain events have a bias for occurring during putative bouts of tonic and phasic REM sleep, respectively (Figure S3C), characterization of both event types is an important step for understanding potential differences in interregional interactions during REM sleep. Furthermore, a recent appreciation for the role of sleep stage sub-states (Chang, Tang et al. 2025) further emphasizes the importance of investigating these events separately. We expect future studies to further dissect the roles of tonic and phasic REM states in memory and cognition, and our finding provides an account of the existence of different events for future reference.

      HFO chains during periods of greater synchrony: While it may seem natural that chaining would occur during periods of high synchrony (theta synchrony here), we show that our result is HFO-specific, especially for spiking activity modulation (Figures 4F-H and Supplementary Figures S5K-N), which makes it novel and important. Previous studies have focused on theta-gamma cross-frequency phase amplitude coupling, which has been proposed as a mechanism where slower theta oscillations temporally organize faster local population activity, thereby synchronizing neural ensembles within and across brain regions (Belluscio, Mizuseki et al. 2012). Here, we present a comparison of HFOs and gamma events and show that while gamma events can also occur in chains (added to Supplementary Figure S4H), possibly due to increased synchrony as the Reviewer stated above, we do not observe theta-modulated population activity aligned to these events (Supplementary Figure S4J), which is a defining property of HFOs that we propose underlies our results. This indicates that our reported results are unique to periods of synchrony associated with HFOs.

      Control analyses for theta periods: Regarding the comment about lack of control for sleep state, we refer the Reviewer to our response to Reviewer 1’s comments. We have performed additional control analyses comparing PFC activity during REM theta periods adjacent to detected HFOs and demonstrate that phasic PFC population activity is largely absent (Figure S5).

      We are aware that it is difficult to dissociate the generation of ripples and HFOs from the underlying brain state, since they are so tightly linked. However, we do expect that our clarifying points in addition to the REM theta state controls provide strong evidence that the results that we present are intrinsic to HFOs and not general reflections of activity during baseline theta activity in REM. Here, we further reiterate a few of the main results that demonstrate this:

      (1) Phasic spiking modulation of PFC population activity associated with HFOs is not present during baseline theta periods (Supplementary Figure S5K).

      (2) Phasic spiking modulation during HFOs is not linked to extracted gamma events (Supplementary Figure S4J).

      (3) Phasic spiking modulation, as well as activity suppression, is strongest during chains of HFOs (Figure 4F, Supplementary Figure S4J).

      (4) PFC-CA1 theta coherence surrounding HFOs increases relative to baseline (Figure 3E, z-scored relative to baseline coherence).

      (5) Assembly peaks are sequentially organized surrounding HFO chains but not during isolated or shuffled (baseline) HFO times (Figure 6C, Supplementary Figure S7E, and Author response image 2).

      (6) A higher proportion of CA1-CA1 pairs are high cofiring during HFOs as compared to baseline periods (Figure 7B), indicating specific CA1 engagement during HFOs (in addition to coherence).

      Lastly, we show that HFOs detected during periods of active behavior (high theta) on the W-Track are not associated with many of the defining features of HFOs in REM sleep, thus demonstrating the specificity of REM HFOs despite similar background theta activity (Figure S11, in response to Reviewer 3 comment #1).

      (2) It is crucial that the study provides REM-specific sleep examples for each of the data sessions, marking tonic and phasic REM and indicating when isolated and chain HFOs are observed. It remains unclear how interspersed these events are. Do some REM episodes just have isolated HFOs and others chains?

      We thank the Reviewer for raising this point and apologize for the lack of clarity. For additional transparency, we now provide example sleep state plots for each animal (Supplementary Figure S1H-I).

      We also provide hypnograms that show bouts of putative phasic REM on top of REM periods (Supplementary Figure S3). Additionally, as a compact way of demonstrating the validity of our separation of putative tonic and phasic bouts, we provide a plot showing the average velocity and spectrogram surrounding putative phasic REM bouts across all putative phasic REM transitions (Supplementary Figure S3A). Regarding the incidence of HFOs in tonic and phasic REM, we refer the Reviewer to Supplementary Figure S3D, where we show that HFO rate is significantly higher during putative bouts of phasic REM, during which they tend to be organized in chains as compared to putative tonic bouts (Supplementary Figure S3E). In line with this, we additionally report that although both isolated and chain HFOs occur in putative tonic and phasic bouts of REM sleep, the proportion of chain HFOs during phasic REM is significantly higher than that of isolated HFOs (Figure S3), indicating a bias for chains to occur in phasic REM sleep, potentially due to stronger theta input. However, since putative phasic REM accounts for <10% of REM sleep in our dataset, consistent with previous reports using similar methods (Mizuseki, Diba et al. 2011), we were unable to restrict spiking analysis to HFOs in phasic REM. We found that chain HFOs during both putative tonic and phasic REM sleep elicited theta modulated population activity in PFC, thus we pooled HFOs across REM states for analysis.

      While a larger proportion of chain events occur in putative bouts of phasic REM sleep as compared to isolated events (Figure S3), chain and isolated events occur in both putative tonic and phasic substates (Proportion of events in putative tonic REM is (1 - proportion in phasic) shown in Figure S3C, right). Thus, the majority of analyses comparing isolated and chain HFOs, especially for spiking data, were pooled across states. Instead of showing example plots for all 22 sleep epochs (22 out of 36 epochs with >5 s of putative phasic REM), we expect that this analysis will be sufficient to illustrate the distributions of isolated and chain HFOs across putative REM sleep substates. We have also removed the qualifying term “phasic REM” from the abstract.

      To clarify this, we have added a statement to the new “Limitations” section of the manuscript on Page 16, Lines 10-20:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states. (Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult. Finally, although our model predicts that distinct cell-type activity profiles shape REM sleep HFO dynamics, we did not record a suDicient number of interneurons to test these predictions directly. Future studies using appropriate behavioral tasks, longitudinal sleep recordings, and cell-type specific opto-tagging will be able to resolve these limitations and further clarify the roles of high-frequency oscillations in REM sleep.”

      We have now added Supplementary Figure S3.

      This is also important because some of the REM sleeps detected and shown in Figure S1H seem unusual. For example, in Animal 1, there are multiple bursts of REM that seem very irregular. In some other sessions, animals occasionally appear to enter REM sleep with very little preceding NREM, which goes against existing literature. It's not clear which ones of these meet the 30 s duration threshold. Are the chain events occurring in these periods?

      In reference to the hypnograms shown in Supplementary Figure S1H (Now Supplementary Figure S1I), these were generated by concatenating all 9 sleep epochs regardless of whether they passed the inclusion criterion of > 30 s of total REM sleep. Furthermore, only a subset of the epochs shown were included for analysis based on a secondary, manual inspection that is performed to confirm inclusion. Sleep state plots (e.g. Supplementary Figure S1H) were visually inspected to further confirm the transition into REM sleep, ensuring absence of noisy T/D ratio or spurious detection due to noisy signals – epochs where microarousals or persistent subthreshold fluctuations in animal movement induced noisy TD ratio increases, and thus inaccurate REM designation, were excluded. We thus used a total of 36 sleep sessions from a possible 90 sleep sessions.

      We apologize for not specifying what portions of data represented by the hypnograms were included. We have now provided updated hypnograms only illustrating the sleep epochs included for analysis (Figure S1I).

      We have now added Supplementary Figure S1.

      Regarding the Reviewer’s point “Are the chain events occurring in these periods?”: Yes, all of the chain events (and isolated events) come from these updated epochs that are now shown. No events in the excluded epochs were included, since we wanted to only analyze the data that came from curated REM epochs that we were confident in.

      (3) The analysis supporting structured reactivation was not generally convincing. Figure 4E does not provide convincing evidence of this. Indeed, reactivation strength is lowest around the time of the event, and appears highest 0.5s before. The example panel seems rather anecdotal. It's also not clear why REM reactivations should be compared to NREM ones here. I could not follow what was done in Figures 4F-H. Why should the first half and second half of an HFO event be correlated?

      We appreciate the Reviewer for raising these concerns. We agree that this section is somewhat dense and at time hard to follow, so we will clarify with further explanations and analyses (see also response to Comment #8 with new reactivation figures).

      In Figure 6B (originally Figure 4E), our intention was to show that if you simply align assembly activation to REM HFOs, a relatively flat response is observed when averaged across all assemblies, in stark contrast to assembly reactivation during NREM cortical ripples. This could suggest a couple of things: 1) There is no real assembly activation in response to REM HFOs or 2) Assembly activation is organized differentially (compared to NREM) surrounding HFOs. What we hypothesized, and then quantified based on observations, is that assembly activity is sequentially organized around HFOs. If this were true, it would suggest a consistent temporal relationship between REM HFOs and assemblies (for example, assembly 1 tends to be active 115 ms after the onset of chains, assembly 2 is active 230 ms after, etc.). Of course, the sequential pattern that we show in Figure 6A may arise trivially, especially since these types of sequential visualization plots can simply arise from noise. Thus, a cross-validation method must be utilized to ensure the sequences are indeed reflective of an underlying computation.

      In the methods section “Assembly sequence detection surrounding HFOs” we explain the splitripple/HFO procedure that we used to compare assembly sequences across two halves of the data, similar to methods used in hippocampal place cell sequence cross-validation (Plitt and Giocomo 2021, Sosa, Plitt et al. 2025). For every split and assembly alignment, a sequence similar to Figure 6A, right is generated based on the first half of aligned data. Then the second half of aligned data is sorted based on the peak indices of the first half of data and the correlation between the peak reactivation indices across all assemblies for the two datasets is calculated. A high correlation indicates high sequence similarity across the two halves of data (not two halves of an HFO as the Reviewer mentioned), suggesting temporal consistency of assembly reactivation. By utilizing this method, we show that assembly reactivation sequences across randomly chosen halves of data are most similar for chained HFOs (Figure 6C and Supplementary Figures S7C-F). We have now provided an additional control analysis that investigates this sequential assembly activity for time-shifted chain events (Author response image 2). As in Figure 6, assembly activity was aligned to the first event in each time-shifted chain. In addition to the analyses in Supplementary Figure S7, this control further indicates that sequential assembly activity is preferentially restricted to HFO chains.

      Author response image 2.

      Assembly sequences surrounding time-shifted HFO chains. (A) Distributions of r values calculated from the Pearson correlation between peak reactivation bins across all assemblies for two randomly chosen halves of the shifted HFO-aligned data. There was no difference between r values for time-shifted chains and shuffled data, indicating no structured assembly activity during periods outside of real HFOs chains.

      To further quantify this, we calculated two additional metrics: 1) the slope difference between fitted lines for the two halves of data for each split and 2) the absolute peak difference between assemblies in the two halves of data. First, for the slope difference metric, the slope of the best-fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data in Figure 6D indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. Secondly, Figure 6E is a quantification of the peak reactivation displacement between the two halves of data. If there is a high probability of small peak differences, as in the real data, this indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data.

      We apologize for the omission of the description of the slope and peak difference metrics that we used in Figures 6D,E. We have added information in the Methods section under “Assembly sequence detection surrounding HFOs” on Page 43, Lines 23-33:

      “Furthermore, the slope and peak differences were calculated as additional metrics of sequence and temporal reactivation consistency between the two halves of data, respectively. For the slope difference metric, the slope of the best fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. For peak reactivation difference, the temporal displacement of the peak reactivation index between the two halves of data was calculated and compared to shuffle. A high probability of small peak differences indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data. Shuffling of assembly strength was carried out as above.”

      (4) The authors argue that the occurrence of HFOs, rather than theta power, is the reason for lower MUA activity, but the analysis for this (Figure S4I) is quite confusing. The left panel actually seems to indicate that MUA is indeed lower when theta power is high.

      We apologize for the confusion regarding this figure and appreciate the Reviewer’s point that periods of high theta power can also appear to be associated with reduced MUA. We agree that the original presentation may not have clearly separated the contributions of theta and HFOs. What we convey with Supplementary Figure S5K (originally Supplementary Figure S4I) and new Supplementary Figures S5L-M is that the observed theta-modulated PFC spiking response (Figure 2A) cannot be solely explained by baseline theta periods (outside of HFOs) in REM sleep. Since theta oscillations are ubiquitous during REM sleep, an important control is to demonstrate that the theta-modulated population activity is not simply a consequence of ongoing theta activity. Thus, we aligned PFC activity to theta oscillations of varied power and show that the fluctuating, theta-modulated population activity is absent, indicating that HFOs associated with the theta oscillation are driving this phasic response.

      We are not solely arguing that the presence of HFOs is the driver of decreased PFC multiunit activity. In both the data and model, we show that the magnitude of theta power detected in PFC (possibly input to PFC) is inversely related to multiunit activity (Figure 9E). Our interpretation is not that theta is unrelated to MUA, but rather that HFO occurrence provides additional explanatory power beyond theta alone. Accordingly, we show directly in Supplementary Figures S5L-M, in response to Reviewer 1’s comment #1, that theta periods in phasic REM not associated with HFOs do not elicit the MUA activity suppression similar to HFOs.

      The phase-alignment performed is hard to follow and is not being applied to HFO periods. If the study is trying to argue that high-theta periods without HFOs in the same recording sessions show lower MUA, then perhaps some sort of shuffle or jitter would be more suitable. For example, in Figure 1B, it seems there are some high-theta periods that don't have HFOs and appear to have higher MUA.

      We expect that the new analyses where we provide additional baseline theta controls for periods adjacent to HFOs in Supplementary Figures S5L-M now clarifies this point. Briefly, alignment to theta phases at different temporal distances from detected HFOs does not exhibit the same fluctuating PFC activity as HFO alignment.

      Regarding the phase alignment procedure that we used for the control analysis, since REM sleep is characterized almost entirely by ongoing theta activity, the control condition was not a separate brain state but rather theta periods outside of HFOs. We therefore needed a systematic way to select comparable theta cycles and phase bins in order to make a valid comparison with HFO-aligned PFC population activity. Since we demonstrated that there is significant phase amplitude coupling between theta and HFOs, thus a theta phase preference of HFOs, we used that specific phase bin across multiple theta cycles to align PFC activity. This phase bin selection was performed separately for each epoch to account for inter epoch and animal variability in phase preference. We reasoned that this procedure would allow for a valid comparison as compared to random alignment, since activity was aligned to similar phases in the baseline theta vs HFO conditions.

      (5) As far as I could tell, the study does not distinguish between putative excitatory and inhibitory neurons in the PFC, but only in the CA1, even though these play very different roles in the model. What is the rationale for not separating these? How are reactivations to be interpreted among interneurons?

      We apologize for the lack of emphasis on this point, which was originally highlighted in Supplementary Figure S2G (now in Supplementary Figure S2F), and for omitting the explanation as to why we did not separately analyze putative excitatory and inhibitory neurons. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with recording primarily from pyramidal neurons (Supplementary Figure S2F, left). We therefore decided to pool and not separate the populations into putative pyramidal cells and interneurons for the spiking analyses presented. To further validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO-aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript, as we did not have enough interneurons to investigate them separately. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      Regarding the interpretation of reactivation in the context of interneurons, we expect reactivation reflects coordinated ensemble activity with excitatory neurons encoding task-relevant information, and inhibitory interneurons shaping timing and neural synchronization. We however did not record enough distinct interneurons to test the predictions of the model, which is now noted in the Limitations on Page 16.

      Relatedly, we find that PFC assemblies detected from pooled data have task relevant representations (Figures 5F-G).

      (6) Are the high-cofiring CA1 neurons generally higher-firing than the other cells? Could this perhaps explain why they behave differently?

      We apologize that this information was not more evident in the manuscript, as it is an important control. In Supplementary Figure S9A, we show that there was no difference in baseline firing rate between low and high cofiring CA1 neurons.

      We have now explicitly referenced this figure in the main text on Page 9, Lines 42-45:

      “Analysis of low and high cofiring CA1 neurons during REM HFOs showed that high cofiring neurons exhibited elevated activity during chained events as compared to low cofiring neurons (Figure 7D), independent of baseline firing rates (Figure S9A).”

      (7) It appears that the decreased firing around HFO's could be a consequence of the stronger firing modulation around these periods, related to time averaging, rather than suppression per se. How does the firing rate compare to other periods with similar modulation that might not have HFOs?

      We thank the Reviewer for raising this important point. We expect that the new analyses, where we provide additional baseline non-HFO-associated theta periods, and theta periods adjacent to HFOs at different temporal distance as controls in also Supplementary Figure S5L-N now clarifies this point. These figures show that the decreased firing rate is specific to HFO chains (Supplementary Figure S5N). Indeed, if the suppression that we observe is related to time averaging, or another analytical artifact, our claims of suppression during HFOs would not be valid. We present a number of results and provide further explanations to support the accuracy of our characterization of PFC population suppression.

      First, event-aligned multiunit activity was quantified as baseline-normalized population firing relative to detected events. For both NREM and REM, activity was normalized by the mean population firing rate during a baseline period within the same sleep state in which events were detected. Values are therefore expressed as deviations from baseline. This normalization allows comparison of relative changes in firing around events within each state. Importantly, values below baseline reflect reductions relative to the state-matched baseline period.

      Second, we show that smoothing activity with a larger gaussian kernel preserves the dip in population activity, consistent with suppression of activity surrounding HFOs (Figure 9C, note that this is a data figure presented in the context of the model). However, as raised by the Reviewer, this normalized measure does not on its own distinguish sustained suppression from transient deviations introduced by event-locked temporal structure in firing.

      Third, to address this, we employed an alternative method to demonstrate that HFO chains tend to occur during periods of PFC suppression (Figure 4G). Briefly, we detected events in PFC where activity fell below a threshold and calculated the probability of HFOs surrounding these “suppression” events. We refer the Reviewer to the Methods section under “Detection of population suppression” where we explain this procedure in more detail. We found that chain ripples, during which the strongest suppression is observed (Figure 4F), are associated with decreases in PFC activity (Figure 4G and Supplementary Figure S6).

      Fourth, we refer the reviewer to Figure S5N, where we compare the firing rates of PFC neurons during HFO chains and theta periods outside of HFOs during putative phasic REM bouts. The observed reduction in firing during true HFO chains compared to non-HFO periods therefore reflects HFO event-specific activity suppression rather than a common occurrence during baseline REM periods.

      Lastly, we performed a control analysis complimentary to Supplementary Figure S5K where we aligned PFC activity to the preferred theta phase of HFOs and investigated how distance from detected HFOs modulates PFC activity (Figure S5). We found that there was no consistent theta-modulated activity aligned to HFO-adjacent theta phases.

      Regarding the final comment, if the Reviewer meant “modulation” as in the theta modulation or suppression observed when PFC population activity is aligned to HFOs (Figures 2A and 4F), we are not aware of any other REM periods where this strong theta modulation or suppression of PFC population activity is present. To our knowledge, we are the first to demonstrate such a modulation of PFC population activity in REM sleep. The closest comparison that we can make is PFC activity aligned to gamma events that are coordinated with HFOs (suppression of PFC), but this is explained only with association with HFOs, as shown in Supplementary Figure S4J.

      (8) The assembly reactivation measure does not control for pre-existing assemblies. The term "activation strength" would therefore be more appropriate.

      We thank the reviewer for this important methodological point. The concern that ICA-based reactivation strength does not, by itself, distinguish behavior-induced reactivation from pre-existing assembly activity/structure is well-taken, and we have implemented several complementary analyses that directly address these concerns.

      First, the interleaved structure of our recordings (8 run epochs interleaved with 9 sleep epochs) allows us to investigate the within-session pre/post assembly strength differences (i.e. each W-Track run session has a preceding (pre) and following (post) sleep session). An increase in assembly strength from pre to post is a hallmark of behaviorally relevant assembly reactivation (Kudrimoti, Barnes et al. 1999, Peyrache, Khamassi et al. 2009). For each run epoch, the same run-derived templates were projected onto the preceding and following sleep epochs, and the pre and post strengths were compared. The distribution of post-minus-pre reactivation differences across all epoch pairs is significantly skewed toward positive values (Figure 6F), indicating that templates from the run epochs are more strongly expressed in the post-sleep epochs. This asymmetry cannot be explained by pre-existing assembly structure, which would predict similar assembly strengths.

      Second, reactivation strength in post-experience sleep increases across the experiment, with templates from later running epochs producing the strongest reactivation in the following post-sleep (Figure 6G). This increase in reactivation strength over time cannot be explained by preexisting assembly structure, which predicts similar assembly expression strength independent of experience.

      Third, the detected assemblies carry behaviorally meaningful structure. Assembly activation maps computed during running exhibit spatially organized "assembly fields" similar to single-cell place fields (Figure 6H), demonstrating that the detected assemblies represent specific spatial locations or task variables rather than behavior-independent states. Pre-existing co-firing structure unrelated to ongoing experience would not be expected to produce spatially tuned, task-locked assembly activation. Furthermore, this spatial tuning of assemblies was verified by comparison with surrogates, where assembly maps were generated using circularly shuffled activation times (1000 shuffles). Assemblies with p < 0.05 (z > 1.65) were considered to have significant spatial structure (Figure 6I).

      These results establish that what we measure is the selective re-expression of behaviorally relevant assemblies in subsequent sleep epochs, consistent with the use of "reactivation" in the established literature (Peyrache, Khamassi et al. 2009, Lopes-dos-Santos, Ribeiro et al. 2013). We have therefore retained the term "reactivation strength" and have added text to the manuscript noting these new results that justify the use of “reactivation” strength.

      We have now added Figure 6.

      We have also added the procedure for the calculation of spatial information to the Methods section under “Spatial information of assembly fields” on Page 44, Lines 8-23.

      (9) Can the study rule out that the rank-ordering in Fig 4C is related to firing rates? Higher-firing rates tend to fire earlier, and lower-firing cells later.

      We thank the Reviewer for raising this interesting point. Here, we assume that the Reviewer meant the baseline firing rates of the neurons, not the intra-HFO firing rates of the neurons. Indeed, when we look at baseline REM firing rates of these PFC neurons, we do find that neurons with higher firing rates tend to fire earlier than low-firing-rate neurons (Author response image 3). This is also true when PFC rank and firing rate are assessed for isolated REM HFOs and NREM PFC ripples (Author response image 3). Similarly, we also observe this relationship in CA1 during SWRs, during which rank order correlation is typically assessed as a method for replay detection. In line with this, a previous study has shown that CA1 neurons with high excitability at animals’ current location tend to initiate replay events (Karlsson and Frank 2009). Furthermore, high-firing-rate, rigid CA1 neurons are more active during SWRs than low-rate, plastic neurons (Grosmark and Buzsaki 2016), and there are distinct populations of neurons in both hippocampus and PFC that are preferentially active during immobility in sleep epochs (Jarosiewicz, McNaughton et al. 2002, Kay, Sosa et al. 2016, Tang, Shin et al. 2017), potentially biasing replay activity during high-frequency events. Similar dynamics may underlie activity during PFC ripples and HFOs in NREM and REM sleep, respectively. The critical point here is in the leave-one-out cross-validation that we implemented to determine sequence similarity—each left out event’s cell rank was correlated with the averaged rank of the template that was generated from all other events. This analysis provides a basis for our claim that there is preserved sequential PFC activity across HFO chains. We did not observe neuron firing consistency during isolated HFOs or during pseudo-HFO chains (coherently shifted chain HFO times), which indicates that REM HFO chains are unique temporal windows during which PFC activity proceeds in a more structured manner.

      Author response image 3.

      Firing rate difference of low and high rank neurons (A) Comparison of baseline firing rates of PFC and CA1 neurons split by average rank across all PFC REM HFOs, PFC NREM ripples, or CA1 SWRs. Baseline rates were calculated separately for NREM and REM sleep.

      Minor:

      (1) It gets confusing that the authors sometimes (but not always) refer to HFOs during NREM as "ripples" but not if they occur during REM. The terminology is inconsistent. When they refer to HFO chains, it seems they now pool between REM and NREM periods, as well as across phasic and tonic REM periods, which is confusing.

      We apologize for the confusion regarding the terminology. In the revised manuscript, we now use NREM ripples exclusively for NREM events and REM HFOs exclusively for REM events. We have removed mixed labels such as “ripple/HFO” except where a collective term is explicitly defined. We also clarified that HFO chains refer to REM events only and revised the relevant text/figure legends to avoid any implication that chain analyses pool NREM and REM events.

      (2) P7 L8: It might be helpful to emphasize "broader temporal distribution".

      We thank the Reviewer for the suggestion. We have updated the text on Page 8, Line 18:

      “Since we observed a broader temporal distribution of activity…”

      (3) P8 L9: What do they mean by spatial? Do they mean the place-fields of these same neurons during a previous task period?

      We apologize for the confusion. The Reviewer is correct. Here, we calculated the spatial rate map correlation between CA1 neurons as a measure of place field similarity during the W-Track session prior to the sleep epoch being assessed.

      For clarification, we have added “rate map” to the text on Page 9, Line 33:

      “…spatial rate map correlation…”

      We have also updated the y-axis label for Figure 7B for clarity.

      (4) P8 L27: What do they mean by "coordinated SWRs"? As opposed to what?

      Here, we are referring to our previous study where we investigated ripples in NREM sleep and showed that ripples and SWRs in PFC and CA1, respectively could either be independent from or coordinated with events in the other region (Shin and Jadhav 2024). A main result in the study showed that CA1 neurons are strongly suppressed during independent PFC ripples and that there was a relationship between activity suppression and reactivation during coordinated SWRs (CA1 SWR-PFC ripple coordination in NREM). We specifically mentioned “coordinated” since these are SWRs that are also coupled with SOs and spindles as compared to SWRs that are independent from PFC ripples (Shin and Jadhav 2024). Overall, we wanted to frame this result in the context of oscillatory coupling and mechanisms of memory consolidation.

      (5) P37 L28 says "we observed a bimodal distribution" but L31 says "unimodal". Which is it?

      We apologize for the confusion. We observed a bimodal distribution for CA1-PFC cofiring in REM sleep, but a unimodal distribution in NREM sleep. Because of these two observations, we decided to split the CA1 population into high and low cofiring neurons based on two different thresholds:

      (1) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) 0.

      (2) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) the mean of the distribution of averaged cofiring values.

      Using two separate thresholds to split high and low cofiring CA1 neurons demonstrates the robustness of the firing rate change result in Figures 7F and Supplementary Figures S9B-D.

      We added a statement that clarifies that the bimodal distribution was seen in REM sleep only on Page 44, Lines 37-43:

      “Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0. Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions.”

      Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which, if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      (1) Distinction from wake HFOs

      The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent a more general theta-associated phenomenon.

      To investigate PFC high-frequency oscillations during running behavior on the W-Track, events were extracted in the same manner as NREM and REM events (Methods). Events during wake were subset by periods where the animals’ velocity was >4 cm/s to provide a comparison of events during periods of high theta. While we were able to detect HFOs during wake that exhibited a similar spectral profile in the high frequency band, we did not observe 1) strong association with gamma or theta oscillations, 2) prominent HFO chaining, 3) HFO aligned theta modulated PFC activity, 4) comparable levels of theta phase amplitude coupling, 5) association with population suppression, or 6) a relationship between peri-event theta power and multiunit activity (Supplementary Figure S11). Many of the defining features of PFC REM HFOs are absent during wake, indicating REM specificity of the results we present.

      (2) Link to memory consolidation

      The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      To address these concerns, we have now added an explicit “Limitations” section in the main text of the manuscript that includes a statement about learning. We have also added Supplementary Figure S1 in the manuscript, which illustrates the performance of all 10 animals on the W-Track task. Finally, we have also included Figures 6F G in the manuscript, showing that PFC assembly reactivation strength during sleep epochs increases during learning.

      Reviewer #3 (Recommendations for the authors):

      Most of my specific comments were related to further clarification that will help the reader's understanding.

      (1) I'd recommend simplifying terminology. Open to debate, but would it not be simpler and clearer to just say NREM HFO vs REM HFO? Obviously, there is a need to mention how NREM HFOs have previously been referred to as cortical ripples, but I'm not sure it is such a helpful terminology to continue for the field, given, as you state, how different cortical ripples are from hippocampal SWRs. If not, I'd at least provide a clearer explanation early in the manuscript, distinguishing NREM ripples from REM HFOs but collectively still calling them 'cortical events'.

      We appreciate the Reviewer’s suggestion regarding terminology and agree that it would be simpler and clearer to use a single term, ripple or HFO, to describe these events. Initially, we had used a unified term (ripples across both states) but ultimately decided to switch to state-specific terminology due to previous comments we received on the manuscript and to emphasize the distinctions between NREM and REM events. Ultimately, we decided on calling them ripples in NREM and HFOs in REM since there is precedence for both terms in each respective sleep state (Khodagholy, Gelinas et al. 2017, Vaz, Inati et al. 2019, Bueno-Junior, Ruckstuhl et al. 2023, Shin and Jadhav 2024), but we do agree that this distinction can be confusing if not clearly stated. Thus, we have added an additional statement in the manuscript on Page 4, Line 44 to Page 5, Lines 1-3 for clarity:

      “Similar criteria were used to detect cortical high-frequency events in NREM and REM states; however, to avoid ambiguity and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout”

      (2) I assume experiments occurred during the light phase, but it would be good if this could be stated explicitly.

      We have now added text specifying that these experiments took place in the light phase on Page 34, Lines 9-12 of the Methods section under “Behavior”:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Page 2 Line 14: Rephrase to make clearer, e.g. 'that have a shift in...'.

      We have rephrased the sentence for clarification on Page 2, Lines 13-15:

      “REM HFO chains also preferentially engage CA1 neuronal populations that demonstrate a shift in their preferred theta-phase from behavior to REM sleep.”

      (4) Page 4 Line 36 - 'and find coherent shifts in TD', this is self-fulfilling. I would rephrase to something like 'resulting in...'.

      We have rephrased the sentence on Page 4, Lines 36-38:

      “We separated NREM and REM sleep stages based on theta-to-delta (TD) ratio in CA1, which revealed coherent shifts in TD ratio across CA1 and PFC at the onset and offset of REM sleep…”

      (5) Figure 1H left - clarify how many animals or multiunits this is based on.

      We have now updated Figure 1 legend to specify the number of animals and epochs included on Page 19, Line 4:

      “REM HFO aligned multiunit activity (MUA) in PFC (n = 10 animals, 36 epochs)…”

      (6) Figure 1I (right), it would be useful to see the x-axis frequency start from 0, since you are cutting the peak in power.

      We thank the Reviewer for this suggestion. We had initially set the frequency limits to 4 and 12 to specifically illustrate the absence of theta-modulated activity during NREM ripples. However, as suggested by the Reviewer, it is informative to expand the frequency range to ascertain the location of the peak frequency for NREM. Indeed, the peak frequency of NREM ripple-aligned PFC activity tends to be lower than 4 Hz, which is consistent with a single peak of activity that lasts <1 s.

      We have now updated Figure 2C with these new panels.

      (7) Figure 5 E/I - use of ** is confusing, it looks like a significance comparing e.g. quartile 2 to 1, but I think this is the correlation significance. I'd move ** to the top right corner and ideally include rho values.

      We thank the Reviewer for this suggestion. We have now added the r values for the correlation to Figures 7E and 8B to resolve any ambiguity. In addition, we updated Figure 7F, right and Figure 5B to maintain consistency across main figures.

      (8) Figure 7D - Why are stimulated and non-stimulated cells so different at baseline? This is not the case for the NREM results.

      The y-axes in Figures 10C,D (Formerly Figure 7) show the fraction of cells with at least one spike per 10ms bin, either for all pyramidal cells or all interneurons. We report in the figure legend that the stimulated cells represent only 30% of the network for any given stimulus (Page 32, Line 18), so at baseline in both the NREM and REM simulations, there are roughly 3x as many non-stimulated cells with a spike per 10ms bin compared to stimulated cells, as these populations are active at roughly equal firing rates per neuron outside of stimulation periods.

      (9) Methods - there is limited info on spike sorting procedure - can you provide a reference with further details (I couldn't find)? Is this method equally valid for identifying PFC units?

      Matclust is a MATLAB-based spike sorting graphical user interface that allows for manual curation of neuron clusters through the visualization of spike waveform amplitude, peak-to-trough, and principal components. Polygons or boxes are drawn around spike data points, and single unit clusters are resolved through refinement in multiple dimensions. It was developed by Mattias Karlsson and was first used in a publication reporting replay of remote experiences in the hippocampus (Karlsson and Frank 2009). Although there is no formal reference, it can be found at https://bitbucket.org/mkarlsso/matclust/src/master/. Other labs have used the software for clustering neurons from cortical areas (Yu, Liu et al. 2018, Proskurin, Manakov et al. 2023), which demonstrates its robustness across multiple brain areas.

      Additionally, examples of clustered neurons in PFC over the course of the experimental paradigm used here can be found in our previous publication (Shin, Tang et al. 2019). In addition to the aforementioned references, we show that PFC neurons can be accurately clustered and that neurons are stable over time, according to a number of cluster metrics.

      (10) Why were there no further analyses of pyramidal cells and interneurons beyond Figure S2?

      We thank the Reviewer for bringing up this important point, which is similar to Reviewer 2, comment #5 above. We repeat our response here. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with primarily recording from pyramidal neurons (Supplementary Figure S2F, left). All results are similar if we exclude putative interneurons. To validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      (11) Looking at Figure S1B, the second to last main block of NREM sleep shown has a clear peak passing TD threshold, but oddly not classed as REM - I can only assume this is due to the duration limits on your classification?

      Yes, this is due to the REM duration threshold that we implement in our sleep scoring algorithm. We used a minimum REM bout threshold criterion of 10 s for inclusion, as in previous reports (Rothschild, Eban et al. 2017, Zhang, Zhang et al. 2020). Additionally, we have included example sleep plots for all 10 animals in Supplementary Figure S1.

      (12) Include details of how head speed was calculated - just based on the 30fps video?

      Yes, the head speed of the animal was determined by tracking the animals’ position and calculating the speed based on cm/pixel values. We have now added more detail on this in the Methods section under “Surgical implant and electrophysiology” on Page 33, Line 44 to Page 34, Lines 1-2:

      “Additionally, the animals’ speed was calculated based on predetermined cm/pixel values and the position displacement between frames captured at 30 fps.”

      (13) Page 31 line 7 - With the reference you cite, they didn't really show tonic and phasic REM can be segregated based on theta frequency - they just defined it as such. You've done it for some, but I'd ensure all references to phasic REM are defined as putative - mostly missed within the discussion - I'd also add this as a brief limitation (without recording of eye movements or PGO waves).

      We thank the Reviewer for these suggestions. We have now ensured that all mentions of “phasic REM” are qualified with “putative” and have added the requested limitation in the Limitations section on Page 16, Lines 10-15:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states.(Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult.”

      We have also modified the wording in the Methods section under “Theta inter-peak intervals during bouts of high and low theta power” on Page 37, Lines 14-15 to indicate that the referenced article simply used the theta frequency-based method to define tonic and phasic REM – not to explicitly separate the two states:

      “Since previous studies have segregated putative tonic and phasic substages of REM sleep based on CA1 theta frequency…”

      References:

      Abdou, K., M. Nomoto, M. H. Aly, A. Z. Ibrahim, K. Choko, R. Okubo-Suzuki, S. I. Muramatsu and K. Inokuchi (2024). "Prefrontal coding of learned and inferred knowledge during REM and NREM sleep." Nat Commun 15(1): 4566.

      Aleman-Zapata, A., R. G. M. Morris and L. Genzel (2022). "Sleep deprivation and hippocampal ripple disruption after one-session learning eliminate memory expression the next day." Proc Natl Acad Sci U S A 119(44): e2123424119.

      Belluscio, M. A., K. Mizuseki, R. Schmidt, R. Kempter and G. Buzsaki (2012). "Cross-frequency phase coupling between theta and gamma oscillations in the hippocampus." J Neurosci 32(2): 423–435.

      Bueno-Junior, L. S., M. S. Ruckstuhl, M. M. Lim and B. O. Watson (2023). "The temporal structure of REM sleep shows minute-scale fluctuations across brain and body in mice and humans." Proc Natl Acad Sci U S A 120(18): e2213438120.

      Cairney, S. A., S. J. Durrant, R. Power and P. A. Lewis (2015). "Complementary roles of slow-wave sleep and rapid eye movement sleep in emotional memory consolidation." Cereb Cortex 25(6): 1565–1575. Chang, H., W. Tang, A. M. Wulf, T. Nyasulu, M. E. Wolf, A. Fernandez-Ruiz and A. Oliva (2025). "Sleep microstructure organizes memory replay." Nature 637(8048): 1161–1169.

      Cheng, S. and L. M. Frank (2008). "New experiences enhance coordinated neural activity in the hippocampus." Neuron 57(2): 303–313.

      Darevsky, D., J. Kim and K. Ganguly (2024). "Coupling of Slow Oscillations in the Prefrontal and Motor Cortex Predicts Onset of Spindle Trains and Persistent Memory Reactivations." J Neurosci 44(43). Davidson, T. J., F. Kloosterman and M. A. Wilson (2009). "Hippocampal replay of extended experience." Neuron 63(4): 497–507.

      Ellenbogen, J. M., P. T. Hu, J. D. Payne, D. Titone and M. P. Walker (2007). "Human relational memory requires time and sleep." Proc Natl Acad Sci U S A 104(18): 7723–7728.

      Foster, D. J. and M. A. Wilson (2006). "Reverse replay of behavioural sequences in hippocampal place cells during the awake state." Nature 440(7084): 680–683.

      Ghosh, M., F. C. Yang, S. P. Rice, V. Hetrick, A. L. Gonzalez, D. Siu, E. K. W. Brennan, T. T. John, A. M. Ahrens and O. J. Ahmed (2022). "Running speed and REM sleep control two distinct modes of rapid interhemispheric communication." Cell Rep 40(1): 111028.

      Grosmark, A. D. and G. Buzsaki (2016). "Diversity in neural firing dynamics supports both rigid and learned hippocampal sequences." Science 351(6280): 1440–1443.

      Helfrich, R. F., J. D. Lendner, B. A. Mander, H. Guillen, M. PaD, L. Mnatsakanyan, S. Vadera, M. P. Walker, J. J. Lin and R. T. Knight (2019). "Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans." Nat Commun 10(1): 3572.

      Jarosiewicz, B., B. L. McNaughton and W. E. Skaggs (2002). "Hippocampal population activity during the small-amplitude irregular activity state in the rat." J Neurosci 22(4): 1373–1384.

      Ji, D. and M. A. Wilson (2007). "Coordinated memory replay in the visual cortex and hippocampus during sleep." Nat Neurosci 10(1): 100–107.

      Karlsson, M. P. and L. M. Frank (2009). "Awake replay of remote experiences in the hippocampus." Nat Neurosci 12(7): 913–918.

      Kay, K., M. Sosa, J. E. Chung, M. P. Karlsson, M. C. Larkin and L. M. Frank (2016). "A hippocampal network for spatial coding during immobility and sleep." Nature 531(7593): 185–190.

      Khodagholy, D., J. N. Gelinas and G. Buzsaki (2017). "Learning-enhanced coupling between ripple oscillations in association cortices and hippocampus." Science 358(6361): 369–372.

      Kudrimoti, H. S., C. A. Barnes and B. L. McNaughton (1999). "Reactivation of hippocampal cell assemblies: effects of behavioral state, experience, and EEG dynamics." J Neurosci 19(10): 4090–4101. Lee, A. K. and M. A. Wilson (2002). "Memory of sequential experience in the hippocampus during slow wave sleep." Neuron 36(6): 1183–1194.

      Leemburg, S., V. V. Vyazovskiy, U. Olcese, C. L. Bassetti, G. Tononi and C. Cirelli (2010). "Sleep homeostasis in the rat is preserved during chronic sleep restriction." Proc Natl Acad Sci U S A 107(36): 15939–15944.

      Lopes-dos-Santos, V., S. Ribeiro and A. B. Tort (2013). "Detecting cell assemblies in large neuronal populations." J Neurosci Methods 220(2): 149–166.

      Mizuseki, K., K. Diba, E. Pastalkova and G. Buzsaki (2011). "Hippocampal CA1 pyramidal cells form functionally distinct sublayers." Nat Neurosci 14(9): 1174–1181.

      Nitsche, M. A., M. Jakoubkova, N. Thirugnanasambandam, L. Schmalfuss, S. Hullemann, K. Sonka, W. Paulus, C. Trenkwalder and S. Happe (2010). "Contribution of the premotor cortex to consolidation of motor sequence learning in humans during sleep." J Neurophysiol 104(5): 2603–2614.

      Peyrache, A., M. Khamassi, K. Benchenane, S. I. Wiener and F. P. Battaglia (2009). "Replay of rule-learning related neural patterns in the prefrontal cortex during sleep." Nat Neurosci 12(7): 919–926.

      Plitt, M. H. and L. M. Giocomo (2021). "Experience-dependent contextual codes in the hippocampus." Nat Neurosci 24(5): 705–714.

      Proskurin, M., M. Manakov and A. Karpova (2023). "ACC neural ensemble dynamics are structured by strategy prevalence." Elife 12.

      Rothschild, G., E. Eban and L. M. Frank (2017). "A cortical-hippocampal-cortical loop of information processing during memory consolidation." Nat Neurosci 20(2): 251–259.

      Shin, J. D. and S. P. Jadhav (2024). "Prefrontal cortical ripples mediate top-down suppression of hippocampal reactivation during sleep memory consolidation." Curr Biol 34(13): 2801–2811 e2809.

      Shin, J. D., W. Tang and S. P. Jadhav (2019). "Dynamics of Awake Hippocampal-Prefrontal Replay for Spatial Learning and Memory-Guided Decision Making." Neuron 104(6): 1110–1125 e1117.

      Siapas, A. G. and M. A. Wilson (1998). "Coordinated interactions between hippocampal ripples and cortical spindles during slow-wave sleep." Neuron 21(5): 1123–1128.

      Simor, P., G. van der Wijk, L. Nobili and P. Peigneux (2020). "The microstructure of REM sleep: Why phasic and tonic?" Sleep Med Rev 52: 101305.

      Singer, A. C. and L. M. Frank (2009). "Rewarded outcomes enhance reactivation of experience in the hippocampus." Neuron 64(6): 910–921.

      Sosa, M., H. R. Joo and L. M. Frank (2020). "Dorsal and Ventral Hippocampal Sharp-Wave Ripples Activate Distinct Nucleus Accumbens Networks." Neuron 105(4): 725–741 e728.

      Sosa, M., M. H. Plitt and L. M. Giocomo (2025). "A flexible hippocampal population code for experience relative to reward." Nat Neurosci 28(7): 1497–1509.

      Stark, E., L. Roux, R. Eichler and G. Buzsaki (2015). "Local generation of multineuronal spike sequences in the hippocampal CA1 region." Proc Natl Acad Sci U S A 112(33): 10521–10526.

      Tang, W., J. D. Shin, L. M. Frank and S. P. Jadhav (2017). "Hippocampal-Prefrontal Reactivation during Learning Is Stronger in Awake Compared with Sleep States." J Neurosci 37(49): 11789–11805.

      Tort, A. B., R. Scheffer-Teixeira, B. C. Souza, A. Draguhn and J. Brankack (2013). "Theta-associated high-frequency oscillations (110-160Hz) in the hippocampus and neocortex." Prog Neurobiol 100: 1–14.

      Valero, M., T. J. Viney, R. Machold, S. Mederos, I. Zutshi, B. Schuman, Y. Senzai, B. Rudy and G. Buzsaki (2021). "Sleep down state-active ID2/Nkx2.1 interneurons in the neocortex." Nat Neurosci 24(3): 401–411.

      van de Ven, G. M., S. Trouche, C. G. McNamara, K. Allen and D. Dupret (2016). "Hippocampal Offline Reactivation Consolidates Recently Formed Cell Assembly Patterns during Sharp Wave-Ripples." Neuron 92(5): 968–974.

      van der Helm, E. and M. P. Walker (2011). "Sleep and Emotional Memory Processing." Sleep Med Clin 6(1): 31–43.

      Vaz, A. P., S. K. Inati, N. Brunel and K. A. Zaghloul (2019). "Coupled ripple oscillations between the medial temporal lobe and neocortex retrieve human memory." Science 363(6430): 975–978.

      Wilson, M. A. and B. L. McNaughton (1994). "Reactivation of hippocampal ensemble memories during sleep." Science 265(5172): 676–679.

      Yang, S. R., H. Sun, Z. L. Huang, M. H. Yao and W. M. Qu (2012). "Repeated sleep restriction in adolescent rats altered sleep patterns and impaired spatial learning/memory ability." Sleep 35(6): 849–859.

      Yu, J. Y., D. F. Liu, A. Loback, I. Grossrubatscher and L. M. Frank (2018). "Specific hippocampal representations are linked to generalized cortical representations in memory." Nat Commun 9(1): 2209.

      Zhang, L. B., J. Zhang, M. J. Sun, H. Chen, J. Yan, F. L. Luo, Z. X. Yao, Y. M. Wu and B. Hu (2020). "Neuronal Activity in the Cerebellum During the Sleep-Wakefulness Transition in Mice." Neurosci Bull 36(8): 919– 931.

    1. eLife Assessment

      This study presents a useful combined analysis of human behavior, pupillometry and EEG measures to probe differences in ambiguity assessment between individuals during value-based decision-making. Using computational models, the authors aim to test for distinct groups of participants that differ in the relationship between evidence accumulation, arousal, and neural decision-making computations. However, at present, the evidence for group differences in the physiological, neural, and behavioral profile of decision-making under uncertainty is incomplete, lacking an appropriate statistical test for group differences, and a more complete assessment of the modeling results on evidence accumulation in decision-making in this study is required to robustly characterize and interpret the results.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript investigates value-based decision-making under risk and ambiguity using a combination of behavioral, pupillometric, and EEG data. Participants are stratified into three "decision styles" (ideal, aggressive, conservative) based on how their choices under known risk align with expected-value optimality. The central claim is that ambiguity aversion is not a uniform bias but reflects heterogeneous internal belief models, and that physiology tracks subjective belief rather than objective task structure. While this is an interesting conceptual question, the evidence is underwhelming given that differences between groups are not tested statistically (but just described), there are clear problems with how the computational models are implemented, and there are serious issues with sampling of participants.

      Strengths:

      The multimodal design (behavior, pupillometry, EEG) and the attempt to link a latent belief parameter to physiological signatures address a question of clear interest.

      Weaknesses:

      Framing and motivation

      (1) The framing conflates two questions that appear distinct. The motivation centers on "ambiguity aversion," but the study's actual aim - how individuals internally represent ambiguous outcomes - seems like a different question. The relationship between these two framings needs to be made explicit, because as written the motivating phenomenon and the studied phenomenon are not obviously the same thing.

      (2) Several of the contrasts the paper sets up against prior literature read as strawmen. The claim that ambiguity aversion is treated as "a single bias or fixed trait that applies uniformly" is presented as the view being overturned, but it is not clear this is a position the field actually holds - it reads as a strawman. Relatedly, the central objective-versus-subjective valuation distinction that the results are built around also reads as a strawman dichotomy rather than a genuine competing account.

      (3) The motivation for the physiological measures is overly broad. The statement linking EEG to control, attention, valuation, uncertainty, conflict, effort, and engagement is so general as to be uninformative - EEG signals have been linked to essentially everything, so this does not constrain the hypotheses or predictions. A more specific, falsifiable rationale is needed.<br /> Design, sample, and grouping.

      (4) The inclusion of the collaborative spacecraft/Apollo task is unclear. It is not explained why this task is included, and its role relative to the core ambiguity question needs justification (this also bears on the leadership analyses; see below).

      (5) The participant numbers do not add up and must be reconciled. The text reports 57 participants, yet the analyses describe three groups of roughly 32 + 32 + 31. The relationship between participants, sessions, and group n's needs to be stated clearly and consistently, because at present the sample description is internally contradictory.

      (6) The rationale for categorizing participants into three discrete groups is not established, and the approach is statistically questionable. Decision tendency appears to be a continuous variable; dichotomizing/trichotomizing a continuous measure is generally discouraged and can manufacture or distort group differences. The authors should justify why discrete groups are needed at all, and ideally show whether there are genuine group differences (e.g., evidence of discontinuity/clustering) rather than an arbitrary split of a continuum.

      Statistics

      (7) Key claims about how ambiguity affects groups differently are made without the appropriate test. To support a claim that the effect of ambiguity differs across groups, the interaction (group × ambiguity) must be shown - group-wise effects reported separately are not sufficient. This is really a key limitation of the current work.

      (8) The methods mentioned that some participants performed multiple sessions, but their data were treated as if coming from separate participants. This is incorrect for several reasons, particularly given the focus on individual differences.

      Belief parameter and terminology

      (9) The term "ideal" is not justified. It is unclear why this group is labelled "ideal" - are they Bayes-optimal, or optimal in some defined sense? If the label implies normativity, that needs to be demonstrated; otherwise it should be renamed.

      Drift-diffusion modelling

      (10) The boundary parameter is fixed (a detail which is hidden in the methods), but this is highly problematic. By enforcing the same boundary value for all participants, the model is forced to capture any variation as drift rate effects. As such, all conclusions about drift rate are not interpretable as they might reflect boundary effects in disguise.

      (11) The DDMs are fit separately per group of participants, which again precludes testing interactions. As with the behavioral analyses, fitting separate models means group differences cannot be properly compared within a single statistical framework, and interactions cannot be assessed. The paper does mention some comparison between groups, but comparing DDM parameter estimates across separately fit models is not valid.

      (12) Overall, the DDM is very complex, and the manuscript does not yet provide enough validation to make the model trustworthy. Given the number of trial-wise covariates entering the drift rate and the per-participant fitting, stronger evidence that the model is identifiable and that its parameters are recoverable/reliable is needed before the conclusions drawn from it can be accepted.

      Methods - EEG and analysis details

      (13) The high-pass filter setting appears very aggressive. The authors should confirm whether this risks removing genuine low-frequency signal of interest, particularly given that delta-band effects are later interpreted.

      (14) There is an apparent inconsistency in the epoching/time-locking. The time-frequency analysis appears to be computed on choice-locked data, yet elsewhere the epochs are described as stimulus-locked. This needs to be clarified and made consistent, as it affects interpretation of the pre- versus post-decision EEG clusters.

      (15) The mixed-effects modelling appears to omit random slopes. The authors should justify the random-effects structure (e.g., why only random intercepts), as this affects the validity of the inference.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Qin and colleagues entitled "Pupil and Neural Dynamics Reveal Belief-Dependent Decision Making Under Ambiguity" examines decision-making under risk and ambiguity using pupillometry and EEG. The study employs a lottery choice task with three levels of ambiguity (zero, low, high). Participants were classified into three groups based on their choice behavior in a condition with risk and no ambiguity: ideal (choosing in line with objective expected values), aggressive (preference for investments), and conservative (preference against investments). The authors then compared behavior, pupil, and EEG results across these groups. The study concludes that individual beliefs about ambiguity are reflected in different behavioral strategies and neural correlates.

      Strengths:

      The combination of behavior, computational modeling, pupillometry, and EEG.

      Weaknesses:

      (1) It is unclear whether group definition is theoretically justified.

      One general concern is that the strategy to form three distinct groups is not clearly motivated. The authors created the three groups, "aggressive", "ideal", and "conservative", based on the zero-ambiguity trials. However, as the authors state: "Ambiguity differs fundamentally from risk at both the physiological level (34; 6) and the behavioral level" (page 4). Under this assumption, it is questionable whether forming groups based on risk preferences is a useful strategy for studying ambiguity. What do we learn about ambiguity processing when group differences are primarily based on risk preferences? Might the present results partly be driven by risk preferences rather than ambiguity preferences? I recommend the following two points: (a) Clearly justify the reasoning behind the group approach; (b) Add an additional continuous analysis approach indicating whether the key results hold independent of the group definition based on risky decision-making.

      (2) k-parameter.

      The authors use the k-parameter that infers the expected high-payoff probability (e.g., page 11). On page 22, this is explained as: "the subjective value term K was assigned according to each participant's internal belief of the high-payoff rate under ambiguity, yielding a participant-specific estimate of expected value under uncertainty." I hope I have not missed anything, but I neither understood the role of this parameter nor how it was computed.

      (3) How were individual beliefs and models computed?

      A related but more general point is that it remained unclear how the authors computed internal beliefs and internal models in the study. The study contains many statements suggesting that the authors measured internal beliefs. For example:

      a) Abstract: "We show that individuals adopt distinct decision strategies that reflect different internal beliefs about unknown outcomes."<br /> b) Page 3: "We then inferred subjective belief parameters that captured how individuals internally interpreted the ambiguous probability mass and examined how these beliefs related to choice behavior, arousal dynamics, and neural activity."<br /> c) Page 16: "Together, these findings show that ambiguity does not evoke a uniform behavioral or physiological response across participants with different decision-making styles; instead, individuals rely on distinct internal models and computational strategies when forming decisions under ambiguity."<br /> d) Page 16: "Taken together, these results show that ambiguity aversion is not a uniform psychological bias, but a set of heterogeneous belief-driven strategies that shape how ambiguity is represented and acted upon."<br /> e) Page 18: "Ambiguity processing, therefore, reflects distinct belief-driven pathways rather than a single canonical mechanism."

      Based on the present data, analyses, and results, I don't think that the authors can draw these conclusions. Which analyses in the manuscript identify these internal beliefs, models, or strategies? How can we dissociate a unified strategy from a heterogeneous set of strategies based on the present results? My feeling is that the k-parameter might be related to this, but as explained above, I did not understand how it was computed and what it is supposed to reflect. The DDM analyses might also be targeted at this. However, it remains elusive how the DDM captures internal beliefs about ambiguity itself. My recommendation is that the authors more clearly explain (a) why the DDM is a useful model to study ambiguity, (b) what the different parameters exactly reflect about ambiguity processing, and (c) how the DDM captures internal beliefs and distinct belief-driven strategies in this context.

      (4) Statistical tests.

      4.1. Figure 2B: The authors summarize the number of participants with significant effects of ambiguity on choice behavior for each group. I recommend a statistical test at the second level that properly assesses the effects of ambiguity and group within a common statistical model. In my opinion, it is not enough to simply count the number of significant tests (from the first level) for each group.

      4.2. Figure 2C: For the analysis of response times, the authors might want to consider reporting the main effects of group and ambiguity.

      4.3. Figure 2D: The text on page 8 states that Figure 2D indicates that "aggressive investors showed no significant pupil modulation by ambiguity...". However, the figure and its caption indicate significant differences between ambiguous and non-ambiguous trials across all groups. Moreover, if the authors want to compare the groups, it is necessary to compare the groups to each other; a test against zero within each group would not be enough to demonstrate any group differences. In my mind, this would also be important for analyses in Figure 3C and D.

      4.4. Strictly speaking, for the statistical tests, it would be necessary to take into account that participants completed multiple sessions (within-subject variance is different from between-subject variance). Currently, each session is treated independently (page 19: "Each individual completed one to three experimental sessions. For data analysis, each session was treated as an independent participant, yielding a total of 108 sessions.")

      (5) Necessary quality control for pupillometry and EEG data.

      The task was performed in a virtual reality environment with a head-mounted display. The task was not isoluminant, and, to the best of my knowledge, participants were not instructed to avoid eye movements. The authors applied a GLM to control for luminance effects in the pupil data. For EEG, they used ICA to remove ocular and muscular artifacts. While these methods are established, they are usually applied to more controlled paradigms optimized for EEG and pupillometry. To demonstrate high data quality despite these issues, it is necessary to present quality-control analyses. Can the authors please indicate how many blinks had to be removed from the data? Could the authors please indicate how many blinks were removed from the data? Can the authors please show trial-level data (after preprocessing) for a few subjects?

      (6) Quality control for the DDM.

      The manuscript lacks systematic posterior predictive checks and parameter recovery for the DDM results. It is important to validate that the model accurately captures the data. Currently, we only see the model parameters, but it remains unclear whether the model performs well on the current data set. Moreover, if the authors aimed to test different strategies using the DDM, it might be useful to perform systematic model comparison.

      (7) Implications of the second experiment with collaborative task remain unclear.

      To me, the link between the main study and the second experiment on leadership and team performance is not obvious. In my opinion, this topic is beyond the scope of the present paper. Linking the two studies more comprehensively based on deeper theoretical grounds would likely be better suited for an independent manuscript.

    4. Author response:

      We read the Assessment as identifying two decisive gaps: (i) claims of group differences are supported by within-group tests rather than by a test of the group X condition interaction, and (ii) the drift-diffusion modeling is not yet validated or fit in a framework that permits group comparison. We agree with both, and we do not defend the current versions of these analyses. The revision will rebuild them rather than supplement them. We also agree with the reviewers that several of our conclusions about “internal beliefs” are currently stated more strongly than the analyses support, and these will be scaled back to what the modeling can carry.

      Below we first note factual errors and internal inconsistencies in the manuscript that the editors asked us to flag promptly, then summarize the planned revisions, then respond to each public review comment in turn.

      (1) Corrections and clarifications for the record

      Reviewer #2 identified one outright error in our text, and in re-checking the manuscript we found several further inconsistencies. We list them here so that they are on the record alongside the first version of the Reviewed Preprint. In each case the reviewers' reading of the manuscript is correct and the manuscript is at fault.

      (a) Aggressive investors and pupil modulation (Reviewer #2, 4.3). Our Results text states that “aggressive investors showed no significant pupil modulation by ambiguity,”. This is incorrect and contradicts our own Fig. 2d, which reports a significant, ambiguous vs. non-ambiguous difference in all three groups, including aggressive investors (T(31) = 2.74, P = 0.0304, Bonferroni-corrected). The figure and its statistics are correct; the text is wrong. The interpretive claim built on it - that aggressive investors show “blunted” arousal - is therefore unsupported and will be removed. The revised manuscript will state the correct within-group result and will test group differences directly rather than by contrasting significant against non-significant within-group tests.

      (b) Within-group versus between-group claims. Relatedly, our Results state that ambiguity-related pupil differences under the subjective model are “no longer significant relative to zero,” whereas the Discussion describes them as no longer differing across groups. These are different claims, and only the first was tested. This conflation runs through several of our physiological conclusions and is the same problem the Assessment identifies. It will be resolved by replacing these statements with explicit between-group and interaction tests.

      (c) Sample description (Reviewer #1, 5). The numbers are not contradictory but are certainly underspecified, and we accept that as written they cannot be reconciled by a reader. To state them plainly: 57 unique individuals each completed one to three sessions, yielding 108 sessions; 7 sessions were incomplete and 5 showed no response variability, leaving 96 sessions with usable behavior; of these, 95 had usable pupil data and 79 usable EEG. The three strategy groups are tertiles of these 96 sessions (32 each), and the smaller Ns in Figs. 2-4 (N = 31, 29, 25) reflect modality-specific exclusions within each tertile. The revision will include a participant/session flow diagram and a table reporting how many individuals contributed one, two, or three sessions, together with per-figure Ns.

      (d) DDM fitting procedure (Reviewer #1, 11). One clarification: the DDMs were fit separately for each participant, not separately per group (Methods, Section 4.6); group comparisons were then performed on the participant-level parameter estimates. We note this only for accuracy of the record. The reviewer's substantive objection is unaffected and we accept it: comparing parameters estimated in independent per-participant fits does not constitute a test of group differences within a common statistical framework, and it cannot test interactions.

      (e) Inconsistent specification of the EEG regressor. The trial-wise EEG covariate is described in one paragraph of Methods 4.6 as 8-13 Hz power averaged over 0-0.5 s, and in the text following Eq. 3 as a 1315 Hz difference over 0.1-0.4 s. The former corresponds to the analysis actually performed. We will correct Eq. 3's description and report the frequency band and window once, unambiguously.

      (f) Time-locking and figure/caption errors. As Reviewer #1 notes (comment 14), Methods describe stimulus-locked epochs (-0.25 to 1 s) while Figs. 2-3 are labeled relative to decision onset over -1 to 2 s. We will state for each analysis whether it is stimulus- or response-locked and harmonize axis labels accordingly. In addition: the Fig. 3 caption reads "N = 25 for ideal and aggressive; N = 29 for aggressive," where the second instance should read conservative; and Section 2.6 cites Fig. 1b and 1c for the leadership and team-performance results, which are Fig. 4e and 4f.

      (g) Delta/theta claims in the Discussion. Our Discussion attributes early delta- and theta-band enhancements to ideal investors. No such effects appear in our own time-frequency results, which report frontal beta-range and parietal alpha/low-beta clusters. These Discussion statements are not supported by the data presented and will be deleted. This also bears on Reviewer #1's comment 13, since it removes the only interpretation that depended on the low-frequency edge of our filter passband.

      (2) Summary of planned revisions

      (2.1) A single statistical framework with explicit interaction tests. All group comparisons will be replaced by unified models that include group, condition, and their interaction, with sessions nested within participants. Choice will be modeled with a generalized linear mixed model of the form invest ~ ambiguity ⨉ group + trial + (1 + ambiguity | participant/session); decision time with the corresponding linear mixed model including random slopes for ambiguity (Reviewer #1, 15). Time-resolved pupil and time-frequency EEG effects will be evaluated using cluster-based permutation tests on the group-by-condition interaction statistic rather than by aggregating within-group tests. Where our claim is that an effect is absent, most importantly, that ambiguity-related pupil differences vanish under subjective valuation, we will support it with equivalence testing and Bayes factors rather than with a non-significant P value, since a null result is not evidence for the null.

      (2.2) Continuous analyses as primary; grouping justified or abandoned. We accept that trichotomizing a continuous decision tendency requires justification that we did not provide. The revision will (i) define a continuous EV-consistency index and re-run every key analysis with it as a continuous predictor, establishing that the principal conclusions do not depend on the split. The “ideal” label implies a normative optimality we have not demonstrated and will be replaced with descriptive labels (EV-consistent, EV-exceeding, EV-shortfall).

      (2.3) Full specification, validation, and reliability of the belief parameter k. We agree the current description of k is inadequate. The revision will give the generative choice model in full, and explicitly state the estimation procedure and parameter bounds.

      (2.4) Rebuilt and validated drift-diffusion modeling. The boundary will be freed and estimated per participant/session, so that variance is no longer forced into the drift term. We will report whether the group effect on baseline drift survives.

      (2.5) Report why repeated sessions are treated as independent samples. We will conduct additional behavioral analyses to show repeated sessions from the same individual are sufficiently independent to be analyzed as unique samples. In particular, we will examine whether within-participant similarity across sessions is greater than cross-participant similarity. These analyses will provide direct evidence for whether sessions can reasonably be treated as separate samples rather than requiring sessions to be nested within participants.

      (2.6) Sharper framing and appropriately scaled claims. We will remove the framing that positions the field as treating ambiguity aversion as uniform and instead situate the work within literature that already documents heterogeneity in ambiguity attitudes. The distinction between ambiguity aversion and the internal representation of ambiguity will be made explicit, with the latter identified as our actual question. Claims about “internal belief models” will be restated in terms of what is estimated. Physiological hypotheses will be stated as specific, directional, falsifiable predictions rather than by appeal to the broad range of processes EEG has been linked to.

      (2.7) Quality control for pupillometry, EEG, and the VR context. We will report blink rates and counts, proportion of interpolated samples, trials and channels rejected, and ICA components removed. We will also report validate the luminance GLM.

      (2.8) The collaborative task. The Apollo Distributed Control Task preceded the Lottery Choice Task by design, so it must be reported as part of the protocol regardless of the leadership analysis. However, we agree that the leadership and team-performance analyses are not sufficiently motivated to carry the interpretive weight currently given them. They will be moved to a clearly labeled exploratory section, removed from the Abstract and framing, and presented without causal or trait-level interpretation.

      (3) Timeline and next steps

      The revisions above require refitting the behavioral, physiological, and computational analyses rather than adding to them, including hierarchical model fitting, parameter recovery, and new quality-control analyses. We therefore anticipate submitting the revised manuscript within approximately three months, and would welcome guidance if the editors would prefer a different schedule. We are content for the first version of the Reviewed Preprint to be published with this provisional response attached.

      We are grateful to both reviewers for the time invested in this manuscript. Several of the problems they identify are ones we should have caught ourselves, and the paper will be considerably stronger for their having been raised now rather than after publication.

    1. eLife Assessment

      This valuable study advances our understanding of the physiological conditions that promote the formation of mitochondrial-derived compartments (MDCs). The genetically amenable yeast system allowed the authors to generate convincing data showing the rapid formation of MDCs under physiologically relevant conditions where the mitochondrial proteome needs to be expanded during metabolic remodeling. The identification of a key role for a yeast AMPK-related protein called Snf1 will be of interest for cell biologists and biochemists.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigated the formation of mitochondrial-derived compartments (MDCs) under metabolic adaptations. They hypothesized that MDCs may play a role in regulating the mitochondrial proteome under these conditions by removing excess and superfluous membrane proteins that may challenge mitochondrial proteostasis. They found that glucose restriction, carbon-source switching, and osmotic stress can stimulate MDC formation. Underlying these stressors is a common signaling pathway that involves Snf1-dependent derepression of mitochondrial biogenesis and rapid synthesis and trafficking of nuclear-encoded proteins into mitochondria. They then showed that MDC formation is attenuated in tom70/tom71 mutants, suggesting that the delivery of these proteins to mitochondria is critical. Data also suggested that HAP4-stimulated mitochondrial biogenesis promotes MDC formation, which is further enhanced by glucose restriction and is suppressed after prolonged adaptation.

      Strengths:

      The genetically amenable yeast system allowed the authors to generate convincing data showing the rapid formation of MDCs under physiologically relevant conditions where the mitochondrial proteome needs to be expanded to accommodate increasing metabolic function. MDCs therefore function to buffer spillovers of outer membrane proteins upon an abrupt protein influx. Overall, the data presented are of high quality. The conclusion is strongly supported by the data.

      I think this is a significant study as (1) it supported MDCs as a physiologically relevant mechanism of mitochondrial proteostasis; and (2) it offers a common mechanistic framework explaining the MDC phenomenon under many other conditions such as TOR inhibition and hydrophobic protein overloading previously published by this group. Although the precise mechanism of MDC formation and how MDC formation contributes to the overall proteostasis of mitochondria remain unknown, as the authors stated in the manuscript, the current work is a clearly identifiable milestone in this specific area of investigation.

      Weaknesses:

      Although the data are overall strong, weaknesses are mainly related to potential misinterpretation of the data.

      (1) I have reservations regarding the interpretation of some results. First, the authors concluded that MDC biogenesis is activated when glycolytic metabolism is altered. I disagree with this. The authors should distinguish between "loss of glycolysis" and "loss of glucose repression". The yeast S288C strains are GAL2 and can ferment galactose. Likewise, glycolysis is also supported by raffinose and sucrose. In a broad sense, these carbon sources do support glycolysis as long as sugar influx is maintained at a high level. However, these alternative carbons do not repress mitochondrial respiration like glucose. It is likely the derepression of mitochondrial respiration (which is stated in some sections of the manuscript) instead of loss of glycolytic metabolism that stimulates MDC formation. This needs to be made clear throughout the manuscript. As such, the statement that "Carbon-source switching" stimulates MDCs is not accurate and needs to be re-interpreted.

      (2) The explanation for the requirement of low glucose levels could be misleading. A complete lack of carbon sources and high concentrations of 2-DG may shut down global protein synthesis, cell cycle progression, and many other processes, including mitochondrial biogenesis. Glucose at 0.02% is not sufficient to cause glucose repression, as only the high-affinity but low-influx transporters are functioning. Under the low glucose conditions, mitochondrial respiration is also derepressed. In this scenario, low glucose simply plays a role in supporting cell growth without causing the repression of mitochondrial biogenesis.

      (3) The idea that MDCs are formed when protein load exceeds the capacity the organelle can accommodate is attractive. Early studies have shown that the mitochondrial compartment is expanded by several folds in volume when yeast cells are switched from fermentative to oxidative metabolism. Perhaps, space expansion takes longer than protein influx increase. It would be interesting to see whether there is a correlation between MDC frequencies and the delay in volume expansion. Long-term adaptation would solve this challenge, as it allows the cell to complete volume expansion.

      (4) HAP4 may primarily activate OXPHOS genes but not some MDC cargo proteins. The requirement for "metabolic remodeling" for full induction of MDC formation may be an overstatement. The authors should either reexamine the proteomic data to see whether known MDC cargos are not subject to HAP4 activation or have this discussed in the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      Price et al. present new work providing insight into the function and mechanisms of mitochondrial-derived compartments (MDCs) in yeast. The Hughes lab previously established that these large ~micron-sized structures are formed under a variety of conditions including amino acid stress (rapamycin, conA, cycloheximide), alterations in mitochondrial metabolites and lipids, or acute expression of specific outer membrane proteins. These stressors lead to the sequestration of outer membrane proteins (that can include mistargeted inner membrane proteins) that extend or tubulate into large multilamellar structures that ultimately target the vacuole in an ATG5/Dnm1 dependent autophagy related pathway for degradation. Initially reported in aging yeast a decade ago, it has now been accepted as a mechanism to remove excess mitochondrial proteins as a pathway distinct from mitophagy or the extraction of stalled precursors from the import translocon.

      In this study, the authors examined additional metabolic transitions they suspected would drive increased mitochondrial protein expression and promote MDC formation. Indeed, they show that glucose-restricted conditions (or a switch to galactose or incubation with 2DG) induced MDCs within 2 hours. This correlated with increased transcription/translation of mitochondrial precursors porin and OM45. Similar results were seen with osmotic shock, a process previously shown to induce mitochondrial gene expression. The metabolic or osmotic shift was shown to activate a yeast AMPK-type kinase called Snf1, which phosphorylates a key substrate Mig1 - an established repressor of mitochondrial gene expression. Loss of these pathways abolished the generation of MDCs under these conditions. As the key novel finding in the study, the authors explored the relationship/requirement for Snf1 and Mig1 using multiple approaches in different backgrounds and employing auxin-inducible degron tools for acute depletion. These data further support the hypothesis that excess mitochondrial outer membrane proteins result in MDC formation to facilitate their removal, at least transiently until the import machinery can adapt to the increased import demand. To test this more directly, they generated an inducible yeast strain to express a canonical transcription factor Hap4 that induces mitochondrial gene expression. In this system, induction of Hap4 expression also resulted in MDC formation. While not all previously reported MDC inducers act through Snf1/Mig1, the common feature is the transcriptional induction of mitochondrial protein expression.

      Strengths:

      The important aspect of this work is that the authors dissected the transcriptional signaling pathway that induces MDCs in a much more physiological metabolic transition, which complements the more common use of chemical compounds. They had previously shown that overexpression of individual outer membrane proteins could lead to MDCs, but here the Hap4 expression offers a new condition to show that the canonical induction of mitochondrial biogenesis leads to MDC shedding. Overall, the data are of high quality, the findings are clear, and the work provides important new insights into the regulation of MDC formation.

      Weaknesses:

      There are a few points that should be addressed.

      (1) MDCs are almost exclusively monitored through GFP-tagged TOM70, and the authors do not show the inclusion of any endogenous cargo. The evidence for their fate in the vacuole is through the appearance of cleaved, free GFP after 6 hours that is dependent on ATG5, Dnm1, Pep4, etc. Can the authors demonstrate the appearance of MDCs without expressing any GFP tags and instead monitor known outer membrane cargoes? In the case of Hap4 expression, the proteomics identifies some very highly induced mitochondrial proteins, and there surely must be some with antibodies that can detect the protein by IF and Western blot.

      (2) There is a very unexpected ~10X increased in a sporulation factor SPO21 upon induction of Hap4. I see no evidence of sporulation, and it's not long enough for stationary phase. Is the increased mitochondrial biogenesis driving a specific metabolic state of these cells that is signaling to other biology?

      (3) It is important to understand the kinetics and stoichiometry of outer membrane loading that drives MDCs, and their transit to the vacuole. This is why it would be highly informative to monitor some endogenous cargoes (previous point). In the review the authors cite (NRMBC, Pfanner lab 2019), it was stated that the import machinery is not generally increased upon metabolic induction of mitochondrial gene expression. Therefore, (pre-MDCs) the field concluded that the import machinery has a very high capacity for the rapid biogenesis of newly synthesized proteins, along with regulation through the phosphorylation of import receptors (ie; the work of Meisenger). Consistent with this, the Hap1 proteomics did not show any increases in the core import machinery, while ETC subunits and a large swath of mitochondrial proteins were elevated over 2-fold (I looked carefully through the Excel sheet). Since MDCs are induced transiently about 2 hours after glucose deprivation, and fully dependent on de-repression of Mig1, the authors are right to imply that this is coupled to the import of newly synthesized proteins.

      However, it seems to me that MDCs are being formed at very early stages of mitochondrial protein expression, not after they have necessarily "overloaded" the outer membrane. The Hap4 proteomics after 3.5hr of induction would suggest that the bulk of the mitochondrial proteins have been successfully inserted (no import failure) and are likely already functional (metabolizing). I'm trying to understand the percentage of the proteins that would be incorporated within MDCs, as the mitochondria appear to handle the bulk of their newly inserted proteins without issue. How can the authors adapt their "free GFP" assay to understand the stoichiometry of the transport of endogenous, newly imported outer membrane proteins to the vacuole?

      (4) As a last theoretical point for discussion: Can the authors exclude that MDCs are not functional or play a signaling role? Given the emerging work on SPOTs (Lena Pernas), and from the new evidence from Craig Thompson's lab that there can be very specific functional mitochondria (oxidizing vs reducing), it is possible that MDCs are not simply there to be degraded. They last at least 3 hours, which is a long time for yeast (budding cycle 90 min, 3 hours in glucose deprivation). Taking the data presented here very objectively, there is no direct evidence that the cargoes within MDVs reflect any failure to import, or that they are damaged in any way. The deletions of Tom70/71 have way too many pleotropic effects and essentially demonstrate only that the MDC cargoes came from the mitochondria. It could be helpful if the discussion also positioned these MDC mechanisms within the context of other aspects of selective mitochondrial-related compartments that have been emerging in the literature.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Price et al. report the physiological conditions and proteins involved in the formation of mitochondria-derived compartments (MDCs), specialized domains exclusively containing outer mitochondrial membrane (OMM) proteins, in budding yeast. Hughes and his colleagues have previously established MDCs as unique multilamellar membrane structures derived from the OMM that arise with both mitochondrial metabolic perturbation and hydrophobic OMM protein load. Whether cells undergo MDC formation in response to physiological changes in mitochondrial biogenesis remains to be explored. In this study, the authors sought to test if glucose restriction, carbon-source switching, and salt stress can induce MDC formation, and found that these situations, which naturally promote acute mitochondrial biogenesis concomitantly with metabolic transitions, trigger MDC induction. Under these conditions, loss of Snf1, an AMP-activated protein kinase that facilitates mitochondrial biogenesis under metabolic stress, almost completely abolished MDC formation. Snf1 induces MDC induction under metabolic stress via phosphorylating (suppressing) Mig1, a transcriptional repressor of mitochondrial biogenesis. Consistent with this idea, loss of Mig1 mostly rescues MDC formation under glucose restriction or salt stress in cells lacking Snf1. The authors further found that acute induction of Hap4, a core activator of mitochondrial biogenesis, is sufficient to trigger MDC formation even without metabolic stress. Finally, cells lacking Tom70 and Tom71, protein receptors of the TOM (translocase of the outer membrane) complex that mediate targeting of hydrophobic mitochondrial proteins, almost failed to form MDCs under glucose restriction. Correctively, the authors propose that MDCs act in the reduction of OMM protein load upon metabolic stress-induced acute mitochondrial biogenesis.

      Strengths:

      The experiments for this study are well-designed, and the resulting data are mostly convincing, with proper controls and significant statistics to support their conclusions. The paper potentially provides new insights into the physiology of MDC formation.

      Weaknesses:

      There are only a few new mechanistic advancements in this paper.

    1. eLife Assessment

      This important study uses dual-color super-resolution microscopy, quantitative clustering analyses and simulations, as well as biochemical approaches to investigate the nanoscale organization of mucins and trans-sialidases on the surface of Trypanosoma cruzi, the causative agent of Chagas disease. The evidence supporting a non-uniform distribution of these molecules into segregated nanoclusters and a more ordered non-clustered population is solid. Overall, this work provides a generalizable analytical framework that can help understand how parasite surface proteins are spatially organized. The main remaining limitations concern the mechanistic basis of the proposed organization and the need to exclude possible effects of sample preparation before fixation.

    2. Reviewer #1 (Public review):

      Summary:

      Escalante et al. employ super-resolution microscopy to achieve a clearer, nanoscale view of how trans-sialidases and mucins are organized on the Trypanosoma cruzi parasite membrane. Comparing the experimental data using clustering analysis with model-based simulations, they report two kinds of organizational states describing the non-uniform distribution of these two proteins: a segregated state where mucins and trans-sialidases form spatially distinct nano-clusters, and a non-clustered state where they share a proposed fibrillar network with more ordered, shorter-than-random separation distances. They also look at the oligomerization states of the two proteins to try and propose a mechanistic basis for the observed distributions.

      Strengths:

      The in-depth analysis of the distributions of both proteins coupled with model-based simulations brings out new insights into organizational principles underlying protein distribution on the membrane surface. The ability to resolve shorter-than-random separation distances even in the non-clustered state is to be highlighted and is a key take-away from this manuscript.

      Weaknesses:

      The authors propose the oligomeric state of mucins compared to the non-oligomeric trans-sialidases as a basis for explaining the distinct organization of these proteins. Although this hints at how segregation may occur, it does not inform us of how the more ordered non-clustered state could co-exist with the clustered segregated state and warrants further investigation.

      Overall, the analytical framework applied in this study to elucidate organizational principles for the non-uniform distribution of proteins can potentially be used in a wide-range of contexts across different organisms and systems. This study also lays the groundwork to understand mechanisms that spatially regulate how trans-sialidases act on their substrates. Going forward, it could be very interesting to look at how different kinds of mucins and trans-sialidases are organized with respect to one-another and amongst themselves. Also, the development of tools to observe the dynamics of these proteins live will likely provide further insights into the mechanism.

    3. Reviewer #2 (Public review):

      The manuscript describes a numerical analysis of the domains of the T. cruzi cell surface containing different proteins. It has the potential to be of great interest.

      I do not have the expertise necessary to comment on the image collection or analysis.

      I have one concern: the amount of manipulation of the cells prior to fixation; these were clearly stated in the methods, which is good.

      My concern is whether these manipulations prior to fixation alter the observations. The 'Labelling sialic acid acceptors' involves >6 centrifugations and >90 minutes incubation in PBS prior to fixation, and the 'immunostaining' protocol involves cells 'extensively washed with PBS' prior to fixation. I would like to suggest that the authors do controls in which they compare the pattern of anti-SAPA staining under four conditions.

      (1) Cells fixed in culture by the addition of paraformaldehyde to 4%, followed by blocking and PBS washes.

      (2) Cells fixed in culture by the addition of paraformaldehyde to 4% and glutaraldehyde to 0.2% followed by blocking and PBS washes.

      (3) Cells fixed by the 'labelling sialic acid acceptors' protocol.

      (4) Cells fixed by the 'immunostaining protocol'.

    4. Reviewer #3 (Public review):

      Summary:

      The authors present an innovative approach to tackle the lateral organization of mucins and trans-sialidases (TS) on the cell membrane of the organism Trypanosoma cruzi. By applying dual-color super-resolution microscopy (STORM), the authors report on a differential nanoscale distribution between mucins and TS on the cell membrane. Moreover, they find that 60% of mucins and TS are organized in nanoclusters with an inter-nanocluster distance following a random distribution. The remaining 40% of both proteins are organized in a non-random manner, and, using simulations, the authors claim that they are organized in rectilinear fibers.

      Strengths:

      The authors use dual-color STORM microscopy to unravel the protein nanoscale organization of mucins and TS on the cell membrane of Trypanosoma cruzi for the first time. They perform a dedicated analysis of the localizations and clustering of both proteins. Moreover, they perform, for every type of analysis on real data, simulations to compare their results for random organization. They also use an analysis approach together with simulations to propose that the lateral organization of both non-clustered proteins are within rectilinear fibers. They also complement their microscopy findings with BN-PAGE. Overall, the use of advanced microscopy techniques, corresponding data analysis and simulations is very solid and remarkable.

      Weaknesses:

      As the authors point out, they do not provide a molecular/biophysical mechanism explaining the non-random lateral organization of mucins and TS (both clustered and individual proteins).

    1. eLife Assessment

      A regulated cell death pathway that intentionally causes ferroptosis has not been previously described. The authors describe a useful finding that the ferroptosis inducers ML162 and erastin induce caspase-5-mediated cleavage of GSDME in mesenchymal-like ovarian cancer cells, resulting in pore formation. The data supporting the main claims are incomplete. This work will be of interest to researchers who study regulated cell death.

    2. Reviewer #1 (Public review):

      Summary:

      This report seeks to understand the mechanisms whereby the ferroptosis inducers ML162 and erastin cause cell death in several tumor cell lines. They present evidence that caspase-5 is activated and required for ferroptosis, but other caspases, including caspase-1 and -4, are not important. Surprisingly, caspase-5 cleaved and activated GSDME, instead of the expected gasdermin target GSDMD

      Strengths:

      The magnitude of effect for triggering ferroptosis by ML162 and erastin is strong, and the strength of inhibition by YVAD is also very strong, making these effects convincing. The lack of effect of DEVD, which inhibits apoptotic caspases, is also convincing. Also, the lack of effect of necrostatin is convincing. These negative results strengthen the positive results seen with YVAD.

      Inhibition by disulfiram is convincing.

      Caspase-5 knockout single-cell clones and the ability to complement these with caspase-5, but not catalytically inactive caspase-5 in Figure 5, is strong data.

      Weaknesses:

      (1) Prior publications have asserted that ferroptosis is caspase-independent. Can the authors repeat some of these experiments directly to reveal whether there was an error in the published work that resulted in missing the phenotype for a caspase in ferroptosis? In my experience, caspase inhibitors sometimes only delay cell death because they are not 100% effective, especially over hours of time. Can the authors repeat the prior experiments to reveal whether this caveat affected previously published data? At the least, the authors should use the z-VAD-fmk and Boc-D-FMK inhibitors to determine whether they give the same effects as YVAD to rule out a very unlikely possibility that these "pan-caspase" inhibitors do not inhibit caspase-5.

      a) The original report describing ferroptosis by Dixon and Stockwell, doi: 10.1016/j.cell.2012.03.042, shows that erastin treatment-induced ferroptosis is not affected by z-VAD-fmk in 3 cell lines.<br /> b) A later report from Dr. Stockwell states in data not shown that a different pan-caspase inhibitor (Boc-D-Fmk) does not block erastin-driven ferroptosis. Doi 10.1016/S1535-6108(03)00050-3<br /> c) An earlier 2008 report from Dr. Stockwell shows that z-VAD-fmk and Boc-D-fmk do not rescue cells treated with RSL-3 or RSL-5 treated cell lines derived from BJ cells. doi 10.1016/j.chembiol.2008.02.010<br /> d) A recent paper shows a delay of ferroptosis after RSL3 treatment by pan-caspase inhibitor Q-VD-OPh. Doi 10.1038/s41418-025-01514-7. The delay was about 8 hours in time, so cells were still dying.<br /> e) Gpx4 knockout cells or erastin or RSL3 treatment are unaffected by z-VAD-FMK. Doi 10.1038/ncb3064<br /> f) I encourage the authors to do more thorough searching of the literature to find more publications that have used caspase inhibitors.

      (2) The authors should discuss how mouse cells can undergo ferroptosis while they do not encode caspase-5, and the evolutionary conservation of caspase-5 in general. If caspase-5 is not encoded by an animal (as is the case with mice), can their cells undergo ferroptosis?

      (3) Disulfiram is not a specific inhibitor. It is a nonspecific inhibitor that modifies cysteine residues of many proteins. This should be described in more detail so the reader can appreciate the strengths and weaknesses of the inhibitor.

      (4) I encourage the authors to assess IL-1β processing by Western blot and show that this is inhibited by YVAD. Because ELISA can detect release of the pro form after lytic cell death by other mechanisms.

      (5) ASC knockdown in Supplementary Figure 3a for two cell lines is not sufficient to draw any conclusions in Figure 3a.

      (6) Caspase-5 can be more specifically inhibited by LEVD inhibitors. Can the authors show that these work as well?

      (7) I would like to see a positive control in Figure 5a to show what a strong caspase signal activity looks like.

      (8) Since caspase-3 is known to cleave GSDME, the authors need to assess whether caspase-3 is also activated, and whether other caspase-3 target proteins are also cleaved. There are many to choose from. Caspase-3 western blots, including with the cleaved caspase-3-specific antibody, are critical. This is in addition to the blot shown in Supplementary Figure 10. Positive controls should be included. It is important to continue to add controls to rule out caspase-3, with more than just negative data with DEVD inhibitors and the western blot in Figure S10.

      (9) The data in Figure 6c are not strong.

      (10) One would expect that any mode of activation of caspase-5 should lead to its proteolytic activity upon its preferred substrates, so LPS should cause caspase-5 to cleave GSDME and not GSDMD. Additional data to strongly activate caspase-5 with LPS should be investigated to see if this leads to GSDME cleavage and pyroptosis via GSDME and not GSDMD.

    3. Reviewer #2 (Public review):

      Summary:

      In the submitted manuscript, Akter et al use a series of ferroptosis inhibitors in mesenchymal-like ovarian cancer cells and discover that the ferroptosis inducers induce cell death that is inhibited by pyroptosis inhibitors, namely YVAD-fmk and disulfiram, which inhibit pore formation by gasdermin D (GSDMD). Remarkably, the authors also saw the release of IL-1β in response to ferroptosis inducers. Unexpectedly, they did not observe the involvement of caspase-1 but rather observed that caspase-5 was activated in response to the ferroptosis inducers. Moreover, they found that caspase-5 directly cleaves GSDME in response to the ferroptosis inducers, establishing CASP5/GSDME as downstream executors of ferroptosis.

      Strengths:

      These findings are interesting because only CASP1 is known to induce IL-1β maturation, and their data suggest that CASP5 rather than CASP1, is responsible for IL-1β activation in the context of ferroptosis inducers. Notably, CASP3 is the only caspase reported to be able to cleave GSDME, so the identification of CASP5 as a driver of ferroptosis in this context is a significant finding. They genetically show that loss of CASP5 and GSDME knockdown inhibits cell death in response to the ferroptosis inducers ML162 and Erastin, which is evidence that they play a role in this context.

      Weaknesses:

      The major findings in this paper are interesting, but the data presented do not robustly support the claims made in this paper. For example, they claim that CASP5 is responsible for the activation of GSDME by cleaving it directly to induce cell death. They try to rule out the involvement of CASP1, ASC, and CASP4 using siRNA targeting these genes, but the knockdowns are incomplete, and the loading controls are inconsistent. They also claim they do not see GSDMD or CASP3 cleavage and activation but use negative data to make that claim. It is unclear if the antibodies used can detect cleaved GSDMD or CASP3 as they do not include a positive control to show that they can indeed detect these activation events if they were occurring. This needs to happen in the same experiment - they need to show in the same experiment with the same lysates that they can detect CASP5, GSDME and IL-1β activation but not CASP1, GSDMD, CASP4, or CASP3 activation. Of course, they should include agonists for positive controls of CASP1, CASP4 and GSDMD activation, which are lacking in the current manuscript.

      Notably, the major evidence supporting a direct role for CASP5 cleavage of GSDME is one Coomassie gel using recombinant CASP5 and GSDME, but there were too many non-specific bands, and the full-length uncleaved protein could not be detected even in the untreated lanes. The authors need to show a gel where the protein can easily be identified and should also include a positive control protein like GSDMD to show the relative cleavage efficiency of GSDME compared to a known substrate. It would also be great to compare this to CASP3-mediated cleavage of GSDME. With recombinant proteins, calculating the catalytic efficiencies would be the best way to ascertain if this is biologically similar to other known substrates.

      The way that ferroptosis is defined, it is caspase-independent, and pyroptosis is defined as gasdermin-mediated cell death. Given that these agents lead to activation of CASP5/GSDME, it would be more accurate to say that these ferroptosis inducers also induce CASP5/GSDME-dependent pyroptosis, as opposed to them being the executors of ferroptosis. This can be a distinct mechanism/pathway from the ferroptosis pathway, as multiple cell death pathways can be initiated in cells. Consistent with this, ferrostatin-1 also inhibited cell death, likely due to inhibition of the ferroptosis signaling cascade. It is unclear if this pathway is upstream of the caspases. How these ferroptosis triggers selectively activate CASP5 and not CASP4 to induce GSDME cleavage is a major unresolved question. Notably, it is also unclear if this biology is specific to the mesenchymal-like cells used in this study or if it expands to other cells.

    4. Reviewer #3 (Public review):

      Summary:

      Akter et al. identify caspase 5 activation and Gasdermin E cleavage as a novel downstream executioner of ferroptotic cell lysis induced by erastin and ML162. These data are novel and very interesting to the wider cell death community.

      Strengths:

      Strengths of the study include the use and validation of findings in several mesenchymal ovarian cancer cell lines, the rigorous validation using small molecule approaches, siRNA-mediated silencing and CRISPR/Cas9-mediated knockouts with re-expression.

      Weaknesses:

      A weakness of the study is the fact that ferroptosis was not induced genetically (GPX4 ko) and, hence, off-targets of the mode of induction cannot be ruled out at this point (e.g. ML162 also targets TrxR1). Moreover, it would be vital to understand at which point in ferroptosis execution caspase 5 is activated in a time-resolved kinetic together with lipid ROS tracing to also obtain hints as to its possible activation.

      Conclusion:

      Despite the weaknesses described, this is a very interesting, timely, and well-executed study with the described limitations. The work provides important mechanistic insights into the interplay between ferroptosis and pyroptosis with possible consequences for inflammatory responses.

    1. eLife Assessment

      This study presents an important empirical analysis of the acoustic space of vocal repertoires across primates and humans, providing new data that challenge the widely held hypothesis that the expansion of the human vocal space, driven by modifications of the vocal tract, was a key prerequisite for the evolution of spoken language. The evidence is convincing in its technical implementation and acoustic measurements, offering helpful comparative insights into voiced vocalizations. However, the theoretical framing and interpretation would benefit from further discussion.

    2. Reviewer #1 (Public review):

      Summary:

      The authors conducted a comparative acoustic analysis of primate vocal repertoires, focusing on the assumption that speech and language evolution required and involved an expansion in the acoustic space of voiced vocalizations from non-human primates to humans. Results challenge this idea. The study compiles and analyzes a large dataset of calls to quantify differences in vocal production space.

      Strengths:

      The study is technically sound, with a solid implementation of acoustic measurements and a valuable new dataset that brings empirical rigor to test a dominant, yet hitherto strictly theoretical, notion about what speech and language evolution entailed. It provides concrete comparative acoustic data across species to disprove that speech and language required an increase in the range of voiced calls, and thus, by extension, of vowels. The approach is methodologically rigorous and directly engages with the relevant data, rather than relying on untested presumptions of what great apes "ought" to be able to do or not.

      Weaknesses:

      The theoretical contextualization should be strengthened and updated, as several aspects contain inaccuracies, most notably by equating voiced calls or vocalizations with speech (overlooking the critical role of consonants, as human languages typically show vowel:consonant ratios of 1:4 or greater) and misrepresenting the premises and current status of the neural (Kuypers-Jürgens) hypothesis.

      The discussion drifts into speculative territory on features like syntax and co-articulation that fall outside the paper's scope and data, and it does not sufficiently engage recent evidence on vocal learning and consonant-like capacities in great apes.

      Minor issues include incomplete sampling justifications, imprecise terminology, and reliance on references that have been critiqued in more recent work.

    3. Reviewer #2 (Public review):

      This study examines the evolutionary context of the emergence of human speech. The authors address the widely held hypothesis that the expansion of the human vocal space, resulting from modifications of the vocal tract, was a key prerequisite for the evolution of spoken language.

      To test this hypothesis, the authors quantified the acoustic space of human speech, non-linguistic vocalizations, and musical vocalizations and compared it with that of nonhuman primates, chimpanzees, bonobos, and chacma baboons.

      The authors found that speech and song occupied significantly less volume in the acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volume of speech and song was not statistically distinct from that of non-human primates. Accordingly, the authors conclude that the evolution of human speech did not depend on an expansion of the human vocal acoustic space.

      I find the analysis presented in this manuscript highly convincing. It is conducted at a contemporary scientific standard, and the results provide strong support for the authors' conclusions. I particularly appreciate that the authors explicitly discuss the limitations of their approach. For example, they acknowledge that MFCCs cannot capture all aspects of acoustic structure.

      I have only three minor comments:

      First, the authors may wish to briefly summarize the main findings of the study by Anikin et al., as it represents the central reference for the present work. A concise summary in two or three sentences would help readers who are not familiar with that study.

      Second, I would appreciate a brief explanation of why the authors chose this particular statistical approach.

      Third, the authors could briefly mention that the Chacma baboon dataset provides a very comprehensive representation of the vocal repertoire of this species, although a small number of rare vocalizations are not included. I am not sure whether a similar limitation also applies to the chimpanzee and bonobo datasets, but if so, it would be useful to mention this as well.

    1. eLife Assessment

      The ability to measure autophagic flux in vivo in response to physiological stresses remains a challenge for investigators in the field; the novel mouse model present in this work is significant and should prove valuable to investigators in addressing shortcomings of existing models. In addition, the development of the microplate reader approach to permit semi-high-throughput analysis of samples is convincing and a significant advance. The application of these approaches to measuring autophagy in multiple tissues is appreciated but raises some questions that need to be answered and highlighted, including what new biological insight has been generated for tissues under study, how overall autophagy versus rates of flux are determined, and how the sex of the animal affects outcomes. The neuronal populations under study should be reassessed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors develop a GFP-LC3-RFP autophagy reporter under the control of the Rosa26 locus to measure autophagic flux in mouse embryos as well as adult tissues. While image quantification is consistently used, the authors also develop a semi-high-throughput assay for measuring autophagic flux using a microplate reader. Additionally, the authors cross these mice with a Cre-inducible Atg5-deletion mouse model, allowing the investigation of how autophagy flux is affected upon loss of Atg5. With this model, they demonstrate that loss of Atg5 leads to an increased ratio of GFP/RFP intensity in multiple tissues, including the brain, revealing that the brain undergoes basal autophagy. They further go on to show that the increase in GFP/RFP intensity upon Atg5 loss is greater in adult tissues compared to their embryonic counterparts. The development of an animal model, along with quantitative tools to measure the model, will have a high impact on the field. However, the analyses from the data presented do not fully justify the conclusions.

      Strengths:

      (1) A mouse model to better measure autophagy.

      (2) The plate-reader-based method to quantify autophagy across tissues.

      (3) Assessment of autophagy in many different tissues.

      (4) Crossing the reporter mouse with the Atg5f/f mouse to assess basal autophagy.

      Weaknesses:

      (1) While the tool is of high impact, there is little new biological or mechanistic insight provided in these studies.

      (2) The quantification and normalization method is unclear, making it difficult to compare across tissues accurately.

      (3) Differential expression across cell types is not well documented or taken into account for comparisons.

      (4) There is no consideration for sex as a biological variable.

    3. Reviewer #2 (Public review):

      Summary:

      The aim of the authors was to measure starvation-induced and basal autophagy in vivo across several tissues and developmental stages. For this, they developed a novel mouse model expressing the GFP-LC3-RFP reporter. They also aimed to provide a more high-throughput method for autophagy flux measurements than assessment by imaging and developed an assay based on a microplate reader.

      Strengths:

      (1) Good validation of the mouse model. The knock-in strategy is well explained and illustrated.

      (2) The model has potential to be applied to a wide range of research questions. The Cre-dependent expression allows for customization of KO timing, which will be beneficial in developmental studies.

      (3) The authors presented consistent findings using two different methods to quantify autophagy, strengthening the robustness of their results.

      (4) The authors demonstrated the validity of the high-throughput method (microplate reader).

      Weaknesses:

      (1) The comparison of neuronal populations in different areas of the brain is not ideal. In the cerebellum, Purkinje cells were chosen, which are rare and not representative of this tissue, as well as functionally very different from the neurons in the hippocampus and cortex that they were compared to.

      (2) The explanation of the GFP-LC3-RFP construct and specifically if/how autophagosome formation can be measured and distinguished from flux could be clearer.

      Conclusion:

      The work presented is thorough, and the authors achieved their goals for this study. The effort used to further investigate unexpectedly high basal levels of autophagy in the brain is well appreciated and adds value to this paper. The conclusions of the authors are mostly very well supported by the data provided. The well-structured description of the results, along with clear figures, allows the reader to comprehend the authors' reasoning in reaching their conclusions.

      The presented mouse model has great potential for a lasting positive impact on the research field of in vivo study of autophagy. The method of utilizing a microplate reader will also benefit future research where semi-high throughput is an advantage. Together, the information provided in this study not only presents new methodology that will allow the investigation of new research questions, but also provides novel information about in vivo autophagy flux at the selected developmental stages that opens up new follow-up research questions.

    1. eLife Assessment

      This manuscript presents valuable experimental results describing the localisation and regulation of casein kinase CK1δ during the cell cycle. During the G2 phase of the cell cycle, the autophosphorylation of the C-terminal tail stabilises and inhibits the kinase which might protect CK1δ in the subsequent G1 phase. The results may be of interest to researchers working on casein kinase function and regulation but due to lack of some controls remain incomplete.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      We appreciate the authors have provided answers to many of the points we raised, and the changes made to their manuscript, which we think strengthen the overall evidence presented. However, we find that some important controls are still missing across experiments.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Antibody cross-reactivity was only tested against CK1ɛ, but should also be tested against CK1α, which is abundant in U2OS cells, and is also known to be involved in cell-cycle regulation.

      b. Fig. 1: Statistical analyses are missing from the analysis.

      c. Fig. 2: No colocalisation analysis shown for figure 2, only some arrowheads pointing to puncta. Appropriate colocalisation statistics are important since for practical reasons, only a few representative images can be shown on the figure.

      d. Fig. 6: Even if the figure is illustrative, it is important to show centrosome staining to visualise CK1ẟ's recruitment to the centrosome in G2/prophase, especially since this information is used to propose the model in figure 7.

      e. For all figures: Please mention the number of independent biological replicates in the figure legends (1, 2, 6). For figure 1, if there are 3 independent biological replicates, the quantification should take all of them into account (as opposed to the data points corresponding to 10 cells), and statistics must be done appropriately, taking those independent replicates into account. Same for the colocalisation analysis in figure 2 once you include it.

      (2) Shortcomings in biochemistry experiments:

      On CalA control, this is not a matter of confirming that CalA treatment works in principle, but rather to confirm that CalA treatment worked in this specific replicate. Aliquots may lose potency (e.g. with freeze-thaw cycles / exposure to light), hence checking for enrichment of phospho-proteins is essential to confirm the treatment was successful in this particular instance. In the worst-case scenario, the company may have sent the wrong compound altogether! A positive and a negative control is the basis for every experiment to make meaningful interpretation. On a separate note, many experiments have control and siRNA or compound treatments on two different gels - this should be rectified as they are meaningless if different exposures have been selected for different immunoblots.

      (4) As the authors mention, the kinase is not fully inactive when tail phosphorylated. Recent research has also suggested that tail-phosphorylated CK1ẟ may show increased catalytic activity for a few select, specific substrates, in the co-occurrence of pT220 (Cullati et al., 2022; Cullati et al., 2024). It is thus tricky to directly infer that phosphorylated CK1ẟ is inhibited, when no positive control for CK1ẟ inhibition was shown in the evidence presented. It would be necessary to either nuance your claim or include a positive control for CK1ẟ inhibition. Please revise statements in the manuscript accordingly.

      (5) It would be important to include statistical analyses for the immunofluorescence data in Fig. 1 and 2.

      (9) The authors mentioned "In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable." Unfortunately, the data from biochemical analyses presented in the eLife publication is uninterpretable due to a lack of loading controls.

      (10) While the data presented in Penas et al. strongly suggests a link between CK1ẟ stabilisation and the APC/C-Cdh1 complex, it is the only study to have shown it. Given that (1) science relies on data reproducibility and (2) your proposed model relies heavily on the relationship between CK1ẟ stabilisation and the APC/CCdh1 complex, it would be appropriate to include the investigations mentioned in our original comment.

    3. Reviewer #2 (Public review):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U20S cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 WT but not a phospho-null mutant strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin and Okadaic Acid) and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      The manuscript has improved overall, but some sections are still inconclusive and require clarification.

      Major comments:

      Figure 1 is inconclusive. CK1 nuclear staining is highly similar in untreated cells and in cells treated with CHX + PF670462. The reduction in centrosomal staining in these cells is barely significant. However, the authors draw very strong conclusions from these data sets. In panel B, the cells appear to have been fixed incorrectly, and the anti-PCNT shows a strong background signal. Not convinced that immunofluorescence is the best approach to look at protein dynamics in vivo.

      In Figure 2, panel B, the authors should co-stain the centrosomes of cells that co-express CRY1 and CK1, as some of these dots may represent the centrosomes.

      Figures 3B, please provide information on the non-phosphorylable CK1a mutant (it is mentioned as a variant in which all serine and threonine residues in the C-terminal tail were replaced by alanine). Specify the number of sites mutated and their exact position. Is this non-phosphorylable CK1a version catalytically active? Treatment of samples with inactivated PPase should be used as a control.

      Strengths:

      The authors reveal that the activity and abundance of dephosphorylated and phosphorylated CK1δ are regulated in a cell cycle-dependent manner. This suggests that these different pools are associated with distinct physiological functions.

      Weaknesses:

      Unfortunately, some of the data are inconclusive, and there is no data/information linking the cell cycle regulation of CK1δ to its function during the cell cycle.

    4. Author response:

      General Statements.

      We thank all reviewers for their careful evaluation and constructive comments.

      Fidel Serrano recently completed a related study on the role of CK1δ in the circadian clock which is published in eLife (https://doi.org/10.7554/eLife.110786.1). During these studies, we became interested in the physiological role of CK1δ autophosphorylation, whose functional significance has remained unclear.

      In the present manuscript, Fidel discovered that the autophosphorylated, auto-inhibited form of CK1δ accumulates specifically during mitosis. These findings provide a physiological context for CK1δ autophosphorylation that has remained elusive for many years. They constitute the central foundation of our model, which further integrates previous findings on APC/C function and activity with the data presented in our eLife study.

      Several suggestions and questions raised by the reviewers concern the regulation of CK1δ in cultured cells, which are predominantly in G1. Many of the corresponding experiments and analyses are already included in our eLife paper. We apologize that this overlap may complicate the review process. However, we are unable to publish identical datasets in both manuscripts.

      We therefore briefly summarize here the findings from the eLife study that are most relevant to the present work.

      - Overexpressed CK1δ is unstable, whereas CK1δ-K38R (catalytically inactive) and CK1δ-R178Q (altered specificity/activity) are comparatively stable (eLife Figs. 3A-E and 5B).

      - Assembly with PER2 stabilizes overexpressed CK1δ (eLife Figs. 4A-D).

      - Treatment with PF670462 similarly stabilizes the kinase (eLife Figs. 3F and 5D).

      - CK1δ binding sites in the centrosomal/Golgi area are present in excess even relative to overexpressed kinase (eLife Fig. 5C).

      The reviewers also raised questions about the PER2-CRY1 nuclear foci.

      Thermodynamically stable/persisting nuclear foci form upon coexpression of PER2 and CRY1. They were preliminarily characterized in the eLife study (eLife Fig. 2). For the purposes of the present work, they constitute a serendipitous and convenient tool for studying interactions of PER with CK1δ. To induce foci formation, we used a stable cell line expressing DOX-inducible mK2-CRY1. mK2-CRY1 accumulates only at relatively low levels because the protein without a binding partner is intrinsically unstable, as confirmed by Western blot analysis (eLife Fig. 4C). Coexpression of PER2 stabilizes mK2-CRY1 and promotes the formation of nuclear PER2-mK2-CRY1 foci.

      These foci contain elevated levels of endogenous CK1δ (eLife Fig. 2E and EV2), which accumulate gradually over a 24-hour period. This observation indicates that, at steady state, a fraction of endogenous CK1δ is continuously degraded in the absence of overexpressed PER2. Because endogenous CK1δ is synthesized at a relatively low rate, the unstable pool is small at any given time and therefore not readily detectable in conventional cycloheximide chase assays against the kinetically stabilized steady-state background of CK1δ.

      In our point-by-point response, we therefore refer the reviewers to the corresponding datasets and analyses presented in the eLife manuscript.

      Point-by-point description of the revisions.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      Involved in various cellular pathways and processes, Ser/Thr kinase CK1δ is thought to be constitutively active and its tail autophosphorylation is suggested as a putative inhibitory mechanism of CK1δ kinase activity. Here, the authors investigate CK1δ's phosphorylation status, in relation to its location and dynamics throughout the cell cycle. Immunofluorescence and biochemistry studies showed that the subcellular distribution of CK1δ is dynamic and in equilibrium between centrosomal and nuclear pools. The authors argue that CK1δ phosphorylation protects it from degradation. Finally, the authors show differences in CK1δ phosphorylation status and location throughout the cell cycle. Combining their findings with the existing literature, they propose a model of CK1δ phosphorylation status, location and dynamics throughout the cell cycle.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Lack of a positive control for centrosomal staining, eg. PLK1 (Fig.1A,C, 2B, 6A-B). This would also allow a colocalisation analyses between CK1δ and a centrosomal marker, which would build a more convincing body of evidence in favour of centrosomal recruitment of CK1δ in baseline conditions. A lack of a negative control for CK1δ IF? Most CK1 antibodies also cross-react with other isoforms.

      As suggested, we show centrosomal staining with a well-studied centrosomal marker pericentrin (PCNT; revised Fig. 1). Consistent with what has been well-described in the literature that CK1δ localizes to the centrosome (Sillibourne et al., 2002; Greer & Rubin, 2011; Greer et al., 2014) and with our observations, we confirm that CK1δ is indeed localized at the pericentral region.

      The CK1δ antibody does not crossreact, this is shown in Fig EV3A of the eLife paper.

      b. Lack of proper image quantification (Fig.1A,C, 2B, 6A-B). To establish a centrosomal recruitment of the CK1δ pools, colocalisation of CK1δ and a centrosomal marker must be quantified using a standard colocalisation quantification method and shown appropriately. Same comments regarding the colocalisation of CK1δ with PER2/mK2-CRY1-positive foci.

      Fig. 1: Colocalization analysis with pericentrin (PCNT) as centrosomal marker is provided.

      Fig. 2: The colocalization analysis of CK1δ (green) and mK-CRY1 (magenta) shows that in all PER2-untransfected cells (50 cells evaluated), CK1δ is concentrated at a single, mK2-CRY-negative spot, corresponding to the pericentrosomal region, which we have established in Figure 1. In contrast, in all PER2-transfected cells (15 cells evaluated) CK1δ localizes to mK2-CRY1 positive nuclear foci. The fields-of-view we have provided are representative of the typical phenotype of CK1δ localization in the presence or absence of PER2/mK2-CRY1-positive foci. Here it should be noted that CRY1 foci formation is strictly dependent of co-expression with PER2, (eLife paper).

      Fig. 6: This figure presents an analysis of fixed cells. Because only a small fraction of asynchronously growing cells is in mitosis at any given time, the number of cells assigned to individual mitotic stages is necessarily low (typically fewer than 10 cells per stage). The purpose of this figure is therefore primarily illustrative: to document and confirm the known subcellular localization of CK1δ during the different stages of mitosis rather than to provide a comprehensive quantitative analysis.

      c. Lack of proper statistical analyses (Fig.1, 2, 6). As immunofluorescence constrains one to only show few images per condition at best, statistical analyses on broader image analysis data are essential to measure the significance of the changes observed on the representative images provided.

      See answer to previous question.

      d. Lack of biological replicate numbers (Fig.1, 2, 6). Was it from 3 independent biological replicates?

      Fig.1 and 2: three independent biological replicates.

      Fig.6 is just illustrative to show the already known localization of CK1δ across mitosis. The localization shown in the four panels is seen in all cells at the respective cell cycle stages.

      e. NOTE: if available, using a confocal microscope would be best to provide optimal Z-axis resolution. This would provide you with more accurate colocalisation data, like CK1δ recruitment to the centrosome or PER2/mK2-CRY1-positive foci.

      Characterization of PER-CRY foci is published in the eLife paper in Fig. 2 and EV2.

      (2) Shortcomings in biochemistry experiments

      Lack of loading controls in Fig.3A-C and 4A-B. Until they are added, no conclusions can be safely interpreted from the experiments. Myc is a good turnover control for CHX across all experiments. Enrichment of phospho-proteins upon CalA treatment?

      Loading controls are provided in this revision. Enrichment of phospho-proteins following CalA treatment has been demonstrated and published in numerous previous studies (Cegielska et al., 1998; Ishihara et al., 1989; Rivers et al., 1998).

      (3) The authors infer from the literature that overexpressed CK1δ is unassembled, without checking it in any experiment. It could be that CK1δ overexpression drives overexpression of its binding partners, in which case most of overexpressed kinase could actually be assembled. Since CK1δ assembly is of great importance to the study's conclusions, it should be confirmed experimentally.

      In our eLife study, we show that overexpressed CK1δ is not fully assembled with stabilizing binding partners such as PER2. In fact, centrosomal/Golgi binding partners remain always available in excess even relative to overexpressed CK1δ, as demonstrated by the MG132-induced increase in centrosomal accumulation (eLife Fig. 5C). However, binding is dynamic and the affinities and concentrations are such that a substantial fraction of overexpressed CK1δ remains unbound and is therefore subject to degradation.

      (4) The authors infer from the literature that phosphorylated CK1δ is inactive - which is not a given. All Western blotting experiments should include confirmed CK1δ substrates like PER2 or DVL3 to confirm that phosphorylated CK1δ is indeed inhibited. As added benefit, blotting for known CK1δ substrates will act as confirmation that the FLAG-tag on overexpressed CK1δ/ε does not impact its kinase function, and that CK1δ/ε inhibitor treatments like PHF670 have indeed worked. Furthermore, because CalA and OA inhibit many phosphatases, inferences on CK1δ/ε activity upon these treatments should be taken catiously.

      The activity of phosphorylated CK1δ is severely attenuated by autoinhibition, although the kinase is not completely inactive. This has been demonstrated repeatedly in the literature and does not require re-establishment in the present study. We previously showed this directly in a PNAS study by Marzoll et al. (https://doi.org/10.1073/pnas.2118286119). At some point, one must rely on established published findings unless there is compelling evidence supporting an alternative interpretation.

      (5) Figure 1 requires further biochemical analyses to support immunofluorescence data. Western blotting analyses showing fluctuating CK1δ levels in nuclear vs. Centrosomal (https://pmc.ncbi.nlm.nih.gov/articles/PMC7618310/) fractions in the different treatment conditions would help illustrate the redistribution of CK1δ pools between nuclear and centrosomal areas. Additionally, adding a cytoplasmic fraction in each treatment condition would help visualise the amounts of unrecruited CK1δ in overexpressed conditions.

      Microscopy is a well-established and widely accepted approach for assessing subcellular localization. In many cases, it is superior to biochemical fractionation assays, which rarely yield completely pure nuclear or centrosomal fractions. The fractionation experiments suggested by the reviewer could provide additional supportive evidence, but they are not required to substantiate the conclusions presented here.

      (6) Figure 2 requires further biochemical analyses to support immunofluorescence data showing colocalisation of CK1δ with PER2, using for example immunoprecipitation to show that pulling down PER2 also pulls down CK1δ, and vice versa. Optimally, an additional Western blotting experiment showing total / phospho-PER2 and CK1δ levels for each condition in cytoplasmic vs. nuclear fractions would consolidate the evidence from immunofluorescence and immunoprecipitation.

      Association of CK1δ with PER2 has been demonstrated extensively in numerous previous studies (Aryal et al., 2017; Narasimamurthy et al., 2018; Philpott et al., 2020; Cao et al., 2021; Marzoll et al., 2022; An et al., 2022) including our own eLife publication.

      (7) The authors consistently interpret data showing CK1δ phosphorylation as CK1δ tail phosphorylation (p.8-11). A band shift in CK1δ signal on a Western blot does not in any way show the tail specifically is phosphorylated. Indeed, although CK1δ tail phosphorylation is widely recognised, several identified sites on other CK1δ domains, e.g.: the kinase domain, can be phosphorylated by CK1δ and other kinases. The authors' conclusions must therefore not mention tail-specific phosphorylation unless domain-specific phosphorylation is established. This could be done by comparing signal from antibodies specifically recognising known phospho-sites on the CK1δ tail against total CK1δ signal in experiments of Fig.3, 4, and 5.

      The electrophoretic mobility shift is caused by phosphorylation of the CK1δ tail. This has been demonstrated in numerous previous studies. For example, we generated a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines. In this mutant, CK1δA, inhibition of phosphatases by CalA no longer induces an electrophoretic shift (Figure 3B).

      (8) The use of anti-FLAG antibody in all figures looking at overexpressed CK1δ is not the most appropriate choice, as the anti-CK1δ antibody works perfectly for Western blotting and immunofluorescence studies. This adds more variables and hinders comparisons with endogenous CK1δ. If CK1δ is successfully overexpressed, wild-type levels should be negligible compared to overexpressed CK1δ levels, so the need for a FLAG antibody is not justified. Using an anti-CK1δ antibody in all experiments instead of anti-FLAG would confer them greater solidity.

      The amount of endogenous CK1δ is limited, and we therefore use endogenous detection only when it is scientifically necessary. FLAG-tagged constructs, in contrast, provide a robust and convenient experimental system. In the absence of evidence that the FLAG tag introduces artifacts or alters the observed behavior of the kinase, we do not consider it necessary to repeat these experiments using endogenous detection alone.

      (9) The authors base themselves off a correlation between CK1δ phosphorylation status and its observed stability to establish a causal relationship between both ('phosphorylation protects the overexpressed kinase from degradation' p. 7; 'tail phosphorylation protects the overexpressed kinase from rapid degradation' p.9). However, no experiments performed suggest there is a causal relationship occurring. This could be done by looking at specific known CK1δ tail phospho-sites using phosphosite-specific antibodies. The authors could assess whether CK1δ stability is impacted when overexpressing wild-type vs. phospho-dead mutant CK1δ.

      In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable.

      (10) The authors consistently link APC/CCDH1 with CK1δ's putative degradation, when none of their study touched on APC/CCDH1 activity or its involvement in CK1δ dynamics. To support these conclusions, fig.5 requires further biochemistry studies. To show APC/CCDH1's interaction with CK1δ, I suggest the authors (1) pull down CK1δ by immunoprecipitation and check for APC/CCDH1, phosphorylated CK1δ, and ubiquitin levels in enriched samples for each cell cycle stage in untreated vs. MG132-treated conditions. To consolidate this, the authors could (2) pull down CDH1 and check for phosphorylated and total CK1δ levels in pulled-down samples. Finally, they should investigate whether reducing CDH1 (3) expression (siRNA knock-down) and (4) CDH1 activity (specific E3 ligase inhibitor) have any impact on CK1δ levels at each cell cycle stage.

      These data were previously published in Penas et al. (doi: 10.1016/j.celrep.2015.03.016). For example, Fig. 5C shows that siRNA-mediated depletion of Cdh1, the G1-specific cofactor of APC/C, stabilizes overexpressed CK1δ, as well as the established APC/C-Cdh1 substrate Cyclin B1.

      (11) Methods section lacks a subsection detailing image acquisition, including the type of microscope used (confocal vs. widefield), magnification used, whether acquisition settings were kept consistent throughout conditions / technical/biological replicates.

      Is provided in the revised manuscript.

      (12) Methods section must include a subsection detailing immunofluorescence image analysis parameters and statistical analyses, e.g.: criteria for centrosomal localisation categories in Fig.1, criteria for categorising cells in different stages of mitosis (Fig.6), overall threshold / criteria stringency.

      Is provided in the revised manuscript.

      Minor comments:

      The Reviewer has raised about 50 minor points, counting the individual remarks and sub-point and sub-sub points. We have addressed a number of these comments where they were scientifically relevant or helpful for improving clarity. Most of the questions raised concern published and generally accepted data. We therefore do not believe that a point-by-point response to every individual minor remark is constructive or necessary.

      (1) PF670 was shown to selectively inhibit CK1δ/ε over 42 common kinases (TOCRIS), however it is not well-characterised regarding the remaining 474 kinases encoded by the human genome. There is thus a possibility for PF670 inhibition overlap between CK1δ/ε and other kinases, which should be mentioned, and the authors' conclusions should he more nuanced as a result.

      (2) CHX is a protein synthesis inhibitor, which means it does not selectively target CK1δ expression, but that of every protein within the cell. This should be addressed and controlled for, if possible. A potential way to go about it would be to knock-down CK1δ using siRNA and compare untransfected controls with select post-transfection timepoints to assess the impact of inhibiting CK1δ expression on total CK1δ levels.

      (3) All unshown Western blotting replicates should be included in the supplementary materials.

      (4) Figure 1:

      a. A-B: experiment is missing PF670 alone and CHX alone conditions to control for direct effects of either compound, versus combined.

      b. A-B: in the text, please address the fact that >50% cells show unclear or no centrosomal pattern in untreated conditions.

      c. C: the authors claim that CK1δ-FLAG levels are decreased in PF670-treated conditions because PF670 inhibits excess CK1δ autophosphorylation, inducing its subsequent degradation. Before this claim is made, a proteasome inhibitor like MG132 should be included within the experiment, to show that this decrease in CK1δ expression under PF670 treatment can be rescued with MG132 treatment. This would plead in favour of excess CK1δ degradation and exclude the possibility that 4h of PF670 treatment simply reduces the rate of CK1δ expression. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ degradation occurs via APC/CCDH1-mediated ubiquitination of CK1δ.

      d. C-D: experiment is missing pre-DOX induction control to confirm that CK1δ-FLAG is indeed being overexpressed.

      (5). Figure 2:

      a. A: should include non-transfected control panels to confirm the success of PER2 overexpression.

      b. B: Although observed in a preprint from the same team, there is no peer-reviewed evidence that establishes mK2-CRY1 as a reliable indicator of PER2 overexpression. 2A shows a correlation between both but does not exclude the fact that PER2 must be included in the imaging of the 2B panels, instead of using mK2-CRY1 as a proxy readout of PER2 expression and location. Appropriate analyses would then be required to show colocalisation of CK1δ with PER2 in the highlighted puncta.

      c. 'PER2-dependent foci': cannot be said of the data unless PER2 dependency has been validated in those images. Please see above point to resolve this.

      (6) Figure 3:

      a. B: the use of kinase-dead CK1δ mutant does not allow to fully separate direct autophosphorylation from phosphorylation by other kinases, unlike what the authors mention: CK1δ kinase activity may be required for the phosphorylation of certain sites by other kinases. In this case, the decrease of phosphorylation in the kinase-dead CK1δ mutant would not only result from inhibited autophosphorylation, but also from reduced phosphorylation by other kinases. Please adjust your conclusions accordingly (p.8).

      b. C: poor visualisation of CK1δ overexpressed condition, especially showing critically reduced signal at the 6min CHX timepoint, compared to its kinase-dead homolog. If this change is present in every replicate performed, please address it in the text. If not, perhaps you may have to display another representative replicate in the figure.

      c. The authors claim that the increased levels of overexpressed with CalA treatment suggest 'that full or partial phosphorylation of the CK1δ tail stabilizes both active and inactive forms of the kinase' (p.8). Importantly, tail phosphorylation was never shown, so this should be corrected. Additionally, it could be that CalA treatment considerably increases the rate of CK1δ expression - which would also match data in fig.4A, since CalA and CalA+PF670 treatments alone drastically increased CK1δ levels. This should be checked by pre-treating cells with CHX before applying CalN, or by treating cells with CHX and CalN simultaneously, and interpreted accordingly.

      (7) Figure 4:

      a. A: CK1δ levels in CHX-free controls from both 1h pre-treatments look much higher compared to untreated controls, which the authors interpret as an indication that 'most of the newly synthetised overexpressed kinase was degraded in untreated cells' (p.9). However other explanations are not explored: since loading controls are not provided, it may be that sample loading in the gel is simply off. Importantly, it is also possible that pre-treatments increased CK1δ expression before CHX application. Please make sure you touch on each

      b. A-B: the authors claim that CK1δ-FLAG and CK1ε-FLAG levels are decreasing with CHX treatment because they are being degraded. Adding a panel with a proteasome inhibitor like MG132 would solidify this argument. Rescue of CK1δ/ε degradation under CHX treatment would show that the loss is indeed mediated by the UPS. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ/ε degradation occurs via its ubiquitination by APC/CCDH1.

      c. B, D: blot in B does not match the quantification trends in D. E.g.: 60min CHX + CalA + PF670 condition shows clearly lower CK1ε signal compared to its 0min CHX control. Please ensure the biological replicate you display on the figure is indeed representative of your results.

      d. A, C: 'Hyperphosphorylated CK1δ remained stable throughout the CHX chase [...], indicating that tail phosphorylation protects the overexpressed kinase from rapid degradation' (p.9). Meanwhile this is true, unphosphorylated CK1δ in the CalA+PF670 treatment condition was also stabilised, showing that CK1δ phosphorylation may not be required for kinase stabilisation. This is an important point and should be addressed in the data interpretation. On another note - and as mentioned above -, tail phosphorylation specifically is not shown and cannot be inferred unless domain-specific phosphorylation is investigated.

      e. B, D: CK1ε-FLAG levels decrease with CHX treatment compared to its baseline in the CalA+PF670 condition, which is not the case for CK1δ-FLAG (A). Thus, the data shown does not support the authors' conclusions 'CK1ε turnover is regulated in a similar manner to CK1δ'. Please address this issue and adjust your conclusions accordingly.

      f. C-D: performing statistical analyses on protein band intensity in different conditions would be interesting to establish the significance of those changes.

      (8) Figure 5

      a. B: non-arrested controls mentioned in the text are missing from the figure.

      b. B: 'the accumulated CK1δ remained dephosphorylated' (p.10). The experiment is missing important positive / negative controls of CK1δ phosphorylation status to conclude whether CK1δ remained phosphorylated or unphosphorylated across conditions. One or the other cannot be concluded from the blot as it is. Using phospho-specific antibodies may also help to visualise phosphorylation status.

      c. B, D-E: blots should include CDH1 phosphorylation levels (hyperphosphorylated CDH1 is inactive), CK1δ substrates (indicators of CK1δ activity), and relevant phosphatase substrates to create a cohesive picture of changes in mitosis and support fig.7.

      d. D-E: please address the differences in CK1δ profile between overexpressed and endogenous CK1δ in G2/M phase.

      (9) Figure 6

      The authors say: 'similar results were observed when using U2OStx cells and staining for endogenous CK1δ'. However, CK1δ's subcellular distribution pattern is different in endogenous vs. overexpressed conditions in the telophase-cytokinesis and post-mitosis stages. In the telophase-cytokinesis stage, endogenous CK1δ seem to form nuclear hotspots, while overexpressed CK1δ is more diffuse. In the post-mitosis stage, overexpressed CK1δ shows a clear polar pattern in the nuclear periphery, while endogenous CK1δ shows a diffuse pattern similar to that of overexpressed CK1δ described in fig.1 as 'unassembled' by the authors. Please address this and adjust your conclusions appropriately.

      (10) Figure 7

      a. The authors infer APC/CCDH1's activity levels or relationship to CK1δ from existing literature only. Since it is a central mechanism of their study, literature-only components are insufficient for a summary figure. To include these elements in the figure, the authors must include investigations of APC/CCDH1's activity levels and involvement in CK1δ degradation at different stages of the cell cycle in their study. Please refer to point 10.

      b. Similar comment for CK1δ assembly status and activity levels. Please refer to point 4.

      c. Similar comment for phosphatase activity levels in different stages of the cel cycle. By observing phosphorylation status in known substrates of established CK1δ phosphatases, one can easily confirm phosphatase activity levels.

      d. Does not consider the fact that some results were different in endogenous vs. overexpressed CK1δ models. Please nuance your claims.

      (11) Reference missing p.12 paragraph 1: 'Yet, overexpression of CK1δ consistently accelerates the circadian clock, implying that kinase availability can influence clock speed. This finding suggests that CK1δ activity may be regulated not only by catalytic mechanisms but also by its spatial availability within the cell'. The facts stated there are not covered in the study's findings and is not referenced with a published study.

      (12) Grammar mistakes / typos to report:

      p.3 paragraph2: 'the kinases undergo futile cycles of phosphorylation and dephosphorylation'

      p.26 Fig.1A legend: 'Endogenous CK1δ was detected by IF to'? Unfinished formulation

      p.8: 'PP1 was previously suggested as a major PPase of CK1δ/ε'

      p.8: 'fewer phosphorylation sites are targeted by other kinases'

      Figure 4C legend: 'Overexpressed unphosphorylated CK1δ [instead of CK1ε] is degraded with a half-life of about 15 min'

      Figure 4E: Ponceau staining

      (13) Nomenclature inconsistencies to report:

      p.4 paragraph2: 'protein PER2', then p.4 paragraph3 'PERIOD2'.

      'CK1δ' used throughout the article's body text, but 'CSNK1D' is used in IF panel legends. The gene name was never introduced in the main text, nor has it been explained in the figure legends, so perhaps go for CK1δ for all mentions, including in figures.

      'FOV' nomenclature is unclear, please define in the figure legend.

      Figure 2B: what are the arrowheads pointing to? - please clarify in the figure legend.

      Mislabelled figures: figure 3 in the text refers to figure 4 in the figure section, and figure 4 in the text to figure 3 in the figure section.

      Figure 6A-B: abbreviated mitosis stages 'Pro' and 'Meta-Ana' should either be defined in the figure legend or put in full writing within the figure.

      Reviewer #1 (Significance):

      The paper will appeal to those working on CK1 biology, including cell cycle and circadian rhythms.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      In this study, authors aimed to address functional links between CK1 activation, subcellular localization and protein stability, which is an important biological question. This manuscript is a follow up of a recent study by the same team, which is currently deposited at Biorxiv and a fraction of data seems to overlap, which is somewhat confusing. Overall, the concept that the stability of CK1d is dynamically controlled across the cell cycle is interesting. On the other hand, how this is functionally connected to the circadian cycle described in the previous study remains unclear.

      The data presented in this study are highly preliminary and lack a number of essential controls, which weakens an otherwise interesting concept.

      Major issues:

      One of the main conclusions of the study is that CK1d is degraded predominantly in the nucleus while the centrosomal pool is protected from the degradation. This concept is interesting but unfortunately, the experimental evidence supporting this model is very limited. Authors previously showed that neither inhibition of proteasome or treatment of cells with PF670462 significantly influenced levels of endogenous CK1d but both treatments promoted accumulation of the tagged and overexpressed FLAG-CK1d. Absence of the phenotype at the level of endogenous protein clearly raises question whether this may be just an artifact of overexpression, tagging or both combined.

      As shown in our previously published eLife study, CK1δ is synthesized at a relatively low rate from its endogenous locus. Free CK1δ continuously shuttles between the cytosol and the nucleus. Although association of CK1δ with the centrosomal/Golgi area is dynamic, nuclear export followed by binding to centrosomal/Golgi structures constitutes the major sink for the kinase at steady state. Consequently, the fraction of unbound CK1δ that is targeted for degradation in the nucleus at any given time is very small and cannot be detected in cycloheximide chase assays against the much larger background of kinase stabilized by association with centrosomal/Golgi binding sites. To reveal that CK1δ is subject to degradation at all, we had to overexpress the kinase. Under these conditions, the fraction of unbound CK1δ increases substantially, allowing nuclear degradation to be detected experimentally.

      The statement that degradation occurs in the nucleus is based on weak data with CK1d-NES construct that was claimed to accumulate at higher levels compared to the wild type CK1d. However, that experiment used quantification of microscopic data in which two distinct regions (nucleus, cytosol) were compared, which is technically challenging. Including a reporter for normalizing for the transfection efficiency would strengthen conclusions of that experiment. Overall, quantification by immunoblotting may be more accurate.

      As suggested, we performed CHX time-course experiments with NES- and NLS-tagged CK1δ constructs followed by immunoblot analysis (shown in revised Fig. 4E).

      Specificity of the microscopy staining was not validated Fig. 1. Confirmation of the staining specificity by RNAi or KO approaches is essential. This antibody from Abcam has been discontinued which makes it impossible to reproduce this experiment.

      The authors should therefore attempt to demonstrate localization using other available antibodies against CK1d.

      The antibody is available through Thermo Fisher in the United Kingdom: https://www.fishersci.co.uk/shop/products/100-ul-mouse-monoclonal-af12g4-casein-kinase/13070202#

      It is specific for CK1δ and does not recognize CK1ε, as demonstrated in our eLife study.

      Furthermore, FLAG-tagged CK1δ detected with anti-FLAG antibodies, as well as endogenous CK1δ detected with the commercial antibody, localize to the pericentrosomal region and relocalize upon PER2 overexpression to mCRY1-containing nuclear foci. Together, we consider these data compelling evidence for the specificity of the antibody staining.

      The authors also failed to demonstrate that the dots represent centrosomes. To do so, they must perform co-staining with a robust centrosome marker.

      As requested, Pericentrin (PCNT) was used as a centrosomal marker and is now provided in the revised Fig. 1.

      This would help classify approximately 50% of the cases that are currently assessed as "potential" but are in fact inconclusive.

      Centrosomal colocalization with PCNT improved the robustness of the quantification.

      The quantification in the experiment is incorrect. Since the percentages are reported, both columns should add up to 100%. In Panel 1B, approximately 10% of the cells are missing, while in Panel 1D, there appear to be about 20% more cells.

      As indicated in our previous figure, the y-axis represents the number of cells, not percentages. We analyzed approximately 100 cells, which may have led to the misunderstanding that the values refer to percentages. As well, we have replaced this figure with a revised Figure 1 which shows that CK1δ co-localizes with PCNT as a centrosomal marker.

      Fig. 2B shows four different fields based on which authors come to conclusion that co-expression of CRY and PER2 promotes re-localisation of CK1d to nuclear foci. This is an interesting possibility but it is hard to conclude without any quantification. What was the fraction of cells that expressed CRY and PER2 that showed this phenotype? What was the fraction of cells that did not show this phenotype although both of these proteins were expressed?

      It would help to label cells expressing PER2 either by expressing it as a fusion protein or by co-expressing a marker protein ideally from the same plasmid.

      All cells in this stable cell line express mK2-CRY1. The cells were transiently transfected with PER2, and therefore only a fraction of the cells received the PER2 expression construct. In our eLife paper, we showed that all PER2-expressing cells stabilized mK2-CRY1 and formed nuclear foci. We demonstrate here that every cell containing such nuclear foci also showed accumulation of endogenous CK1δ within these structures. We did not observe a single cell with PER2-induced nuclear foci that lacked endogenous CK1δ accumulation.

      Cells that did not express PER2 were identified by the absence of nuclear foci and by low levels of mK2-CRY1, which was homogeneously distributed throughout the nucleus, as also shown in the eLife manuscript. In these cells, endogenous CK1δ was concentrated at a single discrete structure that, based on data shown in Fig. 1, corresponds to the pericentrosomal region. Under these conditions, neither CK1δ nor mK2-CRY1 was detected in nuclear foci.

      Fig. 5 suggests that massive phosphorylation of CK1d, which is responsible for its mobility shift on SDS-PAGE, is most likely linked with mitosis. Authors should use established markers to estimate a fraction of mitotic cells in their G2/M fraction. In principle, there are two possible explanations for the doublet observed with CK1d staining. Ether CK1d exists in two pools with different phosphorylation states in mitosis, or perhaps more likely, this fraction contains G2 cells where CK1d is not yet modified and mitotic cells where CK1d is fully phosphorylated. Performing a shake-off experiment yielding a pure fraction of mitotic cells could help to distinguish between these two options.

      As described in the main text and the methods section, the fractions were prepared by mitotic shake-off to further enrich our sample for rounded cells that are loosely attached during mitosis.

      Degradation of CK1d in telophase/cytokinesis when APC/Cdh1 becomes active is not apparent in Fig 6. The signal at mitotic spindle is missing, but there is still plenty of signal remaining in the cells. It is possible that the signal is just redistributed in the cell and the data shown do not support degradation of the protein. Authors could film cells expressing fluorescently labeled CK1d and quantify the signal during progression through mitosis and mitotic exit. The statement that "Following nuclear envelope reformation and mitotic exit, CK1d localized primarily to the single centrosome in each daughter cell" is incorrect. First, authors cannot deduce from the fixed cells whether they have just formed the nuclear envelope and exited mitosis.

      The reviewer is, of course, absolutely correct in the points raised.

      First, we cannot deduce from fixed cells whether they have only recently exited mitosis. This was not our intention. We merely selected cells in G1 and referred to them as “post-mitotic,” without intending to imply that these cells had just exited mitosis. To clarify this point, we changed the previous statement:

      “Following nuclear envelope reformation and mitotic exit, CK1δ localized primarily to the single centrosome in each daughter cell (Fig. 6A, 4th column)”

      to:

      “In G1, CK1δ localized primarily to the single centrosome (Fig. 6A, 4th column),” and replaced in column 4 of Fig. 6A and B the label “post-mitosis” with “G1.”

      Furthermore, we cannot deduce from fixed cells whether, or to what extent, CK1δ is degraded upon mitotic exit. This was neither the intention nor the conclusion drawn from Fig. 6. Rather, we show in Fig. 3 that phosphorylated CK1δ is not degraded, and in Fig. 5 that CK1δ is predominantly hyperphosphorylated during mitosis, leading us to conclude that this phosphorylated pool of CK1δ is stable. In G1, CK1δ is dephosphorylated and unassembled kinase is degraded. We currently have no data regarding the kinetics of CK1δ dephosphorylation, assembly with centrosomal/Golgi structures, versus degradation of unassembled dephosphorylated CK1δ.

      The data shown in Fig. 6 serve merely to illustrate the subcellular distribution of CK1δ, which is consistent with previous reports. In Fig. 7, we present a model that attempts to integrate the new findings reported here together with the data from our recent eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation. Of course, we do not claim that this model does by no means represents a final verdict, and many important questions remain open. However, we believe that the model provides plausible novel concepts and mechanistic ideas that have not been proposed in previous publications and therefore merit publication, as they provide a basis for further investigation and discussion.

      Live-cell imaging:

      Live-cell imaging of fluorescently tagged CK1δ throughout mitotic progression and mitotic exit could, in principle, provide additional insight, but such experiments are technically extremely challenging. Moreover, the central idea of our model is that as much CK1δ as possible is preserved throughout the cell cycle, whereas degradation selectively targets unassembled and potentially harmful kinase. When CK1δ is expressed at physiological levels, which would require tagging the endogenous locus, the fraction of kinase degraded upon mitotic exit is expected to be very small and therefore likely below the threshold for reliable quantification by fluorescence microscopy. Similarly, although overexpressed CK1δ undergoes substantial degradation in G1, the kinase is simultaneously synthesized at a high rate. Hence, quantitative interpretation of overexpressed CK1δ levels during mitotic exit by microscopy (without CHX) would still be difficult.

      Second, the images of interphase cells constantly show multiple dots (probably surrounding the centrosome), which is a pattern that likely corresponds to Golgi rather than a single centrosome.

      The reviewer is correct. Indeed, CK1δ localization to the Golgi apparatus is well established in the literature. In our original wording, we did not explicitly distinguish between Golgi and centrosomal localization, which may have been somewhat misleading. In the revised version, we therefore refer more cautiously to the “pericentrosomal region” rather than strictly to the centrosome.

      The data shown in Fig. 6 serve to illustrate the known subcellular distribution of CK1δ. In the model presented in Fig. 7, we attempt to integrate the new findings reported here together with the data from our eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation.

      It is unclear to which figure points the paragraph "Tail phosphorylation protects CK1δ/ε from degradation". I assume that one figure is missing.

      Figs. 3 and 4 were accidentally swapped, and we apologize for this error. The paragraph in question refers to Fig. 4, which shows the cycloheximide-induced degradation kinetics of CK1δ/ε.

      Minor points:

      CK1 kinase inhibitor PF670462 should not be named as PF670 as this causes confusion. Authors should either use the full name of the compound or just call it as CK1 inhibitor with providing details in the methods.

      PF670 has been changed to PF670462.

      Fig. 3B is discrepant with the figure legend. Figure shows CK1e but legend says kinase dead CK1D-K38R

      The captions to Figs. 3 and 4, as well as the references to these figures in the text, are correct. However, the actual Figs. 3 and 4 were inadvertently swapped during figure assembly. We apologize for this mix-up.

      The authors` interpretation of CK1 involvement in checkpoint is incorrect. The authors state that CK1 activity decreases p53 function promoting recovery, but Inuzuka et al (ref. 51) showed that inhibition of CK1 leads to this outcome.

      We thank the Reviewer for noting this mistake. We are no experts in p53 regulation, which is rather complex. CK1 decreases MDM2 stability and hence enhances p53 function.

      We corrected the statement and placed it in the right context: “CK1 phosphorylation triggers β-TrCP-mediated degradation of MDM2 and activates p53, thereby enhancing p53-dependent responses involved in checkpoint signaling and DNA repair (Inuzuka et al., 2010; Winter et al., 2004). After DNA repair, CK1δ has…”

      Reviewer #2 (Significance):

      It is generally assumed that CK1 is constitutively active, which is likely an oversimplified view; in a physiological context, some degree of regulation can be expected. Demonstrating that there are several pools of CK1 that are differently regulated at the level of protein stability during the cell cycle would be a significant advance in our understanding of CK1 functions.

      Reviewer #3 (Evidence, reproducibility and clarity):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U2OS cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin), and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      Overall, they propose that the activity and abundance of CK1 are regulated during the cell cycle. However, this claim would require several experiments to support it.

      Major comments:

      - As presented, some of the data are inconclusive. Co-staining with a centrosomal marker is required to determine whether or not Ck1 localises to the centrosomes. A large proportion of cells exhibit "potential" (their term) centrosomal staining, so a centrosomal marker is essential before any conclusions can really be drawn.

      In our revised Figure 1, we confirm this centrosomal staining using an antibody against pericentrin (PCNT).

      - Figures 3 and 4 have no loading controls, and these two figures have been mixed up in the text.

      The reviewer is right, we have corrected the Fig. 3 and Fig. 4 mix-up.

      Loading controls are now also provided.

      - A mobility shift on SDS-PAGE does not prove that a protein is phosphorylated. The authors should provide experimental evidence that the mobility shift is really due to phosphorylation. As they are inactivating phosphatases using CalA, it is likely the case, but they should prove it. Furthermore, the authors did not map any phosphorylation sites in this study, so they do not know whether CK1 phosphorylation occurs in the tail (as they assert) or elsewhere.

      We provide data in Fig. 3B showing that a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines does not undergo an electrophoretic mobility shift upon CalA treatment. These results demonstrate that the CalA-induced mobility shift is caused by phosphorylation of the CK1δ C-terminal tail.

      - The figure legends in general are limited and lack crucial information. For instance, in Figures 3C and 3D, how was the half-life of CK1 determined?

      We have adapted the figure caption:

      (C) Densitometric quantification of n=3 Western blots (see A) shown as mean ± SD. Overexpressed unphosphorylated CK1δ is degraded with a half-life of about 15 min. Both CalA and CalA + PF670462 treatments stabilize the kinase. (D) Densitometric quantification of n=3 Western blots (see B) shown as mean ± SD.

      - The CK1 regulatory model presented in Figure 7 is not supported by the data. What experimental evidence, for instance, shows that CK1 is inactive during mitosis? To make this claim the authors should directly assay its activity.

      We show that the majority of CK1δ is phosphorylated and therefore auto-inhibited during mitosis. In the original version, we referred to this fraction as inactive. In the revised manuscript, we refer to the phosphorylated kinase as auto-inhibited.

      Reviewer #3 (Significance):

      This study may be of interest to researchers working on cell cycle regulation.

    1. eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are compelling, and supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

    3. Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are largely supported by strong genetic and biochemical evidence, although some claims regarding the dispensability of fatty acid synthesis pathways remain incomplete. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We kindly ask that the editorial assessment be revised to focus on the major conclusions. Alternatively, we would respectfully suggest revising the second sentence to something akin to “The main conclusions are largely supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve.”

      Public Reviews:

      Reviewer #1 (Public review):

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

      We thank the reviewer for these positive comments.

      Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      To clarify, we conclude that the essential function of ACP in blood-stage P. falciparum parasites includes a critical stabilizing interaction with pyruvate kinase II. Apicoplast ACP presumably still plays a central biochemical role in FASII pathway function, but that role in FASII is dispensable for blood-stage parasites.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages. However, there are currently some important experimental issues/flaws, missing experiments that induced wrong interpretations and thus do not support some important claims of the study, notably for the role of FASII and the interaction between ACP and PKII.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We elaborate on these points and address the reviewer’s critiques below.

      Therefore, at this point, the study is only partial and would require major additions and/or important text edits/revisions before being considered for acceptance.

      We note that the manuscript has already been accepted for publication in accordance with the current eLife publishing model.

      Major points:

      From the graph of P. falciparum growth, we can see that in the lipid-rich condition, where both FabH KO and ACP KO can survive, the addition of mevalonate was essential for the growth of ACP KO. Along with the other evidence (PKII association, DNA levels...), we therefore agree that PfACP is involved in the mevalonate pathway.

      To clarify, our model is that ACP supports IPP synthesis by the apicoplast nonmevalonate/MEP pathway indirectly by stabilizing and thus supporting function by pyruvate kinase II that supplies the pyruvate and NTPs required for IPP synthesis by the MEP pathway.

      The authors claim that the FASII pathway is inactive/not essential in the P. falciparum blood stage. However, the authors have not shown any evidence on whether ACP is or not involved in the FASII pathway during the asexual blood stage.

      To clarify, there is overwhelming data in the prior published literature that we cite (including refs. 13, 14, and 32) to establish that FASII is dispensable for blood-stage Plasmodium growth in vivo in rodent parasites and in vitro culture in human parasites. Prior studies also strongly support a role for apicoplast ACP as the central scaffold for FASII-mediated acyl chain synthesis. However, our and prior studies support the conclusion that essential ACP function in blood-stage parasites is independent of its role in FASII.

      As currently designed, the experiments presented cannot conclude on that point for several reasons. Indeed, it was previously shown that (i) the expression of the protein from the FASII pathway are all present in blood stages and are significantly upregulated in patients that are under under "nutrient starvation" (Daily et al. Nature 2007), (ii) that, growing parasites under similar low lipid conditions in vitro induces an activation/upregulation of FASII, which can be measured by stable isotope precursor labelling and lipidomics (Botté et al. 2013).

      We are aware of these prior studies and cite and discuss the Botté et al. 2013 reference in our manuscript, which provided isotope-labeling evidence to support FASII activity in low-lipid growth conditions for P. falciparum. We note that neither study addresses whether FASII activity is required for growth in low-lipid conditions.

      (iii) that growing the PfFabI KO line under deprived lipid conditions leads to parasite death (Amiar et al. 2020), indicating that the FASII pathway can become critical, if not essential, depending on the host nutritionnal content together correlating patients' data and metabolic adaptation for the same reasons in the related parastie Toxoplasma gondii (Amiar et al. 2020, Krishnan et al. 2020, Liang et al. 2020, Primo et al. 2021, Charital et al. 2024, Dass et al. 2024, Bitew et al. 2025).

      All of the studies cited by the reviewer focus primarily or exclusively on Toxoplasma gondii parasites. We agree with the reviewer that these and other studies provide strong evidence that FASII activity contributes to growth of T. gondii parasites, including roles for apicoplast ACP that appear to differ from what we have unveiled for P. falciparum malaria parasites. We acknowledge and discuss these differences from T. gondii in the final section of the Discussion section and think that exploring these differences will be a fascinating area for future study.

      The Amiar et al. 2020 paper cited by the reviewer is the only study we are aware of that has directly tested the ability of a ∆FASII parasite (in this case, ∆FabI) to grow in low-lipid conditions. We acknowledge that they observed little to no growth of ∆FabI parasites in these conditions. Our growth assays with ∆ACP and ∆FabD parasites indicated a different outcome in which both WT and ∆FASII parasites grew similarly in low-lipid conditions. Our results thus contrast with the prior study. As noted below, the minimal lipid growth conditions explicitly reported in the methods section of the Amiar et al. paper are identical to those used in our study: fatty acid-free BSA, 30 µM palmitic acid, and 45 µM oleic acid (all sourced from Sigma) with daily media changes. Thus, the basis for these differences is unclear and additional follow-up work will be needed to explore and resolve these differences.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

      Here, the authors are expecting to show that FabH (and thus the FASII pathway) is not essential in an experiment that is not designed to be in low lipid conditions but rather in lipid rich conditions: Such high lipid conditions of culture in this study is granted by daily feedings with high fatty acid supplement (30-90 uM palmitic acid and 30-60 uM oleic acid). These fatty acid concentrations were used previously by Mitamura et al. (2005) and Miichi et al.(2007) to replace non-determined supplements such as Serum or Albumax supplement to grant similar growth by a completely controlled culture medium.

      This means the concentrations above do not represent limited fatty acid concentrations, especially not with daily feeding (representing an excess supplied amount of lipids, unlike regular 48h feedings) that allowed the authors to easily reach very high non-physiological parasitaemia of more than 20%!! Amiar et al. previously showed essentiality of FabI in P. falciparum in the limited fatty acid culture at a lower concentration (<30uM 16:0, <45um 18:1), than the Mi-Ichi et al. controlled medium with regular 48 h culture feeding. Therefore, with the current experimental settings, the FAH KO is placed in high lipid conditions, thus preventing any conclusion on its essentiality under low lipid conditions.

      The basis for the reviewer’s statements here is unclear, as this critique and the conditions it describes do not conform to the published conditions reported in the Amiar et al. 2020 paper. The methods section of that study for “Plasmodium falciparum growth assays” explicitly states (page e7):

      “Media was replaced daily, sub-culturing were performed every 48 h when required, and parasitemia monitored by Giemsa-stained blood smears. Growth assays in lipid-depleted media were performed by synchronizing parasites before transferring trophozoites to lipid-depleted media as previously reported (Botte ´ et al., 2013; Shears et al., 2017). Briefly, lipid-rich AlbuMAX II was replaced by complementing culture media with an equivalent amount of fatty acid-free bovine serum albumin (Sigma), 30 µM palmitic acid (C16:0; Sigma) and 45 µM oleic acid (C18:1; Sigma).”

      We used identical culture conditions to those described above: fatty acid-free BSA in place of lipid-rich AlbuMAX, 30 µM palmitic acid, and 45 µM oleic acid. We thus obtained growth results that contrast with the prior study and suggest that additional, future studies will be required to understand and resolve these differences.

      Furthermore, it is too uncertain to conclude that ACP is only essential for the mevalonate pathway.

      Please see our response above that clarifies our model for ACP function in supporting pyruvate kinase II and the many apicoplast pathways that appear to depend on PKII.

      This would be a similar discussion to the Yeh et al. 2011 and the Swift et al., where induced Apicoplast knockout caused parasites to require IPP to survive, but there were always remnant apicoplast vesicles and thus the putative presence of an active FASII in the parasite, where de novo fatty acid synthesis could be maintained.

      It is extremely unlikely that FASII remains active upon apicoplast disruption and loss of the apicoplast genome. The apicoplast-encoded SufB is lost upon apicoplast disruption and can no longer participate in making Fe-S clusters. Without Fe-S synthesis, the apicoplast lipoate synthase (LipA) cannot make lipoate to activate pyruvate dehydrogenase (E2 subunit) and produce the acetyl-CoA needed for FASII activity. There is no experimental evidence that FASII remains active upon apicoplast disruption and loss of the apicoplast genome.

      Amiar et al. (2020) and Krishnan et al. (2020) showed that disruption of FASII and absence of de novo FA synthesis in T. gondii could be compensated by the exogenous supplementation of myristic acid, C14:0.

      As explained above, we acknowledge that FASII contributes to Toxoplasma gondii growth and includes functions that appear to differ from P. falciparum.

      Here, high fatty acid supplementation using commercially available fatty acids may include unexpected fatty acid species such as myristic acid in palmitic acid or oleic acid, since all commercially available fatty acids guarantee only >99% but not 100%. If P. falciparum requires a very, very low amount of myristic acid to survive, the amount of possible contamination, like 1 nM, may be sufficient to maintain their survival. Thus, ACP and FabH might be very important to generate de novo fatty acids within parasites, but this was not shown by the authors.

      As noted above, we used identical culture conditions and commercial sources of defined fatty acids to those reported in the Amiar et al. study. We do not see a basis for the reviewer’s critique that the two studies utilized differing culture conditions. Nevertheless, we agree that future studies are needed to understand and resolve these differences.

      Therefore, the manuscript currently contains incorrect conclusions on the potential essentiality/use of FASII, against current experimental evidence.

      As explained above, we do not see a basis for the reviewer’s critique here or for viewing one study as more or less definitive than the other, as identical culture conditions were used yet contrasting results were obtained for reasons that remain uncertain. Future studies beyond the scope of the present manuscript will be required to fully understand and resolve these differences.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      We either request more solid experimental evidence showing the absence of fatty acid synthesis at low fatty acid conditions by re-doing the growth assay in the lower fatty acid feeding conditions without daily feeding to clarify if the ACP and FabH are essential in the blood stage growth, or not; as well as showing the absence of fatty acid synthesis at low fatty acid conditions using isotope labelled precursor. Without these, the authors cannot conclude on this important point. Alternatively, toning down the text to acknowledge the possibility of FASII being active and critical under certain conditions would be acceptable.

      We note that the Amiar et. al 2020 study cited by the reviewer reported growth assays for WT and ∆FabI NF54 P. falciparum in low-lipid conditions that similarly lacked direct tests of FASII activity by isotope labeling.

      We agree that our results, which utilized distinct ∆ACP and ∆FabD NF54 (PfMev) lines, contrast with the results and conclusions of the Amiar et al. study. We fully agree with the reviewer that future studies, utilizing tandem growth assays and isotope-labeling metabolic flux assays (e.g., mass spectrometry), will be required to fully test and understand the dependence of FASII activity on the lipid content of the growth medium and the functional dependence of P. falciparum growth on FASII activity in low-lipid conditions.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

    1. eLife Assessment

      This important work identifies phlda2 as a specific marker for primordial cardiomyocytes in the adult zebrafish heart and demonstrates their essential role in myocardial morphogenesis and coronary vascularization, but not in heart regeneration. The conclusions are well supported by single-cell transcriptomics, new genetic tools, and cell-specific ablation experiments. The revised version strengthens these findings with analyses of the epicardium, long-term time points showing that the developmental defects persist, and a control confirming that loss of primordial cardiomyocytes alone does not trigger a regenerative response. Overall, the evidence is solid and provides insight into the difference between developmental and regenerative cardiac programs, and this work will be of interest for those studying cardiac development and regeneration.