10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This computational study constitutes an extension to prior work on biophysical calcium-based synaptic plasticity rules, investigating how the regulation thresholds for strengthening and weakening connection help solve learning tasks. This important work presents a significantly simpler solution to the studied problem with potentially broad applicability, the evidence to support the core conclusions is solid.

    2. Reviewer #1 (Public review):

      Summary

      This computational modelling study investigates the functional role of metaplasticity in a calcium-dependent synaptic plasticity rule, using a biophysically detailed model of a striatal projection neuron. The learning rule is deliberately simple: local calcium signals derived from different sources regulate LTP and LTD through separate plasticity thresholds, while dopamine provides a reward-related signal that determines the direction of synaptic updates. The author uses linear and nonlinear feature-binding problems (FBP and NFBP) as learning tasks to illustrate the mechanics of competing LTP and LTD processes in this rule. The central aim is thus to use these classification tasks to understand what metaplasticity in the two calcium thresholds contributes to learning.

      The study identifies complementary roles for metaplasticity in the LTP and LTD thresholds, enabling synapses exposed to competing plasticity processes to converge toward a stable outcome. In reversal-learning tasks, relaxing the conditions for metaplasticity allows previously stabilized synaptic states to become flexible again and permits relearning. Finally, the manuscript investigates whether two separate calcium thresholds are themselves necessary. While a modified rule with a single adaptive calcium threshold can solve the same tasks comparably well, two thresholds allow separate control over synaptic strengthening and weakening.

      Strengths

      A major strength of the study is the simplicity of the learning rule, which makes it possible to dissect the roles of individual components while retaining a connection to experimentally motivated mechanisms of cortico-striatal plasticity. The use of shared features in the FBP and NFBP is useful for exposing the interaction between competing LTP and LTD processes and for showing how adaptation of one plasticity threshold can indirectly permit expression of the opposite form of plasticity.

      The manuscript also provides a much more systematic analysis of the rule than the original version. In particular, the expanded parameter analyses explore both successful and failing regimes for the learning and metaplasticity rates, while the scan over the maximum synaptic weight provides a clear mechanistic motivation for the additional upper LTP threshold.

      Importantly, the single-threshold simulations directly address whether two calcium thresholds are computationally necessary. These additional simulations clarify that the principal advantage of two thresholds is not that they uniquely enable nonlinear learning, but that they permit independent regulation of strengthening and weakening. This considerably strengthens the conceptual conclusions of the study.

      Weaknesses

      The main limitations concern the scope and biological generalizability of the results. The simulations in this paper are based on a particular detailed SPN model and depend on several specific biological and modelling assumptions. Most notably, solving the nonlinear task requires appropriately clustered excitatory inputs, but the learning rule itself does not provide a mechanism for generating these clusters. The manuscript now discusses this limitation explicitly and appropriately identifies structural plasticity as an important direction for future work.

      A second important assumption is the upper calcium threshold that prevents LTP once sufficiently strong supralinear responses are reached. This mechanism plays an important functional role in the model but has not been experimentally established for cortico-striatal synapses. The manuscript appropriately identifies this as a major assumption and discusses potential molecular mechanisms that could implement such an effect.

      Overall assessment

      The revised manuscript is substantially stronger. The conclusions are carefully matched to what is demonstrated by the simulations, and several important alternative explanations and model configurations are tested directly. In particular, the revised framing provides a more specific mechanistic account of what metaplasticity in two separate thresholds contributes within the proposed rule. The work therefore provides a valuable contribution to computational studies of synaptic plasticity and metaplasticity, with convincing evidence supporting its principal conclusions.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript proposes interesting synaptic plasticity rules grounded in experimental data. Its main features are: (1) plasticity depends on local calcium concentration driven by presynaptic activity and is independent of somatic action potentials, (2) the rules incorporate metaplasticity, and (3) they demonstrate how a single neuron could address the feature-binding problem at the dendritic level.<br /> The work extends a previous study (https://doi.org/10.7554/eLife.97274.2), to which the author also contributed.

      The author models two calcium thresholds (LTP/LTD) from two different calcium sources (NMDA/VGCC) and these thresholds are flexible (metaplasticity rule, similar to BCM) which is claimed to be necessary for successful learning of both FBP and NFBP (linear and nonlinear feature binding problem with 1 or 2 patterns). The role of each threshold seems to be is opposite and complementary. One extra condition has been added : an upper threshold for LTP. This threshold serves to stop synaptic strengthening once synapses are strong enough to evoke a plateau. With that, synapses are not strengthened to the maximal value, avoiding strong supralinear integration for irrelevant patterns.

      Strengths:

      The current model implements not only local synaptic plasticity but also metaplasticity and solves the FBP at the dendrite level. Another strong aspect of the model is that metaplasticity in the LTD threshold protects strengthened synapses from weakening. In this way, as the author mentioned, metaplasticity is able to protect learned patterns from being forgotten or weakened and prevent irrelevant patterns from being stored. This is a nice modelling example of metaplasticity being helpful in preventing the catastrophic interference or forgetting (as has been explicitly discussed in a recent article https://doi.org/10.1016/j.tins.2022.06.002). The author might want to briefly mention or emphasize this aspect of the model, which might be interesting also for the AI community.

      Comments on revised version:

      The author improved the paper significantly. He conducted new simulations on relearning. He addressed the problems and limitations, added new simulation results and reasoned better in the discussion and the introduction. Also, the figures are improved.

    4. Author response:

      The following is the authors’ response to the original reviews.

      The goal of the study was to demonstrate the role of the two thresholds in learning, i.e. the role of metaplasticity with a learning rule that is both as simple as possible, but also biologically-based (hence the use of the detailed SPN model, as well). The goal was not to solve the FBP and NFBP with a different rule. However, the text likely gives that impression. Therefore, in addition to addressing the reviewers’ comments, the text has been substantially revised to reflect the following:

      Main results

      (1) Metaplasticity enables synapses that undergo both LTP and LTD to ultimately express just one plasticity outcome (either LTP or LTD).

      (2) Metaplasticity in the LTP threshold allows LTD to be expressed, and vice versa. (That is, the threshold regulating one plasticity process allows the expression of the opposite plasticity process.)

      Reason for using the FBP and NFBP

      The reason why the FBP and NFBP are particularly useful to demonstrate the thresholds’ roles is that the patterns in the tasks share features. (For example, in the FBP, the feature ‘strawberry’ is shared between both ‘red strawberry’ and ‘yellow strawberry’.) The synapses for these shared features experience precisely such competing LTP and LTD processes where the role of metaplasticity becomes evident. Specifically, the role of the LTD threshold is visible in shared synapses that need to be strengthened during learning (such as the shared ‘strawberry’ synapses in the FBP), while the LTP threshold’s role is visible in shared synapses that need to be weakened (there are no such synapses in the FBP, but there are in the NFBP, which is why it is needed in this article). In addition, feature binding is relevant for the striatum (now described in the introduction).

      In the first version of the article there are sentences that are likely confusing and misleading regarding the goal of the study, and these have been removed (lines 98-100 in the tracked changes file).

      The revised article now contains 4 new main figures and 15 new figure supplements prompted by the reviewers’ comments, thus expanding the scope of the study and hopefully clarifying its content, the mechanisms, and the conclusions.

      Finally, the question whether this rule can solve the NFBP with excitatory synapses alone is being addressed in a new, ongoing study, as written at the end of the Discussion section in the first version of this article. The new study tests 100+ SPN models (71 dSPN and 34 iSPN models differing in ion channel composition, with two different morphologies), different locations of the synaptic clusters along the dendrites (from proximal to distal), and two different learning regimes (suprathreshold – where the neuron initially spikes for all patterns, and subthreshold – where the neuron is initially silent for all patterns, as in this study) containing both clustered and distributed synapses in the same setup. Where necessary, preliminary results from this study are attached in the responses.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This computational modelling study addresses the important question of how neurons can learn non-linear functions using biologically realistic plasticity mechanisms. The study extends the previous related work on metaplasticity by Khodadadi et al. (2025), using the same detailed biophysical model and basic study design, while significantly simplifying the synaptic plasticity rule by removing non-linearities, reducing the number of free parameters, and limiting plasticity to only excitatory synapses. The rule itself is supervised by the presence or absence of a binary dopamine reward signal, and gated by separate calcium-sensitive thresholds for potentiation and depression. The author shows that, when paired with a strong form of dendritic non-linearity called a "plateau potential" and appropriate pre-existing dendritic clustering of features, this simpler learning mechanism can solve a non-linear classification task similar to the classic XOR logic operator, with equal or better performance than the previous publication. The primary claims of this publication are that metaplasticity is required for learning non-linear feature classification, and that simultaneous dynamics in two separate thresholds (for potentiation and depression) are critical in this process. By systematically studying the properties of a biophysically plausible supervised learning rule, this paper adds interesting insights into the mechanics of learning complex computations in single neurons.

      As mentioned in the introductory response above, the goal of the article was not to solve the NFBP, but to study the role of metaplasticity. The text has now been substantially revised to reflect that. One of the primary claims that “metaplasticity is required for learning non-linear feature classification, and that simultaneous dynamics in two separate thresholds (for potentiation and depression) are critical in this process” is now removed from the abstract. Instead, the abstract now focuses on the main results regarding the role of metaplasticity, as well as the results which arose from addressing the reviewers’ comments. However, I guess I should note that to be able to study the role of metaplasticity using the FBP and the NFBP, it was first necessary to have a rule that solves these tasks – that allows to later observe the effects of systematically fixing or modifying parameters in the rule. Another aspect which may have added to the confusion about the study’s goal may be due to the NFBP being a complex task requiring supralinear dendritic integration. Because of this, plateaus have received significant attention in the article, possibly diverting from the main focus. This is also now clarified in the article.

      Strengths:

      The simplified form of the learning rule makes it easier to understand and study than previous metaplasticity rules, and makes the conclusions more generalizable, while preserving biological realism. Since similar biophysical mechanisms and dynamics exist in many different cell types across the whole brain, the proposed rule could easily be integrated into a wide range of computational models specializing in brain regions beyond the striatum (which is the focus of this study), making it of broad interest to computational neuroscientists. The general approach of systematically fixing or modifying each variable while observing the effects and interactions with other variables is sound and brings great clarity to understanding the dynamic properties and mechanics of the proposed learning rule.

      It would indeed be great if this rule facilitates similar studies in other neuron types. The rule requires at least qualitatively accurate calcium dynamics, which at the moment are not widely available for other neuron types (but will likely appear in the future).

      Weaknesses:

      General notes

      (1) The credibility of the main claims is mainly limited by the very narrow range of model parameters that was explored, including several seemingly arbitrary choices that were not adequately justified or explored.

      The parameter range has now been expanded to address all comments below.

      (2) The choice to use a morphologically detailed biophysical model, rather than a simpler multicompartment model, adds a great deal of complexity that further increases uncertainty as to whether the conclusions can generalize beyond the specific choices of model and morphology studied in this paper.

      Regarding the plasticity rule and the role of metaplasticity, the conclusions should hold if the calcium dynamics of other (simpler or detailed) neuron models is qualitatively similar (i.e. shows the same monotonicity of “stronger synaptic input – stronger calcium”, and not something like Figs. 12A, B for which the rule was not tested).

      As to whether the NFBP can be solved, that is being addressed in the new study with 100+ neuron models (71 dSPN models and 34 iSPN models, and two morphologies: one for the dSPNs and another for the iSPNs). So far, the results show that the important ingredient is the threshold nonlinearity provided by the plateau potential (both dSPNs and iSPNs exhibit this nonlinearity, irrespectively of where the synaptic cluster is located, see Author response image 1).

      Also, plateaus can be evoked in passive SPN dendrites (as stated in the methods in lines 916-919, by “turning off” ion channels in the computational model as in Fig. S7 of Du et al. (2017) (see list of references provided)). Although plateaus are somewhat modified by active conductances, the main behavior important for the NFBP is given by NMDARs. The most striking contributions of the other ion channels in SPNs are the low resting potential around -85 mV, delayed first spike under current injection and strong inward rectification. Basically, SPNs require more input to be stimulated compared to other neurons that lack these properties. So, other neurons with NMDAR-dependent nonlinearities should in principle be able to solve the same task, but perhaps needing less inputs (provided dendritic integration is sufficiently supralinear – see response to comment (2) under Reviewer #2 (Recommendations for the authors)).

      (3) The requirement for pre-existing synaptic clustering, while not implausible, greatly limits the flexibility of this rule to solve non-linear problems more generally.

      In my view, the limitation of this rule is that there is no mechanism for structural plasticity with which clusters could be “grown”. For example, Hedrick et al. 2022 demonstrate that during learning in the motor cortex, existing functional clusters are “enlarged” by growing new spines. Devising a biologically realistic rule for structural plasticity is a separate research project, which would be very interesting to do.

      In any case, to solve the NFBP with excitatory synapses alone, supralinear voltage elevations are necesssary (Tran-van-Minh et al. 2015). When it comes to NMDAR-dependent supralinear integration, clustering is basically necessary, within a smaller, e.g. 20 micron dendritic stretch, or more loosely, within a single dendritic branch (Losonczy and Magee, 2006; Branco et al. 2010). Also, synaptic clustering is quite present across different brain regions (such as the sensory and motor cortices, pallium, striatum, hippocampus, brainstem, the developing cortex, and across species, e. g. see the introduction of Kirchner and Gjorgjieva, 2021), so I suppose it is not that implausible. Inputs from motor cortex to striatum are also clustered (Hwang et al. 2022).

      (4) In order to claim that two thresholds are truly necessary, the author would have to show that other well-known rules with a single threshold (e.g., BCM) cannot solve this problem. No such direct head-to-head comparisons are made, raising the question of whether the same task could be achieved without having two separate plasticity thresholds.

      Although the goal of the study was not to solve the NFBP, this is a very interesting question related to the question “what is the use of having two separate thresholds for plasticity”? I have added simulations in which the rule is modified to use a single calcium threshold instead, showing that the FBP, NFBP, and reversal learning are solved just as well (Fig. 12 and Figure 12 – figure supplements 1 – 4). The results show that two calcium thresholds allow for separate control over the LTP and LTD processes (lines 687-706).

      Specific notes

      (1) Regarding the limited hyperparameter search:

      (a) On page 5, the author introduces the upper LTP threshold Theta_LTP. It is not clear why this upper threshold is necessary when the weights are already bounded by w_max. Since w_max is just another hyperparameter, why not set it to a lower value if the goal is to avoid excessively strong synapses? The values of w_max and Theta_LTP appear to have been chosen arbitrarily, but this question could be resolved by doing a proper hyperparameter search over w_max in the absence of an upper Theta_LTP.

      Thank you for raising this question – it is closely related to the very important question of when to stop updating the weights. Before providing the response, I should clear up any confusion made due to an error in notation. In this study, , and , meaning that during learning the weights can increase up to twice (up to 100% of) the initial value, and can decrease down to 1% of the initial value. Currently, it says w<sub>max</sub> = 2, and w<sub>min</sub> = 0.01, which is a mistake, and possibly a reason why the chosen values seem arbitrary.

      The reason for introducing the upper threshold θ<sub>LTP</sub> is most easily seen in the parameter scan for w<sub>max</sub> done in the absence of θ<sub>LTP</sub> (Figure 4 – figure supplement 4) – it is very difficult to choose a value for w<sub>max</sub> with which all three tasks will be solved (FBP with linear integration, i.e. distributed synapses, the FBP with supralinear integration, i.e. clustered synapses, and the NFBP). Put differently, if there is a range of w<sub>max</sub> values which solves all three tasks, it is very narrow. So, θ<sub>LTP</sub> is introduced to solve all three cases without specially tuning w<sub>max</sub>.

      The following is a description of the results in Figure 4 – figure supplement 4, which is also present in the article in lines 383-396. The reason why a single preset value for w<sub>max</sub> does not work is that distributed synapses sum only linearly at the soma, while clustered synapses sum supralinearly. This means that to trigger somatic spiking, the distributed synapses need to be strengthened more compared to clustered synapses. Concretely, this means that a is needed to solve the FBP with linear integration (distributed synapses), but it is already too strong to solve the NFBP (Figure 4 – figure supplement 4, column four: top panel shows that FBP with linear integration is solved, while the bottom panel shows the average performance for the NFBP is right at the border – pink line aligns with the dashed line). On the other hand, is enough in the FBP with clustered synapses and the NFBP to trigger glutamate spillover and a plateau potential, which drives somatic spiking (Figure 4 – figure supplement 4, column two). However, it is not enough to solve the FBP with linear integration. Larger values for w<sub>max</sub> increase the performance on the FBP with linear integration, but lower the performance on the NFBP due to spiking for the irrelevant patterns (especially visible once ). As said, in this parameter scan, only the value solves all three tasks, with borderline performance on the NFBP (it is arguable whether the borderline performance means the NFBP is solved). This suggests that if a range of values for w<sub>max</sub> exists which solves all three tasks, it is very narrow.

      With the settings used in the study, the w <sub>max</sub> threshold for triggering glutamate spillover is a number between 1.2 and 1.3 with many decimals – 1.230769231. A possibility could be to program spillover to always occur at , but this is also not a general solution. For example, other dSPN models from the model library in Lindroos et al. (2021) have different excitability, and that would mean a different, specific value of w<sub>max</sub> would be needed for each neuron model.

      θ<sub>LTP</sub> is thus a solution that prevents “over-excitation” in a dendrite. The logic is that synapses in a cluster do not need to be as strong as distributed synapses to be effective in exciting the neuron. θ<sub>LTP</sub> ensures that once plateaus appear, synapses are not strengthened further. This allows using the same value across all scenarios, i.e. avoiding tuning of w<sub>max</sub> or tuning the number of synapses representing the features. (Also, preliminary results show this works well in the new study which tests all dSPN and iSPN models in the model library.) As for the value of θ<sub>LTP</sub>, it can be set anywhere in the “blank” region in Fig. 2C<sub>3</sub> where the supralinear jump in [Ca]<sub>NMDA</sub> occurs. What is important is that [Ca]<sub>NMDA</sub> evoked by a plateau should be above θ<sub>LTP</sub> – then no more strengthening will occur once plateaus can be evoked.

      (b) The author does not explore the effect of having separate learning rates for theta_LTP and theta_LTD, which could also improve learning performance in the NFBP. A more comprehensive exploration of these parameters would make the inclusion of theta_max (and the specific value chosen) a lot less arbitrary.

      I am not sure I have understood this comment properly, but hopefully this response addresses it adequately. In fact, the computer code does use separate variables for the learning rates for θ<sub>LTP</sub> and θ<sub>LTP</sub>, but for simplicity, they are set to the same value. In this response I have included some simulations in which they are set to different values (Author response image 2), but to understand their effect it is better to read the response to the comment under c) first. The metaplasticity rates determine how fast the synaptic weights are stabilized, i.e. how long the time window is where the synapses are “flexible”. When the thresholds reach the calcium amplitudes, the weights are “locked”.

      If the metaplasticity rates are too low, the weights take a longer time to stabilize and solving the tasks takes a long time (Figure 4 – figure supplement 10 from the response to c). If they are too high, weights are quickly stabilized and, in the case of weakened synapses, prematurely so, leading to worse performance on the NFBP due to spiking for the irrelevant patterns. (Figure 4 – figure supplement 11).

      One can use different values for the metaplasticity rates, as in the examples in Fig. 2 , and this affects the speed of learning. A large η<sub>θLTD</sub> = 20 quickly prevents the LTD process (Figs. 2D<sub>1</sub>, D<sub>2</sub>), causing faster synaptic strengthening, but also quickly limiting weakening (Figs. 2B<sub>1</sub>, B<sub>2</sub>) – as a result, after learning, the neuron spiked for an irrelevant pattern (Fig. 2A<sub>1</sub>). Conversely, a large η<sub>θLTD</sub> = 20 quickly prevents the LTP process (Figs. 2C<sub>3</sub>, C<sub>4</sub>), prolonging synaptic strengthening and allowing faster weakening (Figs. 2B<sub>3</sub>, B<sub>4</sub>) – as a result, the NFBP is solved, but takes a longer time. However, this does not remove the need for using an upper threshold θ<sub>LTP</sub>. If I have not understood the comment properly, please let me know so I can address it further.

      (c) Figure 4 Supplements 3-4: The author shows results for a hyperparameter search of the learning rule parameters, which is important to see. However, the parameter search is very limited: only 3 parameter values were tried, and there is no explanation or rationale for choosing these specific parameters. In particular, the metaplasticity learning rates do not even span one order of magnitude. If the author wants to claim that the learning rule is insensitive to this parameter, it should be explored over a much broader range of values (e.g., something like the range [0.1-10]).

      Thank you for bringing this up! I have expanded the parameter scan for both the learning rate η (Figure 4 – figure supplements 7 and 8) and the metaplasticity rate η<sub>θ</sub> (Figure 4 – figure supplements 10 and 11) to cover three orders of magnitude (when taken together with the results already present in the article in Figure 4 – figure supplements 6 and 9).

      Figure 4 – figure supplements 7 and 8: The learning rate η was was tested with the values 0.1 and 20. This shows a known result in other plasticity rules, which is that if the learning rate is too fast or too slow, learning is not successful or is too slow. For η = 0.1, learning takes very long (around 2000 patterns for FBP with linear integration), and for the NFBP is at the borderline score (Figure 4 – figure supplement 7E). As the LTD thresholds rise, they will prevent the slowly decreasing synaptic weights from decreasing sufficiently. For η = 20, the weight updates are too large and no pattern is stored because calcium quickly falls below the thresholds as a result of such large weight updates (Figure 4 – figure supplement 8E). This leaves some intermediate range of learning rates that “works”. (The three values {0.4, 0.85, 1.7} that were tested in Figure 4 – figure supplement 6 are roughly a factor of 2 apart.) Results are described in more detail in lines 409-414 of the revised article.

      Figure 4 – figure supplements 10 and 11: The metaplasticity rate η<sub>θ</sub> was tested with the values 0.2 and 20. This further demonstrates the role of metaplasticity as a “lock” on the learning process. η<sub>θ</sub> determines how fast the weights stabilize (i.e. how long they remain flexible for learning). With the value η<sub>θ</sub> = 0.2 the weights take a longer time to stabilize, because it takes a long time for the thresholds to reach the calcium amplitudes (Figure 4 – figure supplement 10). This affects the learning time on the NFBP only. In contrast, with the value η<sub>θ</sub> = 20, the thresholds quickly reach the calcium amplitudes, and quickly stabilize the weights. While the neuron does learn to evoke plateaus, the weakened synapses are stabilized too early, resulting in spikes for the irrelevant patterns and lowering the performance on the NFBP. (Figure 4 – figure supplement 11). Results are described in more detail in lines 415-429 of the revised article.

      (2) Regarding the similarity to BCM, the author would ideally directly implement the BCM learning rule in their model, but at the least the author could have shown whether a slight variant of their rule presented here can be effective: for example having a single (plastic, not fixed) Cadependent threshold that applies to both LTP and LTD, with a single learning rate parameter.

      - As mentioned above, this is also a very interesting question, related to the important question of “what is the use of having two separate calcium thresholds?”. I have added simulations according to your second suggestion – using a single calcium threshold that applies to both LTP and LTD. A single threshold can solve the FBP, NFBP, and reversal learning equally well (Fig. 12 and Figure 12

      - Figure Supplements 1, 2). This indeed raises the question what are two separate thresholds useful for? Figure 12 – figure supplement 3 shows that when keeping the single calcium threshold fixed, none of the shared synapses stabilize. Comparing this to Figs. 5 and 6, where the LTP and LTD threshold are kept fixed one at a time, it shows that having two separate thresholds allows separate control over the LTP and LTD processes.

      I also tried an implementation of the BCM rule using a nonlinear (quadratic) function in the metaplasticity rule for updating the threshold:

      where c<sub>norm</sub> is a normalization constant set to 80 μM for [Ca]<sub>NMDA</sub> and 20 μM for [Ca]<sub>L-type</sub> (close to the maximal concentrations achieved throughout the article). In the BCM rule, a nonlinear function is necessary to prevent runaway growth (or collapse) of the weights. So, in principle, w<sub>max</sub> and θ<sub>LTP</sub> should be unnecessary with this quadratic metaplasticity rule. However, even though the weights seem to stabilize (around a high value), they are too strong and plateaus are triggered for all patterns (see Author response image 3). It is possible that only a narrow range of parameters in the quadratic function will allow weights to stabilize at appropriate values, similarly to the narrow range in w<sub>max</sub> in the absence of θ<sub>LTP</sub> (as described for comment 1a) above).

      Below are some questions about how to “correctly translate” the BCM rule from its formulation for rate-coded inputs to a formulation for sparsely-coded inputs which are used in this article. I did not address these questions in detail. The following is meant for readers who might be interested in pursuing these questions.

      Questions when translating the BCM rule for rate-coded inputs to a rule for sparsely-coded inputs:

      The similarity of this study’s rule to BCM is only qualitative – a sliding threshold was added that follows an indicator of synaptic activity (in this study it is the amplitude of the calcium concentration). The original BCM rule is defined for rate-based synapses, and its threshold θ is a nonlinear function of the average (synaptic) activity:

      where p and c<sub>0</sub> are positive constants, and c) is the average activity. The threshold θ is updated as the average activity c) changes (the averaging being done over a time interval) [1]. Changes in the rate-coded inputs drive changes in c) continuously, thus changing θ over time.

      (1) What should be chosen as the average activity c̅ ?

      In the BCM model, the average activity, and consequently the threshold θ, are global for the whole neuron (which is represented just by its firing rate), meaning that θ is the same for all synapses. In a morphologically realistic neuron model, what should be the average activity? Especially in the case of supralinear integration occuring in a dendrite, one could also consider average activities in dendrites. In the dSPN model, the calcium concentration decays by the time the next pattern arrives, so one probably needs another, low-pass filtered variable of the calcium concentration (or voltage), with a much longer time constant than that of the calcium decay. (I have chosen to slide the thresholds towards the amplitude of the calcium evoked by a pattern, which could be viewed as an approximation of such a long-lived variable. This is one possible “translation” of this aspect of the BCM rule.)

      (2) How to update the chosen indicator of activity c̅?

      Assuming an adequate indicator is chosen (different from the calcium amplitude that I have chosen), how should it be updated when using sparsely-coded inputs? Should it evolve continuously, or only when the sparsely-coded inputs are active? If updated continuously, it should not decay very fast when synapses are inactive (not activated by a pattern), while still change fast enough to detect changes in synaptic weight occuring due to learning. (I have chosen to slide the thresholds towards the evoked calcium amplitude only when synapses have been activated, so this is also one way of “translating” this second aspect.)

      (1) which in the BCM article is replaced with the average over the input’s probability distribution

      (3) What nonlinear function to use to update the threshold?

      The BCM rule allows any function with p > 1, but how fast the synapses stabilize will likely depend on c<sub>0</sub>, the normalization constant in the nonlinear function, and the power p. The role of the nonlinear function is to ensure stability of the synaptic weights (prevent runaway growth), meaning no θ<sub>LTP</sub> nor w<sub>max</sub> should be needed. However, achieving high performance on all three tasks might still require a specific tuning of these parameters, as w<sub>max</sub> does if no θ<sub>LTP</sub> is used. That is, these parameters will need to be chosen such that the threshold quickly goes up once a plateau appears, so weights are not allowed to strengthen further. Having in mind that there is a very narrow range in w<sub>max</sub> where all three tasks are solved, there might similarly be only a narrow range in c<sub>0</sub> and p.

      Since the goal of the article was to study the role of the two thresholds (not to optimize the rule’s parameters on the three tasks), I have only tried a quadratic function for the metaplasticity rule, and did not search for a region of parameters that could work.

      (3) This paper is extremely similar (and essentially an extension) to the work of Khodadadi et al. (2025). Yet this paper is not mentioned at all in the introduction, and the relation between these papers is not made clear until the discussion, leaving me initially puzzled as to what problems this paper addresses that have not already been extensively solved. The introduction could be reworked to make this connection clearer while pointing out the main differences in approach (e.g., the important distinction between "boosting" nonlinearities and plateau potentials).

      The last paragraph in the introduction now makes this comparison, highlighting the differences in approach, and mainly emphasizing that the goal of the two studies is different. The open question that this study addresses is to pinpoint the roles of the two calcium thresholds (present in dSPNs) in learning. Also see the response to the similar comment (1) from Reviewer #2 (Public review): Weaknesses.

      (4) The introduction is missing some citations of other recent work that has addressed single-neuron non-linear computation and learning, such as Gidon et al (2020); Jones & Kording (2021).

      Thank you for mentioning this. Since the first version of the article caused confusion regarding the goal of the study (which is to study metaplasticity), I avoided focusing on nonlinear computation in the introduction, and instead added these references to the Discussion (in line 850 the section regarding nonlinear computation).

      (5) Figure 1: The figure prominently features mGluR next to the CaV channel, but there is no mention of mGluR in the introduction. The introduction should be updated to include this.

      Thank you for noticing this. It is now included in the introduction in lines 50-53.

      (6) Could the author explain why there is a non-monotonic increase/decrease in the [Ca]_L in Figure 2B_4? Perhaps my confusion comes from not understanding what a single line represents. Does each line represent the [Ca] in a single spine (and if so, which spine), or is each line an average of all the spines in a given stim condition?

      Thank you for noticing this, that information was missing from the figure caption. The line is from a single spine which is placed at a random location on a randomly chosen dendrite. Because each line comes from a different trial with a different number of synapses, the spine it corresponds to is in a different location, resulting in different voltage and different calcium signals when using distributed synapses. (Clustered synapses are placed at approximately the same somatic distance, despite being on a randomly chosen dendrite, so this effect does not appear in Figs. 2C<sub>3</sub>, C<sub>4</sub>). The figure caption has been updated accordingly.

      (7) Row 124 (page 4): L-type Ca microdomains (in which ions don't diffuse and therefore don't interact with Ca_NMDA) is a critical assumption of this model. The references for this appear only in the discussion, so when reading this paper, I found myself a bit confused about why the same ion is treated as two completely independent variables with separate dynamics. Highlighting the assumption (with citations) a bit more clearly in the results section when describing the rule would help with understanding.

      Thank you for pointing this out! This assumption is now clearly stated with citations in lines 174176 when characterizing the plateaus in SPNs, and lines 205-207 when describing the rule (in addition to the existing description in the Methods).

      (8) Row 149 (page 5): The current formulation of the update rule is not actually multiplicative. The fact that the update is weight-dependent alone does not make it a multiplicative rule, and judging by equation (1) it appears to simply be an additive rule with a weight regularization term that guarantees weight bounds. For example, a similar weight-dependent update is also a core component of BTSP (Milstein et al. 2021; Galloni et al. 2025), which is another well-known *additive* rule. An actual multiplicative rule implies that the update itself is applied via a multiplication, i.e. w_new = w_old * delta_w

      For an example of a genuinely multiplicative rule, see: Cornford et al. 2024, "Brain-like learning with exponentiated gradients"). Multiplicative rules have very different properties to additive rules, since larger weights tend to grow quickly while small weights shrink towards 0.

      Thank you for explaining the difference between “additive” and “multiplicative”. I have removed the word “multiplicative” from the description of the rule, as it is not necessary. (In the first version of the manuscript I followed the nomenclature by Gütig et al. (2003), which refers to the weight-dependent update as a multiplicative rule, while an additive rule has no weight dependence in the equation for Δw, which is Eq. 2 in Gütig et al. (2003).)

      (9) Equation 1 (page 5): Shouldn't the depression term be written as: (w_min - w)? This term would be negative if w is larger than w_min, leading to LTD. As it is written now, a large w and small w_min would just cause further potentiation instead of depression.

      Yes, a minus sign is missing in front of the depression term. Thank you for noticing this! It is now corrected.

      (10) In the introduction, the teaching signal is described in binary terms (DA peak, or DA pause), but in Equation 1, it actually appears to take on 3 different values. Could the author clarify what the difference is between a "DA pause" and the "no DA" condition? The way I read it, pause = absence of DA = no DA

      Yes, the “no DA” condition should say “baseline DA”. The DA signal does indeed take three different values. In the striatum, there is a baseline tone of dopamine. The DA peaks are transient increases from this tone, and the DA pauses are transient decreases. The basal dopamine level is now introduced in line 113 of the introduction, so that Eq. 1 should be clearer. In the computer code, DA is in effect a ternary signal, with three values, +1 for DA peak, -1 for DA pause, and 0 for baseline DA.

      (11) Figure 3: In these experimental simulations, DA feedback comes in 400ms after the stimulus. The author could motivate this choice a bit better and explain the significance of this delay. Clearly, the equations have a delta_t term, but as far as the learning algorithm is concerned, it seems like learning would be more effective at delta_t=0. Is the choice of 400ms mainly motivated by experimental observations? On a related note, is it meaningful that the 200ms delta_t before the next stimulus is shorter than the 400ms pause from the first stimulus? Wouldn't the DA that arrives shortly before a stimulus also have an effect on the learning rule?

      These time windows were chosen purely to keep the simulations as short as possible (since a computing cluster was used to run many trials). In reality, two patterns would probably arrive after a longer time window, as an animal should reach for the pattern (and in the case of these fruit patterns, eat it) which would last on the order of seconds. However, the calcium signals that trigger plasticity in these simulations are shorter than that (Figs. 2B<sub>3,4</sub>, 2C<sub>3,4</sub>), and only the amplitude of the signal is used to trigger plasticity. This gives an opportinity to not have to wait for several seconds between two patterns arrive (as might occur in reality), but shorten that interval and so shorten the simulation time. (This is now stated in lines 245-247 in the Results section where the rule is described, pointing to the Methods for more detials.)

      Hence, the dopamine feedback is given at 400 ms after a pattern appears because the calcium amplitudes would have surely occurred by then (even if a plateau were evoked, see Figs. 2B<sub>3,4</sub>, 2C<sub>3,4</sub>). Similarly, the next pattern is given at 600 ms because the calcium signals would have decreased to baseline by then. (This is already explained in the Methods section in lines 1101-1117). The key implicit assumption here is that the calcium amplitude is somehow remembered by the synaptic circuitry. In experiments, dopamine feedback up to 2 seconds after synaptic stimulation can cause LTP, so it is OK to assume this (Yagishita et al., 2014). However, dopamine signaling before synaptic stimulation has no effect (Yagishita et al., 2014). (This assumption is now clearly stated in the Methods in lines 902-906.)

      Lastly, in this rule, if the dopamine feedback coincides with the pattern, then whatever calcium levels were reached at that time would be used to check whether the thresholds are crossed, and these will most likely not be the maximal calcium levels (these are reached some time later, as seen in Figs. 2B<sub>3,4</sub>, 2C<sub>3,4</sub>).

      One last note is that from your comment I got the impression that you may have understood Δt to be the time until dopamine feedback is provided. It is not, it is simply the time step in the simulation (already stated in line 211; now an additional explanation is added in lines 211-214, just in case).

      (12) Figure 4C: How is it possible that the theta_LTP value goes higher than the upper threshold (dashed line)? Equation 3 implies that it should always be lower.

      This is also just a choice of implementation: θ<sub>LTP</sub> evolves independently of θ<sub>LTP</sub>, reflecting an implicit assumption that different molecular machinery implements θ<sub>LTP</sub> and θ<sub>LTP</sub>. This assumption is similar to the findings in Figs. 4, 5 in Ngezahayo et al. 2000, where in weakened synapses the voltage threshold for LTD can increase (to -20 mV) above the voltage threshold for LTP (at -30 mV). (This assumption is now explicitly written out in lines 316-322.)

      (13) Row 429 (page 11): The statement that "without metaplasticity the NFBP cannot be solved" is overly general and not supported by the evidence presented. There exist many papers in which people solve similar non-linear feature learning problems with Hebbian or other bio-plausible rules that don't have metaplasticity. A more accurate statement that can be made here is that the specific rule presented in this paper requires metaplasticity.

      Yes, that is what is meant with that statement – that this specific rule requires metaplasticity. The text has been corrected. Thank you for noticing this.

      (14) The methods section does not make any mention of publicly available code or a GitHub repository. The author should add a link to the code and put some effort into improving the documentation so that others can more easily assess the code and reproduce the simulations.

      The link to the code is given in the Data availability statement. However, this statement might not appear in the manuscript files that you are receiving (it is available online and in the article PDF downloadable from eLife), so I am providing the link here, as well. The Readme.md file now describes the code, and how to reproduce figures in the article.

      https://github.com/danieltrpevski/Plasticity/tree/metaplasticity

      Reviewer #2 (Public review):

      Summary:

      The manuscript proposes interesting synaptic plasticity rules grounded in experimental data. Its main features are:

      (1) plasticity depends on local calcium concentration driven by presynaptic activity and is independent of somatic action potentials,

      (2) the rules incorporate metaplasticity, and

      (3) they demonstrate how a single neuron could address the feature-binding problem at the dendritic level.

      The work extends a previous study (https://doi.org/10.7554/), to which the author also contributed.

      The author models two calcium thresholds (LTP/LTD) from two different calcium sources (NMDA/VGCC), and these thresholds are flexible (metaplasticity rule, similar to BCM), which is claimed to be necessary for successful learning of both FBP and NFBP (linear and nonlinear feature binding problem with 1 or 2 patterns). The role of each threshold seems to be opposite and complementary. One extra condition has been added: an upper threshold for LTP. This threshold serves to stop synaptic strengthening once synapses are strong enough to evoke a plateau. With that, synapses are not strengthened to the maximal value, avoiding strong supralinear integration for irrelevant patterns.

      This is summarizes the learning rule well. What I should probably add is that, as mentioned in the introductory response, the article’s first version may cause some confusion about the study’s goal. The goal was to study the role of metaplasticity in the two thresholds, not to solve feature binding in particular. Nevertheless, feature binding is a very useful task to demonstrate the roles of each threshold because the patterns share features: the synapses for the shared features experience both LTP and LTD, and metaplasticity is necessary to express (or to converge to) just one plasticity state (either LTP or LTD). More specifically, the threshold for one plasticity process (e.g. LTD) is necessary to express the opposite process (e.g. LTP). However, to be able to study metaplasticity using feature binding, a prerequisite is to have a rule that can solve it. (So, in that sense, it was first necessary to show that the rule solves the tasks before exploring the roles of metaplasticity. The article has been revised to clearly state this.)

      Strengths:

      The current model implements not only local synaptic plasticity but also metaplasticity and solves the FBP at the dendrite level. Another strong aspect of the model is that metaplasticity in the LTD threshold protects strengthened synapses from weakening. In this way, as the author mentioned, metaplasticity is able to protect learned patterns from being forgotten or weakened and prevent irrelevant patterns from being stored. This is a nice modelling example of metaplasticity being helpful in preventing the catastrophic interference or forgetting (as has been explicitly discussed in a recent article https://doi.org/10.1016/j.). The author might want to briefly mention or emphasize this aspect of the model, which might be interesting also for the AI community.

      This is a great suggestion, thank you! I have instead emphasized a related aspect of the model: that it can solve the plasticity-stability dilemma (when taking away the closed-loop implementation of metaplasticity). I did not focus on catastrophic forgetting, because in machine learning it has a slightly more specific meaning: forgetting that occurs in sequential learning of different tasks. Instead, I exemplified the plasticity-stability dilemma using reversal learning (the reward policy is switched in the middle of the simulation), which can be viewed as a special case of sequential learning of opposite tasks. So, strictly speaking, there is no demonstration that metaplasticity prevents catastrophic forgetting (in the more general sense), but there is a demonstration that it retains (stabilizes) what has been learned (within the same task). The latter does suggest that metaplasticity could help prevent catastrophic forgetting, as well.

      Weaknesses:

      (1) What is novel in the current paper as compared to Khodadadi et al. eLife 2025? That is not completely clear and should be made clearer. Is it only a minor difference related to the fact that the new learning rule has metaplasticity in both calcium thresholds and is simpler? This seems to be just an incremental increase in knowledge/methods. Can the author defend his paper against this point from the „devil's advocate"? How is the conclusion of the author in the abstract that „metaplasticity in both thresholds is necessary" reconcilable with his previous publication (Khodadadi et al. eLife 2025), in which only metaplasticity in one threshold was successful in solving the nonlinear feature binding problem?

      Thank you for raising this question, since the answer was not clear in the first version. The study may be an incremental increase in methods, but is (in my view) a solid increase in knowledge, giving concrete new insights into metaplasticity (an area that is still comparatively little understood). The two main conclusions are:

      (1) In synapses that are exposed to both the LTP and the LTD process, metaplasticity is necessary to ultimately express just one of them (converge to either LTP or LTD).

      (2) Metaplasticity in the threshold regulating one plasticity process is necessary for expressing the opposite process (metaplasticity in the LTD threshold is necessary for expressing LTP, and vice versa).

      The article has been substantially revised to clearly state these main conclusions (as well as the study’s goal), while a full list of all conclusions is given in points 1-7 the Discussion.

      The conclusion in the abstract that „metaplasticity in both thresholds is necessary" is stated too generally, and it was meant to refer to this rule only. I have now removed it from the abstract, since the study’s goal was not to solve feature binding. Instead, the focus is on the two conclusions above as well as other conclusions obtained from expanding the study with reversal learning and using a single threshold (prompted by the comments from yourself and reviewer #1).

      Lastly, the differences compared to Khodadadi et al. 2025 (eLife) are:

      (1) Yes, the rule is simpler and has metaplasticity in both calcium thresholds, instead of just one in Khodadadi et al. 2025 that operates according to a different mechanism (where the threshold, along with the entire LTP plasticity kernel, moves in a direction opposite of the calcium signal).

      (2) Plateau potentials are used here, versus the “boosting” nonlinearities in Khodadadi et al. 2025.

      (3) The most detailed calcium diffusion model to date for SPNs is implemented in this study (by Dorman et al. (2018)). This was a necessary addition to avoid non-monotonous increases in calcium (shown in Fig. 13A, B). No calcium diffusion between neuronal compartments is implemented in Khodadadi et al., 2025. This is indeed a small methodological difference, but not trivial to implement.

      The relation to Khodadadi et al., 2025 is now stated both in the introduction in lines 137-144, and at greater length in the discussion in lines 856-870.

      Perhaps I should note that the two learning rules (this one and that in Khodadadi et al. 2025) developed in parallel, this one being based on formulations similar to Gütig et al. 2003 by adding metaplasticity, while the rule in Khodadadi et al. 2025 was inspired from the ideas in Schiess et al. (2016) and Urbanczik and Senn (2014) (see list of references provided). The two studies have different goals (and in that sense one was not meant to be an extension of the other):

      - Khodadadi et al. (2025) studies whether the NFBP can be solved by SPNs,

      - this article studies the role of metaplasticity in the two calcium thresholds that exist in dSPNs; it uses feature binding as a task because the role of metaplasticity is exposed precisely by the shared features in the task (they cause shared synapses to undergo competing LTP and LTD, and metaplasticity is necessary to direct synapses into just one outcome); importantly, feature binding is also relevant for the striatum (now described in the introduction in lines 126-136)

      Finally, to obtain the conclusions in this study about the roles of the two thresholds, it was necessary to devise a suitable rule. They cannot be obtained neither with the rule in Khodadadi et al. 2025, nor with any other rule, because none have a formulation with two separate thresholds for LTP and LTD. Phrased more generally, one needs a new/different rule (new/different assumptions) to obtain new/different conclusions, even if the new rule may be related to existing ones. That said, a rule with a single threshold would have reached conclusion 1. above, as is now shown in the article in Fig. 12 and Figure 12 – figure supplements 1 – 4, but SPNs have two thresholds, prompting the use of a rule with two thresholds. Also, although not demonstrated with the cascade model (by Fusi et al. (2005) and its extensions), conclusion 1. should be obtainable with it, and is implied by the dynamics of the cascade model (which uses one threshold to control transitions between strong and weak synapses). More on the cascade model is given below and in the revised article.

      Hopefully that provided a clearer answer. If you have more questions, please let me know so I can try to clarify further.

      (2) As far as I can judge without testing the model, metaplasticity causes thresholds to monotonically increase during systematic pattern presentation, which stabilizes weights and allows pattern separation. Due to the closed-loop nature of the current implementation, where metaplasticity only happens if plasticity happens, this also effectively locks patterns in place. However, flexible learning is an essential mechanism for survival. Imagine a mutation event takes place and bananas suddenly become red and/or strawberries turn yellow. It seems that the current model would be unable to adapt to these new patterns even if rewards were to be shifted. While out of the scope of the study, due to its importance, I feel that pattern shifting/relearning should at least be briefly discussed. How could the model be improved to allow relearning?

      This is a very important question, and I think it is in fact within the scope of the study, so it is now addressed using reversal learning as a task (the reward policy is switched in the middle of the simulation, after the FBP and NFBP are learned). As you pointed out, the closed-loop formulation of metaplasticity effectively locks the patterns in the dendrites, and reversal learning cannot be solved (Fig. 9). On the other hand, if the conditions for metaplasticity are relaxed by allowing any calcium levels to trigger metaplasticity (i.e. no longer have a closed-loop formulation requiring that [Ca]<sub>NMDA</sub> or [Ca]<sub>L-type</sub> be above their thresholds), reversal learning in both the FBP and NFBP is solved (Fig. 10 and Figure 10 – figure supplements 1 – 3). (Note that [Ca]<sub>NMDA</sub> or [Ca]<sub>L-type</sub> still need to be above their thresholds for plasticity to occur.)

      Recommendations for the authors:

      Reviewing Editor Comments:

      While the reviewers were very positive, they highlighted limitations on the strength of evidence. Addressing these points could help revise this assessment.

      Reviewer #1 (Recommendations for the authors):

      (1) Framing the problem as "feature binding" is easy to understand, but it's a slightly narrow view of learning in general. Some mention of how NFBP relates to the XOR problem in the introduction would help you relate your solution to the much broader class of computational problems, since XOR is a more fundamental computational primitive. Non-linear feature learning is very general and relates to all machine learning problems and to computations that occur across every brain region.

      Since the goal of the study was not to solve the NFBP or feature binding (but to study metaplasticity), the relation of the NFBP to the XOR is put in the Discussion instead (to avoid further confusion about the study’s goal). Nevertheless, this connection is clearly made in the introduction of the new study that uses 100+ SPN models.

      (2) Since the requirement for clustered synapses is one of the main limitations of this learning rule, it would help if there were some discussion of how clustering might occur (e.g., whether there are any other related learning rules or mechanisms that might promote clustering, or if you are assuming that this has to occur entirely through developmental wiring).

      This is now discussed in lines 817-827 and related to developing a learning rule that would incorporate structural plasticity.

      (3) Page 5, row 157: Another potentially relevant citation here is Bittner et al. 2017, showing that a dendritic plateau potential in silent CA1 neurons drives place field formation without any prior spiking activity.

      Thank you, that is indeed a very relevant citation here! It is now added in line 222.

      (4) Figure 4C: It would be helpful to add a sentence in the figure legend explaining what the dashed lines represent (even though you also mention it in the main text).

      Added, thank you!

      (5) Across all figures, the author should consider using color combinations that are more colorblindfriendly (e.g., cyan/red). As a R/G colorblind person, I find it slightly difficult to see which line is which in the panels that compare scores in the NFBP.

      The panels showing the NFBP scores have been changed throughout the article, and all other figures have been checked. If there are any more difficult color combinations, please let me know.

      (6) The conventional, widely used abbreviation for dopamine is "DA", not "Da".

      Now changed.

      (7) A few typos I spotted in the paper:

      (a) Row 117 (page 3): missing parenthesis around citation

      (b) Row 173 (page 5): "uner" --> under

      (c) Row 248 (page 7): repeated sentence "...are first seen above..."

      (d) Row 255 (page 7): "NBFP" --> NFBP

      These are now fixed, thank you for reporting them!

      Reviewer #2 (Recommendations for the authors):

      (1) How are synapses distributed in the nonlinear integration case? Are pattern combinations branch-specific, or are features distributed randomly across the whole dendritic tree? Does this matter in any way? In any case, it should be clarified.

      Yes, they are branch-specific. In the NFBP, the feature combinations from Figure 3 – figure supplement 1 (and later from Figure 10 – figure supplement 1) are placed on two dendrites, chosen at random from 8 dendrites, in a 20-micrometre dendritic stretch starting around 120 micrometers away from the soma. 12 different trials are run for each feature combination, meaning that 12 randomly chosen pairs of dendrites were tested for each feature combination. In the FBP, the clusters are placed in one dendrite chosen at random from 8 dendrites, at the same distance from the soma as in the NFBP. This information is now added in the caption of Fig. 3.

      The idea behind placing clusters on the same dendrite is that, after learning, two strong clusters will produce a plateau (e.g. ‘red’ and ‘strawberry’), while one strong and one weak cluster will not (e.g. ‘yellow’ and ‘strawberry’), as in Figs. 3C<sub>2</sub>, 3C<sub>3</sub> (with the weights shown in Figs. 4B<sub>2</sub>, 4B<sub>3</sub>). For this the clusters need to be in the same branch, so that two features are “bound” together “into” a larger, plateau-evoking cluster. Also, a strong and a weak cluster should evoke a sufficiently lower somatic amplitude so that any additional noise will not cause somatic spiking. This is assured by the large voltage jump in the plateau (the all-or-none quality of the plateaus, Fig. 2D<sub>1</sub>).

      If the clusters are distributed randomly on different dendrites, it might or it might not work. Fig. 10 of Oikonomou et al. (2012) shows that dendritic spikes (that individually do not evoke somatic spiking) summate sublinearly at the soma in pyramidal neurons: weak + weak cluster = no spiking (Fig. 10A); strong + strong = no spiking (Fig. 10B); but also strong + strong = spiking (Fig. 10C). To reliably solve the NFBP, a large enough supralinearity is necessary instead, so one should test how the soma integrates dendritic spikes/plateaus in SPNs. As the goal of this study was not to solve the NFBP, this is left for the future (perhaps within the new ongoing study focusing on the NFBP).

      (2) To produce dendritic plateau potentials, the model (as in the previously published model - Trpevski et al. 2023) implements glutamate spillover activating extrasynaptic NMDARs. This is an interesting mechanism, but is there empirical evidence supporting this way of generating plateau potentials? Are synaptic NMDARs not sufficient? If not, which experiments support the role of nonsynaptic ones?

      This is still an ongoing area of research, indicating that glutamate spillover is regulated by reuptake from glial cells, which could even reverse function and excrete glutamate (Rusakov and Stewart, 2021; Malarkey and Parpura, 2014).

      Except for the experiments from Szapiro and Barbour, 2007 showing that climbing fibers in the cerebellum signal to molecular-layer interneurons exclusively through glutamate spillover, there are no other “in-vivo-like” experimental conditions which show that spillover activates extrasynaptic NMDARs (eNMDARs). The strongest other in vitro data come from the following experiments:

      (1) Chalifoux and Carter (2011), where glutamate reuptake by transporters was blocked with TBOA, which promoted NMDA spikes evoked by synaptic stimulation, suggesting that glutamate spillover activates eNMDARs. A similar experiment was done in Suzuki et al. (2008), where, in addition, synaptic NMDARs were blocked, leaving only eNMDARs available to trigger plateaus.

      (2) Glutamate iontophoresis, which ejects glutamate directly into the extrasynaptic space, activates plateaus once the iontophoretic current is strong enough (Oikonomou et al., 2012), with or without blocking glutamate reuptake (Suzuki et al. 2008).

      (3) Repetitive synaptic stimulation is thought to produce glutamate spillover (Suzuki et al. 2008; Oikonomou et al., 2012). For example, two synaptic shocks trigger NMDA spikes, while 5 synaptic shocks trigger plateaus in pyramidal neurons (Oikonomou et al., 2012).

      On the other hand, studies that employ glutamate uncaging at spines, which should activate eNMDARs much less, evoke NMDA spikes instead (Losonczy and Magee, 2006; Branco et al. 2010). Compared to plateaus, these have much smaller amplitudes, and the size of the supralinearity (the voltage jump) is much smaller than in the plateaus obtained with glutamate iontophoresis (compare Fig. 3 in Losonczy and Magee, (2006), Fig. 3B in Branco et al. (2010) to Fig. 4C in Oikonomou et al. (2012), Fig. 6 in Oikonomou et al. (2014)).

      The reason why glutamate spillover was implemented is precisely this robust all-or-none quality of the plateaus, i.e. the large jump in somatic voltage (Fig. 2D<sub>1</sub>) after a threshold level of stimulation (when the stimulus intensity increases in equal amounts). Without spillover, plateaus are graded in amplitude (and duration, Fig. 2B<sub>1</sub> in Trpevski et al. (2023)), and the NFBP cannot be solved (Fig. 6 in Trpevski et al. (2023)).

      (3) The check for dependence on initial conditions of theta_LTP/LTD is missing. Especially for the conditions of the thresholds being fixed. How do you decide what fixed value to use? How about having a corresponding limit for weights and not for thresholds? Why is there still an increase in threshold going on even when weights reach the maximum?

      The initial conditions for θ<sub>LTP</sub> and θ<sub>LTD</sub> are also an important question. The idea behind starting with low values of the thresholds (or fixing them to low values) is to make the synapses flexible for learning. One can view metaplasticity’s role as a “lock” on the weights, locking them once learning is done, and unlocking them (making them flexible) when something needs to be learned. Which signals determine when to lock or unlock the weights is what the metaplasticity rule implements, and is insufficiently understood experimentally. In any case, it makes sense to start with flexible synapses (low thresholds) at the beginning of learning; otherwise, no/little learning will take place (Figure 9 – figure supplement 2, the panels for “thresholded” metaplasticity). Throughout the article I choose the initial values of the thresholds to be lower than the calcium evoked by weakened synapses, ensuring that any calcium signal will trigger plasticity (but higher values also work).

      Instead of doing a scan over the initial conditions of θ<sub>LTP</sub> and θ<sub>LTD</sub>, I chose to show the following:

      (1) An example with initial threshold values higher than the calcium amplitudes: these “lock” the weights (Figure 9 – figure supplement 2, panels for “thresholded” metaplasticity) from the start and no learning can occur.

      (2) An example initialized with high values for the thresholds, but where metaplasticity can be triggered by any calcium amplitudes (no closed-loop in the metaplasticity rule, Figure 9 – figure supplement 2, panels for “relaxed” metaplasticity): here the thresholds adapt and learning can take place afterwards.

      This is supposed to illustrate a mechanism that can “unlock” weights. A scan over the initial conditions of θ<sub>LTP</sub> and θ<sub>LTD</sub> will give the lowest calcium level that the thresholds can be initialized to so that “thresholded” metaplasticity will work. But it is not important to know these precise values to obtain the conclusions in the article, which is why I did not do such a scan. (The above is treated in the article in the section on reversal learning, lines 569-577 and 628-633.)

      Also, when the rule is used for learning, it is not meant to have fixed thresholds, as synapses cannot stabilize and nothing will be stored (Figs. 5, 6, and 8). Fixing the thresholds to a low value was meant to show what happens without any metaplasticity (or if metaplasticity were “dysfunctional”). If in reality a mechanism exists to fix the thresholds to a low value or a high value, fixing would be temporary (until necessary to keep the weights flexible or locked, respectively). 

      I am probably not understanding this part of the comment: “How about having a corresponding limit for weights and not for thresholds?” There are two limits, the maximal and minimal weights in the rule w<sub>max</sub> and w<sub>min</sub>, and are initialized to random values within the interval [0.3, 0.35], but this is probably not what you mean.

      When the weights reach their maximum, the thresholds still increase because they have not reached the calcium amplitudes evoked by the maximal weights. (This is most visible with supralinear integration, where strengthened weights evoke plateaus Fig. 4C<sub>2</sub>-C<sub>4</sub>.) The thresholds move towards the calcium amplitudes with a rate η<sub>θ</sub>, which if made very high (as in Figure 4 – figure supplement 11), will make the thresholds rise as fast as the weights. But this is not good, as it will “lock” the learning processes in the synapses too soon (i.e. shorten the window where they are flexible). Neither is a too low η<sub>θ</sub> good, as then the synapses are flexible for a very long time and learning of the NFBP is prolonged (Figure 4 – figure supplements 10B<sub>3</sub>, 10B<sub>4</sub>, 10E<sub>3</sub>, learning simulations last longer than in Fig. 4).

      (4) Calcium modelling: The author writes that the largest voltage plateaus do not correspond to the largest calcium amplitude because of the plateau's voltage approaching the NMDAR reversal potentials, which might be a problem for the plasticity rule. To solve this, the author tried to implement a monotonic increase of calcium with increasing plateau by implementing axial calcium diffusion and buffering. My question is, is the monotonic increase realistic? Is not the NMDAR reversal effect, described above, a realistic scenario in dendrites?

      It seems to be realistic. The closest experiment that I know of is in Fig. 2C<sub>2</sub>, D<sub>2</sub> in Oikonomou et al. (2012), where 2 synaptic shocks produce an NMDA spike, and 5 shocks produce a plateau. The calcium dye shows a higher response from the plateau, suggesting a higher calcium amplitude (with a monotonic increase). Moreover, in the detailed calcium model for SPNs by Dorman et al. (2018), larger synaptic clusters produce monotonically higher calcium amplitudes. These two pieces of evidence suggest the calcium amplitudes are not affected by the membrane voltage approaching to the reversal potential. Most likely, this happens because of strong intracellular calcium buffering.

      Similarly, the experiments in Figs. 5 and 8 in Losonczy and Magee (2006) with glutamate uncaging at spine heads show a monotonic increase in calcium dye flourescence as the cluster size is increased. However, these experiments evoke NMDA spikes, which might not have approached the NMDAR reversal potential as closely as plateaus do (so, it is theoretically possible that monotonicity does not hold if a plateau were evoked, although it seems unlikely). This is now added to the Methods in lines 957-961.

      (5) Function: What is the actual function of the SPN, and how does it relate to FBP/NFBP learning? SPNs are involved in motor control/decision making and generally in sensorimotor tasks. Rather, use an example for that than a visual stimulus, though I understand it is more easily illustrated. Is feature binding what SPNs do and have to solve?

      It is not known yet what SPNs do precisely. The initial action selection role of the basal ganglia is currently being challenged or replaced by a role in movement initiation and/or invigoration (with action selection being done by the cortex, see e.g. Thura and Cisek (2017)). In any case, SPNs receive convergent input from sensory, motor, limbic and associative areas, as well as thalamic and neuromodulatory inputs (with topographical projections indicating functional specialization). In that sense, feature binding is particularly relevant for the striatum, as diverse features from different areas would arrive there. The basal ganglia participate in non-motor loops as well, so the striatum may have a role in initiation and termination of cognitive processes such as planning and attention, and in regulating emotional and motivated behavior (Purves et al. 2018).

      The role of the striatum and the relevance of feature binding is now stated in the introduction in lines 126-136, and a task with features more suited for the striatum is given in Figure 1 – figure supplement 1 (inspired from similar tasks in Bernklau et al. (2024)). I have kept the example from the visual system in the main text simply because it is widely used throughout the literature and easily recognizable, while clearly stating that other features would be involved in the striatum and pointing to the example in Figure 1 – figure supplement 1. Hopefully this will make a good compromise between ease of reading and relevance to the striatum.

      (6) The population of striatal projection neurons can express dopamine D2 receptors and not only D1 receptors. SPNs with different receptors thus undergo synaptic plasticity according to different rules. Would that lead to similar results? A combination of both populations?

      This is also being tested in the new study, and preliminary results indicate the answer is “yes” (see Author response image 4). Contrary to the dSPNs, the indirect-pathway SPNs (iSPNs), which predominantly express D<sub>2</sub> receptors, are thought to be involved in suppressing competing or related movements. The learning rule is almost the opposite to dSPNs: to trigger LTP, a dopamine pause is needed, and to trigger LTD a dopamine peak is needed (Fig. 4B and Shen et al. (2008)). So, with such a rule, iSPNs will store the irrelevant patterns and spike to them, and be silent to the relevant patterns (which indeed happens after learning in Author response image 4F<sub>2</sub>, F<sub>3</sub>). This aligns well with their role to suppress movements – reaching for the irrelevant patterns will be suppressed, while reaching for the relevant patterns will be disinhibited (the indirect pathway inhibits movement initiation). 

      (7) Other synaptic plasticity & metaplasticity models (possible comparison or discussion): https://pubmed.ncbi.nlm.nih./ https://linkinghub.elsevier.(https://link.springer.com/ For example, in the last model, there is only one modification LTP/LTD threshold, but it is different from the voltage threshold (which would be comparable to calcium thresholds in the current paper - see the voltage threshold different from the modification threshold here: https://doi.org/10.1371/). The modification threshold favors LTP or LTD - depending on the previous spiking activity of the cell and effectively works as an anti-Hebbian firing rate „homeostasis" mechanism. How is the firing rate homeostasis maintained in the current paper/model? Can you explain how this model is related to the previous models of Benuskova and Abraham, and also Clopath (there was also a metaplasticity version of the Clopath model), in terms of firing rate stability? The current model implements not only local synaptic plasticity but also metaplasticity, which is a nice feature, and seems to not only solve the FBP but also maintain firing stability. Is the stability of firing a consequence of the hard bounds for synaptic weights, or is it also a consequence of plasticity and metaplasticity? I like that each synaptic weight is changed independently based on local calcium concentration, with its own LTP and LTD thresholds.

      I will first describe what happens to the firing rate in this model, and then compare with the references you mentioned. In the FBP with linear integration, yes, the firing rate stability is only maintained because of the hard bounds for synaptic weights (Figure 4 – figure supplement 5A<sub>1</sub>, the neuron goes into depolarization block without w<sub>max</sub>). Without w<sub>max</sub>, the weights would grow until [Ca]<sub>NMDA</sub> saturates in the spine (which does not happen for most spines within the simulation time in Figure 4 – figure supplement 5B<sub>1</sub>).

      In the FBP with supralinear integration and the NFBP, the weights are prevented from increasing by the upper threshold for LTP, θ<sub>LTP</sub>. If θ<sub>LTP</sub> and w<sub>max</sub> are gone, synapses also grow without bound (Figure 4 – figure supplement 5B<sub>2</sub>). Regardless of how high the synaptic weight becomes, firing rate stability is ensured by the plateau potential, as the maximally achievable voltage is limited by the NMDAR reversal potential (Figure 4 – figure supplement 5A<sub>2</sub>, a somatic depolarization block does not happen as in Figure 4 – figure supplement 5A<sub>1</sub>). This is part of the plateaus’ function to provide dynamic range compression, as shown experimentally in Figs. 3E, 4D and 12 in Oikonomou et al. (2012). (Of course, such high weights are unrealistic, but they happen in the model without w<sub>max</sub> and without θ<sub>LTP</sub>). This is now described in the text in lines 397-407.

      Regarding the other studies, the second link did not work, so I will refer to Jedlicka et al. (2015), Clopath et al. (2010) and Zenke et al. (2013). (The rule in Jedlicka et al. (2015) is the same as in Benuskova and Abraham, (2007) but with a more detailed neuron model.) In these studies, there is a threshold representing a (weighted) average of the postsynaptic activity (postsynaptic voltage or spiking). In Clopath et al. (2010) and Zenke et al. (2013) it modulates only the amount of LTD, while in Jedlicka et al. (2015) it modulates the amounts of both LTP and LTD that occur due to plasticity. The threshold is global for a neuron (i.e. all synapses use the same threshold), and allows for heterosynaptic effects that compensate increased excitation from synaptic strengthening with stronger weakening in depressed synapses (thus maintaining stable firing rates). (However, the rule in Jedlicka et al. (2015) and Benuskova and Abraham, (2007) has only one threshold, there is no separate voltage threshold. Indeed, the threshold is a low-pass filtered variable of the input spike train, so is similar to calcium concentration.)

      There are no such effects in the current learning rule. Some compensatory effects are visible when comparing supralinear integration with and without an upper threshold θ<sub>LTP</sub>, as e.g. in Fig. 4B<sub>2</sub> and Figure 4—figure supplement 3B<sub>1</sub>. With θ<sub>LTP</sub>, synapses do not increase to w<sub>max</sub>, and without it, they do. As a consequence, the weakened synaptic cluster is weakened less when using θ<sub>LTP</sub> (‘yellow’ synapses in Fig. 4B<sub>2</sub>), and more without θ<sub>LTP</sub> (‘yellow’ synapses in Figure 4—figure supplement 2B<sub>1</sub>). The increased weakening in the latter case is because the strengthened ‘strawberry’ synapses have reached w<sub>max</sub>, and contribute to spiking for the irrelevant pattern (‘yellow strawberry’). As a result, the weakened ‘yellow’ synapses have to decrease more so no spiking for the irrelevant pattern occurs. In effect, the increased strengthening in the ‘strawberry’ synapses drives increased weakening in the ‘yellow’ synapses. However, this is a result of the task structure, i. e. that the patterns share features, and not from the learning rule.

      The above is briefly explained in the Discussion in lines 778-786.

      (8) The last sentence in the discussion is a bold claim: SPN can completely solve NFBP alone. Is that statement too strong, or is it a sign of multiple different mechanisms implementing the same function (i.e., degeneracy)?

      The sentence was not meant to be a bold claim. It currently says:

      “This indicates that ... SPNs might completely solve the task ... Whether this is true will be explored in another study …”

      I have now toned this claim down, as the purpose was only to point to the new study that explores this question (using with 100+ SPN models, in two learning regimes and varying cluster location). Nevertheless, the preliminary results in the new study suggest that SPNs can indeed solve the NFBP (if the plateaus are all-or-none), so I agree that it is a sign of multiple mechanisms implementing the same function.

      Figures/Table:

      (1) Figure 2: The layout of the figure is confusing. B shows distributed input as in upper A, and C shows clustered input as in lower A. This should be illustrated better (e.g., as in Figure 3 or 4). The term cluster size is confusing as well because it also refers to the distributed inputs; better use the number of synapses here. The color choice is difficult, as similar colors are chosen for cluster size and for differentiating linear and nonlinear integration in D.

      Thank you for the detailed comments! The figure has been reworked accordingly.

      (2) Figure 4: C What are dashed lines?

      They indicate the upper threshold for LTP, θ<sub>LTP</sub>.Thank you for noticing this, it is now added in the figure caption, and in other figures where it was missing.

      (3) Figure 5 and Figure 6: B is missing.

      These figures have been reworked to match the revision of the text describing the role of metaplasticity in the shared synapses. Now, the FBP is described first, followed by the NFBP, so the figures have been merged with their figure supplements (where the FBP results used to be in the first version).

      (4) Figure 4 Supplement 1: What do the colors and regions mean in B3? Why show here and not for others?

      The main text in lines 311-312, 331-333, and 358-359 refers to these regions when explaining the dynamics of the calcium, the thresholds and the weights (e.g. the pink region shows that [Ca]<sub>L-type</sub> in the ‘strawberry’ synapses evoked by ‘yellow strawberry’ is much lower than the calcium threshold for ‘strawberry’, protecting these synapses from weakening). They are simply highlighted with different colors, so readers know precisely where to look when reading the main text. The figure caption has been updated to state this.

      Also, the figure caption says that the same regions exist in all panels, although they have not been highlighted. I could have chosen any panel to mark these regions, and I chose B<sub>3</sub> because it looked like it had the least visual clutter.

      (5) Table 1: Typo for LTD

      Now fixed.

      (6) Line 21: „thehsold"

      Now fixed.

      (7) Line 50: Closing parentheses missing.

      Now fixed.

      (5) Line 55: "All metaplasticity models so far have only one modifiable threshold". This is not true as it stands. Multiple sliding thresholds have been proposed before:

      - experimentally by Ngezahayo et al. (2000)

      - multiple pathways reviewed by Abraham (2008)

      - multiple structural LTP thresholds observed experimentally by Ueda et al. (2022)

      - models of synaptic states and cascade models can be considered to employ multiple thresholds (depending on state); e.g., Fusi et al. (2005)

      Thank you for providing these references! The sentence has been changed (it was supposed to refer only to computational models), and the studies have been added in the text where appropriate.

      Importantly, Fusi et al. (2005) is now being discussed throughout the text, as it contains important results and conclusions about metaplasticity which are used as comparison. (In the first version of the article I did not include it because I was not sure about the interpretation of the thresholds – strictly speaking there is only one plasticity threshold (q, determining whether a synapse will switch sign), while the other is a metaplasticity threshold (p, determining whether the synapse will change its metaplastic state). However, the precise number of thresholds is not as important, as the study’s results are relevant to compare with.)

      (9) Lines 117/118: Missing parentheses around reference.

      Fixed.

      (10) Line 140: "a small difference". I feel like "small" is a bit of an understatement here, as from what I can see, the differences are almost a magnitude -- see blue/black traces in Figure 2B4 vs. 2C4.

      Thank you for noticing this. It was meant to refer only to the amplitude of the somatic voltage when the cluster size is small (in Fig 2D<sub>1</sub>). The text has been clarified now.

      (11) I think it's unnecessary to rewrite Eq. 1 in Eq. 3.

      I decided to keep this.

      (12) Lines 181-183: It would probably still be useful to show that NFBP can't be solved by distributed inputs.

      It is now shown in Figure 4 – figure supplement 1, and briefly described in the text in lines 298300.

      (13) Line 210: Typo, double "(the"

      Fixed, thank you.

      (14) Lines 248-249: Typo, double partial sentence.

      Now fixed.

      (15) Line 760: The year in the „STDP rule endowed with the BCM sliding threshold accounts for hippocampal heterosynaptic plasticity" should be corrected to 2007. Also, all other references should be checked for mistakes and missing pages, etc. (see e.g., also line 882).

      Thank you for noticing this! The references have been checked.

      References

      Benuskova, L., Abraham, W.C. STDP rule endowed with the BCM sliding threshold accounts for hippocampal heterosynaptic plasticity. J Comput Neurosci 22, 129–133 (2007). https://doi.org/10.1007/s10827-006-0002-x

      Bernklau TW, Righetti B, Mehrke LS, Jacob SN. (2024) Striatal dopamine signals reflect perceived cue–action–outcome associations in mice. Nature Neuroscience 27(4):747–757. doi: 10.1038/

      Branco, T., Clark, B. A., & Häusser, M. (2010). Dendritic Discrimination of Temporal Input Sequences in Cortical Neurons. Science, 329(5999), 1671–1675. http://www.jstor.org.focus.lib.kth.se/stable/40803162

      Chalifoux JR, Carter AG (2011) Glutamate Spillover Promotes the Generation of NMDA Spikes. J. Neurosci. 31(45):16435–16446. doi:10.1523/JNEUROSCI.2777-11.2011

      Clopath, C., Büsing, L., Vasilaki, E. et al. Connectivity reflects coding: a model of voltage-based STDP with homeostasis. Nat Neurosci 13, 344–352 (2010). https://doi.org/10.1038/nn.2479

      Dorman, D. B., Jędrzejewska-Szmek, J., Blackwell, K. T. (2018) Inhibition enhances spatiallyspecific calcium encoding of synaptic input patterns in a biologically constrained model. elife 7:e38588 https://doi.org/10.7554/eLife.38588

      Du K., Wu Y., Lindroos R., Liu Y., Rózsa B., Katona G., Ding J.B., & Kotaleski J.H. (2017) Celltype–specific inhibition of the dendritic plateau potential in striatal spiny projection neurons, Proc. Natl. Acad. Sci. U.S.A. 114 (36) E7612-E7621, https://doi.org/10.1073/pnas.1704893114.

      Gjorgjieva J., Clopath C., Audet J., & Pfister J. (2011) A triplet spike-timing–dependent plasticity model generalizes the Bienenstock–Cooper–Munro rule to higher-order spatiotemporal correlations, Proc. Natl. Acad. Sci. U.S.A. 108 (48) 19383-19388, https://doi.org/10.1073/pnas.1105933108

      Gütig R., Aharonov R., Rotter S., Sompolinsky H (2003) Learning Input Correlations through Nonlinear Temporally Asymmetric Hebbian Plasticity. J. Neurosci. 23 (9) 3697-3714; DOI: 10.1523/JNEUROSCI.23-09-03697.2003

      Hedrick, N.G., Lu, Z., Bushong, E. et al. (2022) Learning binds new inputs into functional synaptic clusters via spinogenesis. Nat Neurosci 25, 726–737 . https://doi.org/10.1038/s41593-022-01086-6

      Hwang F, Roth R, Wu Y et al. (2022) Motor learning selectively strengthens cortical and striatal synapses of motor engram neurons Neuron 110, 2790-2801.e5

      Jedlicka P, Benuskova L, Abraham WC. (2015) A Voltage-Based STDP Rule Combined with Fast BCM-Like Metaplasticity Accounts for LTP and Concurrent “Heterosynaptic” LTD in the Dentate Gyrus In Vivo. PLOS Comput Biol 11(11):1–24. https://doi.org/10.1371/journal.pcbi.1004588

      Kirchner, J.H., Gjorgjieva, J. (2021) Emergence of local and global synaptic organization on cortical dendrites. Nat Commun 12, 4005. https://doi.org/10.1038/s41467-021-23557-3

      Lindroos R, Hellgren Kotaleski J (2021) Predicting complex spikes in striatal projection neurons of the direct pathway following neuromodulation by acetylcholine and dopamine. Eur J Neurosci. 53:2117–2134. https://doi.org/10.1111/ejn.14891

      Losonczy A, Magee J (2006) Integrative Properties of Radial Oblique Dendrites in Hippocampal CA1 Pyramidal Neurons. Neuron 50: 291-307

      Malarkey EB, Parpura V (2008) Mechanisms of glutamate release from astrocytes, Neurochemistry International 52(1–2): 142-154, doi: 10.1016/j.neuint.2007.06.005.

      Oikonomou KD, Short SM, Rich MT and Antic SD (2012) Extrasynaptic Glutamate Receptor Activation as Cellular Bases for Dynamic Range Compression in Pyramidal Neurons. Front. Physio. 3:334. doi: 10.3389/fphys.2012.00334

      Oikonomou KD, Singh MB, Sterjanaj EV and Antic SD (2014) Spiny neurons of amygdala, striatum, and cortex use dendritic plateau potentials to detect network UP states. Front. Cell. Neurosci. 8:292. doi: 10.3389/fncel.2014.00292

      Purves D, Augustine GJ, Fitzpatrick D, Hall WC, LaMantia AS, White LE, et al. (2018) Neuroscience. 6 ed. New York:Oxford University Press

      Rusakov DA, Stewart MG (2021) Synaptic environment and extrasynaptic glutamate signals: The quest continues, Neuropharmacology 195: 108688, https://doi.org/10.1016/j.neuropharm.2021.108688

      Schiess M, Urbanczik R, Senn W (2016) Somato-dendritic Synaptic Plasticity and Errorbackpropagation in Active Dendrites. PLoS Comput Biol 12(2): e1004638. https://doi.org/10.1371/journal.pcbi.1004638

      Suzuki, T., Kodama, S., Hoshino, C., Izumi, T. and Miyakawa, H. (2008), A plateau potential mediated by the activation of extrasynaptic NMDA receptors in rat hippocampal CA1 pyramidal neurons. European Journal of Neuroscience, 28: 521-534. https://doi.org/10.1111/j.14609568.2008.06324.x

      Thura D, Cisek P. (2017) The Basal Ganglia Do Not Select Reach Targets but Control the Urgency of Commitment. Neuron 95(5):1160–1170.e5. doi: 10.1016/j.neuron.2017.07.039)

      Tran-Van-Minh A, Cazé RD, Abrahamsson T, Cathala L, Gutkin BS and DiGregorio DA (2015) Contribution of sublinear and supralinear dendritic integration to neuronal computations. Front. Cell. Neurosci. 9:67. doi: 10.3389/fncel.2015.00067

      Trpevski D, Khodadadi Z, Carannante I and Hellgren Kotaleski J (2023) Glutamate spillover drives robust all-or-none dendritic plateau potentials—an in silico investigation using models of striatal projection neurons. Front. Cell. Neurosci. 17:1196182. doi: 10.3389/fncel.2023.1196182

      Urbanczik R, Senn W (2014) Learning by the Dendritic Prediction of Somatic Spiking. Neuron 81: 521-528

      Author response image 1.

      Plateau potentials (A1, B1) and their somatic amplitudes in dSPNs (A) and iSPNs (B) when the location of the synaptic cluster is varied. Distally evoked plateaus have a smaller somatic amplitude in both dSPNs and iSPNs, bust still exhibit a nonlinear jump in the somatic voltage. In (A2, B2) all 71 dSPN and 34 iSPN models were tested, color coded with respect to their excitability to a synaptic cluster of 10 synapses.

      Author response image 2.

      Two learning simulations on the NFBP where different metaplasticity rates for LTP and LTD were used. (Top) The metaplasticity rate for LTD is higher, stopping the LTD process prematurely (D1, D2), resulting in less weakening of the synapses (B1, B2) and spiking for one irrelevant pattern (A1). (Bottom) The metaplasticity rate for LTP is higher, preventing frequent synaptic strengthening(C1, C2). As a result, it takes a longer time for synapses to strengthen and trigger plateaus (B3, B4).

      Author response image 3.

      Learning on the NFBP with a nonlinear (quadratic) metaplasticity rule, which implements the BCM rule. (A) Somatic and dendritic voltage before and after learning. (B-D) Evolution of synaptic weights (B), LTP (C) nad LTD thresholds (D). Weights in (B) seem to stabilize around a high value once the LTP threshold reaches the maximal [Ca]NMDA; however, the value is too high and triggers plateau potentials for all patterns.

      Author response image 4.

      Learning to solve the NFBP in one dSPN and one iSPN model. (A, B) Schemes describing the requirements for cortico-striatal synaptic plasticity in dSPNs (A) and iSPNs (B). (C) Dopamine peaks are emitted from the midbrain after the relevant patterns, and dopamine pauses after the irrelevant patterns. (D) Stimulation protocol. (E, F) A single learning simulation in a dSPN (E) and an iSPN (F) model. Before learning the neurons spike to all patterns (due to additional feature-unspecific distributed inputs, shown with the third panels in E3, F3). After learning, the dSPN spikes only for the relevant patterns (E2) and the iSPN only for the irrelevant ones (F2). (E3, F3) The evolution of clustered and distributed synptic weights (In dendrite 1 of the iSPN, the synapses for 'red' experienced some strengtheneing after being weakened, but this does not seem to affect learning in this example.) The neurons in th enew study receive less background synaptic input than in the current article, and additional distributed synapses are needed along with a plateau for somatic spiking, making for a more challenging learning scenario.

    1. eLife Assessment

      This valuable study provides a practical computational framework for inferring latent neural states directly from calcium fluorescence recordings, bypassing the traditional need for a separate spike deconvolution step. The evidence supporting the method is convincing, featuring rigorous validation across multiple latent variable model families (including HMM, GPFA, and LFADS) using both simulated and experimental data. To further strengthen method's generality, it would require further application to a broader range of experimental datasets, such as recordings from different brain regions or using different calcium indicators.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors elegantly combined latent variable models (i.e., HMM, GPFA and dynamical system models) with a calcium imaging observation model (i.e., latent Poisson spiking and autoregressive calcium dynamics (AR)).

      Strengths:

      Integrating a calcium observation model into existing latent variable models improves significantly the inference of latent neural states compared to existing approaches such as spike deconvolution or Gaussian assumptions.

      The authors also provide an open-source access to their method for direct application to calcium imaging data analysis.

      Weaknesses:

      As acknowledged by the authors, their method is dependent on the quality of calcium traces extraction from fluorescence videos. It should be noted that this limitation applies to alternative strategies.

      While the contribution of this study should prove useful for researchers using calcium imaging, the novelty is limited, as it consists of an integration of the calcium imaging model from Ganmor et al. 2016 with existing LVM frameworks.

      Comments on revised version.

      The authors addressed my comments and I have no further concerns.

    3. Reviewer #2 (Public review):

      Summary:

      This compelling study proposes a framework to implement latent variable models using population level calcium imaging data. The study incorporates autoregressive dynamics and latent Poisson spiking to improve inference of latent states across different model classes including HMMs, Gaussian Process Factor Analysis and nonlinear dynamical systems models. This approach allows for a more seamless integration of existing methods typically used with spiking data to apply on calcium imaging data. The authors test the model on piriform cortex recordings as well as a biophysical simulator to validate their methods. This approach promises to have wide usability for neuroscientists using large population level calcium imaging.

      Strengths:

      The strength of this study is the flexibility in the choice of models and relatively easy adaptation to user-specific use cases.

      Weaknesses:

      The weakness of the study lies in its limited validation of biological calcium imaging data. Calcium dynamics in a task-specific context in a sensory brain region might be very different from slower dynamics in a region of integration.

    4. Reviewer #3 (Public review):

      Summary:

      S. Keeley & collaborators propose a computational approach to infer time-varying latent variables directly from calcium traces (e.g., obtained with 2p imaging) without the need for deconvolving the traces into spike trains in a preliminary, independent step. Their approach rests on 1 of 3 families of latent models: GPFA, HMM and dynamical systems - which they augment with an observation model that maps latent variables to fluorescence traces. They validate their approach on simulated data as well as a single real dataset, showing that the approach improves latent variable inference and model fitting, compared to more traditional approaches (although not directly compared with the 2-step one; see below). They provide a GitHub repository with code to fit their models (which I have not tested).

      Strengths:

      The approach is sound and well-motivated. The authors are specialists of latent variable models. The manuscript is succinct, well-written and the figures are clear. I particularly liked the diversity of latent models considered, in particular latent models with continuous (GPFA) vs. discrete (HMM) dynamics, which are useful for characterizing different types of neural computations. The validation on both simulated and real data is convincing.

      Weaknesses:

      The main weakness point that I see is that the approach is tested only on a single real dataset (odor response dataset). The other model fits are obtained from simulated data. While the results are convincing, it would be useful to see the approach tested on other datasets, for instance datasets with different brain areas, different behavioral conditions, or different calcium indicators. This would help assess the generality of the approach and its robustness to different experimental conditions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors elegantly combined latent variable models (i.e., HMM, GPFA and dynamical system models) with a calcium imaging observation model (i.e., latent Poisson spiking and autoregressive calcium dynamics (AR)).

      Strengths:

      Integrating a calcium observation model into existing latent variable models improves significantly the inference of latent neural states compared to existing approaches such as spike deconvolution or Gaussian assumptions.

      The authors also provide an open-source access to their method for direct application to calcium imaging data analysis.

      Weaknesses:

      As acknowledged by the authors, their method is dependent on the quality of calcium trace extraction from fluorescence videos. It should be noted that this limitation applies to alternative strategies.

      While the contribution of this study should prove useful for researchers using calcium imaging, the novelty is limited, as it consists of an integration of the calcium imaging model from Ganmor et al. 2016 with existing LVM frameworks.

      Reviewer #2 (Public review):

      Summary:

      This compelling study proposes a framework to implement latent variable models using population level calcium imaging data. The study incorporates autoregressive dynamics and latent Poisson spiking to improve inference of latent states across different model classes including HMMs, Gaussian Process Factor Analysis and nonlinear dynamical systems models. This approach allows for a more seamless integration of existing methods typically used with spiking data to apply on calcium imaging data. The authors test the model on piriform cortex recordings as well as a biophysical simulator to validate their methods. This approach promises to have wide usability for neuroscientists using large population level calcium imaging.

      Strengths:

      The strengths of this study are the flexibility in the choice of models and relatively easy adaptation to user-specific use cases.

      Weaknesses:

      The weakness of the study lies in its limited validation of biological calcium imaging data. Calcium dynamics in a task-specific context in a sensory brain region might be very different from slower dynamics in a region of integration. The biophysical properties of the data would also be dependent on the SNR of the imaging platform and the generation of calcium indicator being used.

      Reviewers 1 and 2 correctly point out that our method depends on the quality of the upstream calcium trace extraction. As they noted, these traces can vary based on the specific indicator used, the signal-to-noise ratio (SNR) of the recording, and the region of the brain being imaged (such as sensory vs. integrative areas). As Reviewer 1 rightly mentions, this is a universal challenge that applies to alternative strategies as well, rather than a limitation unique to our framework. To address these points, we have added a paragraph to our Discussion section.

      Reviewer #3 (Public review):

      Summary:

      S. Keeley & collaborators propose a computational approach to infer time-varying latent variables directly from calcium traces (for instance, obtained with 2p imaging) without the need for deconvolving the traces into spike trains in a preliminary, independent step. Their approach rests on 1 of 3 families of latent models: GPFA, HMM and dynamical systems - which they augment with an observation model that maps latent variables to fluorescence traces. They validate their approach on simulated and real data, showing that the approach improves latent variable inference and model fitting, compared to more traditional approaches (although not directly compared with the 2-step one; see below). They provide a GitHub repository with code to fit their models (which I have not tested).

      Strengths:

      The approach is sound and well-motivated. The authors are specialists in latent variable models. The manuscript is succinct, well-written, and the figures are clear. I particularly liked the diversity of latent models considered, in particular latent models with continuous (GPFA) vs.

      discrete (HMM) dynamics, which are useful for characterizing different types of neural computations. The validation on both simulated and real data is convincing.

      Weaknesses:

      One advantage … is that one can inspect the quality of the deconvolution step independently from the latent variable inference step. For instance, if the inferred latent variables are not interpretable, how can one determine whether this is due to a poor choice of latent model (e.g., HMM with too few states), or a poor fit of the observation model (e.g., wrong parameters for the calcium dynamics)?

      We agree with the reviewer that integrating the calcium likelihood introduces additional parameters that require careful diagnostics. However, for the vast majority of imaging datasets, there is no simultaneous electrophysiology to verify the deconvolution step. If the final latent states are not interpretable, it remains impossible to determine whether the error originated in the initial spike inference from deconvolution or the subsequent model fitting.

      Our framework addresses this by maintaining the raw fluorescence as the fixed observation. We suggest for those using this model to employ cross-validation using this data to select model parameters. We outline how to do this below, but because the data itself does not change with each model fit, you can compare P(data | λ) across any model configuration. In contrast, different deconvolution methods change the data itself (the spiketimes) making comparison across models impossible.

      Could the authors comment on whether their approach allows for instance to compare different forms of latent models (e.g., HMM vs. GPFA) in terms of model evidence, cross-validated log-likelihood or other model comparison metrics?

      We thank the reviewer for highlighting this. In short: yes. Because our framework integrates the calcium observation likelihood with various latent variable models, we can assess held-out prediction P(data | λ) irrespective of the specific latent structure.

      However, because fitting the LVM requires inferring the latent state z to determine the firing rate λ, proper cross-validation across models involves holding out both neurons and timepoints. A principled approach—which our framework supports—is as follows:

      (1) Train both the latent states z and the model parameters (e.g., the mapping from latent space to observations) on a training portion of the recording.

      (2) On a held-out test segment, withhold a subset of "test" neurons and infer the latent states using only the "held-in" neurons.

      (3) Calculate the likelihood of the observed fluorescence for the test neurons given the inferred rates.

      We clarify this procedure in the revised manuscript. While a comprehensive benchmarking across all possible LVM architectures is beyond the scope of this study, we provide the statistical infrastructure for users to perform such comparisons. Furthermore, we would like to emphasize that while predictive likelihood is a rigorous metric for model selection, the primary utility of these LVMs often lies in the interpretability of the latent states themselves, which can remain biologically informative even if cross-validated performance is not the sole optimization target.

      While it certainly makes sense that models accounting for the full transformation of latent => spikes => fluorescence data should outperform the two-step (1) deconvolution => (2) latent variance inference approach, the amount of improvement is not clear. A direct comparison … would be useful

      We thank the reviewer for this point. Figure 4 was designed to address this comparison directly. By using a biophysical simulator, we generated a pseudo-realistic spiking network with ground-truth latent trajectories governed by a Gaussian Process. This allowed us to explicitly compare our unified approach against the traditional deconvolution-then-Poisson-GPFA pipeline. While a first-order (AR1) calcium likelihood did not show improvement over the two-step deconvolution method in recovering the ground-truth latents, the second-order (AR2) process demonstrated an improvement. Because there are no ground-truth parameters in the model, we use the reconstruction error of the latent values as our primary metric for recovery. These results suggest that when the observation model sufficiently captures the underlying calcium kinetics, the unified approach offers a more accurate estimation of the neural state.

      It would be useful to discuss the possible extension of the approach to other types of data that … have different observation models.

      We thank the reviewer for this helpful comment. We agree that the general framing of the likelihood has potential use in a wider range of data modalities.

      Specifically, all sensors (aside from some voltage sensors) have a rise and decay time in line with our model. Thus the autoregressive (AR) nature of the calcium likelihood we utilize makes the current implementation particularly well-suited for a broad range of fluorescence-based sensors with similar temporal profiles. The specific use and extension would depend heavily on the biological target of the sensor. For example Glutamate, dopamine, and similar indicators can be thought of as having a similar underlying Poisson firing model, as the release of these products is tied to neural firing. Other sensors that might relate to other biological processes, such as hemodynamics (via imaging or ultrasound) or broader neuromodulation (Norepinephrine imaging with nLight) might be more continually varying and therefore would require changing the Poisson with an appropriate alternative, for example a Gaussian Process or similar.

      Voltage imaging is the one exception that may require more complex observation models. However, the challenge in voltage imaging is not the ability to identify individual spikes, but more that the speed and scope of imaging is inherently limited by the speed of the voltage process and signal-to-noise ratios induced by the low quantum efficiency and membrane-bound nature of these indicators. If imaged well, single spikes would be clearly visible and the two-stage likelihood would not be necessary—one could simply use the spike times in a Poisson model just as with electrophysiology. We have added a paragraph in the discussion highlighting these points.

    1. eLife Assessment

      This important study identifies underappreciated experimental factors that influence α-synuclein amyloid polymorphism, with practical implications for the reproducibility of in vitro fibril preparations. Extensive cryo-EM analyses provide convincing evidence that protein purity, monomer preparation, and agitation conditions influence polymorph selection, including substantial effects from small amounts of an N-terminally truncated variant. The study provides a substantial resource for researchers studying amyloid assembly and synucleinopathies, although the proposed nucleation mechanisms need stronger kinetic support and reliable protocols for producing specific polymorphs require further development.

    2. Reviewer #1 (Public review):

      In this work, Frey and colleagues have carried out a very large study of α-synuclein polymorphism as a function of aggregation conditions and sample preparation. They provide valuable insight into the many critical factors affecting α-synuclein polymorphism, thereby illuminating the need for detailed reporting in the literature as well as both rigorous and detail-oriented protocols when working with this protein. Indeed, their observations are in line with the difficulties of reproducing structural outcomes across different laboratories and experiments. The authors must be complemented on their openness about the difficulties experienced and the thoroughness of their work. Efforts like these are going to be crucial to achieve an understanding of the unparalleled structural plasticity of α-synuclein amyloid fibrils. It is particularly notable that the authors have managed to optimize protocols to form single-polymorph aggregation reactions with high reproducibility.

      In this work, the authors focus on the influence of α-synuclein purity in aggregation reactions. They find that using reverse-phase HPLC to purify the protein significantly alters the aggregation behaviour. It is interesting, and rather uncommon in the field, to use reverse-phase HPLC as a final purification step, rather than SEC, which is commonly used in many laboratories. It would be useful to compare this new protocol even more directly and extensively with the commonly used SEC protocol. When mentioning their previously published work, it should be mentioned explicitly how the protein was purified in these previous studies.

      On the point of protein purification, the authors lyophilize their protein prior to storage. In their work, they also find that pre-aggregation oligomer formation alters the aggregation pathway of α-synuclein. While they demonstrate that this can be solved by appropriate filtration, it should be discussed why the lyphilizaiton step was not reconsidered/omitted given that this process is known to facilitate oligomer formation. Do the authors have experience with aggregation studies using α-synuclein that has not undergone lyophilisation and are able to comment on the influence of this step in the protocol?

      The authors make note of several degradation products affecting their aggregation reactions, which is why they employ a much more thorough purification protocol. However, they also point out that some of the degradation products found at the end of their reaction could form during the reaction itself. Unfortunately, they never investigate this further. In particular, it would be very useful to know if different aggregation conditions (pH, salt, agitation) lead to different and characteristic degradation patterns. HPLS/mass spec of the supernatant at the end of each aggregation reaction would have been a very insightful thing to do.

      With respect to degradation products, in this work a NΔ4-variant is produced to mimic a disease-relevant degradation product and indeed it is found to alter the structural outcome even at low relative concentrations (5%). Have the authors investigated the minimal fraction of the NΔ4 variant necessary to still influence the structural outcome of the predominant WT protein? This type of analysis could have significant relevance to disease-related analysis, where several variants (truncations and PTM variants) are present in trace amounts.

      While on this topic, the authors note that in several of their type 5 fibrils, they find unresolved peptide fragments in their cryo-EM structures. Can the authors speculate if these fragments are indeed peptide degradation products or residues wrapping around the fibril originating from the fibril-incorporated protein?

      In this work, the authors have performed an extensive study of α-synuclein polymorphism. However, despite generating what is likely the largest single data set of fibril structures, they perform very little quantitative analysis of their data. It would be interesting to analyse the relative abundance of fibril polymorphs produced in each reaction. Perhaps from the particle-picking data it could be estimated the relative abundance of each polymorph as well as non-resolved fibrils to generate a more nuanced view of polymorphism beyond overall classifications of the resolved structures. Indeed, from such data it could also be studied if certain protofilaments are more prone to pair in asymmetric fibril structures over others or if fibril asymmetry can be attributed to random pairing of protofilaments in accord with their abundance. This latter point is particularly interesting for the type 1 fibrils. Perhaps, the propensity of α-synuclein to form specific symmetries could also be estimated.

      The authors note that pH is a strong factor in determining polymorph selection. This does indeed appear to be the case, but other parameters do not appear to show any clear trend. Have the authors investigated the influence of aggregation parameters (agitation, duration) on structural outcomes systematically or quantitatively, such as with principal component analysis? Indeed, they also find that some fibril types that otherwise are not compatible at the same pH appear to co-exist when the shaking parameter is modified. Are the authors then confident in the claim that pH is a deterministic parameter?

      It is evident from this and other work that amyloid aggregation is highly sensitive to kinetic effects. It is therefore curious that the effect of protein concentration and reaction time has not been systematically investigated. The authors have some data studying dilution series (Figure 6A) and different reaction times (reaction 56 & 57). Could the authors comment on the effects of these two parameters, and might there be more information touching upon this that could be highlighted in this work?

      The authors state that this work likely underreports fibril polymorphisms in samples due to population size or data quality challenges. Could the authors, based on their extensive experience, try to quantify this statement?

      Additionally, the authors point out that the current framework for classifying α-synuclein fibril polymorphism is not sufficient to describe the real complexity of this protein system. However, they do not seem to address some of the recent literature aiming to solve such issues (see Scheres 2026, Connor et al. 2025, Milchberg et al. 2025 & Price et al. 2025).

      Was any biophysical/biochemical analysis performed of the many structures produced here, such as CD spectroscopy, Proteinase K digestion, dye binding or FTIR, which could act as low-resolution structure identifiers and might help to retrospectively explain some findings in the older literature? Such data would be very useful for the vast majority of researchers, who do not have access to cryo-EM.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript describes insights gained during efforts to reproduce disease-relevant alpha-Synuclein (aSyn) fibrils using recombinant protein in vitro. It follows up on a similar article from this team published in 2024. Although the authors have not been able to produce fibrils with the structure of ex vivo fibrils isolated from patients, they share insights gained into which factors influence the formation of specific fibril polymorphs.

      Strengths:

      This is quite an unusual manuscript because it goes into minute detail about sample preparation that are usual just mentioned in the Materials and Methods sections of other manuscripts (if at all). This makes it very valuable for the scientific community working on exactly the problem of reproducing disease-relevant aSyn fibrils in vitro (which will be a major breakthrough in the field). The authors present an impressive array of cryo-EM fibril structures, some of which have not been described before.

      Weaknesses:

      A major concern with the manuscript is that its story and messaging are a bit murky. The authors describe a few new polymorphs, show that some polymorphs (type 1) have small variations, show that sample purity and fragmentation will influence polymorph formation, and present a helical-symmetry mystery. This all reads like a loose collection of findings without any major takeaway. Looking at Table 1, it still seems that the authors do not have a good control over any of these polymorphs. Are they able to make any of these polymorphs reliably? I think the impact of this work could be strengthened if it ended with a reliable protocol for the production of any of the polymorphs described.

      A second major concern is the quality of the aggregation kinetics and their interpretation. Are these kinetics just done once per concentration? The figure caption talks about 'three independent samples' but it is unclear if this refers to the three different concentrations or NΔ4 percentages used, or true repetitions. Looking at the curves themselves, it seems that only one of the conditions was actually done in triplicate, which should be the minimum to draw conclusions. Further, it would have been helpful to characterize the kinetic data quantitatively. Finally, because there are only kinetic data for a fraction of the conditions tested, it is not clear what they add to the overall manuscript. My recommendation is to either remove the kinetics from the manuscript or substantially expand this section.

    4. Author response:

      Reviewer #1 (Public review):

      In this work, Frey and colleagues have carried out a very large study of α-synuclein polymorphism as a function of aggregation conditions and sample preparation. They provide valuable insight into the many critical factors affecting α-synuclein polymorphism, thereby illuminating the need for detailed reporting in the literature as well as both rigorous and detail-oriented protocols when working with this protein. Indeed, their observations are in line with the difficulties of reproducing structural outcomes across different laboratories and experiments. The authors must be complemented on their openness about the difficulties experienced and the thoroughness of their work. Efforts like these are going to be crucial to achieve an understanding of the unparalleled structural plasticity of α-synuclein amyloid fibrils. It is particularly notable that the authors have managed to optimize protocols to form single-polymorph aggregation reactions with high reproducibility.

      In this work, the authors focus on the influence of α-synuclein purity in aggregation reactions. They find that using reverse-phase HPLC to purify the protein significantly alters the aggregation behaviour. It is interesting, and rather uncommon in the field, to use reverse-phase HPLC as a final purification step, rather than SEC, which is commonly used in many laboratories. It would be useful to compare this new protocol even more directly and extensively with the commonly used SEC protocol. When mentioning their previously published work, it should be mentioned explicitly how the protein was purified in these previous studies.

      On the point of protein purification, the authors lyophilize their protein prior to storage. In their work, they also find that pre-aggregation oligomer formation alters the aggregation pathway of α-synuclein. While they demonstrate that this can be solved by appropriate filtration, it should be discussed why the lyphilizaiton step was not reconsidered/omitted given that this process is known to facilitate oligomer formation. Do the authors have experience with aggregation studies using α-synuclein that has not undergone lyophilisation and are able to comment on the influence of this step in the protocol?

      We will include details in our revised version on how samples were prepared in our previously published work. However we do not plan to do a comparison between SEC and HPLC as final purification steps because these are truly orthogonal separation methods. SEC is the gold standard for oligomer removal but not particularly useful for removing degradation products of similar size (which we believe to affect aggregation outcomes). In our hands, the highest purity and reproducibility come from HPLC-purified material followed by a stringent oligomer removal. Because oligomers in the solubilized sample are a concern, SEC as a final post-solubilization/pre-aggregation step might be ideal. However, as we mentioned in the manuscript, SEC dilutes the sample to the point that much of the sample is too dilute for our purposes (aggregation without seeds) and adding an additional concentration step would risk promoting the formation of new oligomers in the concentration device. That is why we resorted to using a 100 kD MWCO filter to remove oligomeric species.

      We considered skipping the lyophilization step and dialyzing the HPLC-purified sample into the buffer of choice but stuck with lyophilization because it provides an easy control over the protein concentration in the solubilized sample.

      To clarify the logic in this choice of sample preparation steps, we will add a section to the revised manuscript listing/explaining our suggested protocols for preparing alpha-synuclein samples for reproducible aggregation experiments.

      The authors make note of several degradation products affecting their aggregation reactions, which is why they employ a much more thorough purification protocol. However, they also point out that some of the degradation products found at the end of their reaction could form during the reaction itself. Unfortunately, they never investigate this further. In particular, it would be very useful to know if different aggregation conditions (pH, salt, agitation) lead to different and characteristic degradation patterns. HPLS/mass spec of the supernatant at the end of each aggregation reaction would have been a very insightful thing to do.

      In retrospect, we agree that this could have been important from the standpoint of understanding how in situ degradation during the aggregation at 37º C could also play a role in polymorph selection. We have begun to save frozen aliquots of our aggregation samples for subsequent MS analyses of the interesting samples in order to be able to address this in the future. However, it was outside the scope of our original search for the PD polymorph.

      With respect to degradation products, in this work a NΔ4-variant is produced to mimic a disease-relevant degradation product and indeed it is found to alter the structural outcome even at low relative concentrations (5%). Have the authors investigated the minimal fraction of the NΔ4 variant necessary to still influence the structural outcome of the predominant WT protein? This type of analysis could have significant relevance to disease-related analysis, where several variants (truncations and PTM variants) are present in trace amounts.

      We did not try lower than 5% because it was our goal to test if impurities at this level (which usually go undetected) could influence the aggregation outcomes. We think that directly relating the precise impurity level in these in vitro experiments to in vivo aggregation would be difficult due to the many factors we do not yet understand that appear to guide in vivo polymorph selection.

      While on this topic, the authors note that in several of their type 5 fibrils, they find unresolved peptide fragments in their cryo-EM structures. Can the authors speculate if these fragments are indeed peptide degradation products or residues wrapping around the fibril originating from the fibril-incorporated protein?

      We will add this speculation to the revised manuscript. The unassigned peptide density most likely originates from the C-terminal residues of the intact chains rather than from a degradation product. The levels of degradation products observed by MS in our other samples were far too low to account for the amount of peptide that would be required for >50% of the fibrils in sample 23 to contain this extra density. Furthermore, in the 5A polymorph the extra peptide density is sometimes present (e.g. sample 52) and sometimes absent (e.g. sample 3), despite the fact that in both samples the coexisting type 5 polymorphs (5m and 5B, respectively) have the peptide bound. Therefore, it appears that subtle differences between 5A polymorphs determine the presence or absence of this density, whereas for 5m and 5B it is consistently present.

      In this work, the authors have performed an extensive study of α-synuclein polymorphism. However, despite generating what is likely the largest single data set of fibril structures, they perform very little quantitative analysis of their data. It would be interesting to analyse the relative abundance of fibril polymorphs produced in each reaction. Perhaps from the particle-picking data it could be estimated the relative abundance of each polymorph as well as non-resolved fibrils to generate a more nuanced view of polymorphism beyond overall classifications of the resolved structures. Indeed, from such data it could also be studied if certain protofilaments are more prone to pair in asymmetric fibril structures over others or if fibril asymmetry can be attributed to random pairing of protofilaments in accord with their abundance. This latter point is particularly interesting for the type 1 fibrils. Perhaps, the propensity of α-synuclein to form specific symmetries could also be estimated.

      We agree that Cryo-EM datasets contain considerable information that could potentially be used to better understand polymorph populations. We have previously used particle counts as an approximate measure of polymorph abundance; however, the biases introduced during particle picking and subsequent curation are substantial, and we therefore do not consider these counts sufficiently reliable for quantitative comparison of polymorph populations.

      Regarding the symmetry of paired filaments, it is clear that the overwhelming preference of all filaments is to pair as symmetric dimers. Among the in vitro polymorphs, type 1 is the most commonly observed to form an asymmetric dimer, either with itself or with the new type 7. However, the number of observations is too small to establish that this represents a statistically meaningful preference. Types 2 and 3 have also been observed to pair in asymmetric fibrils. Overall, we do not think that our dataset is large enough to add statistical weight to previous observations.

      The authors note that pH is a strong factor in determining polymorph selection. This does indeed appear to be the case, but other parameters do not appear to show any clear trend. Have the authors investigated the influence of aggregation parameters (agitation, duration) on structural outcomes systematically or quantitatively, such as with principal component analysis? Indeed, they also find that some fibril types that otherwise are not compatible at the same pH appear to co-exist when the shaking parameter is modified. Are the authors then confident in the claim that pH is a deterministic parameter?

      We did not systematically vary the agitation but in two cases where it was either intentionally or accidentally varied, we found surprising polymorph outcomes. We felt that these observations were worth reporting, but on their own they do not constitute a thorough study. We would rather conclude that pH is a strong selector and can be deterministic under certain conditions: pure sample, no seeds, continuous or intermittent moderate agitation. We will revise the manuscript to make this clearer.

      It is evident from this and other work that amyloid aggregation is highly sensitive to kinetic effects. It is therefore curious that the effect of protein concentration and reaction time has not been systematically investigated. The authors have some data studying dilution series (Figure 6A) and different reaction times (reaction 56&57). Could the authors comment on the effects of these two parameters, and might there be more information touching upon this that could be highlighted in this work?

      We did not collect sufficient data to draw conclusions about the effects of protein concentration or aggregation time, and therefore do not think that a quantitative analysis of these parameters is justified by the present dataset.

      The authors state that this work likely underreports fibril polymorphisms in samples due to population size or data quality challenges. Could the authors, based on their extensive experience, try to quantify this statement?

      Precise quantification would be difficult. The literature has many mentions of amyloids that could not be solved by Cryo-EM due to a lack of twist. In our hands it is very common to have a small subset of non-twisted filaments in a sample and some samples appear to be exclusively non-twisted. Low-abundance or low-quality fibrils are also rather common in our data but also difficult to quantify. Nevertheless, we agree that it would be useful to place a lower bound on this estimate, and we will re-examine our datasets to determine whether this can be quantified in the revised manuscript.

      Additionally, the authors point out that the current framework for classifying α-synuclein fibril polymorphism is not sufficient to describe the real complexity of this protein system. However, they do not seem to address some of the recent literature aiming to solve such issues (see Scheres 2026, Connor et al. 2025, Milchberg et al. 2025 & Price et al. 2025).

      We agree that we should have discussed this literature in greater detail and will do so in the revised manuscript.

      Was any biophysical/biochemical analysis performed of the many structures produced here, such as CD spectroscopy, Proteinase K digestion, dye binding or FTIR, which could act as low-resolution structure identifiers and might help to retrospectively explain some findings in the older literature? Such data would be very useful for the vast majority of researchers, who do not have access to cryo-EM.

      We did not perform these analyses precisely because they are low resolution. Retrospectively, such analyses might have been useful for interpreting past data, but most of these methods (particularly Proteinase K resistance) are difficult to compare between laboratories and work best with side-by-side controls. This is why developing a facile method for polymorph identification is one of our main future research goals.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes insights gained during efforts to reproduce disease-relevant alpha-Synuclein (aSyn) fibrils using recombinant protein in vitro. It follows up on a similar article from this team published in 2024. Although the authors have not been able to produce fibrils with the structure of ex vivo fibrils isolated from patients, they share insights gained into which factors influence the formation of specific fibril polymorphs.

      Strengths:

      This is quite an unusual manuscript because it goes into minute detail about sample preparation that are usual just mentioned in the Materials and Methods sections of other manuscripts (if at all). This makes it very valuable for the scientific community working on exactly the problem of reproducing disease-relevant aSyn fibrils in vitro (which will be a major breakthrough in the field). The authors present an impressive array of cryo-EM fibril structures, some of which have not been described before.

      Weaknesses:

      A major concern with the manuscript is that its story and messaging are a bit murky. The authors describe a few new polymorphs, show that some polymorphs (type 1) have small variations, show that sample purity and fragmentation will influence polymorph formation, and present a helical-symmetry mystery. This all reads like a loose collection of findings without any major takeaway. Looking at Table 1, it still seems that the authors do not have a good control over any of these polymorphs. Are they able to make any of these polymorphs reliably? I think the impact of this work could be strengthened if it ended with a reliable protocol for the production of any of the polymorphs described.

      It is true that the initial results were a collection of findings compiled while searching for the PD polymorph. However, the trends that we saw in these data inspired us to pursue a more systematic approach which was used to show how impurities play a crucial role in polymorph selection even at low levels. We agree that the impact of the work will be enhanced by summarizing the protocols that can be used to obtain the types 1, 2, 3 and 5 polymorphs and will add this to the revised manuscript.

      A second major concern is the quality of the aggregation kinetics and their interpretation. Are these kinetics just done once per concentration? The figure caption talks about 'three independent samples' but it is unclear if this refers to the three different concentrations or NΔ4 percentages used, or true repetitions. Looking at the curves themselves, it seems that only one of the conditions was actually done in triplicate, which should be the minimum to draw conclusions. Further, it would have been helpful to characterize the kinetic data quantitatively. Finally, because there are only kinetic data for a fraction of the conditions tested, it is not clear what they add to the overall manuscript. My recommendation is to either remove the kinetics from the manuscript or substantially expand this section.

      In Figure 6A, the three curves represent three independent samples; in panels B–D, each curve similarly represents one independent sample. We will try to clear up the ambiguity in the revised version. We agree that the kinetic data have limited utility for determining kinetic parameters of the aggregation. The data were collected primarily to determine when the aggregation reactions were complete so that samples could be prepared for cryo-EM. Nevertheless, we think that the kinetic traces provide two qualitative observations that are relevant to the structural results: 1) In panels A and C of figure 6, the type 5 polymorphs are associated with shorter lag times, suggesting that they either nucleate faster, or as we propose, arise from a small amount of oligomeric protein in the original sample. 2) The longer lag phase associated with the NΔ4 construct was reproducible (panels B, C and D), suggesting that its intramolecular self-chaperoning effect is enhanced due to the increased positive charge in its N-terminal region. Future studies will determine whether this same electrostatic change also gives rise to enhanced secondary nucleation.

    1. eLife Assessment

      This useful study reports potential loss-of-function variants, pseudogenes and gene presence-absence variation across multiple chicken genomes, with potential implications for understanding genome evolution and domestication. The evidence for the central claims is unfortunately incomplete, as the inferences of gene loss are not sufficiently robust to account for assembly and annotation artifacts, and, in addition, the analyses can not distinguish between positive selection and relaxed constraint. The overall claim of large-scale gene loss being adaptive and thus being a major driver of chicken evolution and domestication is therefore not sufficiently supported. The area of the study is of interest to colleagues in evolutionary and comparative genomics as well as animal domestication.

    2. Reviewer #1 (Public review):

      Summary:

      The authors have assembled the genome of four local chicken breeds from China and analysed their gene content. They come to the conclusion that thousands of genes present in the current reference genome of chicken have become pseudogenized during chicken evolution and domestication. They argue that their study provides strong support for the importance of the "less-is-more" hypothesis for adaptive evolution.

      Strengths:

      The paper provides medium-quality genome assemblies for four individuals representing four local populations of chickens and analyses their gene content.

      Weaknesses:

      They have not excluded the possibility that the high rate of putative pseudogenes reflects the presence of errors in gene models, in particular in GC-rich microchromosomes that are challenging to assemble correctly. The paper contains no genotype-phenotype analysis, which means that the adaptive significance of a high rate of pseudogenization, if it exists, is unknown.

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to investigate the evolutionary role of gene presence-absence variation and pseudogenization in chicken evolution and domestication. By comparing draft genome assemblies of four indigenous Chinese chicken breeds against the red junglefowl reference genome (GRCg6a) and four PacBio HiFi assemblies, the study proposes that the common ancestor possessed nearly 22,000 genes, and that each domestic lineage independently lost thousands of genes (identifying ~8,000 dispensable genes). The authors conclude that massive loss of function and pseudogenization represent major drivers of chicken evolution under the "less-is-more" hypothesis.

      While the concept that gene loss can drive phenotypic diversification during domestication is compelling, the results do not convincingly support the central conclusions. The scale of reported gene loss and the specific patterns of pseudogenization appear to be potentially driven by well-known genome assembly gaps, annotation errors, and sequencing dropouts rather than genuine evolutionary events, and the authors do not provide enough convincing evidence that this is not the case.

      Strengths:

      (1) The study addresses an important and timely evolutionary question regarding the role of gene loss and loss-of-function variation in animal domestication.

      (2) The inclusion of multiple indigenous Chinese chicken breeds alongside high-accuracy PacBio HiFi assemblies provides a valuable comparative genomic dataset.

      (3) The authors attempt to evaluate pseudogene transcription using RNA-seq data and check for transcript isoforms that bypass candidate loss-of-function mutations.

      Weaknesses:

      (1) Unvalidated gene and pseudogene annotations: Long-read and consensus genome assemblies are known to suffer from residual indel errors that create artificial frameshifts and premature stop codons. The frequency of predicted pseudogenes in this study (~3.5%-4.5%) aligns closely with expected baseline annotation error rates. Although the authors state in their Methods that mutations were validated using short reads, this validation is never quantitatively demonstrated or shown in the results. The fact that most proposed pseudogenes are actively transcribed and lack a paralog strongly suggests that many are intact, functional genes affected by sequencing or annotation artifacts.

      (2) Assembly gaps and GC-bias mistaken for gene loss: The claim that ancestral chickens possessed ~22,000 protein-coding genes and lost thousands of genes in only ~10,000-50,000 years is inconsistent with the evolutionary conservation of avian genomes. The missing genes are enriched for high GC content and preferentially located on microchromosomes. Avian microchromosomes and dot chromosomes are notoriously GC-rich, repeat-dense, and prone to severe assembly gaps in non-telomere-to-telomere assemblies. The reported gene absences reflect assembly fragmentation and coverage dropouts rather than evolutionary deletions.

      (3) Lack of synteny validation: Genuine gene absence requires demonstrating conserved collinear synteny of flanking orthologous genes with an unambiguous sequence deletion at the locus. Relying on sequence alignment or short-read mapping failures across fragmented scaffolds substantially inflates false-positive gene loss calls.

      (4) Positional bias of pseudogenization mutations: The observed concentration of pseudogenization mutations in the terminal 10% of coding sequences (the "bathtub" distribution) is characteristic of alignment boundary artifacts and non-canonical translation start/stop annotations, rather than positive selection to disrupt gene ends. Mutations in the terminal 3' region often produce functional proteins with slightly altered C-termini rather than complete loss of function.

      (5) Inconsistent terminology: The manuscript alternates between identifying pseudogenes as unitary (lacking a functional paralog in the same genome) and evaluating sequence identity against "parental genes," creating substantial confusion regarding whether loci are duplicated paralogs or orthologous reference genes.

    4. Reviewer #3 (Public review):

      Summary:

      The authors reanalyze genome assemblies of four indigenous chicken breeds from Yunnan Province together with the red jungle fowl reference (GRCg6a), and search for genes with disrupted protein-coding sequences. They catalog candidate pseudogenes and missing genes to estimate that 7,993 of the ancestral genes are dispensable. They characterize the positional distribution of pseudogenization mutations along coding sequences, their fixation in breeds, their estimated ages, and their pathway enrichments. Most results are replicated in four independent PacBio HiFi-based chicken assemblies. From the biased position of pseudogenization mutations toward CDS ends, their frequent fixation, and the phylogenetic signal in gene-loss patterns, the authors conclude that large-scale loss of function is a major driver of chicken evolution and domestication, consistent with the "less-is-more" hypothesis.

      Strengths:

      The catalog itself is a substantial resource: the comparison spans multiple closely related genomes, the main patterns are checked in a second, independently assembled set of HiFi genomes, and population resequencing data are used to ask whether the pseudogenization mutations are fixed rather than segregating. The finding that candidate loss-of-function genes are under relaxed purifying selection is well supported. The question of what gene loss contributes to domestication is worth asking, and this is a useful dataset for asking it.

      Weaknesses:

      The finding that genes carrying disruptive mutations are under relaxed selection is not particularly surprising, and the more interesting claim, that the observed patterns reflect positive selection for gene loss, is less certain in my opinion.

      (1) Relaxed purifying selection versus positive selection. The bias of pseudogenization mutations toward the two ends of coding sequences is interpreted as positive or artificial selection, along with elevated dN/dS. But the alternative, that disruptive mutations at gene ends are simply better tolerated, is equally consistent with the results. Several mechanisms would produce this pattern under relaxed constraint without positive selection per se: alternative downstream start sites that rescue 5-prime disruptions; the small fraction of protein truncated by 3-prime disruptions; and the enrichment of disordered regions at protein termini.

      (2) The functional status of the pseudogenes is assumed, not demonstrated. Genome-scale work cannot be expected to validate individual genes, but the language of the paper should reflect the candidate status of these calls. Nearly all predicted pseudogenes (~95%) were reported as transcribed in multiple tissues. It is possible for a nonfunctional coding sequence to retain intact regulatory sequences, but the observation deserves more attention in the paper, particularly because transcripts carrying premature termination codons can be the targets of nonsense-mediated decay, which is not discussed. These remain candidate pseudogenes defined by the presence of a putatively large-effect mutation (e.g. premature stop or frameshift).

      (3) What "missing" means. For a gene to be scored as completely absent, it could be genuinely deleted, or its allele could be diverged enough that annotation and orthology/mapping no longer detect it. These are different phenomena. Related, since most pseudogenization mutations are reported as fixed or nearly fixed in their populations, the history of alleles matters: it is not clear how the authors established which state is derived, and whether the reference sequence assumed to be functional is in fact the "functional" version.

      (4) Limited biological insight into domestication. The main biological interpretation rests on hierarchical clustering of dispensable genes followed by ontology enrichment within clusters (Figure 7a), with narrative connections to breed phenotypes. Only 19.6% of dispensable genes have Gene Ontology assignments, and the phenotype links are speculative. The section on the subspecies origin of the GRCg6a reference is only loosely connected to the loss-of-function story, and the population-genetic analysis supporting it is thin as described.

    1. eLife Assessment

      The observations in this study related to a pleiotropic EPAS1 enhancer that mediates adaptation to hypoxia in adipocytes of Tibetans are a valuable contribution. The data are solid, but additional experiments would strengthen the claims.

    2. Reviewer #1 (Public review):

      In the article, the authors set out to characterize in adipocytes an enhancer, ENH5, of the gene EPAS1, a gene that was found to show strong selection in Tibetan populations. They investigate whether this enhancer contributes to adipocyte response to environmental stress. The authors show that ENH5 is active in preadipocytes and that the Tibetan high-altitude allele confers reduced activity. They then use a mouse ENH5 knockout model to show a hypoxia/thermogenesis responsive phenotype of stronger transcriptional downregulation of aerobic respiration, electron transport chain, and adipogenesis pathways. The authors interpret these findings as evidence that ENH5 conditionally regulates adipocyte energetics and thermogenic response, potentially favoring energy conservation in Tibetans exposed to the demands of high-altitude hypoxia and cold. Overall, the paper presents an interesting potential connection between EPAS1-mediated high-altitude tolerance and energy metabolism; however, more work needs to be done to establish this connection.

      Major comments

      (1) The authors use mouse ENH5 enhancer knockout (ENH5-KO) as the model of the Tibetan EPAS1 locus because the high-altitude allele of human ENH5 has lower transcriptional activity than the low-altitude allele (Figure 1A) and mouse ENH5 (musENH5) has enhancer activity (Figure 1D) in mouse preadipocytes. However, it is an overstatement to claim the functional role of Tibetan ENH5 haplotype only based on these data because musENH5 is neither identical to human ENH5 nor the murine high-altitude haplotype. The title should also be revised to better reflect the function of ENH5, like "An EPAS1 enhancer mediates hypoxic and cold response in mouse adipocytes". The authors should consider in some way to actually show that the Tibetan haplotype in ENH5-KO leads to expression changes. This could be done by inserting the haplotype into preadipocytes via CRISPR (realize this is a tough one) or if they have available cells from Tibetans or some eQTL or other similar datasets. The more closely they can connect this haplotype to EPAS1 expression, the more beneficial it would be for the article. As it stands, they currently have episomal luciferase assays showing reduction of enhancer activity in mouse preadipocytes of a human allele and a complete knockout of the mouse enhancer that doesn't recapitulate the Tibetan haplotype. A bit more work is needed to connect all of these to the Tibetan adaptation. As it stands now, this is all very circumstantial.

      (2) In Figure 2A, the body weight of ENH5-KO normal diet is significantly lower until four weeks in male, and until 11 weeks and 18 and 19 weeks in female than that of WT. These are slight but significant differences between ENH5-KO and WT; therefore, the authors should describe and discuss this and how it could affect their results.

      (3) For the mouse work, it is not clear why the authors did not do cold-exposure or some type of hypoxia experiment for the mice themselves. This will be helpful to support their claim, and if not done, or done without significant differences in the results, the authors should add and mention this. The RNA-seq work, while substantial, again provides circumstantial support.

      (4) The authors used CL316243 as a β3-AR selective agonist to mimic thermogenesis in vitro. In humans, it is not β3-AR but β2-AR that mainly drives thermogenesis (Blondin et al., Cell Metabolism, 32, 287-300. e7). Therefore, the authors should describe the limitation due to the difference in mechanisms of action of thermogenesis between humans and mice, as they use the mouse cells as a human model.

      (5) The authors note in the discussion that Figure 3's CL316243 stimulation intended to simulate a thermogenic reaction to cold temperatures also generated a change in OXPHOS pathways associated with hypoxia, thus making it difficult to separate the contribution of β3-adrenergic/thermogenic effect from an indirect local hypoxia response. The authors could further interrogate this effect by measuring canonical hypoxia-responsive genes, oxygen consumption, or performing an in vivo cold challenge.

      (6) Figure 4: The authors mention that reduced aerobic respiration pathways are evidence for reduced thermogenesis, but it does not directly demonstrate altered thermogenesis correlates. They do not measure heat synthesis, oxygen consumption/respiration, uncoupled respiration, UCP1 protein levels or activity, mitochondrial changes, etc. Their evidence for changed thermogenesis stems from transcriptional changes in energy consumption pathways shared with hypoxia changes. They should tone down their findings.

      (7) In the Discussion, the authors should interpret and discuss their data carefully. For example, in the GSEA analysis, the authors identified enriched gene sets in ENH5-KO. Therefore, the authors should discuss which genes might contribute to each pathway, since some genes show large logFC changes. In the current discussion, there is little mention of their results, and it mostly focuses on prospects.

    3. Reviewer #2 (Public review):

      Summary:

      This study extends previous work on the adaptive EPAS1 locus by examining the pleiotropic activity of the ENH5 enhancer in adipocytes and its potential role in metabolic and thermogenic responses. The authors progress from demonstrating enhancer activity and evolutionary conservation to in vivo phenotyping and environmentally dependent transcriptional responses in primary adipocytes. The work provides an interesting example of how pleiotropic regulatory effects can contribute to the complexity of adaptation, with a single adaptive regulatory locus influencing multiple biological processes in an environmentally dependent manner.

      Strengths:

      The study is well executed and is clearly presented, with a logical experimental progression. A particular strength is the genotype-by-treatment interaction analysis demonstrating that ENH5 loss alters metabolic and adipocyte-associated transcriptional programs following both hypoxia and beta-3-adrenergic stimulation. The convergence of these responses is particularly interesting in the context of regulatory pleiotropy and suggests that selection at the EPAS1 locus may have consequences extending beyond the canonical hypoxia response. The inclusion of negative findings, including the absence of an overt baseline metabolic phenotype and the lack of an additive response to combined stimulation, also provides a balanced presentation of the results.

      Weaknesses:

      The principal conclusions are generally supported by the data, and the weaknesses are relatively minor and primarily relate to the scope of interpretation. The study assesses transcriptional programs associated with thermogenic signaling using CL316243 rather than directly measuring physiological thermogenesis. Nonetheless, CL316243 is a rational and well-established approach for experimentally inducting beta-3-adrenergic thermogenic signaling. In addition, the murine ENH5 knockout is a useful model of reduced enhancer activity that phenocopies the Tibetan ENH5 haplotype, but it is not genetically equivalent to the naturally occurring Tibetan ENH5 haplotype. The authors generally recognize these limitations, and they do not substantially detract from the central findings.

    1. eLife Assessment

      This study provides a valuable and comprehensive in vivo map of proteins in proximity to three vesicle-tethering complexes. The evidence supporting the overall proximity mapping is convincing. However, some aspects of the functional characterization of selected candidates remain incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      The CATCHR complexes are a family of five multisubunit complexes that act in vesicle tethering in several key transport steps. COG and GARP act at the Golgi, EARP acts on early endosomes, EXOCYST acts in transport to the cell surface, and DSL1 acts on the endoplasmic reticulum (ER). The authors have used in vivo proximity biotinylation and mass spectrometry to look for new neighbours of three of these complexes, COG, GARP, and EARP (despite the title, only two of the three are on the Golgi). The authors then follow up two of the hits, CCDC186 and WWOX, with more directed experiments.

      Strengths:

      The strength of the paper is that proximity biotinylation is of a high standard. To avoid overexpression, the authors express TurboID-tagged CATCHR subunits in cell lines from which the genes have been deleted. This allows them to confirm that the tagged proteins are functional and correctly located. Mass spectrometry is used to identify the proteins biotinylated in each cell line with four replicates, and the data are clearly presented in figures and supplementary tables. The authors make a good choice of proteins to follow up, as both CCDC186 and WWOX appear potentially interesting.

      Weaknesses:

      Overall, although the paper is based on a high-quality initial set of data, it seems somewhat incomplete and preliminary. There is undoubted value in presenting a through if descriptive set of in vivo proximity labelling data for a comprehensive overview of proteins or complexes. However, in this case the authors have only addressed three of the five CATCHR complexes, and so it is not a complete overview. It was also somewhat unclear why they examined VPS52 and VPS53, as these are present in both GARP and EARP, which adds some ambiguity, even if they did at least provide useful confirmation of some of the hits with EARP. Of course, a proximity biotinylation analysis does not need to cover all members of a family if it generates substantial biological insight.

      However, the investigation of CCDC186 and WWOX does not provide significant insight into either function or mechanism. CCDC186 has already been identified in C. elegans as a protein involved in dense core vesicle biogenesis (CCCP-1, as the authors acknowledge), and work in C. elegans and mammalian cells has already linked it to EARP function. Less has been published on WWOX, but all that is found here are some small changes in Golgi appearance and glycosylation when it is reduced by RNAi. Thus, the paper falls between two stools: it is neither a comprehensive application of proximity biotinylation to the CATCHR family, nor is it the application of proximity biotinylation to reveal new insight into membrane traffic. I feel that for a broad-interest journal such, it should be one or the other of these.

    3. Reviewer #2 (Public review):

      In this manuscript, Aragon-Ramirez et al. present the first systematic proximity-interaction map of the three Golgi/endosomal CATCHR tethering complexes - COG, GARP, and EARP. They generated hTERT-RPE1 knockout lines rescued with C-terminally TurboID-tagged subunits (COG4, COG6, VPS50, VPS52, VPS53, VPS54) expressed from the COG4 promoter at near-endogenous levels, verified that each construct rescues its KO phenotype, and confirmed expected localization by immunofluorescence. Fifteen-minute biotin pulses followed by streptavidin capture and label-free DIA mass spectrometry, benchmarked against GFP-TurboID, yielded compartment-resolved neighbor lists. The central claim of this work is that each complex sits within a distinct "trafficking module" of coiled-coil tethers (CCTs), SNAREs, SM proteins, Rab GTPases, coats, and homeostasis regulators. COG neighbors nearly all Golgi golgins plus the STX5-SCFD1 fusion machinery, with COG4 and COG6 lobes showing overlapping but non-identical hierarchies - this is offered as in vivo support for a two-lobe model. GARP associates with TGN golgins GOLGA1/GOLGA4, STX16 and partners, TBC1D23, and CLINT1. EARP associates with GRIPAP1, RELCH, RAB11FIP5, and the VPS33B-VIPAS39 (CHEVI) SM complex; a VPS53 MUN-domain mutant loses VPS33B proximity while retaining VPS50 labeling. Two hits are followed up in this study. CCDC186, a poorly characterized coiled-coil protein, localizes to the TGN and, when ectopically anchored to mitochondria, captures ~60 nm vesicles - presented as direct evidence of tethering activity. WWOX, a tumor suppressor with no prior trafficking role, colocalizes with COG8 in the medial Golgi; siRNA knockdown reduces Golgi area, increases HPA and GNL lectin binding (O- and N-glycosylation defects), and displaces COPB2 from the Golgi. Based on these findings, the authors conclude that CATCHRs are not isolated tethers but organizing hubs that assemble compartment-specific tethering-and-fusion modules. Overall, this is a solid piece of work that will be of interest to cell biologists.

      (1) The CCDC186 mitochondrial assay demonstrates sufficiency in a non-native context but includes no loss-of-function work, and the cargo-specificity result carrying most of the interpretation is "data not shown." At a minimum, the authors should show KO/KD and rescue data.

      (2) The WWOX section rests on a single siRNA with no second oligo, no rescue, knockdown shown only by RT-PCR, and localization based entirely on overexpressed myc-tagged protein.

      (3) Proximity of labeling does not distinguish direct from indirect interactions. When referring to "interactions", the authors either need to show recombinant protein binding data, or soften the tone to acknowledge potential indirect interactions.

      (4) The VPS53 MUN-domain mutant design rests on a "manuscript in preparation," and the mutant's localization isn't shown (VPS50 labeling establishes complex incorporation, not correct targeting). The HEK293T WWOX replication is "data not shown." The authors should either show the data or delete these claims.

      (5) References 73 and 78 are duplicates.

      (6) "NZR" is used in the introduction. Should this be NRZ?

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive proximity proteomics analysis of Golgi- and endosome-associated CATCHR complexes, including COG, GARP, and EARP, using near-endogenously expressed TurboID-tagged subunits. The study aims to define how these tethering complexes organize distinct trafficking modules involving coiled-coil tethers, Rab-associated proteins, SNAREs, SM proteins, and other trafficking regulators. In addition to validating known associations, the authors identify CCDC186 as a candidate vesicle tether and WWOX as a potential regulator of Golgi homeostasis and glycosylation. The work provides a broad resource for understanding the spatial organization of CATCHR-associated trafficking networks. However, several conclusions require further experimental and statistical support, particularly those interpreting proximity-labeling data as physical interactions or evidence of discrete functional complexes.

      Strengths:

      The manuscript addresses an important question in membrane trafficking and provides a systematic comparison of the proximity interactomes of COG, GARP, and EARP complexes. The use of near-endogenously expressed TurboID-tagged subunits is a strength, as it may reduce artifacts associated with protein overexpression. The dataset is comprehensive and has the potential to serve as a valuable resource for investigators studying Golgi and endosomal trafficking. The comparative analysis identifies both known trafficking factors and potentially novel regulators, including CCDC186 and WWOX. The functional follow-up experiments add biological relevance to the proteomic findings and extend the study beyond descriptive mapping. Overall, the scope is well defined, and the manuscript proposes an interesting model in which CATCHR complexes act as organizing hubs for vesicle tethering and fusion.

      Weaknesses:

      A major limitation is that TurboID proximity labeling is repeatedly interpreted as evidence of physical interaction or assembly into discrete complexes, although the method primarily reports spatial proximity. The soluble GFP-TurboID control may not adequately account for enrichment caused by membrane confinement and high local protein concentration in Golgi/endosomal microdomains. Some spatial enrichment analyses do not reach statistical significance, weakening claims of compartment-specific labeling. Several low-fold-change SNARE hits are emphasized despite limited enrichment. Additional quantitative imaging is needed, including colocalization analysis of streptavidin labeling with Golgi/endosomal markers and localization of GFP-TurboID. The WWOX knockdown phenotype requires validation by rescue, independent siRNA, or CRISPR-based approaches. Several supporting data, quantifications, figure corrections, and nomenclature revisions are also needed.

    1. eLife Assessment

      This important study investigates how GDNF treatment restores nerve cells in the bowel of a mouse model of Hirschsprung disease, and proposes that this regeneration occurs through diverse resident precursor cells and a non-canonical signaling pathway. The evidence is solid for the ability of the treatment to induce new bowel nerve cells and for a contribution from multiple precursor populations, supported by extensive cell tracing and imaging experiments. This work will be of interest to researchers in developmental biology and peripheral neurobiology.

    2. Reviewer #1 (Public review):

      Gary, Soret et al., present an interesting and ambitious study investigating the cellular and signaling mechanisms by which GDNF induces enteric neurogenesis in a disease model of HSCR. The manuscript combines scRNA-Seq, pharmacological inhibition, and several lineage-tracing strategies. The lineage tracing experiments in particular are technically impressive and clearly represent a substantial amount of work. The figures are generally beautiful and high quality, and the biological question is an important one.

      However, while I found the study interesting, I also think that several of the major conclusions are substantially stronger than what the current data support. In particular, the scRNA-Seq dataset that the study is founded on is difficult to interpret because of the RFP-based enrichment strategy, the lack of WT reference controls, and the very non-stringent FACS gating. Several mechanistic claims are also based primarily on expression patterns or sequencing results.

      Major concerns:

      (1) The G4-RFP based enrichment strategy raises some concerns. The authors note that 88-90% of SOX10+ glial cells (where did this number come from, it doesn't seem to match the range in the bar graph?) and only 70% of GDNF-induced neurons. They then sort in "yield" mode using a non-stringent gating strategy, but do not provide the flow graphs in the supplements for interpretation. Consistent with this, the initial dataset contains multiple non-ENS populations, including lymphoid, myeloid cells etc, suggesting that their strategy, in addition to missing cell populations, additionally includes potential negative cells based on background fluorescence.

      (2) It's not clear to me why the authors did not include a wild-type control for comparison. Due to the RFP-enrichment strategy, it is additionally difficult to compare to integrated with scRNA-seq datasets. The comparison seems important because the central question is not only whether GDNF induces neurons in the mutant colon, but whether the induced cells and associated populations resemble those found in a normal ENS.

      (3) The proportional changes in Figure 2 are hard to interpret. The authors write that GDNF treatment leads to an enrichment in enteric neurons, "at the expense of" SCPs. This seems to be a strange conclusion or writing, unless the authors are suggesting that certain cell populations are additionally depleted by GDNF treatment. I think the authors should provide stronger support, or dial down on conclusions based on cluster proportions, given the sorting strategy.

      (4) Some of the mechanistic conclusions are far too strong for the evidence. For instance, the claim that GDNF signals through GFRa1/2 based on only expression UMAPs is not justified. The authors should at least confirm some of the major findings via smFISH (HCR or RNA-Scope), and dial down the mechanistic claims. The pharmacological data is more supportive but still indirect. This section would be stronger if the authors validated NCAM1/FAK activation in the specific progenitor populations proposed to respond to GDNF.

      (5) The inhibitors are administered during the same P4-8 window as GDNF. Given that the authors report cellular-level responses as early as 6 hours post GDNF treatment, it is important to know whether the inhibitors were already active at the time of GDNF administration. Why were the inhibitors not administered prior to the start time of GDNF treatment?

      (6) The authors draw inferences and conclusions based on IHC images, but do not provide any details on the analysis procedures. For example, in Figure 3, the authors interpret changes in NCAM and FAK expression level. However, it does not appear that there are notable differences in intensity between the two timepoints. There are however, differences in signal coverage. Without knowing what was actually measured and quantified, it is not possible to interpret the results. Please provide the necessary information for all experiments in the methods section.

      (7) The final results section introduces one of the more surprising claims in the paper, that a non-neural crest-derived progenitor contributes to regeneration, but the section ends quite quickly and abruptly after introducing the observation. The authors should provide additional validation or interpretations of the data.

    3. Reviewer #2 (Public review):

      Summary:

      Previous work from this group demonstrated regeneration of the ENS following exogenous GDNF treatment within the aganglionic portion of the bowel in several Hirschsprung (HSCR) murine models as well as colon from HSCR patients (Soret, et. al., Gastro. 2020). Focusing on the Holstein (HolTg/Tg) HSCR mouse model, the authors build upon their prior work utilizing lineage tracing, immunohistochemistry, scRNA sequencing, and pharmacological methods to further delineate the mechanisms and cell types contributing to ENS regeneration in this context. This work adds to the knowledge base regarding natural ENS development as well as ENS generation outside of natural ENS development. Furthermore, this work provides important clinical implications for potential curatives treatments for HSCR disease.

      The conclusions in the paper are overall well supported by the data with the majority of technical and model limitations openly acknowledged. Strengths include the use of scRNA sequencing, several genetic mouse models with ample, well-planned time points within experiments. However, additional analysis, clarity or more detail around experimental protocols, and broader discussion of prior studies would strengthen the manuscript by improving interpretation of the findings and placing them more clearly within the context of existing work in the field.

      First, utilizing the G4-RFP transgene for ENS cell selection and subsequent scRNA sequencing of RFP+ cells, the authors nicely demonstrate the presence of known cell clusters (based on markers and analysis from prior sequencing studies) including Schwann cell precursors (SCPs), enteric glial cells (EGCs) and neuron subtypes in the aganglionic colon of HolTg/Tg GDNF-treated and untreated mice. Further use of this dataset led to the discovery of Ncam1 expression, a GDNF receptor, within all ENS cell clusters, suggesting importance of this protein in the GDNF-induced ENS regeneration phenotype. Interestingly, Ret was limited to some neuron types. On the protein level, temporal and GDNF-treatment changes in NCAM1 expression and downstream phospho-FAK(Y397) were nicely demonstrated via Western blot. Mechanistically, the importance of NCAM1 in ENS regeneration was demonstrated given pharmacologic blockade of downstream FAK signaling via PF-562271 resulted in significantly fewer GDNF-induced neurons. The paper would be strengthened here, however, if the following areas were addressed:

      1) Although the GDNF-induced regeneration appears likely through an NCAM1 signaling mechanism, inhibition of FAK phosphorylation by PF-562271 could affect more than NCAM1 signaling. Blocking NCAM signaling through another method (such as through NCAM1 targeted monoclonal antibodies or small molecules) would support the authors' interpretations further.

      2) A point of ambiguity is the exact cells included in the final scRNAseq analysis. The authors noted they flowed and gated broadly, including RFP negative cells, in the initial set of cells selected for sequencing (given RFP expression is not visualized within all neurons and glia in GDNF-treated HolTg/Tg mice.) Pan-neuronal (Elavl4, Tubb3) and pan-glial (Sox10) marker expression were utilized to select cells for final analysis (unsupervised analysis, pseudotime trajectories, etc.) Here, it is unclear if RFP expression was examined in these cells, either by looking at unmapped reads or if included in the initial alignment. If this is discernable within the authors' data set, this provides an opportunity for an interesting additional analysis to determine if progenitors or neuronal or glia populations differ in the RFP+ and RFP- groups.

      3) The authors presumably sequenced both the submucosal and myenteric plexus ENS components within their scRNAseq analysis. Given the relatively sparse numbers of ENS cells in the submucosal plexus compared to myenteric plexus in the distal colon, this is unlikely to change their analysis or conclusions. However, given their tissue analysis focuses on the myenteric plexus, this difference should be acknowledged.

      The authors go on to further to utilize their scRNA sequencing data set with pseudotime trajectory analysis to determine that SCPs appear to go through an EGC-like state prior to GDNF-induced neurogenesis. This was followed nicely by use of transgenic mice (Dhh-Cre) and immunohistochemistry within their HolTg/Tg model to capture the SCP population and demonstrate in vivo this progression. Additionally, they were able to demonstrate that temporally, SCPs generated neurons much sooner (within 6 hours of GDNF-treatment) compared to neural crest derived EGCs (labeled by Slc18a2-Cre and GFAP-CreERT2 lines).

      Fascinating here as well is that the authors unearthed that most generated neurons following GDNF-treatment appeared to derive from direct cell transdifferentiation (suggested by EdU incorporation investigation at various time points) and that these neurons did not seem to originate from SCPs or EGCs (i.e. a non-neural crest origin.) Acknowledging that their mouse genetic tools and sequencing could have missed a NCC-derived population, the authors go on to label neural crest derivates utilizing the Wnt1-Cre2 transgene in their wildtype and HolTg/Tg mouse model and find that while Wnt1 labels the vast majority of neurons in wildtype mice, nearly 25% of neurons in P20 HolTg/Tg mice appear non-NCC derived. Finally, once again utilizing various genetic mouse lines for lineage tracing and immunohistochemistry, the authors demonstrate NCC-derived neural progenitors (SCPs, EGCs) appear to preferentially give rise to cholinergic neurons where non-NCC derived neurons tend to be more nitrergic.

      4) The high magnification of the majority of images provides the reader with extremely compelling evidence in regard to visible overlap or lack thereof across various reporter lines, antibody markers, and EdU labeling. Unfortunately, this can somewhat create a trade off with the area or number of cells examined which appears somewhat less than the number of cells and/or area examined by many prior studies in the field. The near equal averages across the three mice in each group in many of their experiments makes this less of a concern, but the paper may be strengthened by noting exactly how the 3-11 images per mouse were selected. Were these randomly selected across the tissue, moving from mesenteric to anti-mesenteric? Or caudally to distally?

      5) In line with this, the authors report a near 100% of Wnt1-Cre2 driven reporter expression in wildtype mice but a large portion (~25%) of neurons in HolTg/Tg model appear non-NCC derived (i.e. not Wnt1-Cre2 labeled.) The original Wnt1Cre line (Danielian, et.al., Curr Biol. 1998) has been reported to incompletely label all neural crest derivatives/ENS (Hari, et. al., Development. 2012.; Deal, et. al., Dev Bio. 2021). To my knowledge, the same type of analysis has not been carried out with the same rigor in the ENS with the Wnt1-Cre2 line. It could be -- depending on tissue sampling, area covered via imaging, and if there is any patchy or unique patterning to non-NCC derived neuro regeneration -- any incomplete labeling of neural crest by Wnt1-Cre2 could be missed. Of course, alternative interpretations are that the non-labeled cells are truly non-NCC derived. Utilizing other lineage tracing models for other non-NCC derived lineages may help address this issue, but would require many mouse crosses and additional experiments, and thus likely outside the scope of this manuscript. However, this caveat should be addressed.

      6) A very important and fascinating find by the authors is that the NCC versus non-NCC derivatives appear to preferentially give rise to cholinergic or nitrergic neurons respectively. This will have important implications in treatment options for ENS regeneration not only within HSCR disease, but other ENS disorders as well. A point of discussion the authors missed out on is several previous reports regarding skewing of ENS neurons (typically toward higher nitrergic numbers) in ganglionic and/or hypoganglionic segments of HSCR bowel in murine models and patients (Zaitoun, et.al., Neurogastroenterol. Motil. 2013.; Musser, et.al, CMGH. 2015.; Cheng, et.al., J Pediatr Surg, 2017.; Sukhada, et. al., Front. Cell Dev. Biol. 2022.) Could ganglionic portions of the bowel in HSCR disease contain more of the non-Wnt1 labeled population? Why do proportions in the aganglionic region generate correct proportions following GDNF-induced treatment compared to ganglionic bowel? This does not need to be experimentally addressed in the manuscript, but including in the discussion impresses upon the audience the importance of the authors findings and that these factors need to be further experimentally delineated and considered with any therapeutic interventions.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors aim to build on their previous work to further define the mechanism by which GDNF can induce enteric nervous system migration and differentiation in the Holstein mouse model (Hol Tg/Tg) of Hirschsprung's disease (HSCR). They show that GDNF can cause neuronal differentiation from either Schwann cell precursors (SCP) that migrate from the periphery and/or enteric glia (or neural crest precursors located in the intestine). Surprisingly, the rescue effect of GDNF is mediated through signaling with NCAM1 rather than the canonical receptor RET, with blockade of NCAM1 significantly reducing the number of GDNF-treatment induced neurons in the distal colon of Holstein mice. The authors then perform scRNA-seq data of the colonic ENS from Holstein mice +/- GDNF, and pseudotime analysis of this data indicates the SCP downregulate SCP genes and upregulate classical enteric glial genes on their transition to becoming enteric neurons. Following this, the authors employ Dhh-Cre mice (to label SCP), Slc18a2Cre and GFAPCreERT2 (to label enteric glia) and finally Wnt1Cre2 mice (to label neural crest derived cells) to assess the kinetics and end identity of enteric neurons driven by GDNF treatment. Using these animal models, the authors propose a model by which transdifferentiation of SCP gives rise to enteric neurons first, followed by transdifferentiation of enteric glia; the resulting enteric neurons labelled by Dhh, Slc18a2 or GFAP Cres are all biased to a cholinergic lineage, raising a questions about the source of nitrergic neurons. Finally, the authors postulate that at least some nitrergic neurons come from a non-neural crest lineage in GDNF treated Hol Tg/Tg mice as they are not labelled by a Wnt1Cre2 model, while all ENS cells are labelled in three wildtype Wnt1Cre2Tg/+;R26YFP/+ mice.

      Strengths and Weaknesses:

      This work convincingly shows that NCAM1 is required for GDNF-induced rescue of the Holstein mouse model of HSCR and, along with the NCAM1 staining in human HSCR tissue, highlights the potential of this treatment even in HSCR patients with loss of function RET mutations. The major weakness of this study is in the interpretation of the mouse models; many of these models are not fully penetrant/show mosaicism in other studies (like the DhhCre and Wnt1Cre) or may not only label the specific subpopulation they present (in this case, the GfapCreERT2 and Slc18a2Cre; these genes have been shown in Schwann cells/SCP in other studies, and the expression pattern of these two genes outside of the gut at the relevant developmental timepoints are not shown). While these are unavoidable limitations with our current mouse models, the authors need to clearly acknowledge these limitations in the interpretation of their data (with the DhhCre and Wnt1Cre) or more definitively show the models are only labelling enteric glia (in the case of the Gfap and Slc18a2Cre). In addition, in some figures the authors should add untreated Hol Tg/Tg and ideally WT mice so readers can contextualise the effect of GDNF on Hol Tg/Tg mice. Finally, it appears the statistics throughout the paper are performed on using each ROI taken from n=3-5 animals - this is inappropriate, as images from the same animal are not independent samples. The statistical unit in question for every analysis performed should be 1 animal, thus the statistics should be done on the n=3-5 animals, not the n=3-11 images from 3-5 animals (which leads to an n of 6-55).

      Specific comments:

      In general I find the author's conclusions are relatively well supported, albeit a little overstated for the evidence presented in the later figures and with the need to correct the statistical analysis. For Figures 1-5, I have only minor comments, as follows.

      In Figure 1-2, the Gata4-RFP reporter is used to sort cells for scRNA-seq, and the authors get a large number of non-ENS cells so utilise Elavl4 and Sox10 expression to pull out neurons and glia to analyse. Do other putative neuronal/glial genes pull out the same cells, and can the authors find evidence of non-neural crest derived glia/neurons in this dataset? This would ensure there isn't a bias in this first analysis and strengthen the later claims respectively.

      In Figure 3, the interpretation of increased NCAM1 and pFAK would be aided by adding the MFI for the untreated Hol Tg/Tg and ideally WT mice at p10 to assess if this is a normal developmental increase or induced by GDNF treatment - I would ask the authors add at least untreated Hol Tg/Tg at p10 as authors already present this data as images in figure 3D. Similarly, it would be nice to have neuronal counts of untreated Hol Tg/Tg and WT mice presented in Figure 4c to assess if NCAM1 inhibition completely or only partially prevents the effect of GDNF.

      In the text accompanying Figure 2 and 5, the authors use a variety of markers of assign SCP or enteric glia identity to the clusters, but some of the enteric glia genes can be expressed by Schwann cells/SCP in other contexts - it is worth adding references to the genes chosen to define SCP and enteric glia so readers can understand why those were chosen and the strength of evidence for each marker. For me, with the exception of Dhh I do not find any of these genes definitive as, at least on the SCP/Schwann cell side, they have been shown to express Gfap, Slc18a2, Sox2 and single cell analysis indicates they also express Cpe. The ability for SCP/Schwann cells to express both Gfap and Slc18a2 is important for the interpretation of the upcoming experiments.

      In Figure 6, the authors state in line 213 that "quantitative analysis as a function of time revealed that GFAP+ cells in the distal colon are not all derived from Dhh-expressing SCPs on the first day of treatment"; however, the DhhCre is restricted to specific SCP subsets (Xie, Meng, et al. "Schwann cell precursors contribute to skeletal formation during embryonic development in mice and zebrafish." Proceedings of the National Academy of Sciences 116.30 (2019): 15068-15073) and thus may not be labelling all SCP coming into the intestine. As far as I know there are no models that would allow labelling of all SCP without labelling enteric glia so would require the generation of a new models or techniques that is out of scope for this work - however, the authors should acknowledge the possibility of Dhh negative SCP contributing to the ENS.

      Figure 7 and 8 focuses on the contribution of Dhh+ cells, Slc18a2+ cells and cells expressing GFAP between p3-p8 to the ENS of GDNF treated Hol Tg/Tg mice. The authors first show that by p20 the number of neurons per mm2 is equivalent to those at p60, indicating p20 as an appropriate timepoint to assess the relative contribution of each pool of glia to the resulting neurons. The authors show that neurons derived from a DhhCre are detectable first, then Slc18a2, followed finally by GFAP. This data supports the authors claims that neurons first arise from Dhh+ SCP, but as I raised earlier, both Slc18a2 and GFAP can be expressed by schwann cells later in development - the authors in fact show that GFAP is in fact expressed by Dhh+ cells in figure 6, and showed with pseudotime analysis that SCP undergo a transition to Gfap+ Slc18a2+ glial cells before becoming neurons. The data presented here negates that analysis, or at least suggests some Dhh+ cells skip this step, otherwise the kinetics of neuronal differentiation would align across the 3 lines used. Could the authors address this discrepancy? For example, one addition could be to show that SCP/schwann cells on extrinsic gut innervating nerve fibres do not express Gfap or Slc18a2 at the timepoints used. In addition, their scRNA-seq analysis suggests that Gfap & Slc18a2 are expressed in the same cells, so one would expect similar kinetics in both populations - what is the overlap of GFAP staining with Slc18a2-YFP expression at the earlier stages (say p6 and p10?). Alternatively, the inducible nature of the GFAPCreERT2 means some bona fide GFAP expressing cells may be missed if the tamoxifen dose is not saturating, thus underestimating the contribution of this population to the total neuron count. As such, could the authors please show co-staining with GFAP in some induced GFAPCreERT2 samples to assess the effectiveness of this Cre. As with other figures, it would also be nice to add Holstein negative reporters at p20 as a control to assess the level of contribution each one of these cell lineages has in a normally developing animal compared to GDNF treated Hol Tg/Tg mice.

      In Figure 9, the authors show that neurons derived from labelled cells in the DhhCre-YFP or Slc18a2Cre-YFP skew strongly cholinergic (81%) compared to nitrergic (13%), but only 52% of the total neurons in GDNF treated Hol Tg/Tg mice are cholinergic, while 42% are NOS1+. The authors say that a similar trend is seen in Slc18a2Cre-YFP mice with a 'slightly higher percentage' reflecting Slc18a2 expression in neurons. However, this 'slightly higher' percentage looks to be about 40%, which is the amount expected in GDNF treated Hol Tg/Tg. I think the authors underplay the level of nitrergic neurons that are labelled in the Slc18a2Cre mice and should consider the possibility that Slc18a2+ glia can give rise to the nitrergic neurons seen in GDNF treated Holstein mice. This data also contrasts with prior reports that indicate SCP tend to become NOS1 expressing in a Ret or Ednrb loss of function HSCR model (Uesaka, Toshihiro, et al. "Enhanced enteric neurogenesis by Schwann cell precursors in mouse models of Hirschsprung disease." Glia 69.11 (2021): 2575-2590), a discrepancy worth discussing.

      In Figure 10, the authors assess the contribution of neural crest to the ENS of WT and Hol Tg/Tg mice using a Wnt1Cre2 line, and conclude there is non-neural crest contribution to the ENS of GDNF treated Hol Tg/Tg mice. This is an interesting idea and supported by a few studies the authors mention in the discussion; however, the labelling of neural crest derived cells in Wnt1Cre2 may not be fully penetrant. While this line has a well established leaky phenotype, others have reported it can have patchy recombination in the neural crest too (potentially in a background dependent manner, according to the line information on Jackson laboratories). For example, recent works shows significant mosaicism in the embryonic phase (Gandhi, S., Du, E. J., Pangilinan, E. S., & Harland, R. M. (2024). The Wnt1-Cre2 transgene causes aberrant recombination in non-neural crest cell types. bioRxiv, 2024-11.) This may not be the case in the authors home institution in a FVB background, as they show all HuCD+ and Sox+ cells overlap with the Wnt1Cre2 reporter, but the Hol Tg/Tg mice have aberrant neural crest development - this could interfere Wnt1 expression thus the effectiveness of the Wnt1Cre2, thus leading to non-labelled, but actually neural crest derived cells. It is also possible that those neurons are truly from a non-neural crest background, and the neural crest labelling tools are all flawed in one way or another - as such, I think the authors just need to discuss this alternative possibility in the discussion.

      Finally, Figure 11 shows NCAM1 expression in human tissue resections from HSCR patients. This data bolsters the idea that GDNF treatment could be an effective treatment for HSCR. The methods indicate that 3 tissue samples from HSCR patients were obtained, but staining is only shown for one - could all 3 be included, along with secondary only controls? This is important as human tissue resections are often processed many hours after removal, so are prone to higher background. It would also be interesting to include tissue from non-HSCR patients (the controls used most often in other studies are resections from anorectal malformations which are performed at approximately the same age as HSCR resections) to see if NCAM1 expression is influenced by the absence of an ENS in human patients, but I understand these samples are not always possible to obtain. If not possible, the authors could consider analysing publicly available RNA sequencing datasets to assess if RET is solely expressed in mature neuronal populations while NCAM1 is expressed by glial and neuronal populations like in their mice data.

      Overall, the authors clearly show that GDNF can drive differentiation of neurons from different glial populations through NCAM1 not RET, and supports continued effort to translate this work to the clinic as a potential therapy for HSCR patients. I am not convinced that the relative contributions of each glial subpopulation can be extrapolated from the mouse models used - this is a problem across the field, and will only be rectified with the development of new models, which is out of scope for this work. This work complements recent work showing NCAM1-GDNF signaling promotes neurogenesis in glia in Ednrb-/- mice (Mueller, Jessica L., et al. "Intramuscular enteric glia persist in Hirschsprung disease and undergo neurogenesis in response to GDNF-NCAM1 signaling." Scientific Reports 15.1 (2025): 33200), and will add to our understanding of neural development.

    5. Reviewer #4 (Public review):

      Summary: The manuscript by Gary et al works to reveal processes that generate enteric neurons in the distal colon of Holstein mutant mice that experience aganglionosis and are an established model of Hirschsprung disease (HSCR) following GDNF enema infusion. The primary goal of the study is to determine the how rectal infusion (enema) of GDNF drives neurogenesis in these mutants. The studies focus on determining the molecules that mediate this GDNF effect and identifying the cellular origins of the newly formed neurons. Given that prior work from this group and others has indicated formation of postnatal enteric neurons is feasible, understanding this process is an important step for the field. The data presented partially support the claims of the study. The analysis described do not support unequivocable conclusion that "transdifferentiation" is occurring. To instill confidence in the stated conclusions and make the text appropriately clear, clarifications of the methods and additional information is needed.

      Strengths:

      (1) The authors use in vivo experimentation with multiple mouse lines including Dhh-cre, Slc18A2-cre, and GFAP-creERT2 to trace the production of enteric neurons in the context of GDNF enema in the established Holstein mouse model of Hirschsprung disease over a time line of postnatal days 4 (P4) to P10.

      (2) The authors generate novel molecular profiles of enteric neurons induced in the distal colons of Holstein HSCR mice after GDNF enema induction by flow sorting for fluorescently labeled and adjacent populations with novel gating parameters followed by single cell RNA-sequencing.

      (3) The authors recognize broad expression of NCAM1 amongst the induced neurons in scRNA-seq data and apply experimental approaches to further investigate the role NCAM1 in generation of GDNF-induced ENS neurons.

      (4) The authors apply a widely utilized experimental approach, EdU-labeling, to assess proliferation in the immediate timeframe of GDNF-inducation in an effort to discern the timing of when the newly produced neurons exit the cell cycle.

      Weaknesses:

      (1) The article lacks sufficient information justifying use of the Holstein HSCR mouse model for the study. Authors need be more thorough when introducing the readers to the various HSCR mouse models they consider analyzing and provide an expanded justification for use of the Holstein model in the introduction so the audience can appreciate the rationale behind the analysis presented.

      (2) The text minimally describes the G4-RFP reporter that is a critical element of the analysis and upon which the scRNA-seq profiling hinges. Given the data presented in this submission it's not clear how much of the ENS this G4-RFP line labels.

      (3) The authors state that "NCAM1 is already present in SOX10+ cells before GDNF treatment begins at P4, both within and outside extrinsic nerve fibers (Fig.3b,c)." However, the data shown in Figure 3 panels B and D do not allow one to conclude co-localization of Sox10 with NCAM1 or FAK. Because ENS cells are so closely positioned with one another the signal the authors present could be due to adjacent cells or processes of cells above and below the plane of Sox10+ nuclei.

      (4) The authors use the inhibitor PF-562271 in an attempt to specifically inhibit phospho-FAK[Y397] shown in Figure 4. This particular inhibitor has the known side effect of causing apoptosis in phospho-FAK[Y397]+ cells as shown by Hu et al 2017 Cancer Sci (DOI: 10.1111/cas.13256). Because this compound causes apoptosis in phospho-FAK[Y397]+ cells, the data presented do not prove that the lack of neurons produced in this condition is due to inhibition of phospho-FAK[Y397]+ cells making the neurons versus those phospho-FAK[Y397]+ cells simply dying very early in the process.

      (5) In the text describing the results of Figure 1A compared to Figure 1B the authors have missed an opportunity to elaborate on the spatial distribution of what appear to be small ganglia in the GDNF treated colon of Holstein mutant mice. The schematic shown is rather simple and it's unclear from the text where these ganglia are distributed circumferentially around the gut wall.

      (6) The immunohistochemical labeling for Phox2b shown in Figure 6 is odd. Phox2b is expressed in neuronal progenitors, enteric glia, and ALL enteric neurons as shown in multiple publications. The images shown in Figure 6 offer an outlined region that appears to be a ganglion; however, within that encircled area fewer than half of the cells are labeling with Phox2b by this study.

      (7) The authors conclude that "transdifferentiation" is the origin of the ENS neurons that appear in GDNF-treated Holstein mice. However, the data show simply that most of the newly produced neurons have not recently gone through cell division based on lack of EdU incorporation. Given the data presented, other mechanisms may be occurring and should be considered as possibilities.

      (8) The authors utilized a Wnt1-cre2 transgenic line that has known issues with expression in the male germline and ectopic expression in cells that are dependent upon the reporter line utilized, like the Rosa26-YFP of this study. The methods lack information on whether crosses were performed in such a manner as to avoid issues with male germline activation of this reporter and the potential for ectopic expression cannot be excluded based on the information provided in the study.

      (9) The study lacks data on how GDNF enemas affect enteric neuron density and ganglia distribution in wildtype animals. If the signaling mechanism that produces new neurons in the Holstein model is also operating in wildtype animals, this could be crucial information for investigators interested in neuronal replacement to treat ENS damage resulting from environmental damage, disease, or age.

    1. eLife Assessment

      This work demonstrates an objective way to select parameter values for a quadratic integrate-and-fire model so that its bifurcation diagram matches a specific target diagram, generated from the Wang-Buzsaki model. The method is useful for the field and is presented with convincing evidence. The method is currently limited in its ability to be applied to data, but improves our mathematical tools to treat a rarely studied type of bifurcation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      From a big picture viewpoint, this work aims to provide a method to fit parameters of reduced models for neural dynamics so that the resulting tuned model has a bifurcation diagram that matches that of a more complex, computationally expensive model. The matching of bifurcation diagrams ensures that the model dynamics agree on a region of parameter space, rather than just at specially tuned values, and that the models share properties such as qualitative features of their phase response curves, as the authors demonstrate. A notable point is the inclusion of extracellular potassium concentration dynamics into the reduced model - here, the quadratic integrate-and-fire model; this is straightforward but nonetheless useful for studying certain phenomena.

      Strengths:

      The paper demonstrates the method specifically on the fitting of the quadratic integrate-and-fire model, with potassium concentration dynamics included, to the Wang-Buzsaki model extended to include the potassium component. The method works very well overall in this instance. The resulting model is thoroughly compared with the original, in terms of bifurcation diagrams, production of various activity patterns, phase response curves, and associated phase-locking and synchronization properties.

      Weaknesses:

      It is important to note that the proposed method requires that a target bifurcation diagram be known. In practical terms, this means that the method may be well suited to fitting a reduced model to another, more complicated model, but is not likely to be useful for fitting the model to data.

    3. Reviewer #2 (Public review):

      Summary:

      The authors derive an integrate-and-fire model to describe the dynamics of a more complex Wang-Buzsaki model and compare the two models. A detailed discussion of bifurcation schemes in both models is convincing and allows us to evaluate the simpler model.

      Strengths:

      The idea is interesting, and the mathematical approach appears to be convincing. In addition, differences between the simple and original models are also discussed.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      From a big picture viewpoint, this work aims to provide a method to fit parameters of reduced models for neural dynamics so that the resulting tuned model has a bifurcation diagram that matches that of a more complex, computationally expensive model. The matching of bifurcation diagrams ensures that the model dynamics agree on a region of parameter space, rather than just at specially tuned values, and that the models share properties such as qualitative features of their phase response curves, as the authors demonstrate. A notable point is the inclusion of extracellular potassium concentration dynamics into the reduced model - here, the quadratic integrate-and-fire model; this is straightforward but nonetheless useful for studying certain phenomena.

      Strengths:

      The paper demonstrates the method specifically on the fitting of the quadratic integrateand-fire model, with potassium concentration dynamics included, to the Wang-Buzsaki model extended to include the potassium component. The method works very well overall in this instance. The resulting model is thoroughly compared with the original, in terms of bifurcation diagrams, production of various activity patterns, phase response curves, and associated phase-locking and synchronization properties.

      Weaknesses:

      It is important to note that the proposed method requires that a target bifurcation diagram be known. In practical terms, this means that the method may be well suited to fitting a reduced model to another, more complicated model, but is not likely to be useful for fitting the model to data. Certainly, the authors did not illustrate any such application. Secondly, the authors do not provide any sort of general algorithm but rather give a demonstration of a single example of fitting one specific reduced model to one specific conductance-based model.

      We thank the reviewer for this critical assessment. It is true that we demonstrate our approach using as target the bifurcation diagram of a more realistic model. In principle, the method would be applicable to experimental systems if parameter space is sampled under appropriate experimental conditions, e.g., by recording at different extracellular potassium concentrations. However, this is challenging: generating reliable” experimental bifurcation diagrams” would require tightly controlled experimental repetitions under various conditions, particularly for reconstructing two-dimensional bifurcation diagrams. We hope that this work encourages the development of such methods. We have added a discussion of this point; see the paragraph starting at line 549.

      We have also included a general algorithm (Table I) and provide a second example illustrating the procedure (Supplementary Fig. S2).

      Finally, the main idea of the paper seems to me to be a natural descendant of the chain of reasoning, starting from Rinzel - continuing through Bertram; Golubitsky/Kaper/Josic; Izhikevich; and others - that a fundamental way to think about neuronal models, especially those involving bursting dynamics, is in terms of their bifurcation structure. According to this line of reasoning, two models are “the same” if they have the same bifurcation structure. Thus, it becomes natural to fit a reduced model to a more complicated model based on the bifurcation structure. The authors deserve credit for recognizing and implementing this step, and their work may be a useful example to the community. But the manuscript should have described and cited this chain of works to put the current study in the correct context.

      We have added a paragraph in the Discussion section (starting at line 517) to better situate the manuscript within the relevant literature and to explicitly acknowledge the chain of work.

      Reviewer #1 (Recommendations for the authors):

      Please see my public review. In line with my comments, I recommend that the authors either (a) provide a general algorithm for fitting at least a class of reduced models (i.e., those that satisfy some general assumptions) to a class of bifurcation diagrams, or (b) provide at least one more example of implementing their method. Step (b) would not need to be done to the same degree of thoroughness as the example they provided (e.g., the PRCs and synchrony need not be considered), but to me, this step would be very important if (a) is impractical. Otherwise, the paper should probably be rewritten to de-emphasize the message that this is a general method; instead, this should be a paper about specifically fitting the QIF (with potassium dynamics) to the Wang-Buzsaki model (with potassium dynamics).

      We provide a general algorithm for deriving a quadratic integrate-and-fire model with dependence on a biophysical parameter by fitting the bifurcation structure of a given class I conductance based neuron model near an SNL bifurcation induced by this parameter; see Table I.

      In addition, we provide a second example of the reduction procedure: motivated by Hesse et al. (Nature Communications, 10.1038/s41467-022-31195-6, 2022), we derive a QIF model that captures dependence on temperature instead of potassium concentration; see Supplementary Fig. S2.

      Not surprisingly, I also think it’s essential that the authors describe and cite the chain of works on thinking of neuronal models in equivalence classes based on bifurcation diagrams, and make clear that this paper builds on the ideas set forth in that chain.

      We thank the reviewer for this comment. As mentioned above, we have added a paragraph in the Discussion section, starting at line 517, to acknowledge this chain of work.

      Also, the authors should make clear that their method is not one for fitting a model directly to data, which will require rewriting at least the first paragraph of their Discussion section.

      We thank the reviewer for helping us make our manuscript clearer. To avoid confusion, we have clarified this point already in the Introduction (see lines 52-57) and have included a new paragraph in the Discussion (starting at line 549).

      Other specific corrections are:

      (1) Typos should be fixed, as the paper has several. The first line of the abstract has one (“concentrations” → “concentration”), for starters. “Original model” on pg. 3 is missing “be” in “can defined”. “ceases” → “cease” on pg. 13. “nerons” → “neurons” on pg. 19. “standart” → “standard” pg. 24.

      Done. Additional typos were also corrected.

      (2) The abstract mentions “consequences in networks” in its second sentence. This is misleading because studying network dynamics is not at all the main emphasis of the paper, but rather a corollary application of the main ideas, so some restructuring of the abstract is needed. Similarly, the final abstract sentence overstates somewhat what was done with studying synchronization and should be rewritten more precisely.

      We have restructured the abstract accordingly.

      (3) For readers who are interested in the ideas here but not familiar with the QIF model, it will be very difficult to follow the first paragraph of Results. Elementary aspects of QIF dynamics should be explained here (e.g., what is the saddle-node bifurcation), and a basic figure panel about this should be included in Figure 1.

      We added a supplementary figure (Figure S1) adapted from Izhikevich for readers who might not be familiar with the QIF model.

      (4) Bottom lines of page 3 should be reworded to make clear that the slow variables are averaged over each member of a family of fast subsystem limit cycles. Also, “one limit action potential cycle” is an awkward phrase.

      We rephrased this sentence (see paragraph starting at line 140).

      (5) Text under system (1) – why isn’t c mentioned? Also, references to Figure 3 should be to Figure 2 here. And authors should state what they mean by “target model” and be clear about whether it includes potassium dynamics and/or pump current.

      c scales the parabola corresponding to the branch of fixed points, given by c(I<sub>app</sub> − I<sub>SN,0</sub> − I<sub>pump</sub>) = −a(ν − ν<sub>SN</sub>)<sup>2</sup>. Consequently, it also affects the position of the homoclinic bifurcation: In the previous version of the manuscript, c was inadvertently omitted from the expression for the branch of fixed points; this has now been corrected. See paragraph starting at line 154.

      Figure references have been corrected.

      By target model, we mean the conductance-based model that includes potassium dynamics and a pump current, in our case System 4. We clarified this point at the beginning of the Results section (see paragraph starting at line 116). Throughout the manuscript, we now explicitly indicate when we refer only to its fast subsystem and whether the pump current is included. In particular, note that the parameter derivation shown in Figure 4 is performed on the fast subsystem of the target model, in the absence of the pump current. I<sub>pump</sub> can be considered as a potassium-dependent contribution to the applied current, and can be added a posteriori to the QIF model. See paragraph starting at line 162.

      (6) Next par: is the “saddle-node bifurcation” that with I<sub>app</sub> as bifurcation parameter? Please clarify.

      Yes, it is the saddle-node bifurcation with I<sub>app</sub> as bifurcation parameter. We have clarified this in the manuscript; see the paragraph starting at line 170.

      (7) Bottom pg. 5: does “beyond” mean above? below?

      We meant above (larger values of ). In the text, we have replaced “beyond” with “larger than”. See paragraph starting at line 178.

      (8) Formula for v<sub>r</sub> at top of page 6: Please specify what formula for I<sub>pump</sub> is being used here.

      The formula for I<sub>pump</sub> is given in Eq. 5d. We are using the same formula throughout the paper.

      Note that to clarify the reduction procedure, we derive QIF parameters to match the bifurcation diagram with respect to the applied current of the fast subsystem of the target model when I<sub>pump</sub> = 0. Reintroducing I<sub>pump</sub> produces the same horizontal shift in this bifurcation diagram for both the QIF and Wang–Buzsáki versions.

      We have restructured the paragraph starting at line 178 to clarify these aspects.

      (9) Three lines below this: I don’t understand what “matching...is appreciable” and “in the continuity of...” mean. Please revise and also explain why a closer matching of v<sub>r</sub> to the min voltages in Figure 4d was not used, and exactly how the v<sub>r</sub> that is shown was chosen.

      With “matching...is appreciable”, we meant that values assigned to v<sub>r</sub> should be close to the minimum voltage values reached during spiking. With ”in the continuity of...”, we meant that when .(SNIC case), we choose v<sub>r</sub> by extrapolating the linear fit performed on the values of v<sub>r</sub> assigned when , (homoclinic case).

      When , v<sub>r</sub> was chosen so that the homoclinic bifurcation occurs at the same value of applied current as in the fast subsystem of the target model. This is explained in the paragraph starting at line 178 (see Eq. 2). This criterion also allows the minimum voltage values reached during spiking to be captured reasonably well (compare the green dotted line and the purple solid line in panel e of Figure 4).

      We have rewritten the paragraph starting at line 186 to clarify these aspects.

      (10) Bottom page 6 - reference to Figure 3e should be 4e. Also, the text mentions the shrinkage of spike amplitude, but the figure shows that vth increases over most of the K+ range before decreasing, so a correction is needed.

      We corrected the figure reference.

      The maximal voltage of the limit cycles of the target model’s fast subsystem (upper purple curve in Figure 4e) increases slightly between and , by less than 1mV. It then decreases by about 27mV before the fold of limit cycles. The sigmoidal function vth () allows us to capture this substantial decrease in the QIF model. The small preceding increase is not captured. We reformulated the text to avoid confusion (see paragraph starting at line 209).

      (11) Figure 4d: Why is E<sub>K</sub> plotted here? It should be mentioned in the caption and text. More generally, the caption for Figure 4e should be expanded to mention what the purple curves are, what is the black curve for K < K<sub>SNL</sub>, and what the other structures shown are. Finally, the text describes that theSNIC/SNL/Hom is determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>, so it’s not clear what is I<sub>app,SNL</sub> in the caption - please clarify.

      In conductance-based models, higher weakens the potassium concentration gradient, thereby raising E<sub>K</sub>. The sodium and potassium reversal potentials typically bound voltage oscillations during spiking (see for example Chander and Chakravarthy, PLOS ONE, 10.1371/journal.pone.0048802, 2012), so the minimum voltage of spikes is expected to be higher when is larger. We plotted E<sub>K</sub> in Figure 4d to show that the increase of the reset voltage v<sub>r</sub> at larger reflects this effect in the QIF version of the model. We clarified this in the caption of Figure 4 and in the text (see paragraph starting at line 199).

      Purple curves show families of limit cycles, while black curves, including the one for , show families of fixed points. The green curve shows the linear fit of v<sub>r</sub> from panel d. All these have now been included in the legend.

      In the QIF model, the onset bifurcation (SNIC, SNL, or homoclinic) is indeed determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>. Figure 4e shows the bifurcation diagram of the fast subsystem of the conductance-based model (Wang-Buzsáki). This is a bifurcation diagram with respect to , for a fixed value of applied current. We chose to fix I<sub>app</sub> at its value at the SNL bifurcation, denoted I<sub>app,SNL</sub>. I<sub>app,SNL</sub> is near 0.22 (see Figure 3a).

      (12) Eqn. (2a): Shouldn’t v<sub>SN</sub> depend on potassium like I<sub>SN</sub> and I<sub>pump</sub> do? What is the formula for Ipump there? Why isn’t the RHS of (2b) dependent on v as in the original model? Please clarify.

      For simplicity, we did not include a potassium dependence for v<sub>SN</sub> in the QIF model. Instead, we set it to its value at the SNL bifurcation (Figure 4b). This is explained in the paragraph starting at line 199: “We notice that v<sub>SN</sub> and the normal form coefficient a are relatively conserved in this interval. We fix them to their value at [K<sup>+</sup>]<sub>o,SNL</sub>.”

      The formula for I<sub>pump</sub> is given in Eq. 5d. The potassium dynamics depends on the voltage via the reset rule in Eq. 3d. This allows us to capture the small increments in at each action potential in the original model (see for example Figure 6c,h). We have clarified these two points in the manuscript (see paragraph starting at line 223).

      (13) Figure 3: Please indicate the criticality of the Hopf bifurcations shown.

      The legend of Fig. 3 now indicates that the Hopf bifurcations are subcritical. The same clarification has been added to the following figures as well.

      (14) Bottom pg. 9: Why isn’t there an I<sub>pump</sub> term as in eqn. (6a)? Please clarify. 

      You are correct, the I<sub>pump</sub> term should be included in the equation for the averaged slow subsystem of the QIF model (see paragraph starting at line 243); it was accidentally omitted in the manuscript. Thank you for pointing this out.

      (15) Top pg. 11: It’s important to reference the slow averaged dynamics here, which allows K+ to increase. Also, this first paragraph should already explain that this averaged dynamics is only relevant along the family of FS periodic orbits, not during the recovery when the FS has a branch of stable equilibria.

      The averaged slow subsystem is indeed only relevant along families of limit cycles of the fast subsystem. Along families of equilibria, averaging is not necessary and the standard slow subsystem can be used. We now explicitly define this standard slow subsystem for both Wang-Buzsáki and the QIF model (see paragraphs starting at lines 241 and 628). In panels e and j of Figure 6, both systems are now represented.

      In the paragraph starting at line 273, we now refer to the averaged slow subsystem to explain the overall increase of during bursts (purple curves in Fig. 6e,j), and to the standard slow subsystem to explain the decrease of during quiescent phases (black curves).

      (16) Pg. 11, par 3: This is unnecessarily confusing. Please try to reword and clarify this paragraph.

      We have simplified this paragraph (starting at line 293). The key point is that the reduction to the averaged slow subsystem is not valid too close to the homoclinic bifurcation.

      (17) Pg. 13, end of Scenario 2: Is there any evidence this is a canard effect and not a noise effect? If so, please mention the evidence; otherwise, perhaps take this out.

      What happens there appears to be a noise-induced canard effect: in the beginning of the burst, the system follows a family of stable limit cycles of the fast subsystem, i.e. a stable object. However, at some point, noise induces a transition to a portion of trajectory where the system evolves near the saddle branch, i.e. a repelling object, for a substantial amount of time. Such phenomena have been thoroughly investigated in the literature, and can also be obtained in a deterministic way; see for example Marin et al. (Physical Review E 90, 042718, 2014). Bursting traces similar to the one in Figure 7c are observed experimentally (see, for example, Figure 4c of Marin et al.), which is why we considered it worth mentioning. We have revised the paragraph starting at line 332 to clarify this point.

      (18) I only see 4 curves in Figure 9a,c, but the legend has 5. Are two on top of each other? Please clarify.

      Yes, the curve for = 7.21mM lies beneath the curve for = 5.21mM. This has been clarified in the figure caption.

      (19) Text should note that the QIF iPRC does not develop a negative region at high K+ and high phase, as WB iPRC does.

      We have updated the paragraph starting at line 409 to mention this.

      (20) Pg. 16, line 4: “at the network scale” is cryptic - a more precise phrase would be preferable.

      We have reformulated the sentence to clarify its meaning (see paragraph starting at line 380).

      (21) Pg. 16, line 9: Reordering of words could make this clearer.

      Done (see paragraph starting at line 385).

      (22) Pg. 16: I am confused by line 14 because the big changes in the iPRC in Figure 9c do not align with the spike phase in Figure 9d. Please clarify what is meant here.

      In the QIF model, at a given phase, the iPRC is the inverse of the slope of the voltage trace as a function of phase. Flatter slopes in Figure 9d therefore correspond to larger iPRC values in Figure 9c. This is illustrated in Figure S5. We have revised the paragraph starting at line 393 to make this point clearer.

      (23) Pg. 16: Please clarify what is meant by a “delta synapse”.

      By “delta synapse,” we meant a configuration in which each spike induces an instantaneous voltage jump in the postsynaptic neuron, modeled using the Dirac delta distribution. We have replaced the term “delta synapse” with “pulse-coupled,” which is more commonly used in the literature, and have added a clarification at its first occurrence in the manuscript.

      (24) Discussion, line 2: Delete comma.

      Done.

      (25) Importantly, as noted above, the first par. needs to be rewritten since the presented method won’t work directly from data or from a target model for which most of the parameters, and hence the bifurcation diagram, are not known.

      As mentioned above, we have included a new paragraph in the Discussion, starting at line 549, to clarify this point.

      (26) Pg. 18: ”Originally” → ”Typically”, perhaps?

      Done.

      (27) Note the work of Marder et al. on temperature-related neural variability.

      We thank the reviewer for this comment. We have added two relevant references from the work of Marder and colleagues addressing temperature-dependent neural variability and ionic concentrations in our manuscript (see the sentence starting on line 537).

      (28) Pg. 20: Cut the ”Potassium dynamics and network models” subsection since it does not add anything substantive as written (or else expand it and include it in the subsection below).

      We have expanded this paragraph and incorporated it into the subsequent subsection, as suggested by the reviewer.

      (29) Finally, it’s a bit confusing that the authors refer to the potassium concentration as a slow variable yet have an instantaneous jump in this quantity at reset in their QIF model (i.e., instantaneous is VERY fast). Some explanation about this should be provided. Do they make the general assumption that ∆<sub>K</sub> is small, for example, such that this reset reflects the slow nature of K+ evolution (i.e., during the reset period, K+ would only change slowly, and hence by a small amount)?

      Yes, ∆<sub>K</sub> is chosen to be small, to capture the behavior of the original model (compare for example panels c and h of Figure 6). As a result, in the QIF model, in the same way as in the original model, despite the fact that the dynamics of includes a fast component, on average evolves slowly. By using the averaging method, we can determine whether overall increases or decreases.

      We clarified this in the manuscript, in the paragraph starting at line 243.

      Reviewer #2 (Public review):

      Summary:

      The authors derive an integrate-and-fire model to describe the dynamics of a more complex Wang-Buzsaki model and compare the two models. A detailed discussion of bifurcation schemes in both models is convincing and allows us to evaluate the simpler model.

      Strengths:

      The idea is interesting, and the mathematical approach appears to be convincing. In addition, differences between the simple and original models are also discussed.

      Weaknesses:

      A comparison to experimental data is necessary to support the theoretical work.

      As mentioned above in our answer to Reviewer 1, we demonstrate our method using as target the bifurcation diagram of a more realistic neuron model. Ideally, one would want to derive phenomenological models that capture bifurcation structures obtained from data. However, this is challenging and beyond the scope of the present study. We hope that this work encourages the development of such methods. We have revised the Introduction (see lines 52-57) and added a paragraph in the Discussion (see the paragraph starting at line 549) addressing this point.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is well-structured; however, it appears that it has been edited with less care. Please see comments below:

      (1) Page 2: “A third bifurcation, the saddle-homoclinic orbit (HOM) bifurcation,”: provide a reference for the bifurcation.

      We have added a reference to the book by Izhikevich (see paragraph starting at line 63).

      We have added additional references in the Introduction that we considered helpful.

      (2) Page 3: “while a larger concentrations it is mediated by...”: remove “it”.

      There was indeed a typo in this sentence. The intended phrasing is: “while at larger concentrations it is mediated by...”. We have corrected it accordingly (paragraph starting at line 134).

      (3) Figure 2, caption: “dashed lines for unstable branches”: this is a dotted line.

      Corrected to “dotted lines”. Thank you.

      (4) Page 4: “is smaller than vSN (Fig. 3c),”: this figure panel does not exist, as well as the Fig.3d referred to afterwards. Please correct.

      We intended to refer to Fig. 2. Figure references have been corrected. See paragraph starting at line 154.

      (5) Page 6: “Fig. 3a-d shows ISN,0, vSN and a for [K]+o between 4 and 16 mM”: Fig.3c+d do not exist, please correct. Similar comment to “the absence of pump (Fig. 3e).” on the same page.

      We intended to refer to Fig. 4. Figure references have been corrected (paragraphs starting at lines 199 and 209).

      (6) Page 8: “(panel A)” → ”panel (a)”.

      Done.

      (7) Figure 4e: What is the meaning of the green dotted curve?

      This curve represents the linear fit of the reset voltage v<sub>r</sub> from panel d of Fig. 4, to show that the minimal voltage values of the limit cycles are also well captured. We have added this curve to the legend and included a brief explanation in the figure caption.

    1. eLife Assessment

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants and the authors provide solid evidence for contrasting symbiont-associated phenotypes under controlled conditions: Rickettsiella increases plant damage while reducing wing production and dispersal, whereas Regiella reduces feeding damage and, at some time points, aphid population growth. The stable establishment of these novel symbiont-host associations and complementary experiments spanning individual aphids, whole plants, populations, and mesocosms are notable strengths and the revised manuscript better clarifies the experimental approaches and the context dependence of symbiont effects; however, the underlying mechanisms remain unresolved, horizontal transmission is inferred rather than directly demonstrated, and limited replication and temporal variability constrain some population-level conclusions. Potential applications to pest management therefore require further validation under field conditions. The study will interest researchers working on insect symbiosis, plant-insect interactions, and biologically based pest management.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      It is a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

    4. Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their limitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      [1] Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.

      [2] Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.

      [3] Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.

      [4] James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants. The authors provide solid evidence that the two symbionts can generate contrasting phenotypes: Rickettsiella increases plant damage while reducing wing formation and dispersal, whereas Regiella reduces aphid population growth and feeding damage. The successful establishment and stable transmission of these novel symbiont-host associations, combined with experiments spanning individual, whole-plant, population and mesocosm scales, are notable strengths of the work; however, the mechanisms underlying these effects remain unresolved, evidence for horizontal transmission is indirect, and some population-level conclusions rely on relatively small sample sizes or effects that are not consistently detected across time points, and therefore the potential application of these findings to pest management remains promising but speculative. The study will be of broad interest to researchers working on insect symbiosis, plant-insect interactions, and biologically based approaches to pest management.

      We have made revisions to the manuscript to cover issues raised around sample numbers and mechanisms. We appreciate that we have not been able to finalize the exact mechanism underlying plant damage effects. We note that while comparisons of population performance were limited by the number of populations we could feasibly maintain; sample sizes were substantial in some of the other experiments. We also do substantiate effects through a combination of experimental approaches that start with controlled conditions and then encompass more complex environments.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

      We thank the reviewer for recognizing the strengths of the work and for providing extensive comments.

      Weaknesses:

      There are some major aspects of this paper that I thought could be strengthened. My main concern is that the manuscript feels broader than it is conceptually focused. A wide range of outcomes is measured, which gives the study breadth, but it also makes the central question harder to identify. As written, the paper reads more strongly as a proof-of-principle demonstration of ecologically relevant phenotypes than as a tightly framed test of a specific biological idea.

      The broad range of tests in our work was intentional. Rather than relying on a single experimental approach or scale, we designed the study using multiple complementary approaches to independently evaluate the effects of endosymbionts. Thus, our conclusions are supported across multiple experimental contexts rather than by a single experiment or scale. For example, the effects on plant damage were consistent across different host plants, including wheat and barley (Figures 1 & S1), and across different experimental scales, ranging from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level, using single aphids maintained on individual plants under favorable conditions (Figures S5 & S6), and at the population level under crowding conditions (Figures 3 & 4).

      Despite the breadth of measurements, the study is focused on establishing robust evidence for the contrasting effects of the two endosymbionts on aphid dispersal and plant feeding damage. To help address this concern, we have moved the summary table from the Supplementary Materials to the main text (now Table 1), which provides an overview of the experimental approaches and main findings and should help readers more clearly see how the different experiments are connected.

      A second issue is that the biological basis of the reported phenotypes remains less developed than the phenotypic description itself. The authors make a genuine effort to address mechanism through JA, JA-Ile, SA, and metabolomic profiling, but these analyses only partially explain the main results. The negative result for the canonical defense markers is informative, yet it still leaves a substantial gap between the observed variation in plant damage and the processes responsible for it.

      Our analyses of JA, JA-Ile, SA, and the metabolomic profiles provide some initial insights, but we appreciate that they do not fully explain the differences in plant damage observed between treatments. The primary aim of this study was to evaluate the phenotypic effects of the endosymbionts and their potential for agricultural application, rather than to provide a comprehensive mechanistic explanation. We have accordingly avoided overinterpreting the mechanistic results and now explicitly state that elucidating the underlying biological mechanisms will be an important direction for future research in Discussion section.

      I also think some caution is needed in how the two symbionts are compared. The authors explain why some follow-up experiments were designed differently for Rickettsiella and Regiella, and that rationale is understandable. Still, because the downstream assays were not fully matched, the paper is strongest when each symbiont is interpreted on its own terms rather than as a strict comparison.

      Our initial plant-damage experiment was designed as a first comparison to test whether different endosymbionts can have diverse and contrasting effects on their aphid host population and plant damage. We then investigated the individual phenotypes of each endosymbiont in greater detail, particularly in relation to their potential agricultural applications. Specifically, our results suggest that Regiella may reduce plant damage, whereas Rickettsiella may reduce dispersal. Based on these early findings, some further experiments were conducted with slightly different experimental setups. Nevertheless, many of the experiments conducted for the two endosymbionts were broadly comparable. We designed the additional experiment carried out only with Rickettsiella to test whether reduced alate production observed in our earlier experiments also translated into reduced dispersal at the population level. An equivalent experiment with Regiella was not undertaken because we failed to detect an effect of this endosymbiont on alate frequency in our preceding experiments. We have clarified this rationale in the Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”).

      Overall, I would suggest softening the Significance Statement so that it more clearly reflects what is directly shown here, namely that introduced symbionts can alter plant damage and dispersal-related phenotypes under controlled conditions, rather than implying that the study directly tests management utility in agricultural settings.

      We have done this in the Significance Statement. We appreciate that the current experiments have been carried out under controlled conditions, rather than in agricultural settings. Pending permit approval, we are currently planning to extend this work to contained field settings to test whether effects on plant damage and dispersal ability are also observed under less controlled conditions.

      Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      One thing that I struggled a little with was the rapid spread of Regiella in the shared plant experiments. Possibly this could be attributed to an increased reproductive output (due to faster developmental time, and/or an increase in fecundity) or efficient horizontal transmission. However, the other experiments performed indicate a slight negative impact (Figure 4a) or no influence (Figure 4C, 4D, Figure S6) of Regiella infection on host fitness. Given these other results, it seems that Regiella must spread fairly efficiently between hosts, which comes as a surprise, and there are very few examples of horizontal transmission of Regiella like this in the literature. The manuscript would benefit from a clear and direct demonstration of horizontal transmission, rather than it being inferred indirectly. The similar spread observed in the Rickettsiella mixed cages is less surprising, because there are several examples where this has been demonstrated.

      We agree that the rapid spread of Regiella in the shared-plant experiments cannot be readily explained by host fitness alone, though we have noted fitness benefits of Regiella in a different transinfection in oat aphids (Yu et al., 2025). In a previous study with transinfected green peach aphids and despite a substantial fitness cost, we found that Rickettsiella can spread relatively rapidly in a population and show high stability (Gu et al., 2023) and perhaps transmission of Regiella follows a similar pathway. However, whereas Rickettsiella may spread through plant tissues, Regiella showed relatively low levels of horizontal transmission through this pathway.

      We certainly agree that more work is required to establish the mechanism and dynamics of horizontal transmission in this system. Rather than focusing on mechanism, our objective here was to examine endosymbiont spread where intact plants were available and where there was a mixed aphid population. Note that we also conducted an additional experiment in which Regiella-infected aphids were present at a frequency of only 10% of the initial population, and in this situation Regiella nevertheless still increased in frequency including to a low Cp value in most (8/10) replicates, highlighting its persistence and potential to increase in populations.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Yu et al., A persistent bacterial Regiella transinfection in the bird cherry-oat aphid Rhopalosiphum padi increasing host fitness and decreasing plant virus transmission. Pest Manag Sci 81, 2791-2799 (2025).

      It is also a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

      We conducted the final mesocosm-dispersal experiment specifically with Rickettsiella because our earlier individual-plant experiments had already shown a clear reduction in alate production in Rickettsiella-infected aphids, together with effects on plant damage (Figure S1I) and population growth (Figure 3C). In contrast, we did not detect a significant difference in alate frequency between Regiella-infected and wild type aphid strains in the similar set up experiments (Figure S1I and Figure 4C). We therefore designed the additional experiment to test whether the reduced alate production observed with Rickettsiella also translated into reduced dispersal at the population level. We did not conduct the same experiment with Regiella because there was no difference in alate frequency in our earlier experiments. We have also added this explanation in Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”). We do appreciate however that future experiments on dispersal of Regiella will be worthwhile resources permitting.

      Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      We acknowledge that more replication is always desirable, but we would also argue that significant effects were not marginal and replication was substantial in many cases (e. g. 9-10 replicate plants per damage treatment evaluation). We were also focused on using multiple experimental approaches and scales to independently and repeatedly evaluate the effects of endosymbionts on aphid fitness, wing development, plant damage, and aphid dispersal. Thus, conclusions are not based on a single experiment or experimental scale. For plant damage, for example, we observed consistent effects across different host plants, including wheat and barley (Figure 1 and Figure S1), as well as across different experimental scales, from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level using single aphids maintained on individual plants under favorable conditions (60 replicates per treatment) (Figures S5 & S6) and at the population level under crowding conditions (Figures 3 & 4). These complementary experimental designs allowed us to examine whether the observed phenotypes were consistent across different environmental and population contexts. We did face challenges in high levels of replication of independent aphid strains in population cage experiments but attempted to replicate as much as possible given the resources that were available.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their dlimitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      (1) Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.

      (2) Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.

      (3) Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.

      (4) James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      We agree that qPCR-based measurements of endosymbiont gene copy number have limitations even if they are the standard approach used in most studies. In the current set of experiments, qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples and the only one feasible given the number of monitoring events and samples required. Nevertheless, we acknowledge the limitations of this approach (and have pointed this out ourselves in a recent COIS paper – Hoffmann et al 2026). We now mention this under further work and provide a couple of references (see Discussion).

      Reference:

      Hoffmann, A. A., Yang, Q. and P. A. Ross. Aphid endosymbionts revisited: molecular detection, diversity, and population dynamics. Curr Op Insect Sci (in press). (2026)

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Overall, I found the study interesting and worthwhile, particularly because it shows that novel symbiont associations can generate contrasting phenotypes in an important pest species. My main suggestion would be to sharpen the framing of the paper and to bring the mechanistic and applied discussion into slightly closer alignment with the current evidence base. With those points addressed, I think the manuscript would read as a clearer and more balanced contribution.

      We have rephrased the Discussion around the mechanistic component of the work to sharpen this and link more directly to evidence. For instance:

      “Despite this, our metabolomic analyses suggest that endosymbionts in D. noxia may influence some other aspects of wheat metabolism in a spatially structured way. Within aphid-feeding areas, wheat exposed to Rickettsiella aphids showed increased valine and decreased 3-phenyllactic acid. Changes in valine have been reported in plants responding to herbivory and other forms of stress (43-45), while 3‑phenyllactic acid has been associated with antimicrobial and defense-related activity (46). In non-feeding areas, wheat exposed to Regiella aphids showed reduced urea and increased pantothenic acid. These changes may reflect differences in nitrogen metabolism and allocation (47), and in metabolic processes involving pantothenic acid (48) respectively. However, the present metabolomic data do not establish the functional consequences or causal mechanisms underlying these changes. They indicate that aphids carrying different endosymbionts are associated with some distinct metabolic responses in wheat, including responses that differ between aphid-feeding and non-feeding areas. Our findings are consistent with previous work showing that phloem‑feeding insects can induce changes in plant metabolites in response to herbivory (49, 50) and provide a basis for future studies to determine how endosymbionts influence aphid-induced plant responses and contribute to contrasting plant phenotypes.”

      We have not undertaken a complete reframing of the paper but further emphasized the focus on phenotypic contrasts in a few places including incorporating some changes to the comments below.

      (2) Lines 130-135: Please clarify more explicitly whether the main aim of the paper is to test a specific biological hypothesis about endosymbiont-mediated aphid-plant interactions or to provide a broader proof-of-principle survey of symbiont-associated phenotypes. As it stands, the framing moves between multitrophic biology and pest-management relevance, which makes the central conceptual contribution harder to identify.

      We have rephrased this sentence as “By integrating these factors, we investigate how endosymbiont infection influences aphid fitness and aphid–plant interactions, providing insights into the ecological consequences of novel microbial associations and their potential relevance to sustainable pest management.”

      (3) Line 205: For the metabolomic analysis, please consider adding a formal multivariate test for strain effects, especially for the comparisons shown in Figure 2C and 2D, or otherwise interpret the PCA more cautiously as descriptive rather than inferential.

      We have emphasized the descriptive component and rephrased this sentence as “However, within either area, there was no clear separation among aphid strains (Figures 2C & 2D), suggesting broadly similar metabolomic profiles among strains of the same aphid clone carrying different symbionts.”

      (4) Please clarify how the top and bottom feeding leaves were handled analytically in the analyses, and explain the rationale for collapsing them into a single "feeding area" category. If possible, it would be helpful to show whether leaf position itself influenced the plant-response patterns.

      We combined the upper and lower leaves together to provide a representative measure of the plant-level responses, rather than focusing on responses at a particular leaf position. This approach was consistent with the main objective of our study, which was to investigate whole plant responses to aphids hosting different endosymbionts, rather than differences in responses among different plant parts. In addition, combining the two portions provided sufficient plant material for the GC-MS analysis and helped ensure reliable metabolite measurements from the same material. Because the two leaf positions were combined prior to GC-MS analysis, we were unable to separately test the effect of leaf position on the metabolomic response in this dataset. We have clarified this point in the revised manuscript in Materials and methods section (“Plant defense responses”).

      (5) Line 228: The use of 19{degree sign}C and 25{degree sign}C is not unusual in aphid work, but it would still help the reader if the manuscript stated more explicitly why these two temperatures were chosen in this study.

      We selected 19°C and 25°C because they represent contrasting temperature conditions within the range suitable for Russian wheat aphid development, allowing us to assess whether temperature influences Rickettsiella transmission and population dynamics. In particular, our previous observations indicated differences in the rate of Rickettsiella spread between these temperature conditions (Gu et al., 2023).

      Reference:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      (6) Lines 390-392: It would help to discuss more explicitly how the relatively modest effects in the individual life-history assays relate to the clearer signals seen at the whole-plant and population level.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with endosymbiont infection. In contrast, we also examined aphid performance at the population level (Figures 3 & 4), where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress. Under these conditions, the effects of endosymbiont infection differed from those observed in the individual assays, with Rickettsiella-infected aphids showing greater population growth and Regiella-infected aphids showing reduced population growth.

      These results suggest that the effects of endosymbionts on aphid fitness are context-dependent and may become more pronounced as population density increases and density-dependent stress develops. Thus, relatively modest effects observed at the individual level where experiments are often undertaken may not translate to differences at the population level, and plant level effects may subsequently influence feeding pressure and plant damage. We have added this perspective in the Results and Discussion sections.

      (7) Line 678 & 688: In both whole-plant experiments, please explain how 3 and 4 replicate plants were selected.

      The replicate plants were randomly selected from the available plants for each treatment to minimize potential selection bias. We have clarified this procedure in the revised manuscript in Materials and methods section.

      (8) Lines 686-690 / Figure 4: In the Regiella whole-plant experiment, the Methods state that 16 plants were established per treatment and that 4 plants per treatment were removed at each time point (days 4, 8, 12, and 16). However, in Figure 4B-D, day 12 appears to include 5 data points. Please clarify this apparent mismatch between the described sampling scheme and the data shown in the figure.

      We thank the reviewer for pointing out this and we have corrected this mistake. We initially established 16 plants for the wild type and 17 plants for the Regiella treatment. Four plants per treatment were originally planned to be sampled at each time point (Days 4, 8, 12, and 16). However, because Day 12 was a key time point at which an obvious difference in plant damage was observed between the treatments, we selected one additional plant each treatment for measurement at Day 12, resulting in five data points for this treatment at that time point. The remaining plant was therefore measured at Day 16. We have clarified the sampling procedure in the revised Materials and methods section and figure legend.

      (9) Figure 2A and Figure 4A: These schematics are helpful overall, but the brown supporting sticks stand out quite strongly and may make the panels a little harder to interpret at first glance. I wonder whether they could be simplified, made less prominent, or replaced with photographs of the actual setup if those are available.

      We have revised Figures 2A and 4A to simplify the supporting structures and reduce their visual prominence.

      Reviewer #2 (Recommendations for the authors):

      Some of the statistics were not entirely clear to me, particularly the tests reported in the results which differ from what is described in the figure legends:

      (1) Lines 306-308: "Total nymph numbers were higher on Rickettsiella aphids from Day 21 to Day 25" and indicates that this is based on GLM testing, but the figure does not indicate statistical significance, and the figure legend states independent sample t-tests were used. Similar comment for the following paragraph and corresponding figure.

      The GLMs were used to test the overall patterns in nymph numbers across the relevant time periods, including Days 21–25, rather than testing each time point independently. We also conducted independent-sample t-tests to assess differences between treatments at individual time points. We have clarified this distinction in the Statistical section.

      (2) Line 43: Does not seem like the appropriate reference (reference is on plant virus transmission, not salivary toxins).

      We thank the reviewer for pointing this out. We have removed this reference and replaced it with reference 34 (Luna et al., 2018) that directly supports the statement regarding aphid salivary toxins.

      Reference:

      Luna et al., Bacteria associated with Russian wheat aphid (Diuraphis noxia) enhance aphid virulence to wheat. Phytobiomes J 2, 151-164 (2018).

      Reviewer #3 (Recommendations for the authors):

      Minor editorial comments:

      (1) Figure S10 - the figure legend needs improvement as the current version does not help the reader understand the figure. Please also include a key.

      We have changed the figure legend with reference to our aim, and also explained use of the Cp values. “Figure S10. Rickettsiella Cp values in (A) routine screening of laboratory Rickettsiella colonies and (B) the mixed cage experiment assessing endosymbiont frequency changes over time at 19 °C and 25 °C. The dark red area represents overlap between the 19°C and 25°C experiments. Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA. This figure shows the typical range of Cp values observed in our laboratory Rickettsiella -infected aphid colonies. We used this range as a reference for identifying aphids that acquired Rickettsiella through horizontal transmission, as horizontally infected aphids generally showed much higher Cp values than vertically infected aphids.”

      (2) Define Cp.

      We have explained it as “Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA.”

      (3) Supplemental Information: Line 113 - T is missing from Table.

      This has been added.

      (4) Move Table S2 to the main paper - this table provides a useful summary of the work.

      This has been moved.

      Main Manuscript:

      (1) Line 145: Serratia was not detected at G0 - was it detected later? It seems possible that titer could be very low to begin and increase in later generations; please clarify.

      We have previously monitored the aphid populations for the presence of Serratia transinfected from the same donor resource across subsequent generations, and Serratia was not detected at any later generation. We also did not detect Serratia in the other aphid species we examined, including green peach aphids (Gu et al., 2023 & 2025) and oat aphids (Yang et al., 2026). Therefore, we believe that the absence of Serratia at G0 was not due to a very low initial titer followed by an increase in later generations but instead that Serratia had been lost from the aphid population.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Gu et al., Transinfections of the endosymbiont Rickettsiella viridis in different Myzus persicae (Hemiptera: Aphididae) clones show consistent deleterious effects and stable transmission. J Econ Entomol 118, 1544-1552 (2025).

      Yang et al., A Rickettsiella transinfection in Rhopalosiphum padi reduces fitness and alate production but not plant virus transmission. Pest Man Sci 82, 3894-3906 (2026).

      (2) Line 206: "different aphid strain" - I learned from the manuscript that a single clone of D. noxia is found in Australia. Further, from my reading of the manuscript, I understand that one clonal isolate was propagated and then infected with symbionts. I think it is important to reword this sentence so that it is clear that the aphid genetic background is held constant, and that the only differences here are the presence or absence of the different symbionts. My reaction to this sentence was that you are working with the same aphid strain hosting different symbionts.

      We have clarified it by adding “the same aphid clone carrying different symbionts” after the different aphid strains.

      (3) Measuring symbiont "density" is a tricky thing to do; I explain this above in the public review. I suggest considering some revisions to the manuscript to be sure that you accurately reflect what has been measured and what can reasonably be inferred from using qPCR to measure gene copy number.

      We agree that qPCR-based measurements of symbiont gene copy number have limitations. In this experiment, we had a relatively large number of samples, and qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples. While we acknowledge the limitations of this approach, the relative differences in gene copy number can still provide an indication of variation in endosymbiont abundance and infection status among treatments. Unfortunately, other approaches remain challenging given resource and expertise limitations.

      (4) I think that you may be undervaluing the results of the mixed infection experiments; I find them to be compelling. To me, the data suggest that the symbionts increase aphid fitness.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single wheat plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with symbiont infection. However, we also examined aphid performance under population-level conditions, where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress (Figures 3 & 4). Under these conditions, the effects on fitness were different from those observed in the individual assays. We therefore agree that our results suggest that symbionts can increase aphid fitness under some conditions, but that this effect may be context-dependent and can differ under population-level conditions where density-dependent stress occurs. We have also added this information to our Results section to make this clear to readers.

      (5) Lines 288-291: This sentence doesn’t make sense to me. What "minor fitness costs" are being referred to? If the infected lines are increasing in representation relative to the uninfected lines, that suggests that there are not fitness costs, but fitness advantages under the experimental conditions.

      We have rephrased it to “minor fitness effects” which we refer to the fitness test under favourable conditions.

      (6) The section that begins on line 293 - I find this part to not be particularly robust and suggest dropping it from the paper.

      We appreciate the reviewer’s concern regarding the robustness of this section. We included this experiment to monitor changes in aphid population over time and to help explain the differences in plant feeding damage observed between aphids carrying different endosymbionts. Importantly, we conducted this experiment using whole wheat plants to evaluate whether the patterns observed in our other experiments were also evident. We believe that these results provide important complementary evidence for interpreting the differences in plant damage among the aphids hosting different endosymbionts and therefore are relevant to the overall conclusions of the study. For this reason, we prefer to retain this section in the manuscript.

      (7) Line 298: three plants per time point - I do not think this sample size is reflected in the methods of the paper.

      We have mentioned in the method part with “At day 14, three replicate plants were randomly selected from the available plants for both treatments and we counted the total number of nymphs, alate adults, and apterous adults. This was repeated again at Days 21 and 25”.

      (8) Line 304: "the frequency of alates decreased as plant damage increased" - this seems to be counterintuitive!

      As plant damage increased, the total aphid population also increased, resulting in an increase in the absolute number of alates. However, the frequency (proportion) of alates decreased because the increase in the total aphid population was greater than the increase in the number of alates.

      (9) Line 404: "horizontal transmission through plant tissue and/or transfer via aphid contact or honeydew" - what evidence is there that this happens? I have not kept on top of the literature with respect to transmission of secondary symbionts in aphids, but back when I was very familiar with that literature, the data did not support transmission by any of those routes. If there is now evidence supporting transmission by these routes, please cite it here.

      Previous studies have provided experimental evidence that horizontal transmission of aphid-associated endosymbionts can occur through plants. For example, plant-mediated transmission has been demonstrated for direct detection of secondary endosymbiont in the plant tissue including Hamiltonella defensa (Li et al., 2018), Rickettsia (Shi et al., 2024) and Serratia symbiotica (Pons, et al., 2019a). Our previous research also demonstrates that Rickettsiella endosymbionts were detected in uninfected aphids after feeding by infected aphids regardless of physical contact (Gu et al., 2023). Endosymbionts have also been detected in aphid honeydew (Darby and Douglas, 2003) and some primary transmission route appears to be horizontal, through honeydew (faeces) and host plant phloem (Pons, et al., 2019a & 2019b; Perreau et al., 2021). Appropriate references have been added to the revised manuscript in the Discussion section.

      References:

      Li et al., Plant-mediated horizontal transmission of Hamiltonella defensa in the wheat aphid Sitobion miscanthi. J Agric Food Chem 66, 13367-13377 (2018).

      Shi et al., Rickettsia transmission from whitefly to plants benefits herbivore insects but is detrimental to fungal and viral pathogens. mBio 15, e02448-23 (2024).

      Pons et al., Circulation of the cultivable symbiont Serratia symbiotica in aphids is mediated by plants. Front Microbiol 10, 764 (2019a).

      Darby and Douglas, Elucidation of the transmission patterns of an insect-borne bacterium. Appl Environ Microbiol 69, 4403-4407 (2003).

      Pons et al., New insights into the nature of symbiotic associations in aphids: infection process, biological effects, and transmission mode of cultivable Serratia symbiotica bacteria. Appl Environ Microbiol 85, e02445-18 (2019b).

      Perreau et al., Vertical transmission at the pathogen-symbiont interface: Serratia symbiotica and aphids. mBio 12, e00359-21 (2021).

      (10) Line 511: revise to "to measure their relative densities relative to a host gene".

      We have revised it.

      (11) Throughout the manuscript, please replace "five aged-matched" with an accurate description of the aphids used in the experiment. Please pay particular attention to the figure legends. Simply state e.g. "five 10-day-old apterous females" etc.

      This has been replicated in the main manuscript and figure legends.

      (9) Line 727 - lowercase t for Tests.

      This has been added.

      (10) Line 738 - what happens when you don’t exclude the early time points? Do your significant results go away? Also, please define what is meant by "early time points".

      When the early time points (Day 14 for Rickettsiella and Day 4 for Regiella) were included in the analysis, the GLM still showed a significant effect of strain on alate production for Rickettsiella (F<sub>1,12</sub> = 26.434, P < 0.001). We excluded these early time points from the analysis presented in the manuscript because, at these stages, aphid population sizes were still similar between treatments. We also conducted independent-sample t-tests at individual time points and observed differences at the later time points, when aphid population sizes began to diverge. We therefore considered the later time points to be more informative for assessing fitness effects and population sizes under increasing crowding conditions. We have now clarified this as “The earliest time points in the experiments for Rickettsiella (Day 14) and Regiella (Day 4) were excluded” in the manuscript.

      (11) In the legends of all figures, please be explicit about sample sizes.

      We have added the relevant information about replicate number or sample sizes in the main and supplementary figures.

      (12) Line 1010 - replace "each leave" with "each leaf" - there was at least one other place, I think in the supplemental information, that leave was used instead of "leaf".

      We replaced this.

    1. eLife Assessment

      This valuable study provides a systematically curated atlas of non-canonical open reading frames translated across normal human and mouse tissues by uniformly integrating nearly 400 ribosome-profiling datasets with independent mass-spectrometry evidence for ncORF-encoded peptides. The convincing analyses connect ncORF evolutionary age and coding constraint with translation level, tissue distribution, and co-translation with canonical coding sequences, offering a broad comparative view of how non-canonical translation changes over mammalian evolution. Although the functions of most ncORF-encoded peptides remain unknown, the scale of their detection across tissues, together with their proteomic and evolutionary signatures, suggests that a substantial subset may have biological functions.

    2. Reviewer #1 (Public review):

      [Editors' note: This revised version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the substantive comments raised in the previous round of review. The qualifications and limitations are now more clearly acknowledged in the revised manuscript.]

      Summary:

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Comments on previous version:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

    3. Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Weaknesses:

      Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

      (1) Bias and representations of data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

      (2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TErelated mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

      (3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated. Comparisons with other noncoding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

      (4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in drosophila, worms, mouse, and human. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied, for their functions and corss-species conservations. The authors should explicitly show what is new here in their analyses.

      (5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

      (6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

      (7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

      Comments on revisions:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We thank Reviewer #1 for the constructive comments and recognition of the value of our ncORF atlas. We have addressed the key concerns by strengthening the analyses and clarifying the framing and limitations of our conclusions. We appreciate the reviewer’s support for publication and believe these revisions have further improved the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

      (2) Some analytical methods and standards were not clearly presented in the manuscript.

      We thank Reviewer #2 for the positive assessment of our comprehensive dataset and standardized analytical framework. We have clarified the analytical methods and criteria throughout the manuscript and better defined the scope and limitations of our bioinformatics-based analyses. We appreciate the reviewer’s constructive suggestions, which have helped improve the clarity and rigor of the manuscript.

      Recommendations for the authors:

      Reviewing Editor:

      We have evaluated the revision together with the reviewers' second-round assessments and your responses. The reviewers agree that the manuscript has improved in clarity and that the standardized integration of large-scale Ribo-seq datasets provides a valuable resource for the field. However, several important concerns remain insufficiently resolved. In multiple cases, the revision relies primarily on acknowledgment or reframing of limitations rather than additional analyses or clearer methodological justification, leaving the evidential support for several conclusions incomplete.

      Because the study is entirely computational, all analytical procedures, criteria, and thresholds should be explicitly defined and adequately justified to meet the expected standard of technical rigor. In particular, key components of the analytical framework require clearer description, including the definitions and criteria used for co-translation and ncORF classification. The limitations of the dataset should also be discussed more explicitly, especially regarding the heterogeneity of public Ribo-seq datasets, technical factors influencing detection sensitivity, and the interpretation of variable detection frequencies across samples.

      In addition, several conclusions remain largely descriptive, and the distinction between novel findings and confirmation of previous observations should be clarified more carefully. Conclusions should be framed strictly within the limits of the presented data and positioned appropriately relative to prior work, with suitable citation to avoid overstating novelty.

      We therefore ask the authors to refine the technical descriptions, ensure that all methods and analytical criteria are presented unambiguously, and expand the Discussion to clearly articulate the limitations of the dataset and analysis. The conclusions should also be revised to reflect an appropriately cautious interpretation of the findings.

      The primary strength of this study lies in the scale and standardization of the dataset as a community resource. Given this substantial resource value, we believe the manuscript could become suitable for publication provided that the issues outlined above are addressed clearly and transparently. With these revisions, the work will provide a useful foundation for future studies in this area.

      We therefore invite you to submit a final revised version addressing the points described above.

      We thank the Editor for the careful assessment and constructive guidance. In the final revision, we have clarified all key methodological definitions and analytical criteria, expanded the Discussion of dataset and detection limitations, and revised the conclusions to avoid overstatement. We believe these changes improve the technical rigor, transparency, and overall value of the manuscript as a community resource.

      Reviewer #1 (Recommendations for the authors):

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We appreciate the reviewer’s support for publication.

      Reviewer #2 (Recommendations for the authors):

      While the authors have made commendable efforts to revise the manuscript and address previous concerns, the revised version still falls short of fully resolving several critical issues regarding data interpretation and methodological transparency. I recommend the following modifications before the manuscript can be considered for publication:

      (1) Although the authors have annotated the detection frequencies of individual sORFs in the revised

      Supplementary Table 3 and Fig. S1B, the biological and technical implications of these data require further clarification:

      (a) As the data demonstrates, even the most abundant sORFs were detected in no more than half of the samples. If the authors attribute this low detection rate to technical limitations (e.g., batch effects, sequencing depth, or threshold stringency), this must be explicitly discussed and annotated in the main text to prevent readers from misinterpreting this as low biological penetrance.

      We thank the reviewer for this constructive comment. We have added a paragraph to the Discussion explicitly addressing the potential technical factors underlying the variable detection frequencies to avoid overinterpreting these frequencies as biological penetrance.

      (b) The authors did not fully address my previous query regarding tissue specificity. Given the diverse and complex origins of the analyzed ribo-seq datasets, it is crucial to know whether any of these sORFs are tissue-specifically translated. The authors should analyze and state whether certain sORFs are exclusively detected in specific sample categories (tissues/organs), and whether this translation pattern aligns with the tissue-specific expression of their corresponding host transcripts.

      We thank the reviewer for raising this important point. Strict tissue-exclusive translation is difficult to establish from heterogeneous public Ribo-seq datasets, as gene expression is inherently stochastic and most genes have some probability of being expressed across tissues, although expression levels may vary substantially between tissues. Moreover, failure to detect an ncORF in a given tissue may reflect low expression or insufficient sequencing depth rather than true biological absence. We therefore avoid making definitive claims about tissue-specific translation and instead quantify variation in ncORF expression across tissues using a tissue specificity index.

      (2) Echoing the concerns raised by Reviewer #1, I remain concerned that the observed uORF-CDS cotranslation might be a computational artifact or false positive. The current Methods section lacks sufficient detail on how "co-translation" was strictly defined and quantified. I strongly recommend that the authors include a schematic diagram (e.g., in Figure 6 or supplementary figures) that explicitly details their analytical strategy, statistical thresholds, and evaluation criteria for defining co-translation. As was pointed out, it is well-established that uORFs typically exert an inhibitory effect on the translation of the main CDS. To validate the accuracy and robustness of their analytical pipeline, the authors should use their collected dataset to demonstrate the prevalence and nature of this canonical inhibitory phenomenon. Successfully capturing this expected repression would serve as a crucial positive control for their methodology. While the authors provided a theoretically acceptable mechanistic model in the text to reconcile co-translation with uORF-mediated repression, this hypothesis currently lacks literature support. The authors must cite relevant prior studies that support this specific regulatory dynamic to strengthen their argument.

      We thank the reviewer for this important comment. We have further clarified the definition, detection criteria, and statistical framework for co-translation in the Methods, and have made the complete analysis code publicly available to facilitate reproducibility and independent evaluation. We have also added relevant literature supporting the proposed regulatory interpretation clarifying the relationship between our observations and the established inhibitory effects of uORFs.

    1. eLife Assessment

      This study presents valuable findings regarding cardiac and autonomic effects of seizures and epilepsy, with relevance to sudden unexpected death in epilepsy (SUDEP). They present solid evidence that genetic deletion of the potassium-chloride co-transporter in hypothalamic corticotropin-releasing hormone (CRH) neurons exacerbates bradycardia and enhances autonomic disturbances in a mouse model of temporal lobe epilepsy. This work will be of interest to neuroscientists working on epilepsy, the HPA axis, and autonomic control.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript entitled "Autonomic reflex plasticity associates with time-dependent SUDEP susceptibility in a murine model with hyperactive stress circuits" by Dr. Saunders and colleagues combined a traditional mouse model of SUDEP, ventral intrahippocampal kainite (vIHKA) injection, with a genetic model of chronic hyperactivity of central corticotropin-releasing hormone (CRH) neurons (Kcc2/Crh) that further increases the risk of SUDEP in the weeks following seizure.

      Strengths:

      Their results show during spontaneous seizures Kcc2/Crh mice had more pronounced reflex-like ictal bradycardias compared to WT controls that notably occurred prior (~10 sec) to seizure termination and had greater autonomic disturbances compared to WT controls, including a pronounced serotonin-mediated Bezold Jarisch reflex. These results show chronic hyperactivity of central corticotropin-releasing hormone (CRH) neurons (Kcc2/Crh) increased autonomic disturbances and risk of SUDEP in a kainic acid model of epilepsy.

      Weaknesses:

      This study could be improved with a more thorough assessment of heart rate, blood pressure and breathing during and following the seizures, and in particular the fatal event. It is unclear if the bradycardias were spontaneous, or a result of preceding central or obstructive apneas, oxygen desaturations, hypercapnia, arrhythmias, or other possible triggers.

      Considerable prior work in the literature suggests SUDEP could be mediated, in some patients, by a burst of parasympathetic activity to the heart. Were the heart rate changes in these animals during seizures inhibited or blocked by atropine, or atenolol? The injection of the 5HT agonist phenylbiguanide into the right jugular is not a selective approach for activating the Bezold Jarisch Reflex (BJR) which is caused by increased activity of intracardiac sensory neurons (generally activated with ischemia or a combination of low preload with high contractility). The results should be interpreted more cautiously, as a response to systemic administration of phenylbiguanide only.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors set out to evaluate the role of hypothalamic pituitary axis hyperactivity on cardiac and autonomic changes during epileptogenesis and following seizures in a mouse model of temporal lobe epilepsy. Epilepsy is very common. It can frequently result in death from sudden unexpected death in epilepsy, or SUDEP. SUDEP is thought to be at least in part due to seizure related cardiac and autonomic instability. Increased stress states are well known to be comorbid with epilepsy. This comorbidity is thought to increase the risk of SUDEP. Here the authors hypothesized that a mouse model of heightened stress in which there is hyperactivity of the CRH neurons in the hypothalamus would demonstrate exaggerated cardiac and autonomic effects of seizures and epilepsy.

      Strengths:

      For the chronic stress model, they employed the Kcc2/Crh mice that have a genetic deletion of the potassium chloride cotransporter in CRH neurons. They treated these mice and their wild type littermates with intra hippocampal kainic acid or saline, as epileptic and sham-treated animals respectively. The assessed cardiac activity, blood pressure, baroreflex, and the Bezold-Jerisch reflex during epileptogenesis. This in general is an interesting study. They make some interesting and potentially important observations regarding heart rate and blood pressure in seizures and epilepsy.

      Weaknesses:

      While the revised manuscript is much improved, there are still some concerns that should be addressed.

      (1) The low-pressure baroreceptor responses they show in Figure 4 are somewhat confusing. Should they not be seeing a reflex increase in heart rate when blood pressure is lowered with sodium nitroprusside? It would be helpful if they could describe whether the control responses were as expected or not, and if not, why? This makes it difficult to assess changes seen in the different genotypes and conditions.

      (2) It does not seem appropriate to label the assessments associated with Figure 5 as the Bezold Jarisch Reflex. This reflex involves bradycardia, hypotension, vasoconstriction, and hypopnea. They seem to be only looking at the cardiac component, which is likely mediated though peripheral 5HT3 receptors. Did they measure blood pressure and breathing? Can they include these? If they can only comment on HR, then the discussion should reflect this.

      (3) In Figure 1B, it would be helpful to show some short (e.g., 0.5 sec) snippets of ECG traces that exemplify the changes in HR (spikes/second).

      (4) The day 21 examples given in Figure 1B, do not seem to be representative of the data depicted in Figure 1C.

      (5) From the top panel examples in Figure 2A it looks like there might be greater EEG suppression following seizures in the Kcc2/CRH mice. Was this consistent? It might be worth looking into.

      (6) Can the authors include scale bars for the top panels in Figure 2A?

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents valuable findings regarding cardiac and autonomic effects of seizures and epilepsy, with relevance to sudden unexpected death in epilepsy (SUDEP). They present solid evidence that genetic deletion of the potassium-chloride cotransporter in hypothalamic corticotropin-releasing hormone (CRH) neurons exacerbates bradycardia and enhances autonomic disturbances in a mouse model of temporal lobe epilepsy. However, the evidence that this deletion produces chronic hyperexcitability of the hypothalamic-pituitary-adrenal axis was incomplete, leaving a mechanistic gap. This work will be of interest to neuroscientists working on epilepsy, the HPA axis, and autonomic control.

      We thank the editors and reviewers for their feedback. Although the loss of Kcc2 from CRH neurons in Kcc2/Crh mice has been confirmed (Melon et al., 2018) and leads to HPA axis hyperexcitability in response to stress or seizures, it does not chronically drive HPA axis hyperactivity in unstressed conditions (Basu et al., 2024).

      The following details have been added regarding the Kcc2/Crh model. We now describe the Kcc2/Crh as “hyperreactive” rather than “hyperexcitable/hyperactive” throughout the manuscript.

      -In the Introduction: “This loss of Kcc2 in CRH neurons has been previously confirmed and shown to cause an exaggerated HPA axis response to stress that is absent in baseline conditions (Basu et al., 2024; Melon et al., 2018)”

      -In the Discussion: “This aligns well with lack of elevated plasma corticosterone at baseline in Kcc2/Crh mice, compared to WT, because elevated PVN<sup>CRH</sup> neuron activity should otherwise increase this signal (Basu et al., 2024).”

      -In the Discussion: “Most notable, our model utilizes a developmental strategy to knock out Kcc2 from CRH neurons, which has been confirmed previously (Melon et al., 2018).”

      Public Reviews:

      Reviewer #1 (Public review):

      This study could be improved with a more thorough assessment of heart rate, blood pressure and breathing during and following the seizures, and in particular the fatal event.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2. Overall, pronounced bradycardia that occurred near seizure termination was followed by recovery of HR to pre-ictal baseline in the early post-ictal period (30 sec). In Results: “Independent of genotype, HR recovered to baseline levels during the immediate post-ictal period (0-10 sec, Fig. 2G; 10-30 sec, Fig. 2K).

      In pilot work, we determined that HR during spontaneous seizures were fundamentally different than HR during status epilepticus (see Author response image 1). Therefore, it was critical for us to examine HR during spontaneous seizures. We previously published that Kcc2/Crh+KA mice have a rate of 1-2 seizures per day. To limit additional stressors and seizure provocation, we employed radio telemetry (over tethered systems) and opted out of carotid instrumentation for BP as well as restricted environments of plethysmography chambers for these assessments of HR. After determining that heart rate was different between our mouse lines, we tested whether this change was mediated by central circuits regulating HR (ie: baroreflex, Bezold-Jarisch reflex) (Fig. 3, 4, 5).

      Author response image 1.

      In Discussion: “Although the present study includes only non-fatal seizures, our report of HR during spontaneous seizure events supports work suggesting physiological events during non-fatal seizures predict SUDEP risk (Lamrani et al., 2023; Ryvlin, Nashef, & Tomson, 2013; Schuele et al., 2011). It remains to be determined if ictal events during fatal and non-fatal spontaneous seizures are different and future studies could help clarify any distinctions. Our examination focused on HR (and not respiration or BP). Whether exaggerated BJR-mediated HR response co-occurs with greater magnitude BJR-mediated apnea and hypotension remains to be determined.”

      It is unclear if the bradycardias were spontaneous or a result of preceding central or obstructive apneas, oxygen desaturations, hypercapnia, arrhythmias, or other possible triggers.

      Our work demonstrates the occurrence of ictal bradycardia whereas identifying precipitating factor(s) and interaction(s) of this phenomenon will require alternate approaches. Normal activation of hypoxic and hypercapnic ventilatory responses would be expected to increase HR. However, whether these chemoreflex circuits undergo remodeling in Kcc2/Crh mice remains unknown. Obstructive apnea via laryngospasm can cause reflex bradycardia, but we did not record airflow or respiratory EMG in these studies to determine the existence of obstructive apnea. More testing is merited.

      In Discussion, “…Additional seizure-related disturbances such as central or obstructive apneas may contribute to BJR activation. Hypoxia, which could result from ictal apnea, is known to increase excitatory neurotransmission to cardiac vagal motor neurons that cause vagal bradycardia (Griffioen et al., 2007) and induces platelet activation (Tyagi et al., 2014) which is considered the main source of circulating serotonin for the BJR. As such, hypoxia resulting from apnea may increase likelihood of exaggerated BJR during seizures. Consistent with this…”

      Considerable prior work in the literature suggests SUDEP could be mediated, in some patients, by a burst of parasympathetic activity to the heart. Were the heart rate changes in these animals during seizures inhibited or blocked by atropine or atenolol?

      With this study targeting spontaneous seizures we were unable to test acute pre-treatment with atropine or atenolol. We did observe reduced mortality in Kcc2/Crh+KA mice that underwent chronic parasympathetic blockade via osmotic minipump of methylscopolamine (Figure 5).

      In Discussion:

      “Chronic inhibition of vagal parasympathetic motor output (the driver of BJR reflex bradycardia) improved mortality by 10% in Kcc2/Crh mice. Although this improvement provides some hope for patients at high risk for SUDEP with no treatment options, additional avenues of investigation are needed to more directly link seizure-related bradycardias to vagal parasympathetic motor output.”

      The injection of the 5HT agonist phenylbiguanide into the right jugular is not a selective approach for activating the Bezold Jarisch Reflex (BJR), which is caused by increased activity of intracardiac sensory neurons (generally activated with is chemia or a combination of low preload with high contractility). The results should be interpreted more cautiously, as a response to systemic administration of phenylbiguanide only.

      BJR can be experimentally triggered with intravenous infusion of various compounds including veratrum alkaloids (Cramer, 1915), 5HT (Fozard 1983), or 5HT3R agonists (Verberne & Guyenet, 1992). We now specifically refer to BJR in our study as that induced by PBG (a 5HT3R agonist), as others have done (Yamano et al., 1995; PMID: 8786638) and acknowledge endogenous BJR activation in the discussion.

      Added to the results: “As dysfunction of serotonergic signaling is implicated in the pathophysiology of SUDEP (Richerson & Buchanan, 2011), we investigated the Bezold Jarisch Reflex (BJR) (Fig. 5), a cardioinhibitory reflex that is reliably triggered experimentally by activation of cardiopulmonary vagal afferents containing serotonin type 3 receptors (5HT3R) (Fozard 1983; Yamano et al., 1995).”

      Added to Discussion: “Although our report is the first to link BJR to SUDEP, serum serotonin levels are elevated following generalized seizures (Murugesan et al., 2018), likely via release from activated platelets (Cloutier et al., 2018). This surge in serum serotonin could lead to endogenous activation of BJR, as bolus intravenous infusion of serotonin reliably triggers BJR experimentally (Fozard 1983; Whalen et al., 2000).”

      Reviewer #2 (Public review):

      Some of the conclusions may be a bit overstated as is and would benefit from more discussion and perhaps additional data.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2.

      The Discussion now includes more details regarding respiration, BP, obstructive apnea, properties of the BJR, and distinctions in seizure type based on whether evoked or lethal.

    1. eLife Assessment

      This important study provides new insights into the patterns of organelle inheritance in the protozoan parasite Toxoplasma gondii. The authors introduce an innovative dual-labeling approach to distinguish maternally inherited from de novo synthesized organelles, representing convincing evidence that different organelles follow distinct inheritance fates during parasite replication. Future studies will be needed to determine whether the residual body functions as a central recycling hub, as the current data are also consistent with alternative models.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting and metabolic compartments) are divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to better understanding the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites. In particular, this study shows that the residue body is a region of the cell syncytium that organelles can be actively transported from. Therefore, it is a space that can actively contribute to the segregation of the late segregating micronemes and rhoptries.

    3. Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of toxoplasmosis. Parasite invasion of host cells, intracellular replication, and subsequent egress, which results in destruction of the infected cell, are central to pathogenicity. This manuscript focuses on understanding how maternal resources, specifically cellular organelles, are shared between daughter parasites during cell division. Many organelles are present as a single copy, making their division and inheritance essential for successful replication. In T. gondii, our understanding of how organelles are divided during cell division remains limited, and this study helps address this important knowledge gap.

      Strengths:

      The major strength of this study is the use of a Halo-based pulse-chase assay to characterize patterns of organelle inheritance and to monitor protein synthesis, turnover, and movement. This approach will be of considerable interest to the field. Using this method, the authors identify three major modes of organelle inheritance:

      (1) Organelles present in multiple copies (such as micronemes and rhoptries) are partitioned between daughter parasites, with additional contributions from newly formed vesicles. Newly synthesized and pre-existing material remain as distinct populations within the cell.

      (2) Single-copy organelles, such as the Golgi and apicoplast, are expanded through the incorporation of newly synthesized material before division.

      (3) Cytoskeletal structures are synthesized de novo during each round of cell division.

      These findings provide a more refined understanding of organelle inheritance and demonstrate that secretory organelles are not generated entirely de novo during each round of division, as was previously thought.

      The paper places particular emphasis on the fate of maternal micronemes and rhoptries during division. The data show that (1) during division in wild-type cells, maternal micronemes and rhoptries are detectable in the residual body (RB); however, the majority of these organelles are localized within the parasite body, either at the apical or basal ends of the daughter parasites (Fig. 6). (2) In the absence of the myosin motor MyoF, micronemes and rhoptries accumulate in the residual body and are not properly trafficked to the daughter cells. Upon restoration of MyoF protein levels, these organelles redistribute to the daughter cells, although in an uneven manner.

      Weaknesses:

      The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. While the authors propose that the RB acts as a central hub for recycling both organelles, the current data do not fully support this conclusion.<br /> The model that microneme and rhoptry recycling is RB-dependent relies largely on the MyoF depletion phenotype and the limited detection of maternal organelles in the RB of wild-type parasites. Alternative models remain plausible, including direct trafficking to daughter cells, with RB accumulation upon MyoF depletion reflecting impaired trafficking rather than an obligatory RB-dependent recycling pathway, as now discussed by the authors.

    4. Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughter-derived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strength:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are convincing. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Weaknesses:

      (1) In addressing the question of residual body participation in sorting of organelles, a clear definition of this structure is required including when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. The authors' definition is as follows: 'The RB originates from the collapse of the maternal parasite during daughter cell budding and occupies the space previously occupied by the mother cell.' As such, a clear marker of the mother cell 'collapse' is required, but such a marker is not identified or used in the study to separate what might be considered an active part of the mother cell during early daughter formation, and the residual body. This might seem like moot a point, but it would help to give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? The authors elegantly show that MyoF is necessary for segregation of micronemes and rhoptries into daughters, and that MyoF depletion leads to accumulation of these organelles within the residual body. Moreover, restored expression of MyoF can then recover these organelles. This clearly demonstrates the activity of the residual body as part of the syncytium space that participates in the maintenance of the vacuole. But does it imply that this space necessarily handles all inherited micronemes and rhoptries as a 'trafficking hub'? My concern with the lack of a clear definition could provide some misinterpretation or overinterpretation of the contribution residual body.

      We thank the reviewer for raising this conceptual point. We agree that, in the absence of a molecular marker that uniquely defines the nascent RB, the precise transition between posterior maternal cytoplasm and a morphologically distinct RB cannot be determined during early daughter formation. We have therefore clarified our terminology in the revised manuscript and define the RB operationally as the posterior compartment/connection between daughter parasites. We also avoid assigning early posterior trafficking events unambiguously to a fully formed RB. We further agree that the current data do not establish that all inherited micronemes and rhoptries must transit through the RB. We have therefore revised the Results and Discussion and softened terminology such as “central trafficking hub” and now conclude that the RB represents an important dynamic compartment in organelle recycling and redistribution, without implying that it is an obligatory intermediate for every inherited secretory organelle.

      (2) A further, remarkable conclusion is that maternal micronemes are evenly segregated into daughters through an active process for 'balanced microneme inheritance'. The proportion of maternal micronemes is quantified up to the 8-cell stage and shown to be not significantly different between cells. But would this result be expected with random assortment at this stage? The authors model the probability of a 32-cell stage vacuole occurring with each daughter having within 0-3 maternal micronemes and this is considered unlikely. However, the authors neither present the modelling for the 8cell stage or show quantification of 32-cell vacuoles. They do show some images of large vacuoles, but it is not possible to determine the distribution of maternal micronemes in these images. A regulated process of segregation would require a complex mechanism where some form of microneme counting would be required to create the proposed balance. It is, therefore, important to have strong data supporting such a hypothesis, but this is not currently presented.

      We agree that the previous wording implied a mechanistic conclusion beyond what can be established from the present dataset. We have therefore revised the manuscript so that the relatively even distribution of maternal micronemes at the 8-cell stage is presented as an observation that is consistent with a non-random or regulated partitioning process, rather than evidence for an established microneme-counting mechanism. The 32-cell model is now presented as supportive rather than definitive evidence, and we explicitly acknowledge that the quantitative experimental dataset was obtained at the 8-cell stage. 

      Reviewer #2 (Public review):

      (1) The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. The authors strongly argue that the RB is a central hub for recycling micronemes and rhoptries; however, this conclusion is not fully supported by the data. For example, the authors state…Thus, the model that all microneme and rhoptry trafficking is RB-dependent is based primarily on the MyoF depletion phenotype (which results in RB accumulation) together with the observation that a relatively small amount of maternal microneme and rhoptry material is detectable in the RB of wild-type parasites. Although the authors' interpretation-that recycling is RB-dependent-is one possible explanation, alternative models are not discussed. For example, an alternative possibility is that the majority of micronemes and rhoptries are trafficked directly from the apical end of the mother parasite to the daughter cells without passing through the RB. In this scenario, only a subset of the organelles would enter the residual body, perhaps reflecting imperfect trafficking efficiency rather than an obligatory recycling step. Loss of MyoF would impair this trafficking pathway, resulting in the accumulation of secretory organelles within the RB. In other words, RB accumulation could be a consequence of MyoF depletion rather than evidence that all trafficking in wild-type parasites normally proceeds through the RB. This alternative interpretation seems particularly relevant for the rhoptries, given that the authors themselves state that "M-RON2 was integrated into daughter rhoptries prior to mother cell collapse and formation of the RB."

      We agree with the reviewer that the current data do not establish obligatory transit of all maternal micronemes and rhoptries through the RB. We have revised the manuscript throughout to make this distinction explicit. In particular, we now emphasize that maternal MIC2 can be directly observed entering the RB, whereas most maternal RON2 is incorporated into daughter rhoptries before mother-cell collapse and formation of a morphologically distinct RB. This observation leaves open the possibility that a substantial fraction of maternal rhoptries is transferred directly from the mother to developing daughters. We have also revised our interpretation of the MyoF-depletion phenotype. The accumulation of maternal MIC2 and RON2 in the RB following MyoF depletion demonstrates that MyoF is required for efficient redistribution of both organelle populations, but does not by itself demonstrate that both normally follow an identical spatial route through the RB. We now explicitly state that their precise trafficking routes and timing may differ.

      We nevertheless retain the conclusion that the RB is a dynamic compartment involved in organelle recycling because maternal MIC2 can be directly observed entering and leaving this compartment, and material accumulated there following MyoF depletion can subsequently be redistributed after restoration of MyoF.

      (2) Figure S10C. To determine whether microneme degradation occurs in the RB, the authors quantified the fluorescence intensity of individual micronemes in control parasites and following auxin washout, showing that after redistribution the fluorescence intensity of individual vesicles is unchanged. However, this is not the appropriate analysis to address the question being asked. To conclude that micronemes are not degraded, the authors would need to quantify the total fluorescence intensity within the entire vacuole. For example, if half of the micronemes were degraded, the remaining micronemes would be expected to retain the same fluorescence intensity as those in the control parasites. Thus, unchanged fluorescence intensity of individual vesicles does not exclude the possibility that degradation has occurred.

      We agree with this criticism and have revised the interpretation of the experiment accordingly. We no longer conclude that the analysis excludes microneme degradation. We now explicitly acknowledge that analysis of individual recovered micronemes cannot exclude degradation of a fraction of the total microneme population during RB retention. Thus, the experiment supports preservation of MIC2 signal in the recovered organelles but is no longer presented as evidence that no microneme degradation occurs.

      Reviewer #3 (Public review):

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body. 

      We agree and have incorporated this limitation into the revised interpretation. Our live imaging demonstrates that maternal MIC2 can enter the RB and subsequently redistribute to daughter parasites, but it does not establish that every individual maternal microneme follows this route. We thank the reviewer for highlighting this distinction, which has helped us clarify the model presented in the Results and Discussion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Carruthers and Sibley, 1997, is still the cited ref for Tic20 in the apicoplast despite the authors saying that would correct this to a van Dooren publication.

      Corrected

      (2) In Figure 2B three biological replicates were used however there are no error bars shown. The only reason for performing replicates is the observe the variance in the data, so if this is not shown the replicates are effectively meaningless. I strongly advice that error bars are given to indicate this seeing that the data in this figure form the basis of the major conclusions of the study. It might be necessary to show this in supplemental forms with fewer proteins if the error bars are too difficult to see in the combined figure.

      We agree and have corrected Figure 2B to display the variability between the three independent biological replicates. Error bars now represent the standard deviation.

      (3) Line 136: Can you conclude that these inheritance patterns are 'organelle-specific' when each organelle is only sampled with one or two proteins. Isn't it better to conclude that these are protein-specific, with the hypothesis that they might represent the orgnalle as a whole. I imagine that some proteins in organelles such as the apicoplast have shorter half-lives than others, and therefore some apicoplast proteins might behave like 'Group 3' proteins.

      We thank the reviewer for raising this point. We agree that individual proteins within the same organelle may differ in their turnover kinetics and that analysis of one or two markers cannot establish that every molecular component of an organelle behaves identically. However, we do not think that describing the observations exclusively as protein-specific inheritance would fully reflect the biological process investigated here. The proteins analysed are established markers of defined organelles, and our conclusions are based not only on changes in fluorescence intensity, but also on the localization, morphology, partitioning, and spatial relationship between maternally inherited and newly synthesized organelle populations.

      This is particularly evident for micronemes and rhoptries, where maternal and de novo material remain spatially separated and individual organelles can be followed during inheritance.

      In addition, the microneme phenotype observed with MIC2 was confirmed using AMA1, MIC4, and MIC8.

      We therefore retain the terminology of organelle inheritance, while acknowledging that individual proteins within a given organelle may exhibit different turnover kinetics and that the markers analysed may not represent the behaviour of every molecular component of the organelle.

      (4) Line 222: It is an odd phrase to suggest that the Golgi, ER etc 'bypass' the residual body, which suggests an active avoidance mechanism. Would the authors also conclude that the nucleus 'bypasses' the RB? Moreover, the ER and mitochondria are actually known to be present in the RB forming continuous organelles between daughters in a vacuole. So again, this might be an overstatement that mispresents how the RB participates in vacuole functions.

      We agree and have removed the term “bypass.” The revised text now states only that we did not observe comparable accumulation of the analysed Golgi, ER, or apicoplast markers in the RB during inheritance. This avoids implying an active avoidance mechanism and is compatible with the known continuity of ER and mitochondria through the RB.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      Figure 8F and Video S6: The authors should specify the time point after IAA washout at which live imaging was initiated. Does time 0 in the video correspond to the point at which IAA was removed?

      We have clarified this in the Results and Methods. Auxin was removed after 24 h of replication, and live imaging was subsequently initiated. Time 0 in Figure 8F and Video S6 corresponds to the first acquired frame after auxin washout.

      Figure S10 should read auxin, not auxine.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      I have no further suggestions. Congratulations to the authors for a lovely study.

    1. eLife Assessment

      This important study shows that nutrient resorption efficiency in the widespread wetland grass Phragmites australis varies with population origin and remains largely stable under experimentally imposed salt stress, supporting limited short-term plasticity and a role for genetic differentiation. The findings suggest that predictions of wetland nutrient cycling under increasing salinization should account for intraspecific variation and phylogeographic composition, the evidence is compelling, based on a common-garden experiment involving 110 genotypes, paired control and salinity treatments, and convergent metabolomic, ionomic, and whole-plant evidence confirming substantial physiological stress. The revised manuscript clarifies the partial overlap between ecotype and phylogeographic lineage and appropriately qualifies the modest explanatory contribution of latitude, conclusions remain restricted to one species, one growing season, and a single moderate salinity treatment; an ecotype-specific nitrogen resorption response also indicates that plasticity is not entirely absent and the physiological mechanisms and responses to chronic, stronger, or multigenerational salinity exposure remain unresolved. The study will interest researchers working on plant functional ecology, nutrient cycling, and wetland responses to global change.

    2. Reviewer #1 (Public review):

      Summary:

      This study demonstrates that nutrient resorption efficiency (NuRE) in Phragmites australis is genetically canalized rather than plastic to salt stress. Using 110 genotypes in a common garden, the authors show that intraspecific variation in NuRE is explained by phylogeographic lineage, ecotype, and latitude, not by effective salinity. Element specific regulatory strategies further reveal how N, P, and K resorption are differentially controlled. At the population level, this is an important study that fundamentally advances our understanding of plant functional trait evolution and its implications for ecosystem nutrient dynamics under global change.

      Strengths:

      This study is the first to demonstrate genetic determination of a key nutrient conservation trait under effective salt stress in a widespread macrophyte, directly testing the 'plastic acclimation versus inherent conservatism' paradigm in a non-nutrient stress context. The experimental design is rigorous: each genotype was paired across control and salt treatments, and multilevel stress effectiveness (metabolomics, biomass, Na accumulation) was confirmed before evaluating NuRE. The large sample size of a macrophyte and dual classification (phylogeography + ecotype) allow robust disentangling of genetic versus plastic sources of variation.

      The analysis comprehensively tests three resorption control hypotheses using appropriate SMA regression, revealing element specific and condition dependent patterns. The latitudinal gradient and variation partitioning provide strong evidence that genetic origin and geographic context outweigh short term plasticity, with important implications for predicting ecosystem nutrient cycling under global change. This study provides a clear empirical demonstration that a key nutrient conservation trait can remain homeostatic under non nutrient stress, and that intraspecific variation is primarily a product of population differentiation rather than short term plasticity.

      Weaknesses:

      First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study's main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermine the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

      Comments on revised version.

      The author carefully revised the parts that might cause confusion.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript shows that nutrient resorption efficiency in Phragmites australis is largely genetically canalized rather than plastic under effective salt stress, with variation mainly associated with phylogeographic lineage, ecotype, and latitude.

      Strengths:

      The common-garden experiment with 110 genotypes, paired control and salt treatments, multi-level stress validation, and element-specific analyses provides compelling support for the central conclusions.

      Weaknesses:

      The generality is limited by the focus on a single species and a single growing season, and some mechanisms are inferred indirectly.

    4. Author response:

      The following is the authors’ response to the original reviews.

      The major revisions include:

      (1) Conceptual framing and scope: defined canalization in the Introduction and clarified that our conclusions are restricted to limited NuRE plasticity during one growing season under a single moderate salinity treatment.

      (2) Methods and classification: clarified the substrate composition and elemental measurements, specified that the ecotype analysis included only Chinese populations, and explained the partial association between ecotype and phylogeographic group.

      (3) Interpretation: expanded the discussion of K resorption and inverted nutrient limitation and tempered the interpretation of latitude and the substantial unexplained variation.

      (4) Robustness and presentation: added Supplementary Figure S7 showing that carbon standardization did not alter the main conclusions, added significance symbols to Table 1, corrected the unit in Figure 2b, and revised repetitive wording in the Discussion.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      (R1-P1) First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long-term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study’s main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      We agree that our evidence is limited to the absence of NuRE plasticity during one growing season under the imposed salinity treatment and does not resolve chronic or multigenerational responses or their physiological basis. We therefore revised Discussion 4.1 to delimit the canalization inference, identify the proposed mechanisms as untested, and specify the longer-term and mechanistic studies needed to distinguish among them.

      Discussion 4.1, fourth paragraph, inserted immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied ”.

      “Our inference of canalization is therefore limited to the absence of a plastic NuRE response during one growing season under the imposed salinity treatment. Chronic, more severe, or multigenerational salinity exposure may produce acclimatory, epigenetic, or transgenerational responses that cannot be evaluated here. Moreover, although altered phloem loading, disruption of senescence-associated remobilization, and reallocation towards osmotic adjustment are plausible explanations for the observed response, we did not directly measure these mechanisms. Long-term experiments combined with targeted molecular and transport measurements are needed to distinguish among these possibilities.”

      (R1-P2) Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermines the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

      We agree and have clarified all three evidential limits. First, the resorbed N: P and N: K analyses are now described as indirect evidence consistent with nutrient-limitation control, not as a causal test; direct nutrient-addition experiments would be required for causal inference. Second, metabolomics is identified as validation of physiological stress rather than a genotype-specific mechanistic analysis. Third, because ecotype and phylogeographic group are partly associated, they were fitted in separate models. The genotype random effect accounts for paired measurements but does not remove confounding between the classification schemes; ecotype differences are therefore interpreted as complementary rather than independent evidence.

      Discussion 4.2, third paragraph.

      “Within this context, the consistent ‘inverted’ nutrient limitation pattern (i.e., a slope significantly >1 for the relationship between log Resorbed N:P and log Green N:P) provides indirect evidence consistent with nutrient-limitation control at the intraspecific level in P. australis, but it cannot establish causal nutrient limitation. Direct factorial nutrient-addition experiments would be required to determine whether the observed resorption patterns are driven by the relative limitation of N, P, or K.”

      Discussion 4.1, first paragraph, inserted immediately after the sentence ending “providing a robust foundation to evaluate NuRE responses”.

      “In this study, metabolomic profiling was used primarily to confirm that the salinity treatment induced broad physiological stress, rather than to resolve genotype-specific metabolic mechanisms underlying NuRE variation. Integrating metabolite profiles with genotype-level NuRE responses would be a valuable direction for future mechanistic research.”

      Methods 2.4.

      “Phylogeographic group and ecotype were analysed in separate linear mixed-effects models because ecotype classifications were available only for Chinese populations and were partly associated with phylogeographic structure. Each model included salinity treatment and either phylogeographic group or ecotype as fixed effects, with genotype fitted as a random effect to account for the paired experimental design in which each genotype was exposed to both control and salt conditions.”

      Discussion 4.4, first paragraph, inserted immediately after the sentence ending “governed by geographic origin (phylogeographic group and ecotype)”.

      “The separate-model approach avoids including the two correlated classification schemes as simultaneous independent predictors, but it does not fully disentangle deep phylogeographic history from recent habitat-associated differentiation. We therefore interpret the ecotype analysis as complementary evidence of habitat-associated differentiation rather than as an effect independent of phylogeographic history.”

      Reviewer #2 (Public review):

      (R2-P1) The experiment covers only one growing season, with salinity applied in June and measurements in December. While the stress is clearly effective, longer-term or multi-year stress might reveal acclimation or epigenetic effects that are not captured. Given the author team’s expertise in parental and transgenerational effects in clonal plants, this limitation is particularly relevant and warrants more thorough discussion in the manuscript.

      We agree. This concern overlaps with Reviewer #1’s temporal-scope comment. We have revised Discussion 4.1 to state explicitly that our inference is restricted to the absence of a plastic NuRE response during one growing season. We also acknowledge that chronic or multigenerational exposure could induce acclimatory, epigenetic, or transgenerational responses that were not captured by the present design.

      See the full revised text under R1-P1 above.

      (R2-P2) The salinity treatment uses a single moderate level of 10 ppt, which does not allow assessment of whether more extreme stress might trigger a plastic response. A dose-response design across a gradient would have provided stronger inference about the threshold at which NuRE canalization might be overcome. Additionally, the ecotype analysis in Figure 4 applies only to Chinese populations, as classification was not available for non-Chinese populations, which should be stated more explicitly in the Results.

      We agree with both points. We have revised Discussion 4.1 to acknowledge that the single 10 ppt treatment does not exclude the possibility of a plastic NuRE response at higher salinity or along a broader dose-response gradient. We therefore frame the identification of a possible response threshold as a priority for future experiments.

      We have also revised Results 3.2 to state explicitly that the ecotype analysis in Figure 4 included only Chinese populations because ecotype classifications were unavailable for non-Chinese populations.

      Discussion 4.1, fourth paragraph, at the same revision point as R1-P1: immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied”.

      “Because only one moderate salinity level (10 ppt) was tested, our results do not exclude plastic NuRE responses at higher salinity or along a dose-response gradient. Future experiments should determine whether a threshold exists beyond which the apparent stability of NuRE is overcome.”

      Results 3.2, second paragraph, inserted immediately after the sentence reporting the ecotype effects and ending “(Figure 4; Figure S5)”.

      “It should be noted that the ecotype analysis here was based only on Chinese populations, as ecotype classification was not available for non-Chinese populations.”

      (R2-P3) The variation partitioning shows latitude as a significant predictor, but the R<sup>2</sup> values are relatively low, indicating that much variance remains unexplained. The manuscript should avoid overinterpreting latitude’s explanatory power and more openly acknowledge the role of unmeasured factors. The interpretation of slopes greater than 1 for the resorbed N:P versus green N:P relationship, labeled as “inverted limitation”, also needs further explanation regarding its functional significance.

      We agree that the original wording overemphasized latitude. Results 3.4 and Discussion 4.3 now describe latitude as the largest contributor among the measured predictors while emphasizing its modest individual R<sup>2</sup> and the substantial unexplained variation. We also expanded Discussion 4.2 to explain that slopes greater than 1 indicate disproportionate recovery of P or K relative to N and to present nutrient balance, P conservation, K mobility, and constraints on N remobilization as non-exclusive hypotheses rather than established mechanisms.

      Results 3.4, first paragraph, replacing the sentence beginning “Furthermore, variation partitioning analysis indicated that latitude”.

      “Variation partitioning indicated that latitude had the largest individual contribution among the measured predictors, but its contribution was modest for N, P, and K resorption (individual R<sup>2</sup> = 0.092, 0.057, and 0.094, respectively; Table 1). The very small contributions of green-leaf P concentration (R<sup>2</sup> = 0.006) and stoichiometry (R<sup>2</sup> < 0.001) further indicate that most variation in P resorption was associated with factors not represented in the present models.”

      Discussion 4.3, first paragraph.

      “Although latitude explained more variation than the other measured predictors, its individual contribution remained modest. The substantial unexplained variance indicates that additional climatic, edaphic, demographic, or genetic factors also contribute to NuRE variation. We therefore interpret latitude as a significant but limited correlate of NuRE rather than as a dominant determinant.”

      Discussion 4.2, first paragraph.

      “Functionally, slopes greater than 1 indicate that changes in green-leaf N:P or N: K are accompanied by disproportionate changes in the corresponding resorbed ratio, consistent with relatively greater recovery of P or K than of N across the observed nutrient gradient. This pattern may contribute to maintaining internal N:P:K balance during regrowth. It may also reflect stronger conservation of P, the high mobility of K, or constraints on the remobilization of N retained in structural or metabolic compounds. Because the experiment did not include nutrient additions or direct measurements of remobilization costs, these explanations remain hypotheses.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) Clarify the concept of “canalization” early. Consider adding one sentence in the Introduction or Discussion explicitly defining canalization in this context (i.e., NuRE varies among populations but this variation is genetically determined and shows little plasticity to salinity). This will help readers less familiar with evolutionary biology terminology.

      We thank the reviewer for this helpful suggestion. We agree that canalization should be defined when it is first introduced. We have therefore added a sentence to the Introduction, immediately after introducing the evolutionary canalization hypothesis, to clarify how this concept is used in our study.

      Introduction, fourth paragraph.

      “Here, canalization denotes genetically based differences in NuRE among populations, coupled with limited phenotypic plasticity of this trait under the short-term salinity treatment.”

      (R1-R2) Discuss potassium more deeply. In the Discussion (section 4.2 or 4.4), speculate on why K resorption lacks concentration control. Does Na<sup>+</sup> accumulation functionally substitute for K in osmotic adjustment, thereby decoupling resorption from green leaf K concentration?

      We appreciate this suggestion and have expanded Discussion 4.2. We now propose partial functional substitution of K<sup>+</sup> by accumulated Na<sup>+</sup> during osmotic adjustment as one possible explanation for the weak coupling between green-leaf K concentration and K resorption. We explicitly frame this explanation as tentative because it was not directly tested in the present experiment.

      Discussion 4.2, second paragraph.

      “One possible explanation is the partial functional substitution of K<sup>+</sup> by Na<sup>+</sup> during osmotic adjustment. Na<sup>+</sup> can replace part of the nonspecific vacuolar osmotic function of K<sup>+</sup> in plants and has been reported to become a major osmoticum in P. australis from higher-salinity habitats (Wakeel et al., 2011; Zhao et al., 1999). Increased Na<sup>+</sup> accumulation may therefore reduce reliance on K<sup>+</sup> for osmotic adjustment, potentially contributing to the weak relationship between green-leaf K concentration and K resorption efficiency observed here.”

      Wakeel, A., Farooq, M., Qadir, M., & Schubert, S. (2011). Potassium substitution by sodium in plants. Critical Reviews in Plant Sciences, 30(4), 401–413. https://doi.org/10.1080/07352689.2011.587728

      Zhao, K. F., Feng, L. T., & Zhang, S. Q. (1999). Study on the salinity-adaptation physiology in different ecotypes of Phragmites australis in the Yellow River Delta of China: Osmotica and their contribution to the osmotic adjustment. Estuarine, Coastal and Shelf Science, 49(Supplement 1), 37–42. https://doi.org/10.1016/S0272-7714(99)80006-7

      (R1-R3) Address the potential confounding of ecotype and phylogeography. The dual classification (Figure 3 vs 4) is elegant. However, Chinese ecotypes (freshwater, coastal, inland saltmarsh) may be partially confounded with phylogeographic lineages. A brief sentence explaining how the analytical approach (separate models, random effects) helps separate deep evolutionary history from recent local adaptation would strengthen the interpretation.

      We agree. This concern is addressed in detail under R1-P2. We clarified why phylogeographic group and ecotype were fitted in separate models, what the genotype random effect accounts for, and why the ecotype results cannot be interpreted independently of phylogeographic history.

      See R1-P2, Locations C1 and C2.

      (R1-R4) In section 2.1, specify whether the soil mixture ratio (2 soil: 1 peat moss: 1 river sand) is by volume or by mass.

      We thank the reviewer for identifying this ambiguity. The ratio was based on volume, and we have revised Methods 2.1 accordingly.

      “The plants were planted in barrels (total volume 25 L; top diameter 32.5 cm, bottom diameter 28.4 cm, height 38.5 cm) containing 20 L of a substrate composed of soil, peat moss, and river sand in a 2:1:1 volume ratio (Figure S1).”

      (R1-R5) In the Materials and Methods section, the element potassium (K) is described twice. Specifically, in line 176, the sentence “Besides C, N and P, other eight elements (K, Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” should be revised to “Besides C, N, P and K, other seven elements (Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” to avoid redundancy.

      We thank the reviewer for identifying this redundancy. We have corrected the sentence in Methods 2.2.

      “Besides C, N, P and K, seven additional elements (Cu, Zn, Fe, Mn, Mg, Si, and Na) were quantified in leaf tissues.”

      (R1-R6) Citations should be formatted and ordered alphabetically or by year of publication.

      We thank the reviewer for pointing out this issue. We standardized multi-reference citations, reordered the reference list alphabetically by first-author surname, and verified correspondence between in-text citations and reference entries.

      Full-manuscript citation and reference-list audit.

      Reviewer #2 (Recommendations for the authors):

      (R2-R1) In the discussion of the “inverted” nutrient limitation, elaborate on why P and K are recovered more relative to N. Consider whether this reflects a strategy to maintain an optimal N: P: K ratio or arises from higher costs or lower availability of N.

      We agree. This issue is addressed in detail under R2-P3, where we explain the meaning of slopes greater than 1 and present the possible functional mechanisms as hypotheses rather than established explanations.

      See R2-P3, Location C.

      (R2-R2) Expand Table 1 to include confidence intervals or p-values for individual effects. The extremely low R<sup>2</sup> for concentration and stoichiometry on P resorption should be more explicitly noted as evidence for the dominance of latitude.

      We agree and have expanded Table 1 by adding significance symbols (*) to the individual R<sup>2</sup> values. Significance was assessed for the corresponding fixed effects in the full linear mixed-effects models using Type III tests with Satterthwaite’s approximation for degrees of freedom. For P resorption, latitude had the largest individual contribution and was significant (R<sup>2</sup> = 0.057, p = 0.005), whereas the contributions of green-leaf P concentration and stoichiometry were very small and nonsignificant (R<sup>2</sup> = 0.006 and < 0.001, respectively). We revised the Results to emphasize this contrast while acknowledging that latitude explained only a modest proportion of the total variation.

      See R2-P3, Locations A and B, for the corresponding Results and Discussion revisions.

      (R2-R3) Standardize the abbreviation to “NuRE” throughout, correcting the use of “NRE” in the introduction.

      We thank the reviewer for identifying this inconsistency. We standardized the abbreviation to NuRE throughout the manuscript.

      (R2-R4) Acknowledge in the discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response, justifying future dose-response experiments.

      We agree. This limitation is addressed under R2-P2, where we state that a single 10-ppt treatment cannot exclude plastic responses at higher salinity or along a broader dose-response gradient.

      See R2-P2, Location A.

      (R2-R5) In the Results section, explicitly state that the ecotype analysis in Figure 4 is based only on Chinese populations, not all 110 genotypes, because ecotype classification was not available for non-Chinese populations.

      We agree. This clarification is provided under R2-P2, where Results 3.2 is revised to state that the Figure 4 ecotype analysis includes only Chinese populations.

      See R2-P2, Location B.

      (R2-R6) Consider adding a supplementary figure comparing raw and carbon-standardized NuRE values to show whether the correction altered main conclusions.

      We agree and have added Supplementary Figure S7 comparing raw and carbon-standardized NuRE. The two estimates were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999), and analyses using raw NuRE retained the same conclusions for phylogeographic group, ecotype, salinity, their interactions, and latitude. Thus, carbon standardization slightly shifted the absolute values without altering the main conclusions.

      Methods 2.3.

      “Raw NuRE was calculated without carbon standardization and compared with carbon-standardized NuRE using Pearson correlations; the main linear mixed-effects analyses were also repeated using raw NuRE.”

      Results 3.4, inserted after the existing paragraph reporting the latitude effects and referring to Figure 6 and Table 1.

      “Raw and carbon-standardized NuRE were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999; Figure S7), and analyses using raw NuRE did not alter the conclusions for phylogeographic group, ecotype, salinity, their interactions, or latitude.”

      Supplementary Figure S7 legend

      “Figure S7 Comparison of raw and carbon-standardized nutrient resorption efficiency (NuRE) for N, P, and K. Points represent genotype-by-treatment observations (n = 206), coloured by treatment. Grey dashed lines indicate the 1:1 relationship, and black lines show ordinary least-squares fits. Pearson’s r and the mean standardized-minus-raw difference (Δmean, percentage points) are shown.”

      (R2-R7) In Figure 2b, check the y-axis label; the text reports mg/kg, but the axis shows g/kg, which needs correction.

      We thank the reviewer for identifying this unit discrepancy. We corrected the y-axis label in Figure 2b from g/kg to mg/kg so that it matches the units reported in the text.

      (R2-R8) In the Discussion, rephrase the sentence “This indicates that NuRE is a conservative trait…” to avoid repetition with the Results summary, for example, “This finding highlights the conservative nature of NuRE.”

      We thank the reviewer for this helpful wording suggestion. We have rephrased the sentence to avoid repetition.

    1. eLife Assessment

      Dohi et al. asked what role the dorsal hippocampus and medial prefrontal cortex play during different temporal epochs of a memory-guided navigation task, a long-standing question for neuroscientists studying hippocampal-prefrontal contributions to working memory. This useful study used a delayed, cue-guided T-maze task in mice and reported impaired choice accuracy when silencing occurred early in the central-arm run but not during the delay. However, the evidence for the study's claim is incomplete in its current form: the silencing windows are not duration-matched across epochs, key negative findings rest on three to four mice without power analysis or reported effect sizes, no non-mnemonic control task distinguishes disrupted memory-guided behavior from a general action-selection deficit, and no neural recordings accompany the causal manipulations to verify the manipulations' assumed mechanistic effects. The reported perseveration also reflects an increased directional bias rather than repetition of the previous choice, and the language should be revised accordingly.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors trained mice to perform a memory-guided navigation task, in which they must navigate to a previously cued arm after a delay period. They optogenetically inhibited the dorsal hippocampus or mPFC (targeting PL) in different task epochs. They found that for both regions, inactivating at the beginning of the navigation epoch impaired performance and induced mice to revert to habitual side biases. Inactivating during other epochs, including a delay period before the navigation phase, had little or no impact on behavior. The relationship between trial duration and behavioral performance was differentially impacted by hippocampal and mPFC inactivation, suggesting that the nature of the deficits was somewhat different.

      Strengths:

      The effects of perturbations are robust across animals and generally convincing. The lack of effect at some task epochs serves as a nice internal control. The finding that hippocampus and mPFC inactivation produced subtly different effects is interesting.

      Weaknesses:

      The simplicity of the behavior makes it difficult to resolve how exactly the hippocampus and mPFC contribute to working memory. Also, the language does not always reflect the trends in the data: the authors claim that optogenetic perturbations cause mice to repeat previous choices, but the data show that perturbations increase the likelihood of choosing a preferred side (which is left for most mice). A side bias is not the same as choice repetition. This has implications for interpreting the nature of the behavioral effects.

    3. Reviewer #2 (Public review):

      Summary:

      The study uses transient optogenetic silencing of the dorsal hippocampus or prefrontal cortex in mice using a delayed response task in a T-maze, with a temporal delay of 1s after a visual cue, followed by a central stem run period before choice execution. Silencing of either the dorsal hippocampus or the dorsal prefrontal cortex is executed for varying time periods during the stem/ central arm running epoch or the 1s temporal delay epoch by targeting either PV+ or Dlx interneurons in a block-wise or random-trial design. The main result reported is that silencing of either region during long periods of stem running impaired choice behavior, and silencing during the temporal delay period did not have an effect.

      Strengths:

      The major strength of the study is using the optogenetic silencing strategy to target different temporal periods of the task.

      Weaknesses:

      (1) A major weakness of the study is the lack of a balanced design with equal time periods of silencing during the stem running period and temporal delay period in many of the animals, which precludes any conclusion about distinct functional roles of the regions during these two phases of the task. The central question of this study is not new, with many previous studies investigating distinct and overlapping roles of hippocampus and prefrontal cortex in spatial working memory tasks and memory-guided navigation, using inactivation of one or both regions, crossed inactivation approaches, as well as targeting direct and indirect connections between the regions (PMIDs: 20074655, 9030646, 17045348, 30179661, 27511010, 10491611, 26017312, etc.), in addition to several physiology studies. A key extension for the current study would have been to show a distinction between roles in the temporal delay period after the cue and the stem-running working memory period. However, the inactivation period during stem running shows effects only for long inactivation periods, 2s initial periods and later ~1.6s periods (run after 0.8s, >2/3 total running periods), whereas the temporal delay period inactivation is 1s for the large majority of animals, which is a clear mismatch in inactivation periods, obviating this conclusion of distinction.

      (2) It is not clarified why such short delay periods were used compared to long periods of ~10s in T-maze spatial alternation tasks with delay, and whether the temporal delay period of 1 sec is strictly distinct from the spatiotemporal delay period during stem running in terms of short-term memory function. The choice of time windows for inactivation needs to be better justified, which currently appears to be rather random (2s initial running period, run after 0.8s, run after 1.6s; for an average reported running period of ~2.4-2.5s). The ideal design clearly would have been to use a 2s temporal delay period so that the inactivation time in this epoch matched the initial 2s run period. Only a subset of 4 animals were run with a longer temporal delay period, and too with only with hippocampal inactivation (Figure 4d). The main conclusion of distinction between temporal delay and stem running delay periods is therefore not adequately tested for the prefrontal cortex, and the statistics in terms of number of animals for this important control for hippocampal inactivation are also not comparable to the main experiment.

      (3) The mixture of PV-Cre animals (5 animals), Dlx targeting (2 animals), and one WT animal is also suggestive of a fragmented approach, and inactivation efficacy cannot be assumed to be similar for different animals. Importantly, there is no physiological evidence for confirmation of suppression in the optogenetic experiments, even in exemplar animals.

    4. Reviewer #3 (Public review):

      The authors sought to determine when the dorsal hippocampus and mPFC are causally required during a delayed spatial working memory task. Using temporally precise optogenetic silencing in mice performing a delayed cue-guided T-maze task, they tested the effects of transient perturbations during distinct behavioral epochs. Contrary to the common view that these regions are primarily required during the delay period to maintain working memory representations, they report that silencing during the delay had little effect on performance, whereas perturbation during the early phase of central-arm traversal consistently impaired performance. The authors conclude that hippocampal and prefrontal contributions to memory-guided behavior are dynamically engaged during active navigation rather than passive maintenance of information.

      The study has several strengths. The behavioral paradigm is carefully designed to dissociate cue, delay, and movement epochs, allowing temporally specific causal manipulations. The systematic comparison of multiple task epochs represents a major strength and provides compelling evidence that the behavioral effects of perturbation are epoch-dependent. The authors also include several important controls, including stimulation during multiple task phases, a longer-delay condition to dissociate task epoch from elapsed time since cue presentation, and analyses of movement trajectories and perseverative behavior that provide additional insight into the nature of the behavioral deficits. Together, these experiments convincingly demonstrate that transient dorsal hippocampal and mPFC perturbations have markedly different behavioral consequences depending on when they occur within a trial.

      The evidence is generally solid and supports the primary finding that perturbations during early navigation produce larger impairments than perturbations during the delay period. However, some aspects of the broader interpretation are less well supported. Most notably, the study lacks a non-memory control task, such as a visually guided version of the maze, making it difficult to determine whether the observed deficits specifically reflect disruption of memory-guided behavior or more general impairments in action selection, behavioral flexibility, or movement planning. The observed increase in perseverative responding and delayed commitment to a turn are consistent with either interpretation. In addition, several experimental conditions rely on relatively small numbers of animals, limiting confidence in some negative findings, particularly for the longer-delay and later-run manipulations. Finally, while the Discussion proposes that hippocampal-prefrontal circuits become engaged during the transformation of stored information into action, this mechanistic interpretation remains speculative because no neural recordings accompany the causal manipulations.

      Overall, the authors achieve their primary aim of demonstrating that the behavioral consequences of dorsal hippocampal and mPFC silencing depend strongly on task epoch. The data convincingly support the conclusion that these structures are more vulnerable to perturbation during early navigation than during the brief delay period used in this task. The broader conclusion that these findings redefine when hippocampal-prefrontal circuits support working memory should be interpreted more cautiously, as alternative explanations involving action selection or behavioral state remain plausible in the absence of additional control tasks.

      The findings are potentially important because they challenge the common assumption that hippocampal and prefrontal contributions to delayed-response tasks are centered on delay-period maintenance. Instead, the work supports the idea that these circuits may be recruited when remembered information is translated into goal-directed behavior. This framework is broadly consistent with recent distributed models of working memory and provides an interesting perspective that may help reconcile previous studies reporting effects during different task phases. The behavioral paradigm and temporally precise perturbation approach should also be useful for future studies aimed at dissecting the dynamic contributions of hippocampal-prefrontal circuits during memory-guided behavior.

    5. Author response:

      On the eLife Assessment. We appreciate the assessment’s recognition that this study addresses a long-standing question concerning hippocampal and prefrontal contributions to working memory. We think the broader significance lies in constraining how causal manipulations in working-memory tasks are interpreted. Working memory encompasses many processes distributed across a trial and showing that a region is required for a delayed-response task does not establish when its contribution is necessary. The observation that hippocampus and mPFC are required while navigating to a goal, but not during stationary delay, challenges a common assumption and, in our opinion, has implications across the broader field of working-memory research.

      We agree that the original manuscript did not adequately report effect sizes or convey the uncertainty associated with some smaller samples. However, the dataset does include duration-matched stationary and running conditions, including a hippocampal long-delay control matched to early-running stimulation in both duration and elapsed time after cue onset. The new effect-size and within-animal analyses support a robust hippocampal epoch difference, while the corresponding mPFC comparison is less precisely estimated and should be interpreted more cautiously.

      We therefore think the central finding remains well supported: hippocampal function, and potentially mPFC function, is required during the active navigation period of this memory-guided task but not detectably during the stationary delay. This does not establish the specific computation disrupted during running. Neural recordings would certainly provide further insight into the underlying mechanism, but their absence does not detract from the value of the behavioral result itself.

      Reviewer #1 (Public review):

      The simplicity of the behavior makes it difficult to resolve how exactly the hippocampus and mPFC contribute to working memory.

      The task was designed to combine the temporal precision of cue-based delayed-response paradigms with the behavioral richness of freely moving navigation. Few tasks combine a fixed cue-presentation period, an explicit delay, and a subsequent navigation phase involving extended running. This structure creates well-defined behavioral epochs that can be targeted with temporally precise perturbations, allowing us to ask when hippocampal and mPFC contributions are required within an ongoing memory-guided behavior. How these regions contribute is the harder question, and one we are pursuing next. Identifying when perturbations disrupt behavior is an important step toward understanding how these regions support memory-guided navigation

      Also, the language does not always reflect the trends in the data: the authors claim that optogenetic perturbations cause mice to repeat previous choices, but the data show that perturbations increase the likelihood of choosing a preferred side (which is left for most mice). A side bias is not the same as choice repetition. This has implications for interpreting the nature of the behavioral effects.

      We agree that our results do not clearly distinguish a directional bias from a tendency to repeat the previous choice. To examine whether mice consistently favored a particular direction, we compared their side preferences during silencing across sessions. Mice generally favored the same side across silencing conditions, although some switched direction in individual sessions (Author response image 1a). Within sessions, the preferred side was maintained from no-stimulation to stimulation trials in 13 of 19 cases and reversed in six (Author response image 1b). These observations are consistent with a directional preference that can sometimes reverse during stimulation, and cannot be disambiguated from perseveration. In the revision, we will describe the effect as increased directional bias and revise the language concerning choice repetition and perseveration throughout the manuscript.

      Author response image 1.

      Silencing increases directional bias. (a) Fraction of choices made to the right in each session, for every mouse (rows) and each silencing condition (symbols). Open symbols, no-stimulation trials; filled symbols, stimulation trials from the same session; blue and red denote a left or right preference during silencing. Filled symbols falling predominantly on the same side of 0.5 within a row indicate that a mouse generally favored the same direction across silencing conditions, although some mice switched direction. One session per mouse and condition, hippocampal silencing only; T2 and T6 did not perform the 0.8 s condition. (b) The same sessions expressed as signed bias, from no stimulation to silencing. Black lines mark the six sessions in which the preferred side reversed; grey lines the thirteen in which it was maintained.

      Reviewer #2 (Public review):

      A major weakness of the study is the lack of a balanced design with equal time periods of silencing during the stem running period and temporal delay period in many of the animals, which precludes any conclusion about distinct functional roles of the regions during these two phases of the task. The main conclusion of distinction between temporal delay and stem running delay periods is therefore not adequately tested for the prefrontal cortex, and the statistics in terms of number of animals for this important control for hippocampal inactivation are also not comparable to the main experiment.

      We agree that matching stimulation duration is an essential control and recognize that the relevant comparisons and statistics were not sufficiently clear in the original manuscript. Four hippocampal conditions used closely matched stimulation durations of 2 s: Cue+Delay, long delay, early running, and running with a 0.8 s onset (Author response image 2a). Crucially, the long-delay control matched both stimulation duration and elapsed time after cue onset to the early-run condition, while the mouse remained stationary.

      Author response image 2b–c shows the estimated impairment and its 95% confidence interval for each condition. In the hippocampal experiments, the duration-matched stationary conditions showed effects close to zero, whereas the running conditions showed large impairments. Despite the smaller sample, the upper confidence limit for the long-delay impairment was approximately 10 percentage points, substantially below the observed early-running impairment. These estimates establish that despite the smaller sample, the data support a lack of effect compared to early running.

      Author response image 2.

      Duration-matched stimulation produces different behavioral effects across task epochs. (a) Stimulation timing relative to cue onset. Numbers within bars indicate calculated median stimulation duration in seconds; black ticks indicate door opening. Bold labels identify conditions with 2 s stimulation. (b-c) Mean impairment in choice accuracy for hippocampal and mPFC manipulations. Impairment is accuracy during baseline minus accuracy with stimulation, in percentage points. Error bars show 95% confidence intervals across animals; numbers indicate mice. Open circles denote single-animal observations, and arrows indicate confidence intervals extending beyond the plotted range.

      For mPFC, the duration-matched Cue+Delay condition likewise showed an effect close to zero, whereas early-running stimulation produced substantial impairment. However, the long-delay condition included only one mouse. The later-running effects in both regions were also less precisely estimated. We will distinguish these limitations from the more informative stationary-condition results.

      To directly test whether the duration-matched effects differed across epochs, we next compared impairment within the same mice, including only animals tested in both conditions (Author response image 3). For the hippocampus, every mouse showed greater impairment during early running than during either duration-matched stationary condition, and the confidence intervals for both paired differences excluded zero. These within-animal comparisons support an epoch-dependent effect that cannot be explained by stimulation duration alone. The mPFC comparison showed the same direction of effect, although the confidence interval for the Cue+Delay versus running difference narrowly included zero. We will therefore distinguish the stronger evidence for the hippocampal epoch difference from the more limited evidence for mPFC.

      Author response image 3.

      Within-animal comparisons of duration-matched stimulation effects. (a–b) Impairment during Cue+Delay or long-delay stimulation compared with early-running stimulation for hippocampal (a) and mPFC (b) manipulations. Points represent individual mice, lines connect observations from the same mouse, and black bars indicate mean. Annotations report the mean paired difference in impairment (running minus stationary), and its 95% confidence interval calculated using the t distribution. Positive differences indicate greater impairment during running. The single-mouse mPFC long-delay comparison is descriptive. Blue indicates stationary epochs and orange indicates running.

      The central question of this study is not new, with many previous studies investigating distinct and overlapping roles of hippocampus and prefrontal cortex in spatial working memory tasks and memory-guided navigation, using inactivation of one or both regions, crossed inactivation approaches, as well as targeting direct and indirect connections between the regions (PMIDs: 20074655, 9030646, 17045348, 30179661, 27511010, 10491611, 26017312, etc.), in addition to several physiology studies.

      We agree that hippocampal and prefrontal contributions to spatial working memory have been extensively studied. That’s precisely why we find these results impactful when placed in the rich context of the field. The requirement for these regions in delayed working memory tasks has often been interpreted in terms of their contributions during the delay period. However, a requirement during a delay-based task does not itself demonstrate a requirement during the delay. Previous manipulations have not isolated delay periods of waiting from the subsequent navigation within a trial.

      Our experiments extend this work by separately targeting cue presentation, the delay, and different portions of navigation within the same task. This allows us to test whether the behavioral consequences of perturbation depend on the particular epoch in which it occurs. This distinction is important: knowing that a region is required for a memory-guided task does not establish when its contribution is needed. Identifying those periods constrains how we interpret the deficits produced by longer-lasting inactivation. We believe that this is an important result that should be considered when interpreting these broader findings. We will ensure that the appropriate literature and discussion are included in the revision.

      It is not clarified why such short delay periods were used compared to long periods of ~10s in T-maze spatial alternation tasks with delay, and whether the temporal delay period of 1 sec is strictly distinct from the spatiotemporal delay period during stem running in terms of short-term memory function.

      The 1 s delay was chosen to maintain reliable task performance, as some mice could not perform the task with longer delays. Indeed, one reason the longer-delay condition includes fewer animals is that two mice could not perform reliably with the 3 s delay (one additional mouse was not tested in this condition). Delays on this timescale have also been used in rodent cued delayed-response tasks, including a 0.5 s delay in Kopec et al. (2015), a 1.3 s delay in Guo et al. (2014), and a 1.2 s delay in Inagaki et al. (2019). Like these tasks, our paradigm requires mice to remember an externally presented cue specifying the upcoming response, rather than their own previous arm choice as in spatial alternation. It therefore combines spatial navigation with a cued delayed-response requirement, and the delay durations tolerated in alternation tasks are not necessarily directly comparable. We will clarify this rationale in the manuscript.

      We agree that the stationary delay and the subsequent run both require retention of information after cue offset. Our experiments distinguish these behavioral epochs, but do not establish that they involve separate short-term memory processes. The different effects of perturbation suggest that the contribution of these regions changes as the animal moves from waiting to navigating. Determining what accounts for this change is an important direction for future work.

      The choice of time windows for inactivation needs to be better justified, which currently appears to be rather random (2s initial running period, run after 0.8s, run after 1.6s; for an average reported running period of ~2.4-2.5s).

      We sought to target different portions of the central-arm run. Given the typical traversal time of approximately 2.4–2.5 s, stimulation onsets at 0, 0.8, and 1.6 s sampled the beginning, middle, and later portions of the run. Our setup allowed precise control of stimulation timing, and, as shown in Figure 4b, these onset times correspond approximately to the start, middle, and end of the central arm. Each condition used a nominal 2 s stimulation window, truncated if the mouse reached the choice point sooner. The windows therefore overlap, with the later-onset condition generally producing shorter stimulation. We will clarify this rationale and the distinction between stimulation onset and duration in the revised manuscript.

      The mixture of PV-Cre animals (5 animals), Dlx targeting (2 animals), and one WT animal is also suggestive of a fragmented approach, and inactivation efficacy cannot be assumed to be similar for different animals. Importantly, there is no physiological evidence for confirmation of suppression in the optogenetic experiments, even in exemplar animals.

      The Dlx animals were included to improve regional specificity through local viral expression and to test whether the behavioral effects were consistent across targeting approaches. Both approaches produced comparable impairments during running, supporting their inclusion in the same analysis. We will clarify this rationale in the manuscript.

      Optogenetic activation of inhibitory interneurons is an established approach for suppressing local principal-cell activity, with physiological validation in previous studies (Guo et al., 2014; Li et al., 2019, Zutshi et al., 2022). In our experiments, running-period stimulation produced robust behavioral impairments that were consistent across animals and targeting approaches. Furthermore, comparisons across epochs were performed within animals, using the same preparation and stimulation parameters. Differences in efficacy between animals therefore cannot readily account for the observed epoch dependence. Although direct recordings would establish the magnitude and spatial extent of suppression in our preparation, the central behavioral finding is supported by these within-animal comparisons.

      Reviewer #3 (Public review):

      The findings are potentially important because they challenge the common assumption that hippocampal and prefrontal contributions to delayed-response tasks are centered on delay-period maintenance.

      We thank the reviewer for describing our findings as “potentially important” and for highlighting their implications for how hippocampal and prefrontal contributions to delayed-response tasks are understood. We appreciate the constructive suggestions and address the public comments below.

      Most notably, the study lacks a non-memory control task, such as a visually guided version of the maze, making it difficult to determine whether the observed deficits specifically reflect disruption of memory-guided behavior or more general impairments in action selection, behavioral flexibility, or movement planning. The observed increase in perseverative responding and delayed commitment to a turn are consistent with either interpretation.

      We agree that leaving the cue on throughout the trial would provide an important control for distinguishing memory-specific effects from broader effects on action selection or movement planning. The senior author is currently setting up a new laboratory, so implementing this control may take some time. We hope to include it in the revised manuscript.

      The running-period deficit indicates a disruption of processes that enable the animal to act on a remembered cue. Whether this reflects disruption of memory itself, movement planning, or another component of translating the cue into a choice requires further clarification but does not detract from the observed dependence on task epoch. Uncertainty about the mechanism of the running-period deficit also does not change the observation that the same manipulation produced no detectable impairment during the stationary delay, when the cue was absent and still had to be remembered. This will be clearly discussed in the revision.

      In addition, several experimental conditions rely on relatively small numbers of animals, limiting confidence in some negative findings, particularly for the longer-delay and later-run manipulations.

      We agree that small sample sizes limit the interpretation of some negative findings. We now report animal-level effect estimates and 95% confidence intervals for each condition (Author response image 2b–c; Author response table 1).

      Of the nine conditions with no detectable impairment and more than one mouse, seven had confidence intervals that excluded effects as large as the observed mean early-running impairment in the same region. This included the hippocampal long-delay condition, despite its smaller sample. These results argue against similarly large impairments in these conditions, although smaller effects remain possible. The two later-run conditions remained too uncertain to exclude such impairments. These and the two single-animal conditions are marked in Author response table 1 and will be interpreted cautiously.

      Author response table 1.

      Animal-level impairment estimates and uncertainty

      Finally, while the Discussion proposes that hippocampal-prefrontal circuits become engaged during the transformation of stored information into action, this mechanistic interpretation remains speculative because no neural recordings accompany the causal manipulations.

      We will clarify in the Discussion that the proposed transformation of stored information into action remains speculative. Nevertheless, we believe this is an exciting possibility raised by our findings that warrants further investigation. Neural recordings would help test this interpretation but are beyond the scope of the current paper. We plan to explore this question in future work.

      References

      Guo ZV, Li N, Huber D, Ophir E, Gutnisky D, Ting JT, Feng G, Svoboda K (2014). Flow of cortical activity underlying a tactile decision in mice. Neuron 81:179–94. doi:10.1016/j.neuron.2013.10.020. PMID 24361077.

      Kopec CD, Erlich JC, Brunton BW, Deisseroth K, Brody CD (2015). Cortical and subcortical contributions to short-term memory for orienting movements. Neuron 88:367–77. doi:10.1016/j.neuron.2015.08.033. PMID 26439529.

      Inagaki HK, Fontolan L, Romani S, Svoboda K (2019). Discrete attractor dynamics underlies persistent activity in the frontal cortex. Nature 566:212–217. doi:10.1038/s41586-019-0919-7. PMID 30728503.

      Li N, Chen S, Guo ZV, Chen H, Huo Y, Inagaki HK, Chen G, Davis C, Hansel D, Guo C, Svoboda K (2019). Spatiotemporal constraints on optogenetic inactivation in cortical circuits. eLife 8:e48622. doi:10.7554/eLife.48622. PMID 31736463.

      Zutshi I, Valero M, Fernández-Ruiz A, Buzsáki G (2022). Extrinsic control and intrinsic computation in the hippocampal CA1 circuit. Neuron 110:658–673.e5. doi:10.1016/j.neuron.2021.11.015. PMID 34890566.

    1. eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRSIPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. The analysis of reciprocal daughter-cell pairs provides compelling evidence for SCE events, providing evidence consistent with CDK1-TTF2-TRAIP mediated cell-cycle regulated CMG helicase disassembly and fork cleavage at unreplicated regions.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9-induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs.

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

    3. Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed large-scale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations.

    4. Reviewer #3 (Public review):

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are « genetically silent ». Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly « permissive » for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRISPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole-genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. However, the evidence supporting the proposed involvement of under-replicated region/replication termination-zone resolution and TRAIP/URR-like pathways is currently incomplete and could be strengthened with an increased number of reciprocal daughter-cell pairs and by genetic or molecular perturbation, or alternatively, this can be addressed by changing the discussion.

      We appreciate the editor’s and the reviewers’ recognition of Cas9-induced SCE as an important previously invisible repair outcome and of RDCP analysis as a notable feature of the study. As proposed, we incorporated the two recent Science studies and now tone down the Discussion of our RDCP observations as consistent with, rather than definitive evidence for, the TRAIP-dependent pathway. We also detail how our single-cell genomic observations complement and extend these two studies in terms of the biological significance of the CDK1-TTF2-TRAIP axis.

      Specifically, these changes are in:

      Discussion. We changed

      “Two recent studies revealed how the CDK1-TTF2-TRAIP axis is cell-cycle regulated to trigger mitotic CMG helicase disassembly and fork cleavage: one study showed a two-fold SCE reduction in mouse ES cells (Fujisawa and Labib, 2026), while the other showed that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions (Can et al., 2026). Our observation of the "WWC-or-WCC/deletion pair" signature in wild-type cells provides, to our knowledge, the first genetic evidence of linking a deletion with SCE and revealing both W and C unreplicated template strands present in the reciprocal daughter cell, consistent with this mechanism at single-cell genomic resolution (illustrated in Fig.3), although this is limited by the observation of only one RDCP.”

      While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells.

      Fig.3. Title and legends. We changed

      “Haplotype-aware analysis of observed RDCP (Pair 4, chr1) shows SV patterns at the SCE junctions consistent with the predicted RDCP signature of SCE mediated by URRs or replication termination zones (green shaded area), although the lagging strands, rather than the leading strands, must be resolved to generate these mitotic breaks.”

      Reviewer #1 (Public review):

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest. The evidence for structural complexity associated with some induced SCEs is intriguing, but the mechanistic interpretation should either be tested directly or presented more cautiously.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs. Given that potential, the current manuscript would benefit greatly from any experiments characterizing this sub-population: are these cells in a particular cell cycle state, experiencing changes in gene expression, or do they have other unique biological properties?

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

      Thank you very much for this assessment.

      Weaknesses:

      The number of informative RDCPs is limited, and the mechanistic interpretation of the "WWC-orWCC/deletion" signature is more suggestive than definitive. In particular, the manuscript invokes (even though only in the Discussion section) URR or replication-termination-zone resolution and discusses TRAIP-dependent CMG unloading, nuclease cleavage, and polymerase theta-mediated joining, but these pathway components are not directly tested herein. A more conservative conclusion that some Cas9-associated SCEs coincide with structural alterations is more appropriate, particularly in the Discussion and Conclusion. For example, the statement that this work provides "direct genetic evidence" for a URR-type mechanism is overstated unless supported by additional experiments or a more extensive analysis of alternative models. Similarly, while the authors explain the limitations of acute Cas9 disruption of LIG3, LIG4, XRCC1, and XRCC4, the manuscript should clarify what biological questions this experiment can and cannot answer.

      Please see response to the eLife Assessment as this is a common point raised by multiple reviewers.

      Additionally, we clarified what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) by adding the following text in the “Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell” section:

      “One additional complication is that, although Cas9 RNP achieves >90% knockout efficiency in bulk assays and is therefore used as a substitute for siRNA, sorting BrdU-labeled cells in the subsequent G1 enriches for cells that escaped frameshift editing, particularly for essential genes; thus, 100% knockout in 90% of cells is not equivalent to 90% knockdown in every cell, representing a unique challenge for single-cell assays.”

      Reviewer #1 (Recommendations for the authors):

      (1) Temper the mechanistic claims about URR/TRAIP-type resolution.

      The RDCP data support the conclusion that some Cas9-associated SCEs are accompanied by structural alterations and may arise through non-classical mechanisms. However, claims about TRAIPdependent CMG unloading, URR resolution, or polymerase theta-mediated joining should be framed as a model unless directly tested.

      We cited mechanistic dissection of the CDK1-TTF-2-TRAIP axis, which was published since the review of the paper. While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells but qualified that this observation is in only one RDCP. We further tempered our claims by changing “direct evidence” to “consistent with” as we did not perturb genes involved in these processes.

      (2) Clarify the impact of the small RDCP sample size.

      The manuscript would be strengthened by explicitly stating how many total RDCPs were analyzed, how many SCE events were informative, and how much confidence can be placed in the estimated fraction of SCEs associated with SVs. A short table summarizing RDCP counts, SCE counts, copy-neutral events, and SV-associated events would be helpful.

      A total of 15 RDCPs were recovered from close to 4,000 single cells analyzed across all conditions. We added Tab.S3 detailing SCEs in all 15 RDCPs in addition to SCE and SV breakdown in Tab.S2 (originally Tab.S1). The new Tab.S3 is cited in the “RDCP analysis reveals large-scale SVs on chromosomes with induced SCE, as well as structural alterations at Cas9-induced SCE junctions” section.

      (3) Provide more detail on the "rescued" SCE calls.

      Because the central conclusions rely on SCE detection, the criteria for breakpoint R-based calls versus rescued calls should be explained clearly in the main text or methods. It would be useful to know how sensitive the main conclusions are to the inclusion or exclusion of rescued calls. Is this laid out in greater detail in an additional manuscript?

      We previously included an “On-target SCE identification” section in the Methods, where we described in detail the rescue of on-target SCEs missed by the initial breakpoint R calls. We also depicted calls and calls+rescues for all the conditions in Fig.S1B.

      We agree with the reviewer and now expanded the description of rescued SCE calls in the main text (under the “A single Cas9 DSB induces potent local SCE” section). In brief, 50-92% of SCEs (typically >70%) across the four single-targeting sites were directly called rather than rescued, with the exception of LIG4, where only 30% were direct calls. This is because LIG4 is located only 6 Mb from the telomere and is therefore particularly prone to missed breakpointR calls in low-coverage cells. On-target rescue at individual sites is self-contained in this manuscript because our lab primarily focuses on spontaneous SCE, for which there are no expected SCE sites. However, the rescue methodology was previously implemented in the original development of sci-L3-Strand-seq to identify SCEs at centromeres.

      (4) Clarify the biological interpretation of the repetitive-target enrichment.

      The high-SCE subset analysis is interesting, but the manuscript should explain whether these cells have evidence of higher RNP uptake, altered cell-cycle state, greater DNA damage, or lower sequencing quality. If these possibilities cannot be distinguished, the text should state this clearly.

      We agree with the reviewer. The high SCE subset does not have lower sequencing quality by coverage or background (0.3% coverage for high-SCE vs. 0.28% coverage overall, and 3% background for both high-SCE and overall). However, our current data do not allow us to distinguish among biological explanations for the elevated SCE. The original manuscript acknowledged this limitation (“Whether this reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment, remains to be determined.”). To make this limitation more explicit and to address the possibility of sequencing quality raised by the reviewer, we have revised the text as follows: “The high-SCE subset did not show evidence of lower sequencing quality, based on either sequencing coverage (p=0.17) or background SCE levels (p=0.13). However, we cannot distinguish whether the elevated on-target SCE reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment.”

      This question may be better explored by future co-assays with sci-L3-Strand-seq; currently we cannot enrich for cells with high SCEs to characterize the molecular features of this subset of the cells using other omics approaches.

      (5) Reconsider the framing of the DNA repair gene targeting experiment.

      The current data do not strongly test whether LIG3, LIG4, XRCC1, or XRCC4 regulate Cas9-induced SCE, because functional protein loss is delayed and essential-gene targeting introduces selection. This section may be better framed as a negative/control observation rather than as a pathway analysis.

      We agree and please refer to Public Reviews for a single-cell assay-specific explanation.

      (6) Consider including some additional control experiments, for example, Cas9 without sgRNA, nontargeting sgRNA, or mock-transfected cells to make sure that some phenotypes (for example, cell-cycle arrest) directly result from DNA cleavage rather than from the transfection procedure.

      We thank the reviewer for this suggestion. We have carefully considered these additional controls but have chosen not to add further experiments. Our existing Cas9 nickase experiments provide a control that directly addresses whether the observed arrest is attributable to DSB formation rather than RNP delivery/transfection. Both the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9, but neither produced the cell-cycle arrest observed following DSB induction by wild-type Cas9. Thus, these experiments control for Cas9 RNP delivery while altering the nature of the DNA lesion and support the interpretation that the observed arrest is associated specifically with Cas9-induced DSBs rather than the transfection procedure itself.

      (7) Figure 1: the fonts should be increased. The majority of the labels are impossible to read in a printed copy of this manuscript.

      We thank the review for pointing this out. We enlarged Fig.1 fonts.

      (8) Figure 1C. The pileup plots should be described and interpreted in a clear way. In its present form, it is unclear how the interpretations and conclusions are made.

      We added explanation of the pileup analysis immediately following mentioning the Fig.1C pileup: “We next examined … SCEs using genome-wide pileup analysis (Fig.1C, Fig.S1B), in which we plot the total number of SCEs detected across all single cells within each 1 Mb window.”

      Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed largescale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Thank you very much for this assessment.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations. The language and logic in the paper can be improved, and some of the claims seem incorrect. For example, the abstract reads "A single Cas9 cut at a unique genomic locus led to strong local enrichment of SCE at the break site, reaching up to 41% in the same cell cycle and 17% in the subsequent division, indicating that DSB repair frequently engages non-local inter-sister repair." The evidence that only a single Cas9 cut was made is lacking (see my earlier comment); it is not clear how local enrichment or non-local inter-sister repair are defined.

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We also agree that novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome these limitations, perhaps by using vfCas9 but more importantly, if new approaches to turn off Cas9 are developed. We have added a brief discussion on this limitation (in the Limitation section) and the resulting constraints on extrapolating mechanisms of DNA instability and repair from the observed genomic rearrangements. We thank the reviewer for pointing out the distinction between “a single Cas9 cut” vs. "Cas9 targeting of a single genomic locus." We went through the manuscript and revised where cutting only once was implied. We also explicitly acknowledge the possibility of multiple rounds of cutting at the same sites.

      We thank the reviewer for pointing out that “non-local repair” is a non-standard term. We use it operationally to distinguish repair confined to the broken chromatid (e.g., fill-in synthesis or end joining in cis) from repair involving exchange between sister chromatids. We have added a schematic (Fig. S1A) illustrating this distinction and cited this figure immediately before where we operationally defined SCE as a “reciprocal strand switch between sister chromatids, without implying a single mechanistic pathway.” This distinction is important because, particularly for two-ended Cas9 DSBs, an SCE-like outcome could potentially arise through either HR-mediated crossover or NHEJ of DNA ends across sister chromatids; the latter may involve different genetic requirements from classical NHEJ at least in end-tethering. We have revised the manuscript to define “non-local repair” explicitly at its first use.

      Reviewer #2 (Recommendations for the authors):

      References to relevant earlier studies using Strand-seq to study SCEs are missing (PMID: 27185886 and PMID: 29348659).

      We thank the reviewer for pointing this out. We added these references in the 3rd paragraph of the Introduction where we briefly review Strand-seq methods.

      Reviewer #3 (Public review):

      Summary:

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are « genetically silent ». Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly « permissive » for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

      Strengths:

      This is an interesting paper that molecularly explores sister chromatid exchanges, which represent an important challenge in molecular biology since they are genetically silent.

      Thank you very much for this assessment.

      Weaknesses:

      A complexity of the current paper is that it heavily relies on a recently published paper (Chovanec et al 2026, NAR) describing the powerful but complex technique sci-L3-Strand-seq. Knowledge of this paper is a prerequisite to understanding the current manuscript because no reminder is provided. In addition, the current manuscript presents the use of the sci-L3-Strand-seq technique in the study of SCE after Cas9-induced DSBs, while a companion study is referred to several times for containing results about SCE in XRCC1 KO. At some point, one questions the relevance of splitting the use of sci-L3-Strandseq in different papers instead of making a single integrated one.

      We appreciate this concern. The original sci-L3-Strand-seq study is an extensive methodology paper that establishes and validates various computational framework, whereas the companion study focuses on the genetic regulation of spontaneous SCE. The experimental designs and biological questions of the companion study and the present work are therefore distinct, although we draw on selected results from the companion study where they provide useful comparisons and contrasts between spontaneous and Cas9-induced SCE. The present study addresses a distinct biological question, the response to Cas9-induced DSBs, and we therefore believe that combining these studies would make the resulting manuscript unnecessarily broad and obscure their different biological questions.

      We nevertheless agree that the present manuscript should be understandable without requiring detailed knowledge of either paper. We have therefore added a brief description of the sci-L3-Strandseq approach (3rd paragraph of Introduction, Fig.S1A legends, and Fig.1B legends) and clarified the relevant methodological concepts where they are first introduced. We hope to improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract

      "Identical sisters ": redundant

      "non-local" inter-sister repair: the meaning is not clear. Do the authors refer only to "inter-sister" and therefore "non-local" is redundant, or do they imply something specific by "non-local", in which case it needs to be clarified?

      "237 repetitive targets": at least a slight description of this target is needed. Is it a "random" repeat, a satellite sequence, a sequence related to a transposable element ?...

      We agree with the reviewer that “identical sisters” is technically redundant. However, we have retained “identical” here to emphasize the distinction between sister chromatids vs. homolog, as SCE is sometimes misconstrued as exchange between homologs and as potentially causing loss of heterozygosity. We prefer the slight redundancy here for conceptual clarity.

      We thank the reviewer for pointing out that the meaning of “non-local” was unclear. As discussed in our response to Reviewer #2, we use “non-local repair” operationally to distinguish repair confined to the broken chromatid in cis from repair involving exchange between sister chromatids. We have added a schematic (Fig.S1A) illustrating this distinction and explicitly define the term at its first use in the revised manuscript. Please see our response to Reviewer #2 above for the detailed rationale.

      We thank the reviewer for asking us to clarify the nature of the 237 repetitive targets. The sgRNA targets an Alu sequence and was selected from a larger screen of >20,000 sgRNAs targeting repetitive sequences occurring at >200 genomic sites. In that screen, cellular toxicity did not simply scale with the number of predicted target sites; we therefore selected this sgRNA because its intermediate phenotype allowed us to introduce a large number of programmed DSBs without either minimal perturbation or excessive loss of cells. Thus, the 237-site guide was not an arbitrarily selected Alu-targeting sgRNA. The full repetitive-element screen is beyond the scope of the present study, but we have clarified in the Abstract that these 237 sites are Alu targets and added a brief description of the guide selection in the Methods.

      (2) Introduction:

      "non-local outcome / non-local repair processes": The use of "non-local" is not standard and is obscure for the reader. Specify if it has any meaning or remove it.

      Please see our response above to both Reviewers #2 and #3 regarding our definition and use of “nonlocal repair”.

      The authors mention that replication through a DSB generates four broken ends. However, in case the DSB is reached by one replication fork before the converging one, there are only three broken ends for at least the time required for the converging fork to reach the DSB from the other side. This may influence the repair outcome.

      We agree with the reviewer. If one replication fork encounters the DSB before the converging fork, a transient three-ended intermediate can exist before the second fork reaches the break. This temporal asymmetry could influence repair pathway choice, including engagement of HR, end joining, or BIR-like repair. Our assay captures the resulting SCE outcome but cannot distinguish the order in which replication forks encounter the DSB or the repair pathway engaged at these intermediate stages. We have revised the text (2nd paragraph of the Introduction) to clarify that four broken ends represent the eventual configuration after replication through the DSB, rather than necessarily a simultaneous intermediate.

      (3) Results

      Cell cycle arrest experiment: it seems that a control condition with no Cas9 is missing to conclude better about what looks like a G2-M arrest, but that is not clearly mentioned.

      Please see our response to Reviewer #1, Recommendation 6, regarding additional controls for the cell-cycle arrest experiment. Briefly, the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9 but did not produce the cell-cycle arrest observed following DSB induction, providing an internal control for RNP delivery/transfection and supporting the association of the arrest with Cas9-induced DSBs. We would also like to clarify that the observed cell-cycle arrest is primarily a G1/S, rather than G2/M, basing on the FACS signal (see revised Fig.1B legend). This is consistent with the strong G1/S checkpoint in mammalian cells and the predominantly G1 cell-cycle distribution of BJ-5ta cells.

      Note that the font size in Figure 1 is too small for readability.

      We have enlarged the font sizes throughout Figure 1 to improve readability.

      Figure 1B, D10A and H840A conditions:

      The authors mention that nicks can be converted into DSBs through the passage of the replication fork, but do not see any cell cycle defect in the conditions tested. Is it possible that the absence of effect results from the fact that the analysis is done prior to nicks being converted into DSBs? This remark notably applies to the 237 target sites experiment. It seems that controlling for cell cycle delays for longer times is needed to conclude clearly about this aspect. In case a clear absence of cell cycle delay is observed in the 237 target sites in the Cas9 nicking condition, this would suggest that replication born DSBs behave differently from "classical" two-ended DSBs and do not trigger cell cycle arrest.

      We agree that the timing of nick conversion during replication could contribute to the absence of a detectable cell-cycle delay under the conditions examined. However, extending the duration of Cas9 nickase treatment or labeling would not necessarily resolve this question, because persistent Cas9 activity permits repeated rounds of nicking across successive cell cycles, making it difficult to relate a later cell-cycle phenotype to a defined replication-born lesion. More generally, we believe that the relationship between replication-associated nicks, SCE formation, and cell-cycle progression is better addressed in the context of spontaneous SCE, which is the focus of our companion study. The present study is focused on SCE following programmed Cas9-induced DSBs, and analysis of replication-born nick lesions would require precise temporal control (ideally vfCas9 nickases that can be turned off) of individual nicking events relative to replication, for which an appropriate experimental system is not currently available to us.

      Figure 1C should mention somewhere the genomic location of the four targets to clearly show that they correspond to the four major SCE peaks. In addition, there is no legend for the vertical pink stripes. Finally, it might be wise to keep the same y-axis scale for better comparisons.

      Figure 1 overall: it might be wise to clearly show a no Cas9 condition to clearly set the SCE baseline and show that it is independent of Cas9 induction. As of now, it is not clear whether the non-targeted SCE comes from a specific cleavage of Cas9 or not. Such an aspect could benefit from putting Figure S1C in the main Figure 1. Alternatively, results from Chovanec et al 2026 (NAR) should be better restated because the reader does not necessarily have them in mind.

      We thank the reviewer for these suggestions. The expected Cas9 target positions were already indicated by vertical bars in Fig. 1C; however, we agree that this was not sufficiently clear. We have therefore revised the figure legend to explain that the vertical bars indicate the expected Cas9 target positions. The Chovanec et al. (2026, NAR) study focused entirely on spontaneous SCE, which we simultaneously map here as the background signal, rather than the on-target SCEs induced by Cas9. We hope that explicitly identifying the target locations in the revised legend makes this distinction clear and ensures that prior knowledge of the NAR study is not necessary to interpret Fig. 1C.

      We have retained the individual y-axis scales because the magnitude of SCE enrichment differs substantially among conditions. Using a common y-axis scale would make several of the on-target SCE peaks difficult to visualize.

      The section « Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell » is questionable in the results section for the following reasons:

      (i) The DNA repair genes are used here as target sites for Cas9 cleavage, but are not the object of the study, but may be the object of a companion paper. This aspect is slightly misleading.

      (ii) As first mentioned in this section, there is evidence strongly suggesting that inactivating DNA repair genes will not affect SCE, and this is what the authors observed.

      (iii) As an alternative, one could put the emphasis on the fact that the effect of Cas9-mediated inactivation of DNA repair genes (ie LIG3) starts to be detectable only in the washout condition ie after at least one cell cycle. But in this case, this is addressing the role of DNA repair genes in Cas9-induced SCE, which is not the point of the current paper.

      We thank the reviewer for this comment and agree that the original framing of this section could give the impression that these experiments were intended to test the functions of the targeted DNA repair genes in SCE. This was not our intent. Rather, these genes provided defined genomic target sites for Cas9 cleavage, and the primary purpose of the experiment was to characterize SCE associated with Cas9-induced DSBs at these loci.

      As discussed in our response to the Editor Assessment, there are important limitations to using these experiments to infer the consequences of loss of the targeted proteins, including the delay between Cas9 cleavage and depletion of pre-existing protein and, and particularly in the single-cell assay, selection for cells that escape disruptive editing at essential genes. We have added text to the Results explicitly describing these limitations.

      We therefore agree with the reviewer that the delayed effects observed under the washout condition should not be interpreted here as establishing a role for individual DNA repair genes in Cas9-induced SCE. We have revised the section title to “Cas9 targeting of DNA repair gene loci did not immediately alter overall SCE frequency per cell” to clarify the scope of this experiment and to avoid implying that testing the functions of the targeted DNA repair genes is a major objective of the present study.

      The conditions in Table 1 need to be homogenized and better explained:

      - 237 sites and 237 cuts are used: homogenize?

      We thank the reviewer for spotting this. We revised both to be “237 sites”.

      - May explain better the rationale for putting BrdU simultaneously with Cas9 or after 24 h and a wash.

      For the single-targeting sites, we observed more SCE when BrdU was added simultaneously with the Cas9 for the same 24 hours, compared to adding BrdU in the subsequent division after a wash. Therefore, for the 237 sites, we analyzed both conditions.

      - Typo in the text: 237cuts_24ws_40BrdU instead of 237cuts_24ws_BrdU

      We apologize for the lack of clarity in these labels and have substantially revised the Table 1 legend. In brief, the “40” is not a typo. In the 237 sites experiments, wild-type Cas9 considerably prolonged the cell cycle. Therefore, rather than labeling with BrdU for 24 hours as in the other conditions, we extended BrdU labelling to 40 hours in the last two conditions to allow more cells to progress into the subsequent G1 for successful Strand-seq analysis. We clarified that “237 sites 24 + 16hrs BrdU” refers to the condition in which Cas9 RNP and BrdU were added simultaneously. After 24 hours of Cas9 RNP treatment, BrdU labeling was continued for an additional 16 hours (a total of 40 hours of BrdU). The “237sites 24ws40BrdU” condition is the corresponding washout condition, in which Cas9 RNP was removed after 24 hours and cells were then labeled with BrdU for 40 hours post-washout.

      - 237 cuts: Are some sites more enriched in SCE than others?

      Yes, some sites showed greater SCE enrichment than others. We tested whether this variation correlated with chromatin accessibility but found no significant association. This was not unexpected, as the sgRNAs predominantly target Alu elements.

      - Table 1: There is a difference between 237 sites 24ws24BrdU and 237cuts24ws40BrdU, with a significant enrichment of on-target SCE for the latter condition only. Could the increase in SCE rise even more with longer BrdU exposure? In other words, does the low enrichment in SCE at target sites in the 237 sites experiment result from a non-optimal timing for the analysis?

      Yes, this is possible. We did not systematically test additional treatment or labeling durations. A 24-hour Cas9 RNP treatment is typically used for Cas9 RNP-mediated knockout experiments, and we therefore initially used this duration to assess gene-editing outcomes. For Strand-seq, BrdU labeling is ideally limited to approximately one cell division. Because BJ-5ta cells have an approximately 24-hour cell cycle, extending BrdU labeling substantially beyond 40 hours could allow some cells to undergo a second round of replication and become double-labelled. We therefore did not extend BrdU labeling beyond 40 hours. Thus, the lower enrichment in the 24ws24BrdU condition may in part reflect the timing of the assay.

      - Figure 2 / RDCP analysis:

      Interpretation of this figure relies exclusively on the 2026 NAR paper from the authors. This, at least, should be mentioned to help the reader understand it. Once the legend restates, this figure misses clear identification of the SCE and other genomic rearrangements. For readability, maybe the full genome should be kept for the supplementary data, and only the rearranged chromosomes should be kept in the main figure so that the rearrangements are clearly visible and annotated.

      We thank the reviewer for this suggestion. To make the Strand-seq plots interpretable without relying on our 2026 NAR paper, we have added an explanation of Strand-seq orientation in the third paragraph of the Introduction and in Fig.S1A. We have revised Fig.2 legends to improve readability. We have retained the whole-genome view because Strand-seq data are conventionally presented in this format and it provides important genome-wide context for interpreting the observed events. The rearranged chromosomes and events were annotated in Fig.S3.

      (4) Discussion

      - Most DSB never formed or did not produce SCE: how to understand this better? What would be the argument in favor of one or the other possibility?

      We agree that these are two possible explanations that cannot be distinguished by the current experiment. The absence of an SCE at a targeted site could reflect either inefficient DSB formation or repair of a DSB through a pathway that does not generate an SCE. Distinguishing these possibilities would require direct measurement of cutting efficiency at individual target sites, which was beyond the scope of this study.

      - The conclusion about the effect of the Cas9 nickases needs to be toned down as long as the proper timing for SCE analysis has not been performed (see comment above).

      We agree and have toned down this conclusion by specifying that no significant on-target SCE enrichment was detected under the conditions tested and acknowledging that we cannot exclude SCE formation at other time points (Discuss, first paragraph).

      - As much as possible, avoid the use of non-conventional acronyms like URR.

      We agree and have reduced the use of non-conventional acronyms where possible. We have retained URR (under-replicated region), as the term appears seven times throughout the manuscript, but have ensured that it is clearly defined at first use.

      - The discussion about the RDCP analysis in the second paragraph of the discussion should refer to Figure 3.

      Thank you for pointing this out. We added this reference to Fig.3

    1. eLife Assessment

      This study presents a valuable tool for comparing immune receptor data across multiple samples while properly accounting for statistical uncertainty and receptor similarity. The evidence supporting the tool is solid overall. However, some key concerns regarding potential sequencing artifacts in one of the validation datasets and an unclear strategy for false discovery rate control in the proposed framework remain to be addressed.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors present ClustIRR, a computational tool that analyzes multiple TCR repertoires together, instead of one at a time. It builds one shared similarity graph across all the repertoires, then finds communities (CJs) on that graph that can be compared across samples. Building one shared graph across repertoires, rather than clustering each repertoire on its own, is a real improvement over existing tools. However, there are a few concerns that need to be addressed.

      (1) The Introduction motivates ClustIRR by contrasting it with beta-binomial regression, Fisher's exact test, and scCODA/tascCODA (lines 62-77), but none of these are actually run on the same data in the Results. Could the authors include a direct comparison, e.g., applying a standard beta-binomial or Fisher's exact test to the Dataset 1 CJ occupancy matrix, to show the Bayesian model gives a lower false-positive rate or better-calibrated intervals than the alternatives it's positioned against?

      (2) Prior predictive and posterior predictive checks are both reported as "(data not shown)" (lines 644, 728). Since the manuscript's central claim is rigorous, uncertainty-aware inference, it would help to include these diagnostic plots, along with Rhat and ESS values, in the supplement rather than stating they were checked.

      (3) With 7,505 CJs tested simultaneously for differential occupancy in Dataset 1 alone (Figure 1 legend), what is the expected false discovery rate under the non-overlapping-HDI criterion used throughout? A short discussion of multiple-comparisons correction, or an argument for why it isn't needed under this framework, would strengthen the statistical claims.

      (4) Line 115 states the Dataset 1 joint graph produced 10,301 CJs, of which 3,038 were singletons, leaving 7,263 non-singleton CJs. The Figure 1 legend reports 7,505 CJs used for the β modeling. Could the authors clarify how these two numbers relate - whether some singletons were included in the model, or a filtering step was applied that isn't described in Methods?

      (5) The Methods section states that archival pretreatment tumor tissue was available for four patients (Pt4, Pt32, Pt36, Pt38; line 531), but the T+/T- DCJ analysis in Fig. 3B-C is only shown for Pt4. Was this analysis attempted in the other three patients? Extending it, even partially, would substantially strengthen the claim that contracting DCJs are enriched for tumor-infiltrating TCRs, which is currently based on a single patient.

      (6) Dataset 1 was generated by deliberately stimulating T cells with EBV or MART1 antigen, so recovering EBV/MART1-annotated CJs from VDJdb is closer to a positive control than a blinded validation. Do the authors have, or could they obtain, any independent confirmation (e.g., tetramer data or an orthogonal cohort) for the CJs with large β that lack VDJdb annotation (orange dots, Figures 1B-C)?

    3. Reviewer #2 (Public review):

      Summary:

      The study confronts a major obstacle in repertoire analysis. Given that individual TCR sequences are diverse and sparsely detected across repertoires, identifying sample-specific enrichment of TCRs based on sample-to-sample comparison of exact clonotype sequences can be intractable. This paper attempts to build on insights that sequence-similar TCRs can share antigen recognition, such that aggregating similar sequences derived from multi-sample joint-graph communities (i.e. clusters of tightly connected nodes) could reduce sparsity and boost signal.

      Strengths:

      The study is a well-motivated effort to address a need in the field. The paper takes a unique approach. The core method is modeling community occupancy with a hierarchical Dirichlet-Multinomial model that attempts to account for the high level of overdispersion present in repertoire sampling, a technique that has been previously applied to compositional microbiome data.

      The authors apply this framework to both single-cell paired-chain and bulk single-chain TCR data. They develop a set of examples from public and synthetic data, with the most promising real-world application shown in reanalysis of longitudinal data during treatment of cancer patients with checkpoint inhibitors.

      The manuscript is well structured and cogent. The authors are to be commended for contributing a well-documented, open-source R/Bioconductor package and for providing the underlying analysis datasets in a well-organized repository. In the joint graph construction step, the authors opt to use an existing implementation of the BLAST algorithm on CDR3 sequences, which is a slightly odd choice since it ignores potential contributions of other V-gene germline-encoded CDRs, but the authors also envision that their statistical package could be extended to include community graphs developed with other established TCR clustering tools. This will allow others to potentially explore the utility of Bayesian hierarchical Dirichlet-Multinomial models for differential occupancy analysis of immune receptors under varied clustering criteria.

      The methods described here were demonstrated on relatively small datasets from 2-5 samples, and future work is likely needed to extend the joint graph differential occupancy concept to larger datasets. The authors are transparent about this and some of the other limitations in their current tool, most notably the computational cost of graph construction based on an all-versus-all sequence alignment to construct a joint sequence graph and the challenge of Bayesian parameter estimation as the number of subgraph entities scales with input data size. Since efficient approximate methods exist to find edges between similar text strings, the underlying idea of applying uncertainty-aware statistical inference to subgraph communities is promising, and the paper advances its primary goal.

      Weaknesses:

      The paper proposes the utility of the joint-graph community occupancy framework through three examples. I discuss potential weaknesses apparent in each separate example in turn.

      (1) Weaknesses in Example 1

      A broad weakness of the first results section ("Detecting EBV- and MART1-antigen reactive T cell communities from single cell datasets") is its reliance on a single vendor-generated dataset generated by the company ParseBio with no published experimental protocols and limited, if any, prior peer review. For reasons I will explore in greater detail below, the EBV-sample data may be particularly prone to chimeric pairings that confound the authors' primary analysis goals, and, at the very least, may not reflect physiologically realistic conditions for identifying antigen-reactive TCRs in other contexts.

      Let us first consider Dataset 1 in more detail. The authors compare 2 antigen-stimulated repertoires with 3 unstimulated controls. To improve on single clonotype-level comparisons, the authors propose comparing the cell count aggregated across cells within joint graph communities constructed from paired CDR3 sequences across all the samples. Thus, one of the most relevant questions one hopes the authors answer in this section is whether the resolved graph communities are made up of many distinct clonotypes (i.e., are they polyclonal), allowing the method to function as intended by aggregating across multiple clones with putative shared antigen-reactivity.

      Supplementary Figure 1B shows the size of all the communities with callouts for the putative EBV-expanded communities. The authors listed communities strongly enriched in the EBV-stimulated sample as e1, e2, e3, and e5. Each contains {greater than or equal to}100 clonotypes, and the authors note they contain at least one clonotype with a CDR3 sequence matching an EBV-annotated clone in VDJdb - a database of TCRs with some experimental evidence of epitope-reactivity. Community "e2" is notable for its remarkable size, including 1,339 unique clonotypes. At first glance, this seems promising for a method attempting to boost signal through community detection. However, it is worth re-investigating the individual clone sizes and sequences within this extraordinary community.

      In the EBV-associated community "e2", a look at the data provided by the authors on the paper's GitHub repository suggests a single clonotype (clonotype_9; TRAV12-3 CATQGSNDYKLSF / TRBV9 CASSTGQVATNEKLFF) comprises 26,917 cells. As such, it makes up 29% of the sample with a total of 90,588 cells. If one clone supplies most of a graph community's cell counts, the community-level posterior estimate of β (Figure 1B) effectively tracks a single-clone estimate, and the premise of borrowing statistical power across a polyclonal expansion in this example is hard to assess.

      There is also considerable evidence to believe that the apparent mega-polyclonality of cluster e2 may be partially an illusion, stemming from an artifact of this hyper-expanded clone's massive size and the experimental method used to assign TCRα-TCRβ chain pairings. In fact, the same α-chain CDR3 (CATQGSNDYKLSF) appears in ~1,219 clonotypes paired to distinct β chains, generating much of e2's remarkable 1,339-clonotype count. I believe two features warrant caution here. First, a single clonotype making up ~29% of total T cells in the sample is highly unexpected in ex vivo repertoires, suggesting intense non-physiological expansion conditions unlikely to generalize to other settings. That is, one would almost never expect to see a signal this strong.

      Second, one TCR-α chain paired to ~1,220 distinct and diverse β chains in one sample is also biologically unexpected given what we know of the best-characterized epitope-specific responses for EBV, including to the well-known HLA-A*02 EBV BMLF-1 epitope, which recruits a tetramer-stained repertoire with conserved CDR3 motifs in both chains and at least some constraint on favored V-gene/α-β pairing (See Extended Data Figure 5 in Dash et al., Nature 2017). This raises the concerning possibility of barcode collision and mispairing against a hyperexpanded clone in the ParseBio split-pool method, unlikely to be robust to a clone occupying a third of the sample. Most of the ~1,219 β chains paired to the dominant α have a cell count of 1, further raising the concern of artifactual pairing versus genuine convergence. The second largest community "e1" also seems to suffer from the same issue, with a TCRβ sequence from one super clone making up 7% of the sample potentially being artifactually over-paired to >500 rare single-cell-count TCR α chains.

      Taken together, these observations suggest that the authors' first positive control example passes but probably for the wrong reason since at least some of the "antigen-specific communities" are strongly anchored by a single hyper-clone. This could be remedied by repeating the same type of analysis on an ex vivo single-cell repertoire following more modest stimulation or natural infection (e.g., yellow-fever vaccine, influenza, or SARS-CoV-2 single-cell TCR datasets). I would advise future work using an alternative data source with better-documented experimental protocols, given the concerns above.

      (2) Weaknesses in Example 2

      Example 2 explores the application of Bayesian methods for identifying sample-enriched joint graph communities found in longitudinal data from many participants at two time points and longitudinal data from 1 person (Pt4) at 5 time points. A potential weakness of Case Study 2 is that it yields limited additional biological insight compared to what was previously shown by the authors of the underlying input data. Previously, Formenti et al. 2018 showed that the number of expanded clones after treatment strongly reflects responder status in this cohort, greater in patients with CR/PR versus SD or PD (See Figure 2b of the study). This 2018 primary analysis showed that tracking individual clonotypes was sufficient to reveal biological insight without the need for the computational demands of constructing a massive sequence similarity graph, finding communities on joint graphs, or Bayesian statistical inference. Thus, the impact of Case Study 2 in proving the unique utility of ClustIRR is somewhat diminished.

      It is not clear how this study's result is more "robust" than the original analysis. Perhaps the authors could further clarify what is learned from the uncertainty-aware approach that could not be learned from exact clone tracking in time series.

      Thus, example 2 shows that a complex method recapitulated the finding of a much simpler method for analyzing longitudinal TCR data where a strong signal of expansion was already present at the single-clonotype level. Since the Dirichlet method is sensitive to absolute counts, the large expanding clone in each community at 22 days may alone have carried most of the signal, which is not fully explored.

      In this section, the authors make an interesting observation that CDR3β detected in both PBMC and patient-matched tumor samples were enriched in contracting communities (7/14) versus expanded communities (1/41). The authors do not indicate whether a similar enrichment of tumor-infiltrating lymphocytes (TILs) matched Day 0-22 contracting communities in the other 4 patients with tumor-matched samples, which, if consistent, would strengthen their finding.

      (3) Weaknesses in Example 3

      The final example explores "convergent repertoire differences between species." This is intriguing in principle but less informative due to the use of synthetic data with somewhat predictable properties that some may reasonably consider baked in by the data-generating process. That is, some of the results might be guaranteed by the way the data is constructed using the OLGA/IGoR generative model. For instance, the authors observe a positive correlation between CJ community size and Pgen of constituent clonotypes, stating: "This indicates that CDR3 sequences with high Pgen are statistically more likely to be generated, leading to convergence of similar sequences into public CJs." I may be mistaken, but this conclusion is almost guaranteed by the way the data-generating OLGA model outputs more similar high-Pgen sequences and fewer lower-Pgen sequences.

      An interesting finding in this section is shown in Figure 4B, where community-level aggregation allowed for discrete clustering of human samples away from mouse samples that was not possible by comparing cosine similarity of a sparse clonotype occurrence matrix, a result that would be higher impact if it could be shown to separate real repertoire samples from humans with differential serology, vaccination status, or HLA backgrounds.

      (4) Weaknesses in General

      More generally, one aspect of the method that seems under-emphasized is the fact that many of the nodes in a multi-sample joint sequence similarity graph may have no edges. These zero-degree nodes would probably frequently occur in only one sample but be absent in other samples. It is not strongly emphasized in the paper how the model would infer whether such a singleton found in only one sample in the graph could be reliably inferred to be sample-specific enriched (see, for example, the large single node in Figure 3D, the orange node labeled "GQYF" in the far-right position of the lowest row in panel D). Presumably the number of cell counts represented in this single-sequence node is so great at sampled timepoint Day 22 as to yield a statistically strong signal in the multinomial model; however, the authors may wish to comment on how, for such singleton sequences, the power to detect sample-specific enrichment differs from prior single-clonotype-based methods.

      With any large effort to find statistically significant features from a large candidate set examined all at once, a reader might be concerned with the potential for false discovery. Throughout, the authors seem to assign statistical significance when the 95% high-density interval (HDI) of the posterior estimate excludes zero or when the 95% HDIs of two features do not overlap. There is little discussion of how this implicitly handles multiplicity adjustment via shrinkage, which the authors could address more directly and explain more clearly to a broad audience, including many non-statisticians, who will read this paper.

    4. Reviewer #3 (Public review):

      Summary:

      Analysis of immune receptor repertoires (IRR) needs to take into account the underlying diversity of the repertoires analysed, and the limitations inherent to the technologies used to measure IRRs: limited sampling depth relative to total number of cells and clonotypes, and experimental noise. In this work, Kitanovski and colleagues present ClustIRR. ClustIRR proposes to improve the analysis of immune receptor repertoires, specifically TCRs in the presented applications, by performing two steps: (1) consistent and comparable sequence clustering across repertoires, to account for sparsity of sampling, and (2) estimation of sequence clusters of interest using a Bayesian approach. They showcase ClustIRR performance in 3 scenarios: detection of antigen-specific T cells in a peptide-stimulation, detection of T cells responding to immunotherapy in the context of lung cancer, and analysis of mouse and human T cell repertoires.

      Strengths:

      (1) The cluster occupancy calculation presents an important conceptual framework which would be of interest and useful to the TCR repertoire field as it smoothly integrates information over a set of experimental conditions or time points. The application to longitudinal TCR sequencing datasets is particularly interesting, and could easily be extended to BCR sequencing datasets. Moreover, the calculation can be performed with any user-defined grouping of TCR clones, which allows for usage of other existing methods as the user wishes.

      (2) The results of the human and mouse repertoires provide a very insightful argument for the use of metaclones as opposed to single clones for analysis of repertoires compared to single clones.

      Weaknesses:

      (1) While ClustIRR provides an interesting framework to analyse TCR sequencing datasets, it is not clear whether ClustIRR can perform more informative sequence clustering than state-of-the-art methods. A comparison of obtained clusters with existing methods would provide a useful benchmark. Moreover, computation time scales quite fast with the number of sequences included. This is a major limitation, as the authors state that a time of 2.5 hours is required for clustering of ~100,000 sequences, a number of clones that can easily be reached when analysing multiple samples together.

      (2) The authors claim in the abstract that ClustIRR is integrated with gene expression data. However, in the results presented, the gene expression and TCR sequencing data are analysed separately, and the results are simply correlated. No real integration in the analysis exists for these two data types. The claim should be removed from the abstract. Moreover, the differential gene expression section, while it presents interesting results, lacks clarity and transparency. Presentation of the data in more transparent ways (such as showing violin plots or clustering on the UMAP) would increase the strength of the claims.

      (3) The score calculation does not seem to be normalized to take into account the underlying diversities of CDR3a and CDR3b. While still a useful metric for sequence clustering, I worry about the impact of the lack of normalisation on the conclusion that the "alpha chain drives functional convergence through germline bias". The observation that clustering is mostly driven by the J gene is known and expected (https://pmc.ncbi.nlm.nih.gov/articles/PMC5553937/), as the J gene has lower diversity and greater overlap with the definition of CDR3. Because CDR3a is lower diversity generally from CDR3b, it will likely dominate the sequence similarity signal. Thus, it will appear that Ja drives the signal. The authors do try to address this by looking at clusters driven by CDR3b similarity. However, a low number of CDR3b-driven clusters is consistent with the similarity definition. I wonder if instead the appropriate control for this analysis would be to run the same analysis disregarding CDR3a sequence altogether, and quantify whether similar species-specific DCJ are identified when only CDR3b similarity is used? Absence or reduction of species-specific clustering would confirm that the effect is driven by the CDR3a sequence.

    1. eLife Assessment

      This meta-analysis provides a valuable contribution by integrating findings from dozens of macaque electrophysiology studies to reconcile discrepancies in choice probability and identify factors that robustly influence choice signals in the visual cortex. Such systematic synthesis across studies is rare in this field, and the analysis provides solid evidence for several conclusions, including the cross-study consistency of the relationship between choice probability and sensitivity and the distinctiveness of V1.

    2. Reviewer #1 (Public review):

      This meta-analysis addresses long-standing questions about the reliability and interpretation of choice probability in macaque visual areas, and provides some important findings (e.g., the cross-study consistency of the CP-sensitivity relationship, V1 distinctiveness). However, the evidence for several claims is incomplete: the analysis does not consider the statistical dependence of data points from the same studies and monkeys, and both the bistable-stimulus effect and the stimulus-duration effect rely on interpretive assumptions.

      Strengths:

      The paper's transparency about its own limitations is a genuine strength. Several sections of the paper and the supplement report null results (task exposure, lapse rate, eccentricity) rather than omitting them. This kind of self-scrutiny is uncommon in meta-analyses and substantially increases confidence in the parts of the analysis that do hold up.

      Weaknesses:

      (1) No mixed/hierarchical statistical models for nested data. The paper considers 150 data points from 59 studies and treats them as independent samples, though many share monkeys and brain areas. This reduces the confidence in the reported p-values. A standard way of dealing with this would be to use mixed-effect models with random intercepts rather than OLS.

      (2) Evidence for one of the main findings in the abstract ("First, CPs were higher in tasks involving bistable percepts, reinforcing the link between CP magnitude and subjective perception.") is weak. This effect relies entirely on studies using bistable rotating cylinder stimuli performed in a single lab (lines 702 - 711). I would suggest making this more explicit in the abstract / discussion and in Figure 8b,c.

      (3) The interpretation of main drivers of CP is unclear. The discussion summarizes the 4 main drivers of CP as "four systematic drivers of this variability: neuronal sensi758tivity, brain area, stimulus duration, and task type." The independent contribution of task type is however, questionable. In line 623, it is stated that the difference between coarse and fine discrimination can be entirely explained by the difference in sensitivity (explained possibly by differences in optimizing the stimuli). Again, detection tasks (line 658) show a trend for higher CP because most studies used tailored stimuli from single-recording experiments. Bistable task: see point 2. Thus, all "task effects" can be attributed to confounds, and the claim of the "four drivers of CP" should be revised.

      (4) The paper could be improved by a Discussion that synthesizes the results in a concise manner. Now it seems more like a re-iteration of the results. Overall, I appreciate that the paper is thorough and discusses many of the caveats. However, those are somewhat buried in the long subsections, and I fear that the quick reader may walk away with a stronger impression of "four robust independent drivers" than the text, read carefully, actually supports.

      (5) Datapoints are not weighted according to their standard error (common practice in meta-analysis is inverse-variance weighting). The concern is that underpowered studies with high variance (e.g., due to a low number of recorded neurons) have the same impact as well-powered studies, and this may change some of the estimates. For example, Supplementary Figure 4 shows that mean CP values decrease with statistical power of the study, consistent with the concern. If SEMs are not available, could the authors show the robustness of the results by weighting by sample size as a partial check?

      (6) Interpretation of feed-forward vs. feedback origin of CP [Disclosure: I am an author of Wimmer et al. 2015.]. This paper presents a mechanistic network model of area MT and a decision area that decomposes CP into two components with distinct time courses, and, directly relevant to Section 2.6, shows how a combination of an early feedforward and a late feedback component can produce a roughly time-invariant (flat) CP. This is a specific, quantitative instance of the "sustained plateau" pattern the authors themselves note is inconsistent across studies (lines 480-486) but don't develop further. Engaging with this model in Section 2.1/2.6 would let the authors contrast their duration-effect interpretation against an explicit dynamical model rather than the generic feedforward/feedback dichotomy in Figure 7a.

      (7) Reaction-time experiments. I am worried that differences in CP in RT vs. fixed duration tasks (Supplementary Figure 12) could have an influence on the main regression analysis (because RT experiments are mostly from detection tasks, and because RT experiments presumably include less of a post-decision period). Could this factor be included in the main analysis?

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript by Pletenev et al. provides a meta-analysis of 59 published studies on decision-related activity in the macaque visual cortex. The work does not contain original research material, but by conducting extensive analyses and compilations of previously published data, it provides several new insights not already conveyed in recent reviews on this topic.

      Strengths:

      The work is scholarly and helps organize and integrate a broad set of findings. A difficulty in such an undertaking is making sure that the original studies are accurately characterized and tabulated. The lead authors should be commended for the rigorous approach they've taken; the inclusion of many of the authors of the original studies as co-authors provides additional reassurance that trends visible across studies reflect an accurate quantification of what each study has shown.

      Weaknesses:

      My sole scientific concern is on the 'duration' section (lines 487 to 536). Specifically, the authors relate the dependence of choice probability (CP) on time to two competing (but not mutually exclusive) views of how CPs arise-the 'feedforward' vs 'feedback' views. The basis for the predictions in this section was not clear to me. In particular, it was not obvious that the predictions fully considered all the relevant factors. For instance, did the feedforward predictions consider how the response covariance depends on duration and how this would affect CP values (Equation 1)? (See for example Figure 4 of Kang and Maunsell, 2012, JNP.)

      I would suggest either explaining the basis of the predictions much more carefully or moving the predictions to a supplementary section where they can be presented in more detail (i.e., more convincingly). Alternatively, since the section is inconclusive in the end (there is not strong evidence in favor of FF or FB), the authors could mention the theoretical predictions in passing only, i.e., much more briefly, just to make the reader aware that the different theories can provide predictions about the duration dependence.

    4. Reviewer #3 (Public review):

      Summary:

      This study presents a comprehensive meta-analysis of choice probability (CP), a classic metric of the relationship between single-neuron responses and an animal's perceptual judgment. The authors compiled data from 59 macaque electrophysiology studies and identified several factors that consistently influence CP magnitude, including neuronal sensitivity, brain area, stimulus duration, and task type. These results provide evidence that helps settle several long-standing debates about how CP should be interpreted.

      Strengths:

      This work is a rare example of meta-analysis in macaque electrophysiology, focusing on an important and long-debated metric, choice probability (CP), in the study of sensory and decision-making mechanisms. CP has been measured across many studies, but its interpretation remains contentious because it depends on numerous task and recording factors in addition to sensory and decision-making mechanisms. Individual macaque studies also typically include few subjects, and CP effects are generally small, which prevents any single study from drawing strong conclusions about general patterns. The authors identified an ideal use case for meta-analysis and combined fragmented findings from individual primate studies into a coherent picture of which factors matter most for CP and which do not. This work can also serve as a guide for future meta-analyses of macaque electrophysiology data.

      Weaknesses:

      While the survey and discussion of CP's interpretation are comprehensive, the paper would benefit from a clearer conclusion on why measuring CP remains important and what future directions could make CP more useful for revealing sensory and decision-making mechanisms.

    5. Author response:

      We thank the Editors and Reviewers for their encouraging evaluation and constructive feedback. We are glad that they recognized the value of this systematic synthesis in reconciling disparate findings across many studies on an important question, and in providing new insights that help address long-standing debates around the interpretation of CP values. They also acknowledged our rigorous data curation involving many original study authors, and our transparency in reporting limitations and null results.

      To address their constructive recommendations, we will provide additional hierarchical regression analyses where possible, state some limitations more explicitly, and condense the Discussion section to improve focus and readability.

      Regarding the hierarchical nature of the data, we distinguish three potential hierarchical levels: studies, monkeys, and neuronal samples. Because individual monkeys contribute roughly one observation per study and cannot be tracked across publications, animal-level random effects are statistically unidentifiable. We will address the rare cases where identical neuronal pools were evaluated across multiple task conditions by providing a sensitivity analysis restricted to one data point per unique neuronal sample. At the study level (median 2, range 1–7 observations per study), we will present linear mixed-effects models with random study intercepts.

      On the bistability findings, we agree with Reviewer 1's concern about the limited number of studies. We noted in the Results and Discussion that this effect currently relies on rotating-cylinder paradigms and emphasized the need for CP to be quantified with other forms of bistable stimuli. We will make this limitation explicit in the Abstract and Figure 8 caption. We will revise the Discussion to emphasize the three primary drivers (neuronal sensitivity, brain area, and stimulus duration) and treat the bistable stimulus effect separately as a distinct finding.

      Regarding Reviewer 1's concern about reaction-time experiments: because reaction-time (RT) paradigms are heavily confounded with detection tasks in the existing literature (89% of detection tasks are RT tasks, and 67% of RT tasks are detection ones), including RT as a separate factor introduces near-complete collinearity. We will make this constraint and the inability to statistically disentangle them explicit in the main text.

      Addressing Reviewer 2's concern regarding the CP–duration predictions, we agree that they depend on specific modeling assumptions. While our predictions—for the feedforward framework in particular—reflect standard models from the literature, we will explicitly acknowledge that alternative feedforward assumptions—such as duration-dependent response covariance—could alter the expected relationship.

      In response to Reviewer 3’s comments on the rationale and future utility of measuring CP, we will revise the Discussion to emphasize that while the interpretation of CP has evolved from feedforward readout to include feedback mechanisms, our results confirm that CP remains a robust neural correlate of subjective perception. Although the exact mechanisms linking CP to perception remain unresolved, this ambiguity does not justify abandoning the metric; rather, CP remains an indispensable tool, provided it is supplemented with additional analyses as outlined in our recommendations.

    1. eLife Assessment

      This study presents an open-source reinforcement learning framework for the real-time, closed-loop optimization of spatiotemporal electrical stimulation in engineered neuronal networks. Using single-spike-resolution activity as continuous feedback, the authors provide solid evidence that their platform can identify stimulation patterns that drive specific activity motifs within a structurally constrained four-node circuit. While this reproducible system offers a valuable and accessible tool for interacting with biological neural networks in an adaptive manner, further validation is needed to determine how well these stimulation strategies generalize to larger, unstructured, or more conventional network architectures.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript presents several reinforcement learning (RL) approaches to control activity in small, ring-shaped neural networks. The paradigm is to find temporally and spatially patterned stimuli applied to axons that maximize the length of activity propagation along the network. Several RL versions were compared. Some RL-designed stimulation patterns worked better than others.

      Strengths:

      The work is technically and statistically very sound, and important controls were done that are missing in some earlier work along this line of research. The statistics and technical solutions seem solid, though I am no specialist in RL. The figures are well done and informative, with matching quality of the captions. The work presented is of value mainly for someone who wants to build a good closed-loop stimulation system to experiment with neuronal networks in-vitro.

      Weaknesses:

      The manuscript appears undecided about whether it wants to be about RL control, about a technical implementation of long-term stimulation in vitro, or about the properties of neuronal networks and interaction with them. The introduction is well written, comprehensive and insightful, focusing on biological aspects. Methods are then extensively about RL algorithms, without explaining why several were used or why these in particular, but with specialist language hard to understand for neuroscientists. The results then quantitatively compare the performance of the RL but do not really explain what this teaches us about neuroscience or what we learn about the networks beyond that they can be stimulated for longer propagation patterns. Extensive supplementary material almost advertises the hardware built by the team. Some figures suggest, though I'm not 100% certain about this, that different algorithms find different optimal stimulation patterns in the same network - which I find puzzling. What then does this tell us about the stimulation patterns and the variability of the responses? There is some speculative mechanistic reasoning, but no data to support this further. The conclusions hardly address the initial motivation of the manuscript. To me, it is not clear what can be learned that was not, in one way or another, presented previously, with as well as without RL.

    3. Reviewer #2 (Public review):

      The authors developed a reinforcement learning framework to identify stimulation patterns that produce a predefined activity pattern in neuronal networks cultured on a microelectrode array. The main contribution of this study lies in the development and validation of the experimental framework, rather than in providing new insight into neuronal response properties or the biological mechanisms underlying network activity.

      A major strength of the work is that the framework is implemented using open-source software. Electrical stimulation through microelectrode arrays is already widely used, and the proposed approach is therefore likely to be of interest to researchers seeking to automate the exploration and optimization of stimulation parameters. The authors also provide the software, hardware configuration, and experimental data, which substantially improves transparency and should facilitate reproduction and further development of the system by other laboratories.

      The experiments provide a useful proof of concept showing that the framework can search for stimulation patterns associated with the predefined task in the tested neuronal network. The evidence is solid for demonstrating the feasibility of the approach within this specific experimental configuration. In particular, the closed-loop integration of stimulation, recording, evaluation of the neuronal response, and subsequent selection of stimulation patterns is clearly implemented and experimentally tested.

      An important limitation is that the framework was evaluated using a relatively small and highly structured network. Four stimulation electrodes were positioned around a single network, and stimulation was delivered to microchannels in which axons were concentrated. This configuration is well suited to the initial demonstration, but it remains uncertain whether the same approach will perform similarly in conventional monolayer dissociated cultures, larger networks, or systems with different numbers and spatial arrangements of electrodes. Additional validation across a broader range of network structures and experimental configurations would therefore be needed to establish the general applicability of the framework.

      Overall, this study presents a valuable and reproducible methodological framework for the closed-loop optimization of electrical stimulation in cultured neuronal networks. Its likely impact lies primarily in providing an accessible experimental and computational platform that can be adapted for studies requiring systematic exploration of stimulation patterns, although the extent to which the current findings generalize beyond the tested network configuration remains to be determined.

    4. Reviewer #3 (Public review):

      Summary:

      This study integrates living neuronal circuits with a real-time reinforcement-learning framework to enable closed-loop optimization of directional firing sequences at single-spike resolution. By using the spatiotemporal structure of neuronal firing as feedback, the system adaptively searches for stimulation patterns that increase a predefined propagation reward. The work represents an exciting step toward systematic and adaptive control of living neuronal networks.

      Strengths:

      The study presents a cutting-edge closed-loop platform that combines neuronal cultures, high-temporal-resolution electrophysiology, and reinforcement learning with millisecond-scale latency. The ability to optimize directional activity propagation at single-spike resolution is particularly novel. More broadly, the work provides a compelling framework for interacting with biological neural networks in an adaptive rather than purely predefined manner.

      Weaknesses:

      Several aspects of the analysis and interpretation require further clarification. In particular, the definition and computation of directional propagation and reward are not always clear. The generalizability of the optimized stimulation patterns across cultures also remains unclear.

    5. Author response:

      eLife Assessment

      This study presents an open-source reinforcement learning framework for the real-time, closedloop optimization of spatiotemporal electrical stimulation in engineered neuronal networks. Using single-spike-resolution activity as continuous feedback, the authors provide solid evidence that their platform can identify stimulation patterns that drive specific activity motifs within a structurally constrained four-node circuit. While this reproducible system offers a valuable and accessible tool for interacting with biological neural networks in an adaptive manner, further validation is needed to determine how well these stimulation strategies generalize to larger, unstructured, or more conventional network architectures.

      We thank the editors and reviewers for their kind and insightful words, as well as their constructive feedback and assessment. We agree that in its current stage, the manuscript does not convincingly argue, that specific stimulation strategies applied to one network generalise well to other network architectures. We intend to address this problem by adjusting the focus of the manuscript to be more on the platform than on any specific neuroscience claim, as we believe that this is the more useful angle for the community at large. We also intend to discuss the transferability of our results between the presented culture system and other neuroscience systems.

      Based on the reviewer comments, we will further give a more approachable introduction to reinforcement learning to make the concept more accessible to a wider audience. We will also motivate the algorithm selection in more depth.

      In the following, we will discuss the comments put forward by the reviewers and how we intend to adapt our manuscript to address them. In this provisional response, we will only focus on major concerns. Minor comments, where we are following the suggestions made by the reviewers directly, will not yet be discussed.

      Reviewer #1 (Public Review):

      The manuscript appears undecided about whether it wants to be about RL control, about a technical implementation of long-term stimulation in vitro, or about the properties of neuronal networks and interaction with them. The introduction is well written, comprehensive and insightful, focusing on biological aspects. Methods are then extensively about RL algorithms, without explaining why several were used or why these in particular, but with specialist language hard to understand for neuroscientists. The results then quantitatively compare the performance of the RL but do not really explain what this teaches us about neuroscience or what we learn about the networks beyond that they can be stimulated for longer propagation patterns. Extensive supplementary material almost advertises the hardware built by the team. Some figures suggest, though I’m not 100% certain about this, that different algorithms find different optimal stimulation patterns in the same network - which I find puzzling. What then does this tell us about the stimulation patterns and the variability of the responses?

      We thank the reviewer for pointing out these issues with our manuscript. The goal of our work is the hardware framework. We plan to highlight this more in the revised version. The RL part will be revised with less specialist language and will get less emphasis in the next version of our manuscript. We will also discuss why different algorithms seem to find different solutions, which can be traced back to having different networks and agents finding different local maxima.

      By clearly setting the scope of our manuscript on the hardware aspects and revising the RL part to be seen more as one possible closed-loop control paradigm, we further intend to address the concerns put forward by the reviewer regarding our focus on RL in other parts of the manuscript.

      Reviewer #2 (Public Review):

      An important limitation is that the framework was evaluated using a relatively small and highly structured network. Four stimulation electrodes were positioned around a single network, and stimulation was delivered to microchannels in which axons were concentrated. This configuration is well suited to the initial demonstration, but it remains uncertain whether the same approach will perform similarly in conventional monolayer dissociated cultures, larger networks, or systems with different numbers and spatial arrangements of electrodes. Additional validation across a broader range of network structures and experimental configurations would therefore be needed to establish the general applicability of the framework.

      We thank the reviewer for their feedback. They are right to point out that the manuscript here focuses purely on small and highly structured networks. With the presented experiments, our manuscript does not and cannot make any biological claims about how neurons communicate. However, with this manuscript, we also do not intend to do so, as such an endeavour would be out of scope. In the next version of our manuscript, we will better highlight the focus of our manuscript, which lies on the hardware framework itself. Furthermore, we will discuss in the outlook the generalisability and scalability concerns raised by the reviewer in more detail.

      Reviewer #3 (Public Review):

      Several aspects of the analysis and interpretation require further clarification. In particular, the definition and computation of directional propagation and reward are not always clear. The generalizability of the optimized stimulation patterns across cultures also remains unclear.

      We thank the reviewer for their insightful feedback. We agree with them and will discuss (1) the currently implemented reward based on directional propagation of activity and (2) the limitations and aspects influencing reward algorithm selection in more detail in our revised version of the manuscript. We further believe that we cannot make any claims about generalisability between cultures or when changing the experimental paradigm.

    1. eLife Assessment

      This important study provides convincing evidence that somatosensory relay nuclei remain engaged during attempted hand movements following chronic cervical spinal cord injury, including in individuals with severe motor impairment. The combination of functional and quantitative MRI with physiological controls supports the main findings and may have implications that are of importance in understanding sensorimotor processing after spinal cord injury. However, the interpretation of this activity as specifically reflecting top-down, and particularly corticocuneate, signaling may be overstated. The origin and anatomical route of the observed activity appears to remain less firmly established.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigated somatosensory processing along the afferent somatosensory pathway (cuneate nucleus, thalamus, S1) in a group of spinal cord injury patients and a group of controls. They propose that reduced motor function in SCI patients would reduce bottom-up activity; thus, recorded activity in SCI patients would reflect top-down modulation of overt or attempted movements.

      Strengths:

      (1) Strong methods.

      (2) Experimental and control groups.

      (3) Strong writing.

      (4) Results well presented.

      (5) Appropriate statistics.

      Weaknesses:

      Some results (or lack of) cast doubt about the ability of the used technique (3T fMRI) to detect the desired effects (bottom-up vs top-down activity).

    3. Reviewer #2 (Public review):

      Summary:

      This study addresses a question that has been essentially inaccessible in humans: whether the early somatosensory relay nuclei are engaged by anything other than peripheral drive. Using functional and quantitative MRI in individuals with chronic cervical spinal cord injury, the authors show that the cuneate nucleus and VPL are robustly engaged during overt or attempted hand movement, and that this engagement persists in a participant with complete hand paralysis and no detectable muscle activity. They further report structural degeneration of the cuneate nucleus that is unrelated to the preserved task-evoked activity, and interpret the residual activity as reflecting top-down corticocuneate signalling.

      Strengths:

      The work is well conceived and clearly written, and the demonstration that the cuneate nucleus and VPL are robustly engaged during (attempted) hand movement after chronic cervical spinal cord injury is, to my knowledge, novel at this level of anatomical resolution. The finding is convincing. The dissociation between preserved task-evoked activity and marked structural degeneration of the cuneate nucleus is a highly interesting result with clear implications for rehabilitation. The manuscript is straightforward to follow, and the imaging protocol is carefully executed.

      Weaknesses:

      My main reservation concerns the inferential step from "not peripheral" to "corticocuneate". The data establish the former convincingly; the latter may not.

      (1) Attribution of the observed activity to corticocuneate projections.

      The central claim rests on an argument by elimination: because bottom-up drive is excluded in PT01, the residual activity must be top-down and, by extension, corticocuneate. Two distinct gaps should be addressed. First, "top-down" is not equivalent to "direct corticocuneate". Descending influence could reach the cuneate nucleus through multiple indirect pathways. Second, the activity observed at the three levels (cuneate, VPL, S1) need not be serially propagated, since layer 6 corticothalamic projections, for example, could drive VPL independently of any cuneate contribution. I would ask the authors to either provide evidence bearing on the routing, or to consistently use a route-neutral term (e.g. "descending" or "top-down") and reserve "corticocuneate" for the discussion of candidate mechanisms.

      (2) Afferent input arising above the lesion level.

      The EMG control in PT01 was restricted to hand and forearm muscles. Musculature innervated above C4 (cervical paraspinals, trapezius, and to a variable extent the shoulder girdle) remained available to this participant, and attempted hand movement is frequently accompanied by increased proximal co-contraction, postural stabilization, and altered respiratory effort. Afferent to the upper cervical cord is known to project to the ipsilateral cuneate nucleus, and its activity would produce lateralized, ipsilaterally dominant cuneate input - that is, precisely the pattern reported. This alternative is not excluded by the present control and should be addressed directly, ideally with proximal EMG in PT01 (and, if possible, in the other participants), or at minimum with an explicit discussion. Relatedly, the authors recorded respiratory and cardiac signals: please report whether respiratory volume or heart rate differed between movement and rest blocks, and between groups, since the dorsal medulla lies adjacent to cardiorespiratory nuclei.

      (3) Functional significance of the preserved top-down signal.

      The discussion establishes that top-down input persists but says relatively little about why it should. If the principal role of descending input to the cuneate nucleus is the gating of incoming afferent traffic, then in the absence of afferents there is nothing left to gate, and one might have expected the signal to be lost. Its persistence is the most interesting aspect of the finding and deserves fuller discussion. Candidate accounts the authors may wish to consider include: an efference copy or predictive signal delivered to a comparator that no longer receives its input, in the framework the authors already invoke (references 27, 28); engagement of the non-lemniscal outputs of the dorsal column nuclei (e.g., cuneocerebellar, cuneo-olivary projections); attempted movement engages motor imagery and attention, in which case the relevant question becomes what distinguishes these from movement-related gating. A related interpretational point: in behaving primates, movement-related modulation of cuneate transmission is bidirectional and includes prominent suppression (refs. 7/12). Note also that BOLD increases are compatible with increased inhibition, so they do not indicate facilitated throughput.

    4. Author response:

      Reviewer #1:

      Some results (or lack of) cast doubt about the ability of the used technique (3T fMRI) to detect the desired effects (bottom-up vs top-down activity). 

      As the reviewer notes, 3 T MRI alone cannot partition the relative contributions of bottom-up and top-down signals, since both are present during movement in an intact system. This reflects the premise of our study design but is also an important caveat when interpreting the group-level results, which characterise the net task-related response. Our central inference therefore focuses on the persistence of activation in PT01, in whom hand movement was absent. PT01 therefore lacks bottom-up signals, and any observed activity must be driven by top-down processes. In the revised manuscript, we provide measures of signal quality and activation magnitude to better characterise the sensitivity of these measurements and will draw on existing evidence that our approach resolves task-specific responses within these nuclei.

      Reviewer #2:

      (1) Attribution of the observed activity to corticocuneate projections: My main reservation concerns the inferential step from "not peripheral" to "corticocuneate". The data establish the former convincingly; the latter may not. The central claim rests on an argument by elimination: because bottom-up drive is excluded in PT01, the residual activity must be top-down and, by extension, corticocuneate. Two distinct gaps should be addressed. First, "top-down" is not equivalent to "direct corticocuneate". Descending influence could reach the cuneate nucleus through multiple indirect pathways. Second, the activity observed at the three levels (cuneate, VPL, S1) need not be serially propagated, since layer 6 corticothalamic projections, for example, could drive VPL independently of any cuneate contribution. I would ask the authors to either provide evidence bearing on the routing, or to consistently use a route-neutral term (e.g. "descending" or "top-down") and reserve "corticocuneate" for the discussion of candidate mechanisms. 

      We agree with the reviewer that our previous attribution of the observed brainstem effects to corticocuneate processing was speculative and should have been presented as such. Our findings support a non-peripheral, top-down contribution but do not allow us to attribute this descending influence to a specific anatomical route. We have therefore revised the manuscript throughout to use “top-down” processing as a more route-neutral term, reserving the corticocuneate pathway for discussion of possible candidate mechanisms. We have also clarified in the revised discussion that activity observed across the cuneate nucleus, VPL, and S1 does not necessarily imply serial propagation through these structures.

      (2) Afferent input arising above the lesion level: The EMG control in PT01 was restricted to hand and forearm muscles. Musculature innervated above C4 (cervical paraspinals, trapezius, and to a variable extent the shoulder girdle) remained available to this participant, and attempted hand movement is frequently accompanied by increased proximal co-contraction, postural stabilization, and altered respiratory effort. Afferent to the upper cervical cord is known to project to the ipsilateral cuneate nucleus, and its activity would produce lateralized, ipsilaterally dominant cuneate input - that is, precisely the pattern reported. This alternative is not excluded by the present control and should be addressed directly, ideally with proximal EMG in PT01 (and, if possible, in the other participants), or at minimum with an explicit discussion. Relatedly, the authors recorded respiratory and cardiac signals: please report whether respiratory volume or heart rate differed between movement and rest blocks, and between groups, since the dorsal medulla lies adjacent to cardiorespiratory nuclei.

      The reviewer raises an important point. To test whether proximal muscle activity could account for the cuneate response in PT01, we collected additional EMG data during attempted hand movements, focusing on muscles innervated above the lesion level. These included the anterior (AD) and middle deltoid (MD), upper (UT) and middle trapezius (MT), and cervical paraspinals (CP), along with two of the original distal recordings (thenar eminence, TE; extensor digitorum, ED). Nonetheless, we found no significant difference in EMG activity in any recorded muscle during attempted left- or right-hand movement versus rest (Supplementary Figure 3A and 3B). To confirm that the chosen electrode montage could detect proximal muscle activity, we further instructed the participant to perform left or right shoulder shrugs during the same session. This produced a clear increase in activity during movement across the trapezius, cervical paraspinal and middle deltoid recordings (Supplementary Figure 3C and 3D).

      In addition, we analysed the cardiac and respiratory recordings acquired during fMRI to determine whether movement-related physiological changes could explain the observed brainstem activity. Importantly, any physiological change in heart rate during movement is global and therefore cannot explain the hand-dependent lateralisation of the cuneate response. Furthermore, cardiac and respiratory nuisance regressors were included in all first-level models. Heart rate showed a small but significant increase during movement compared with rest (controls: +0.31 bpm; SCI: +0.69 bpm; main effect of condition F(1,33) = 10.22, p = 0.003, η<sup>2</sup></sub>p</sub> = 0.24, BF<sub>10</sub> = 8.44), whereas respiratory volume per time (RVT) did not differ between conditions (F(1,33) = 0.35, p = 0.56, η<sup>2</sup></sub>p</sub> = 0.01, BF<sub>10</sub> = 0.24). Given the autonomic consequences of cervical injury, we also tested whether these changes differed between groups. Neither measure showed a Group × Condition interaction (heart rate: F(1,33) = 1.41, p = 0.24, BF<sub>10</sub> = 0.56; RVT: F(1,33) = 1.60, p = 0.22, BF<sub>10</sub> = 0.61), suggesting that they cannot account for the group differences we report.

      We have added the additional EMG control and the cardiorespiratory analyses to the Supplementary Material of the revised manuscript. Together, these controls suggest that neither proximal muscular nor cardiorespiratory factors explain our findings.

      (3) Functional significance of the preserved top-down signal: The discussion establishes that top-down input persists but says relatively little about why it should. If the principal role of descending input to the cuneate nucleus is the gating of incoming afferent traffic, then in the absence of afferents there is nothing left to gate, and one might have expected the signal to be lost. Its persistence is the most interesting aspect of the finding and deserves fuller discussion. Candidate accounts the authors may wish to consider include: an efference copy or predictive signal delivered to a comparator that no longer receives its input, in the framework the authors already invoke (references 27, 28); engagement of the non-lemniscal outputs of the dorsal column nuclei (e.g., cuneocerebellar, cuneo-olivary projections); attempted movement engages motor imagery and attention, in which case the relevant question becomes what distinguishes these from movement-related gating. A related interpretational point: in behaving primates, movement-related modulation of cuneate transmission is bidirectional and includes prominent suppression (refs. 7/12). Note also that BOLD increases are compatible with increased inhibition, so they do not indicate facilitated throughput. 

      We thank the reviewer for their comment and agree that the persistence of this descending signal despite profound loss of peripheral input is a very interesting aspect of the findings, and that our manuscript will benefit from a more extended discussion of this result. We will expand on this and the candidate accounts raised in the revised discussion. We will also clarify that movement-related modulation of cuneate processing may include both facilitation and suppression, and that our finding of increased BOLD activity does not necessarily imply facilitated sensory throughput.

    1. eLife Assessment

      This valuable study examines the relationship between simultaneous LC single-unit recordings and pupillometry, both within and across baseline and evoked epochs, finding that cross-epoch relationships are unreliable. The evidence is solid and suggests that care should be taken when interpreting pupil responses as reflecting LC activity. This manuscript will be interesting to basic and clinical researchers working in cognitive and decision neuroscience as well as computational psychiatry.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses a valuable dataset of simultaneous LC single-unit recordings and pupillometry in awake monkeys to examine one aspect of the relationship between LC activity and pupil diameter: whether baseline LC activity predicts evoked pupil and whether baseline pupil predicts evoked LC. These relationships are largely absent in the present dataset, and the authors conclude that this type of prediction is not reliable.

      Strengths:

      This is a valuable dataset of simultaneous LC single-unit recordings and pupillometry in awake monkeys, the within-modal and within-epoch analyses are sound. The new results further caution the use of pupil diameter to infer LC activity.

      Weaknesses:

      There is an obvious rationale for asking whether baseline LC relates to baseline pupil or evoked LC relates to evoked pupil. It is also obvious to ask about the relationship between baseline and evoked LC activity, and between baseline and pupil responses. However, the rationale for the cross-modal and cross-epoch analysis is not clear. Why should one expect baseline LC to predict evoked pupil, or baseline pupil to predict evoked LC? What is the biological importance of such predictions? These analyses simultaneously change the modality and the temporal domain.

      Furthermore, some recent studies, which are not referenced in this manuscript, showed variability in the relationship between pupil and LC within the same epoch, and that tonic and phasic pupil responses could be regulated by different inputs to LC. Given that the pupil-LC coupling is readily imperfect and potentially controlled by different cellular and circuit mechanisms, it is not surprising that their cross-epoch relationship is even more variable.

      One main analysis that correlated baseline pupil with evoked LC showed statistically different results between the two monkeys, raising the question of to what extent a general claim for the cross-modal cross-epoch analysis can be made.

    3. Reviewer #2 (Public review):

      The manuscript by Thompson and Gold reported that, although baseline and evoked LC activity were positively correlated with baseline pupil size and evoked pupil dilation, respectively, there were no reliable cross-epoch relationships between LC activity and pupil size - that is, between baseline LC activity and evoked pupil dilation or between baseline pupil size and evoked LC activity. A major strength of the study is its large dataset, comprising recordings from more than 100 LC single units collected across 83 recording sessions in two monkeys. The authors further strengthened their conclusions by performing several control analyses to rule out potential confounds, including whether the findings were influenced by (1) the use of residuals in the main analyses, (2) the possibility that the auditory stimulus failed to evoke the full dynamic range of LC activity and pupil responses, and (3) nonlinear cross-epoch relationships between LC activity and pupil size.

      These findings are important because they remind researchers to exercise caution when assuming that non-illuminance-mediated fluctuations in pupil size can reliably serve as a proxy for LC activity under all experimental conditions

      With that being said, I have two suggestions that I believe would further strengthen the manuscript.

      (1) Previous studies have suggested that the LC is organized into functionally distinct subpopulations (e.g., PMID 28920933 and PMID 40770025). It is therefore possible that a subset of the LC neurons recorded in this study does not participate in controlling pupil size. I encourage the authors to discuss this possibility in the Discussion. In addition, this possibility could be tested by repeating the cross-epoch relationship analyses using only LC units that exhibit a strong correlation with pupil size (the darker points in the last column of the top two rows of Figure 1?).

      (2) The manuscript would benefit from a clearer visualization of the analyses addressing the possibility that the auditory stimulus failed to evoke the full dynamic range of LC activity and pupil responses. A supplemental figure illustrating the distributions or ranges of LC activity and pupil responses, together with the corresponding control analyses, would help readers better understand this important point.

    1. eLife Assessment

      This study addresses a key question concerning whether neurofeedback can enhance the neural representation of a selected speaker during competing continuous speech and whether such enhancement translates into behavioral benefits. The solid findings provide important insights into the online modulation of auditory attention and the extent to which selective listening can be voluntarily controlled. Although the observed effects were relatively small and did not consistently persist beyond the feedback period, the study advances understanding in an under-explored area and provides a helpful foundation for future research on sustained learning and behavioral transfer.

    2. Reviewer #1 (Public review):

      Summary:

      The authors asked whether neurofeedback during competing continuous speech can help to modulate the attention-related N1-component in the temporal response function (TRF), which is an event-related-response-like estimate of the phase-locked EEG activity following the envelope. The research question is relevant because it asks to what degree the strength of attention can be controlled beyond the binary decision to attend or ignore something, and whether this control is beneficial for the behavioral outcome.

      Strengths:

      (1) Sample size of 56 participants.

      (2) Control group with sham feedback.

      (3) Novelty: Under-explored field of neurofeedback in selective speech tracking.

      (4) Pragmatic and reasonable methodological decisions.

      (5) Transparent results not hiding the fact that effect sizes are small.

      Weaknesses:

      Besides some need for clarification, I could only find one methodological weakness, which the authors discuss anyway:

      (1) Overall, speech tracking-based neurofeedback may lead to more robust results, because the N1-extraction does not have to be handcrafted and all components would be taken into account. As the authors state, the P2-component has been related to effort, and this may provide more "room to play" for voluntary modulation.

      The following "weaknesses" are related to the impact of the results:

      (2) Non-translating effects to post-training trials, neither neurally nor behaviorally.

      (3) Neurofeedback-related Modulation of N1

    3. Reviewer #2 (Public review):

      Summary

      This manuscript investigates whether neurofeedback based on the N1 component of the temporal response function can be used to modulate neural responses during selective attention to continuous speech. Participants listened to two competing audiobooks and were instructed to attend to one of them. In the neurofeedback group, trial-by-trial N1 responses were converted into visual feedback, whereas the sham-feedback group received replayed feedback from other participants. The authors found a significant interaction between group and block for the N1 response to target speech over a small fronto-central cluster, with larger N1 responses during feedback blocks in the genuine neurofeedback group but not in the sham group. No neurofeedback effect was found for the distractor response. The authors also reported exploratory post-training effects and an association between changes in N1 and speech-comprehension performance at right fronto-central electrodes.

      Overall, the study is conceptually interesting and novel. The online, trial-by-trial estimation of neural responses from continuous speech is an attractive development for auditory neurofeedback, and the inclusion of a randomised sham-feedback group is an important strength. However, the manuscript provides stronger evidence for modulation of a neural response during feedback than for learning or training of selective attention. Some aspects of the analysis and interpretation also require further consideration/clarification, particularly the use of group-specific N1 time windows, the spatial confound between target and distractor streams, the absence of artefact correction in the signal used for feedback, and the relatively weak behavioural evidence.

      Strengths

      The main strength of this study is its novel use of neurofeedback during continuous competing speech. Rather than providing feedback based on a general measure of brain activity, the authors targeted a specific neural response associated with selective auditory attention. The online implementation is technically impressive, allowing neural responses to be estimated from 22-second speech segments and converted rapidly into feedback. The inclusion of a sham-feedback group is another important strength, as it helps distinguish effects of genuine neurofeedback from nonspecific effects such as task engagement or motivation. The relatively large sample for a neurofeedback study and the use of natural continuous speech also increase the robustness and ecological relevance of the work.

      Weaknesses

      (1) The effect was present during the feedback blocks but did not increase across training blocks, and the post-training effect was not found at the same electrodes used for feedback. The evidence therefore supports online modulation more strongly than learning or lasting self-regulation, and claims about successful training or persistent learning should be interpreted cautiously.

      (2) The offline analysis used different N1 time windows for the neurofeedback and sham groups. This introduces a potential bias in the group comparison because the dependent measure was defined differently between groups. Confirmation of the main result using a common, independently defined N1 window would strengthen the evidence.

      (3) The target speech was always presented from the front and the distractor from behind. Differences between target and distractor responses therefore cannot be attributed entirely to attention because spatial location is also different. This limits the interpretation of target-versus-distractor differences and may also contribute to the weaker reliability of the distractor response.

      (4) Feedback blocks always contained two speakers, whereas half of the baseline trials contained only one speaker. Since the presence of a distractor altered the neural response, it is important that the neurofeedback comparison is based on acoustically matched multi-speaker baseline trials. If this were not the case, differences between baseline and feedback could partly reflect differences in the acoustic condition rather than neurofeedback.

      (5) No artefact correction was applied to the signal used for online feedback. Because feedback was derived from fronto-central electrodes, eye or muscle activity could potentially contribute to the measured signal. An offline demonstration that the main neural effect remains after appropriate artefact control would strengthen the interpretation that the effect reflects neural modulation rather than systematic changes in non-neural activity.

      (6) The behavioural evidence is weaker than the neural evidence. There was no significant overall improvement in speech comprehension in the neurofeedback group. The reported behavioural effect is instead based mainly on associations between changes in the neural response and changes in comprehension, and some of these effects were weak before the whole-scalp analysis. Therefore, these findings are better interpreted as exploratory associations rather than evidence that neural modulation directly caused improved comprehension.

      (7) Adding the distractor did not significantly reduce comprehension performance. The absence of a significant distractor effect on comprehension suggests that the listening condition may not have produced a strong behavioural cocktail-party difficulty in this sample. The large spatial separation between speakers and the nature of the behavioural task may have reduced sensitivity to distraction, which could limit the strength of the conclusions regarding improvement of speech understanding in challenging listening conditions.

      (8) The study is described as double-blind, but the manuscript provides limited detail on how blinding was maintained, particularly when the experimenter manually checked the N1 estimate. In addition, participants' belief in or perceived control over the feedback was not formally assessed. These factors make it difficult to determine how effectively expectancy or motivation-related effects were controlled.

      Overall assessment:

      This study provides a valuable methodological and conceptual advance by showing that online, trial-by-trial neurofeedback can modulate the neural response to attended speech during competing speech. The evidence is solid for an immediate neural effect during feedback, and it was supported by comparison with a sham-feedback group. However, it remains incomplete for broader claims about learned self-regulation, persistent effects, distractor suppression, and improved speech understanding.

    1. eLife Assessment

      This study presents a fundamental molecular resource, offering subtype-specific insight into the composition of ribosome-associated protein complexes in the developing cerebral cortex. The evidence is compelling in terms of data quality and is strongly supported by the results, given the rigorous technical execution. While primarily descriptive in nature, this resource will be of great use to the field.

    2. Reviewer #1 (Public review):

      This work provides a valuable toolkit for endogenous isolation of projection neuron subtypes. With further validation, it could present a solid method for low-input ribosome affinity purification using a ribosomal RNA (rRNA) antibody. The experimental evidence for the distinct ribosomal complexes is limited to this method and indirect support from complementary analyses of pre-existing data. However, with additional experimental data to support the specificity of ribosomal complex pulldown and confirmation of the putative ribosomal complex proteins of interest, the study would provide compelling evidence for translation regulation of neuronal development through compositional ribosome heterogeneity. This work would be of interest to neuroscientists, developmental biologists, and those studying translational networks underlying gene regulation.

      Strengths

      (1) This in vivo labeling of specific projection neurons and ribosomal rRNA affinity purification method accommodates a low input of <100K somata per replicate, which is useful for the study of neuronal subtypes with limited input. In principle, this set of techniques could work across different cell types with limited input depending on the molecule used for cell type labeling.

      (2) The authors are also able to isolate endogenous neurons with minimal perturbation up to the point of collection, preserving the native state for the neuron in vivo as long as possible prior to processing.

      (3) This study identified over a dozen potential non-ribosomal proteins associated with SCPN ribosomal complexes, as well as a ribosomal protein enriched in CPN.

      Limitations

      (1) In this study, the authors address the advantages of their ribosomal complex isolation method in SCPN and CPN against RPL22-HA affinity purification. While this does show more pull down of the ribosomal RNA by the Y10B rRNA antibody, the authors claim this method identifies cell-type specific ribosomal complex proteins without demonstrating a positive control for the method's specificity. There are very limited experiments to truly delineate how "specific" this method is working and whether there could be contamination from other complexes bound by the antibody. I see this as the major limitation that should be addressed. To boost their claims of capturing cell-type specific ribosomal complexes, the authors could consider applying their rRNA affinity purification pipeline to compare cell-types with well-characterized ribosome-associated proteins, like mouse embryonic stem cells and HELA cells. The reviewer can completely appreciate the elegance in the neural characterization here, but it seems there needs to be a solid foothold on the specificity of the method, perhaps facilitated by cell types that can be more readily scaled up and tested.

      (2) The authors followed up on their differentially enriched ribosomal complex proteins by analyzing ribosome association of these proteins in external datasets. While this analysis supports the ribosome-association of these proteins, there is limited experimental validation of physical association with the ribosome, much less any functional characterization. The reciprocal pulldown of PRKCE is promising; however, I would recommend orthogonal validation of several putative ribosomal complex proteins to increase confidence. Specifically, the authors could use sucrose gradient fractionation of SCPN and CPN, followed by western blot to identify the putative interaction with the 80S monosome or polysomes. This would also provide evidence towards the pulldown capturing association with mature ribosome species, which is currently unclear. This experiment would provide substantial evidence for the direct association of these non-ribosomal proteins with subtype-specific ribosomal complexes.

      (3) The authors state interest in learning more about the differences underlying translational regulation of projection neuron development. This method only captures neuronal somata, which will only capture ribosomes in the main cell body. There are also ribosomes regulating local translation in the axons, which may also play a critical role in axonal circuit establishment and activity. These ribosomal complex interactions may also be rather transient and difficult to capture at only one developmental stage. Therefore, this method is currently limited to a single developmental snapshot of ribosomal complexes at P3 within the main cell body. It would be exciting to see extended utility of this method to sample neurites and additional developmental stages to gain further resolution on the developmental translation regulation of these projection neurons.

      Likely impact of the work on the field, and the utility of the methods and data to the community

      The authors introduce a unique pipeline of techniques to identify cell-type specific ribosomal complex compositions. With more validation, there is certainly potential for those studying neuronal translation to leverage this method in limited primary cells as an alternative to existing methods that do not rely on ribosomal protein tagging, such as ARC-MS (Bartsch et al., 2023), RAPIDASH (Susanto and Hung et al., 2024), and RAPPL (Nature Communications, 2025).

      Comments on revised version.

      We thank the authors for their thorough response to our comments. The revised manuscript satisfactorily addresses most reviewer comments through clarification and expanded discussion, although we believe some important limitations remain. In particular, we continue to view experimental validation of the identified ribosome-associated proteins as an important component of introducing this methodology to the field, rather than work that falls beyond the scope of the study. The authors have acknowledged that these experiments are future directions, and it is clear they plan to pursue additional validation outside of the current manuscript. Given their emphasis that the primary contribution is the development of a methodological framework, we believe the work may be appropriately framed as a Tool and Resource article. Such positioning would better align reader expectations with the manuscript's strengths as a technical advance while recognizing that further validation and functional characterization will be needed in future studies. Despite these limitations, I believe the manuscript makes a valuable methodological contribution.

    3. Reviewer #2 (Public review):

      Summary:

      The study by presents a sophisticated molecular dissection of ribosome-associated complexes (RCs) in two well-defined cortical projection neuron subtypes (ScPN and CPN) during early postnatal development. The authors develop and optimize an rRNA immunoprecipitation-mass spectrometry (rRNA IP-MS) workflow to recover RCs from FACS-purified, retrogradely labeled neurons, achieving remarkable subtype specificity and biochemical resolution. Through proteomic profiling, they reveal both shared and distinct ribosome-associated proteins between ScPN and CPN, with a focus on non-core RC components and their potential functional relevance. The work advances our understanding of cell-type-specific translation regulation, moving beyond the transcriptome to explore the proteome-level complexity in neuronal subtypes.

      Strengths:

      This work stands out for its technical sophistication and innovation. The authors combine retrograde labeling, FACS purification, and an optimized rRNA IP-MS approach (low input) to isolate ribosome-associated complexes from highly specific neuronal subtypes in vivo, a challenging issue that they execute with impressive rigor. The methodological pipeline is both elegant and well controlled, yielding high-quality, reproducible data. The depth of proteomic coverage is remarkable, with nearly all known cytoplasmic ribosomal proteins identified, along with hundreds of ribosome-associated proteins (RAPs), including translation factors, chaperones, and RNA-binding proteins. The analysis not only reveals shared components between ScPN and CPN RCs but also uncovers subtype-specific differences in associated proteins.

      Particularly notable is the integration of this new proteomic dataset with previously published transcriptomic and ribosome footprinting data, which helps to validate the specificity and relevance of the findings. Overall, the clarity of the writing, the robustness of the data, and the transparency of the methods make this a strong and compelling contribution.

      Weaknesses:

      Despite the depth and high quality of the dataset, the study remains descriptive. While the identification of subtype-specific RC components is intriguing, the current version of the manuscript does not explore their functional roles or biological consequences of their alterations. There is no perturbation, causal testing, in vitro or in vivo manipulation to demonstrate whether these proteins are necessary for ScPN or CPN identity, specific axonal targeting, metabolism or synaptic function.

      One important point that is also highlighted by the authors in their discussion and that is critical to establish the subtype specificity of the identified protein. One important point highlighted by the authors in the discussion - and critical for establishing the subtype specificity of the identified proteins-is that some ribosomal complexes may be specialized for specific developmental stages, rather than exclusively for the subtype-specific needs of projection neuron development. The work presented here provides a valuable starting point for further investigation into such RC specialization. However, it will be essential to determine to what extent these RCs exhibit true subtype specificity, independently of their temporal maturation context.

      As a result, key mechanistic insights remain a bit speculative. Although several of the identified proteins have known roles in processes like synaptogenesis or metabolism, their relevance to the specific neuronal subtypes under study is not experimentally addressed. That said, given its rich content and the comprehensive early postnatal dataset, the manuscript represents an extremely valuable resource for the community. While primarily exploratory, it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work provides a valuable toolkit for endogenous isolation of projection neuron subtypes. With further validation, it could present a solid method for low-input ribosome affinity purification using a ribosomal RNA (rRNA) antibody. The experimental evidence for the distinct ribosomal complexes is limited to this method and indirect support from complementary analyses of preexisting data.

      However, with additional experimental data to support the specificity of ribosomal complex pulldown and confirmation of the putative ribosomal complex proteins of interest, the study would provide compelling evidence for translation regulation of neuronal development through compositional ribosome heterogeneity. 

      This work would be of interest to neuroscientists, developmental biologists, and those studying translational networks underlying gene regulation.

      Strengths

      (1) This in vivo labeling of specific projection neurons and ribosomal rRNA affinity purification method accommodates a low input of <100K somata per replicate, which is useful for the study of neuronal subtypes with limited input. In principle, this set of techniques could work across different cell types with limited input, depending on the molecule used for cell type labeling.

      (2) The authors are also able to isolate endogenous neurons with minimal perturbation up to the point of collection, preserving the native state for the neuron in vivo as long as possible prior to processing. 

      (3) This study identified over a dozen potential non-ribosomal proteins associated with SCPN ribosomal complexes, as well as a ribosomal protein enriched in CPN.

      We appreciate the reviewer's thoughtful and detailed review. We especially appreciate the positive evaluation of its strengths, including the use of rRNA affinity purification to access ribosomal complexes in low-input neuronal subtypes in vivo with minimal perturbation, and the resulting identification of distinct ribosomal complexes in SCPN and CPN with associated non-ribosomal proteins. We are also pleased by the recognition of its significance in advancing our understanding of neuronal subtype-specific post-transcriptional gene regulation. We have carefully addressed the limitations below.

      Limitations

      (1) In this study, the authors address the advantages of their ribosomal complex isolation method in SCPN and CPN against RPL22-HA affinity purification. While this does show more pull-down of the ribosomal RNA by the Y10B rRNA antibody, the authors claim this method identifies cell-type-specific ribosomal complex proteins without demonstrating a positive control for the method's specificity. 

      There are very limited experiments to truly delineate how "specific" this method is working and whether there could be contamination from other complexes bound by the antibody. I see this as the major limitation that should be addressed. To boost their claims of capturing cell-typespecific ribosomal complexes, the authors could consider applying their rRNA affinity purification pipeline to compare cell types with well-characterized ribosome-associated proteins, like mouse embryonic stem cells and HELA cells.

      The reviewer can completely appreciate the elegance in the neural characterization here, but it seems there needs to be a solid foothold on the specificity of the method, perhaps facilitated by cell types that can be more readily scaled up and tested.

      We thank the reviewer for the opportunity to further clarify how our experimental design addresses the question of specificity of ribosomal complex pulldown. The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or non-specific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.

      (2) The authors followed up on their differentially enriched ribosomal complex proteins by analyzing the ribosome association of these proteins in external datasets. While this analysis supports the ribosome-association of these proteins, there is limited experimental validation of physical association with the ribosome, much less any functional characterization.

      The reciprocal pulldown of PRKCE is promising; however, I would recommend orthogonal validation of several putative ribosomal complex proteins to increase confidence. 

      Specifically, the authors could use sucrose gradient fractionation of SCPN and CPN, followed by a western blot to identify the putative interaction with the 80S monosome or polysomes. This would also provide evidence towards the pulldown capturing association with mature ribosome species, which is currently unclear. This experiment would provide substantial evidence for the direct association of these non-ribosomal proteins with subtype-specific ribosomal complexes.

      We thank the reviewer for the feedback and suggested future directions for candidate validation. We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.

      We appreciate the reviewer's suggestion of sucrose gradient fractionation for polysome profiling. We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph.

      (3) The authors state interest in learning more about the differences underlying translational regulation of projection neuron development. This method only captures neuronal somata, which will only capture ribosomes in the main cell body. There are also ribosomes regulating local translation in the axons, which may also play a critical role in axonal circuit establishment and activity. These ribosomal complex interactions may also be rather transient and difficult to capture at only one developmental stage. Therefore, this method is currently limited to a single developmental snapshot of ribosomal complexes at P3 within the main cell body. It would be exciting to see the extended utility of this method to sample neurites and additional developmental stages to gain further resolution on the developmental translation regulation of these projection neurons.

      We thank the reviewer for the opportunity to further highlight the foundational significance of this work in translational regulation of projection neuron development. Here, we identified subtype-specific differences in ribosomal complex composition within somata at a critical developmental time window for two PN subtypes. This work provides foundation for future investigation of local translation and its regulation in growth cones and axons. It will also enable direct comparison with potential future results from axons and growth cones, developmental subcellular specializations at axon tips that implement pathfinding and circuit formation (such work is not yet feasible due to exceptionally low available input). Our lab has significant ongoing work regarding subtype-specific axon and growth cone biology, including recent investigations of growth cone-localised RNA and protein molecular machinery that regulate circuit formation, maintenance, and function of distinct cardinal PN subtypes (Poulopoulos*, Murphy* et al. Nature 2019; Engmann*, Hatch* et al. Nature Prot 2022; Itoh et al. Cell Rep 2023; Veeraraghavan*, Engmann* et al. Nature Neurosci 2026; Durak*, Kim* et al. bioRxiv 2023; Veeraraghavan*, Tillman* et al. bioRxiv 2025; Tillman et al. bioRxiv 2026). Combining subtype-specific growth cone purification with ribosomal complex investigation represents a logical and exciting future direction. We appreciate the reviewer's encouragement of this future line of investigation and now highlight it in the revised Discussion.

      Likely impact of the work on the field, and the utility of the methods and data to the community:

      The authors introduce a unique pipeline of techniques to identify cell-type-specific ribosomal complex compositions. With more validation, there is certainly potential for those studying neuronal translation to leverage this method in limited primary cells as an alternative to existing methods that do not rely on ribosomal protein tagging, such as ARC-MS (Bartsch et al., 2023), RAPIDASH (Susanto and Hung et al., 2024), and RAPPL (Nature Communications, 2025).

      Reviewer #2 (Public review):

      Summary:

      This study presents a sophisticated molecular dissection of ribosome-associated complexes (RCs) in two well-defined cortical projection neuron subtypes (ScPN and CPN) during early postnatal development. 

      The authors develop and optimize an rRNA immunoprecipitation-mass spectrometry (rRNA IPMS) workflow to recover RCs from FACS-purified, retrogradely labeled neurons, achieving remarkable subtype specificity and biochemical resolution. Through proteomic profiling, they reveal both shared and distinct ribosome-associated proteins between ScPN and CPN, with a focus on non-core RC components and their potential functional relevance. The work advances our understanding of cell-type-specific translation regulation, moving beyond the transcriptome to explore the proteome-level complexity in neuronal subtypes.

      Strengths:

      This work stands out for its technical sophistication and innovation. The authors combine retrograde labeling, FACS purification, and an optimized rRNA IP-MS approach (low input) to isolate ribosome-associated complexes from highly specific neuronal subtypes in vivo, a challenging issue that they execute with impressive rigor. The methodological pipeline is both elegant and well-controlled, yielding high-quality, reproducible data. The depth of proteomic coverage is remarkable, with nearly all known cytoplasmic ribosomal proteins identified, along with hundreds of ribosome- associated proteins (RAPs), including translation factors, chaperones, and RNA-binding proteins. 

      The analysis not only reveals shared components between ScPN and CPN RCs but also uncovers subtype-specific differences in associated proteins. Particularly notable is the integration of this new proteomic dataset with previously published transcriptomic and ribosome footprinting data, which helps to validate the specificity and relevance of the findings. Overall, the clarity of the writing, the robustness of the data, and the transparency of the methods make this a strong and compelling contribution.

      Weaknesses:

      Despite the depth and high quality of the dataset, the study remains descriptive. While the identification of subtype-specific RC components is intriguing, the current version of the manuscript does not explore their functional roles or the biological consequences of their alterations. There is no perturbation, causal testing, in vitro or in vivo manipulation to demonstrate whether these proteins are necessary for ScPN or CPN identity, specific axonal targeting, metabolism, or synaptic function. One important point highlighted by the authors in the discussion - and critical for establishing the subtype specificity of the identified proteins - is that some ribosomal complexes may be specialized for specific developmental stages, rather than exclusively for the subtype-specific needs of projection neuron development. The work presented here provides a valuable starting point for further investigation into such RC specialization. 

      However, it will be essential to determine to what extent these RCs exhibit true subtype specificity, independently of their temporal maturation context. As a result, key mechanistic insights remain a bit speculative. Although several of the identified proteins have known roles in processes like synaptogenesis or metabolism, their relevance to the specific neuronal subtypes under study is not experimentally addressed. 

      That said, given its rich content and the comprehensive early postnatal dataset, the manuscript represents an extremely valuable resource for the community. While primarily exploratory, it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.

      We thank the reviewer for their excellent summary and for their very positive assessment of our work. We are pleased that the methodological rigor, proteomic depth, and integrative analyses were well-received. We thank the reviewer for highlighting that our “work presented here provides a valuable starting point for further investigation into such RC specialization” and that “it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.” This is exactly how we view this contribution – as a foundation to share with broader field so such functional investigation and investigation of developmental dynamics can be pursued by multiple groups in the broader related field.

      We again thank the reviewer for the very positive and insightful comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) As listed in the limitations, I would recommend that the authors consider applying their rRNA affinity purification to additional cell lines to confirm the specificity of the method as a positive control, where just demonstrating the technology may be easier to carry out than with more limited samples.

      As we noted in our response above to Limitation 1: “The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or nonspecific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.”

      (2) Also, as listed in the limitations, I recommend that the authors provide orthogonal experimental evidence for the putative SCPN and CPN ribosomal complex proteins of interest (e.g., sucrose gradient > western blot).

      As we noted in our response to Limitation 2 above: “We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.”

      Regarding sucrose gradient fractionation (for polysome profiling), we also noted in our response to Limitation 2: “We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph”.

      (3) The authors are interested in preserving the native state of the projection neurons to isolate ribosomal complexes; therefore, they may consider using biotin-conjugated CTB (Thermo Fisher) sorting through anti-biotin MACS columns (Miltenyi Biotech) as opposed to FACS in the future. While it does not allow for the same visualization as using a CTB-FP, this slight pipeline adjustment could help save time and physical processing of the projection neurons, helping preserve their endogenous state.

      While we appreciate the reviewer’s constructive suggestion for this theoretically alternative approach, we respectfully submit that this approach is unlikely to effectively isolate projection neurons from the living brain as effectively as FACS approach employed here. We considered this approach. We respectfully offer that CTB standardly enters neurons by binding GM1 gangliosides at the axon terminal, after which it is internalized and undergoes retrograde transport to the soma, several millimeters or more away depending on the projection. By the time CTB reaches the soma, we further respectfully offer that it is standardly fully internalized and is no longer surface-exposed for the theoretically suggested antibiotin capture for MACS. We used magnetic bead-based pull-down for the ribosomes themselves and find those molecular approaches very beneficial. Despite our use and openness to magnetic-conjugate-based separation approaches, we judge that fluorophore-conjugated CTB and FACS remain the most efficient and feasible approach for isolation of these exceptionally polarized projection neurons while maintaining cell viability. 

      (4) While the authors provided a comparison of overall protein detection levels between SCPN and CPN, I recommend an additional analysis and potential normalization for the average core ribosomal protein abundance across samples. I will note that this is more accessible when samples are prepared using TMT-labeling methods, which might be considered for future experiments.

      We thank the reviewer for these suggestions. We appreciate the opportunity to address them together, to further clarify our deeply considered choice of MS-based proteomic analysis and corresponding primary normalization approach. To complement our initial normalization approach, we have now also implemented the reviewer’s suggested approach of normalization to the average intensity of core ribosomal proteins. Notably, these new results agree with those from our initial normalization approach. This insightful suggestion has further strengthened the paper’s results and interpretation. 

      Here, we employed label-free quantification (LFQ), wherein samples are assayed sequentially rather than simultaneously, to most rigorously establish which proteins are truly present in some neuronal subtypes but absent in others. As the reviewer is aware, this capability to determine absence vs. presence of MS-detectable peptides distinguishes LFQ from approaches that assay samples simultaneously, such as isobaric tandem mass tag (TMT) labeling, which standardly offers advantages in relative quantification studies. Advances in sample preparation, instrumentation, and data analytical algorithms (from recent developments in single-cell proteomics and related approaches) now enable application of LFQ in quantitative differential analysis, even in the ultra-low-input regime. We now include both detailed discussion of these points and relevant citation in the text. 

      As the reviewer is also aware, in LFQ, normalization is crucial to ensure the protein quantification is accurate and comparable across sequential runs. For quantitative differential analysis of proteins detected in both subtypes, we have implemented a primary normalization approach employing “median-of-ratios” normalization across all detected proteins for robustness against outliers and technical variability. This primary approach results in similar overall distributions of protein abundances across CPN and SCPN samples (Figure S2A), providing confidence in quantitative comparison between SCPN and CPN in the ultra-lowinput regime. 

      Following the reviewer’s suggestion, to further ensure rigor of identification of differential proteins, we have also implemented a second normalization approach, rescaling each sample to the average intensity of its core ribosomal proteins alone (new Figure panels S2B, C). Subsequent differential analysis reveals equivalent CPN > SCPN enrichment of RPS30/eS30, GUCY1A1, and CELF3 (proteins identified as CPN-enriched with the primary normalization approach). These three proteins rank among the five proteins with lowest p-values, though false-discovery-rate-corrected significance is reduced. This confirmation by a second normalization approach further strengthens the findings.

      We have included clarifications in the main text, added the second normalization approach and subsequent analysis in both the main text and Figure S2. In addition, we have now noted in the discussion that TMT labeling with correspondingly appropriate normalization approaches might better define relative quantitative differences between functional candidates present in multiple subtypes.

      Recommendations for improving the writing and presentation:

      (1) I recommend this as a Tools or Resource article, seeing as the biological conclusions are limited.

      We respectfully submit that this work investigated biological questions and identified biological answers beyond pure development of Tools or offering a dataset as a Resource. We further respectfully submit that the question of differential neuronal subtype-specific translation of shared transcripts has become an emerging area of interest in regulation of precise neuronal and circuit development, maintenance, and function, as well as the neurobiological basis of disease. This has been quite hard to study, and this paper brings a first level of answers to that biological question. Of course, the biological results of this paper are not the complete answer, but as with all biological discovery papers, it provides a foundation for many further studies by multiple labs. 

      (2) In the rationale for studying ribosomal complex machinery, it may be helpful to say that ribosome composition and associated proteins that are present in the cytoplasm provide a way in which ribosomes can tune translation rapidly. This is especially important, seeing as this affords post-mitotic neurons the opportunity to remodel and repair by using readily available proteins while also avoiding the energetic demands of producing new ribosomal complex proteins.

      We thank the reviewer for this insightful comment and fully agree. We have now added this rationale to the Introduction. 

      (3) Is there a need for a CTB injection control? Does GM1 binding affect translation pathways? Please list citations, if possible.

      We respectfully submit that a CTB injection control is not necessary. As noted in our response to Limitation 1 in the Public Review portion, both PN subtypes underwent retrograde labeling with CTB in this work. While it remains unknown whether CTB-GM1 binding affects translation, potential effects of CTB are expected to apply equivalently to both subtypes and would therefore not confound these between-subtype comparisons. Please also see our response to recommendation 4 immediately below, in which we further clarify that retrograde tracing with CTB is a long-standing, well-accepted method shown to cause minimal damage and not interfere with continued neuronal development.

      (4) Do the traced/labeled PNs keep developing normally? In other words, does retrograde tracing inhibit proper PN development? Please list citations, if available.

      We thank the reviewer for encouraging us to further clarify that these are longstanding and well-accepted methods in the field, found to cause minimal damage and not to interfere with continued neuronal development. These and related retrograde labeling methods have led to the identification of the field's cardinal regulatory genes and molecules of axonal connectivity. This includes substantial work from our own lab (PMID in parentheses): Arlotta*, Molyneaux* et al. Neuron, 2005 (15664173); Lai*, Jabaudon* et al. Neuron, 2008 (18215621); Molyneaux*, Arlotta* et al. J. Neurosci., 2009 (19793993); Galazo et al. Neuron, 2016 (27321927); and more recently Sahni et al. Cell Rep., 2021a (34686320); and Sahni et al. Cell Rep., 2021b (34686337). These methods have also employed by other groups, such as Bin Chen (e.g. McKenna et al. PNAS, 2015 (26324926)) and Marta Nieto (e.g. De León Reyes et al. Nat. Commun., 2019 (31591398)). We have now clarified this in the text and included references for the benefit of the readers.

      (5) It may be helpful to mention that RPS30 associates with immature ribosomes during biogenesis (PMID: 25706898).

      We thank the reviewer for the opportunity to further clarify background knowledge of RPS30. As the reviewer is aware, RPS30/eS30 is definitively part of the mature 80S ribosome, as established, e.g., by cryo-EM of human 80S ribosomes (PMID: 25901680). Like multiple other ribosomal proteins, RPS30/eS30 also associates with immature ribosomes during biogenesis (PMID: 25706898). Intriguingly, RPS30/eS30 is produced by cleavage of a fusion protein comprising ubiquitin-like FUBI and RPS30/eS30, with cleavage recently identified as a late step in cytoplasmic 40S maturation (PMID: 34318747). As noted in the text, we confirmed that the peptides used for RPS30/eS30 identification appropriately map only to the amino acid sequence of RPS30/eS30 and not FUBI. We now mention that it is a core component of the mature 80S ribosome and its immature ribosomal association. 

      (6) Please clarify how the rRNA-IP is pulling down mature ribosomes. If not, this should be incorporated into the discussion.

      We thank the reviewer for raising this interesting point. As the reviewer notes, it is well established in the ribogenesis field that ribosomes are continuously produced and therefore exist at various stages of maturation. Our protocol removes a major source of immature ribosomes by subjecting FACS-purified cells to two centrifugal spins that remove the nucleus, the site of ribogenesis and early steps of maturation. In pilot experiments, nuclear removal was confirmed by the absence of a contaminating genomic DNA peak on Bioanalyzer electropherograms of total RNA extracted from input samples (without genomic DNA removal) immediately prior to rRNA-IP. However, ribosomes also undergo cytoplasmic maturation steps, and various functional states of ribosomes have been found to be present in the cytoplasmic fraction. For these reasons, we have referred to what we pulled down as "ribosomal complexes" throughout the manuscript. We now explain this nuclear/immature ribosome depletion and cytoplasmic ribosomal enrichment in the text.

      (7) Would have been interested to see some discussion of the most enriched CPN RAPs or why these might not exist in most replicates (inter-subtype heterogeneity?)

      The reviewer asks an interesting question. As the reviewer is aware, when considering a single sample in isolation, absence of mass spectrometry-based detection is not definitive proof of absence, especially not within this work’s ultralow input regime. One of us (B. Budnik) has substantial experience with ultra-low-input samples across multiple cell types outside of the nervous system, and identifying a protein in three of four identical samples is not uncommon. We therefore used detection in three or four samples as an indicator of presence, while absence across all samples was taken to indicate true absence.

      That said, the reviewer is correct that further diversity and heterogeneity within both CPN and SCPN subtypes additionally might be involved. This is an interesting question for future research, and we have added relevant text to the Discussion. 

      (8) It may be helpful to mention that RPL22, while not stoichiometric, is known to have extraribosomal functions and can pull down independently from assembled ribosomes (PMID: 17381311, PMID: 28575669. Moreover, RPL22/eL22-3xFLAG has been previously used as a control for ribosome affinity-based pulldowns and could be added to citations for Figure S1 (PMID: 28625553, PMID: 28625553).

      We thank the reviewer for this helpful recommendation. We have added relevant text and citations. 

      Minor corrections to the text and figures:

      (1) Want to confirm that in Figure 2, P adj is <0.1 is correct? 

      Yes

      (2) It is not necessary to show MS spectra in Figure 2.

      While we understand the spectra are not strictly necessary, we respectfully submit that they enhance the figure and aid readers in assessing data quality. 

      Reviewer #2 (Recommendations for the authors):

      To strengthen the impact and interpretation of the authors' findings, we encourage consideration of the addition of functional validation experiments for at least one (or more) of the ribosomeassociated proteins that are differentially enriched in ScPN. This could include genetic manipulation (e.g., knockdown or overexpression) to test whether these proteins influence subtype-specific features or neuronal function. Even a limited set of perturbation experiments, such as targeting PRKCE, which is particularly interesting due to its known role in synaptogenesis, would help move the study from descriptive to a more mechanistic nature of the work.

      There appear to be no issues related to data availability, ethics, or compliance, assuming all raw proteomic data and associated code for differential analysis are made publicly available.

      We again thank the reviewer for this encouragement and highlighting that our work provides the foundation for further functional and developmental dynamic investigations. We view this work as providing that foundation for further investigation by multiple labs in the broader fields, beyond the scope of this paper.

    1. eLife Assessment

      This important study addresses how listeners learn the statistical properties of acoustic spaces, combining well-designed psychophysics in virtual rooms with non-invasive brain stimulation. The evidence is solid: the behavioral results are strong and show that speech understanding is best in rooms with everyday levels of reverberation, and the stimulation data are consistent with a contribution of dorsolateral prefrontal cortex, though the effects are modest and the spatial precision of TMS is inherently limited. The work will be of interest to researchers in auditory neuroscience and psychoacoustics, helping to understand how the brain copes with reverberant environments.

    2. Reviewer #3 (Public review):

      Summary:

      This manuscript presents a well-designed study examining human adaptation to room acoustics, building on prior work. The psychophysical results are convincing and add meaningful knowledge to our understanding of reverberation learning. The transcranial magnetic stimulation (TMS) component shows a role of prefrontal cortex in this listening task, targeting dorsolateral prefrontal cortex (dlPFC). Cautious interpretation of the TMS results is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, and the limited ability of TMS to precisely target dlPFC in individual subjects. A surprising and interesting finding in the study is that listeners performed the speech recognition task more poorly in anechoic conditions than in those with naturalistic levels of reverberation. This is likely due to contributions of spatial release from masking provided by reverberation acoustics, which may counteract the detrimental effects of reverberation on speech perception itself. Overall, the experiments are well performed and clearly presented, improving our understanding of how the brain copes with reverberant environments.

      Strengths:

      (1) Well designed acoustical stimuli and psychophysical task.

      (2) Comparisons across room combinations is well conducted.

      (3) Virtual acoustic environment is impressive and applied well here.

      (4) Timely study with interesting behavioural results.

      (5) Causal evidence of a role for dlPFC in reverberation learning.

      Weaknesses:

      (1) Poorer performance in anechoic environments than some reverberant environments suggests and interplay of spatial release from masking and speech intelligibility that are not fully unpicked here. This could be controlled in future experiments, for example by comparing monaural and binaural listening conditions.

      (2) Lack of evidence for targeting TMS to dlPFC in individual participants. This is simply a limitation of the technique which the reader should keep in mind.

      (3) Most interesting effect of TMS is a null result compared to a weak statistical effect for "meta-adaptation"

    3. Reviewer #4 (Public review):

      The authors use a d' defined for 2-alternative forced choice experiments, but their data are 4-alternative (for color) and 8-alternative (for number) forced-choice. So, the d' is not computed correctly. For mAFC experiments, the authors should use the Hacker-Ratcliff (1979) method, also defined in chapter 10 of the Macmillan & Creelman textbook.

      Normalization of the stimuli was arbitrary, and consequently the unexpected improvement in performance in reverberation compared to anechoic condition is still not explained. A natural normalization across different environments is to take the direct portion of the BRIR (and HRTF for the anechoic condition) and make sure that that is scaled identically across the different simulated rooms (with the reverberant tails scaled naturally). That corresponds to the situation when the sources are emitting the sound at the same level in each environment. The current study scaled the overall levels. As a minimum, it should be reported how this scaling boosted/attenuated the targets and maskers in each environment.

      The potential that the listeners are tuning to individual voices, as opposed to rooms, has not been eliminated. The authors suggest that a lack of interaction with different voices is evidence that that is not the case. This is not correct: lack of this interaction just means that there are no differences in tuning between the voices. But that does not mean that the same amount of tuning is happening for each voice, as observed in previous studies. Unless the authors provide a follow-up data with randomly varying voices within each trial, these claims should be tuned down.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of an experiment that demonstrates a disruption in statistical learning of room acoustics when transcranial magnetic stimulation (TMS) is applied to the dorsolateral prefrontal cortex in human listeners. The work uses a testing paradigm designed by the Zahorik group that has shown improvement in speech understanding as a function of listening exposure time in a room, presumably through a mechanism of statistical learning. The manuscript is comprehensive and clear, with detailed figures that show key results. Overall, this work provides an explanation for the mechanisms that support such statistical learning of room acoustics and, therefore, represents a major advancement for the field.

      Strengths:

      The primary strength of the work is its simple and clear result, that the dorsolateral prefrontal cortex is involved in human room acoustic learning.

      Weaknesses:

      A potential weakness of this work is that the manuscript is quite lengthy and complex.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how listeners adapt to and utilize statistical properties of different acoustic spaces to improve speech perception. The researchers used repetitive TMS to perturb neural activity in DLPFC, inhibiting statistical learning compared to sham conditions. The authors also identified the most effective room types for the effective use of reverberations in speech in noise perception, with regular human-built environments bringing greater benefits than modified rooms with lower or higher reverberation times.

      Strengths:

      The introduction and discussion sections of the paper are very interesting and highlight the importance of the current study, particularly with regard to the use of ecologically valid stimuli in investigating statistical learning. However, they could be condensed into parts. TMS parameters and task conditions were well-considered and clearly explained.

      Weaknesses

      (1) The Results section is difficult to follow and includes a lot of detail, which could be removed. As such, it presents as confusing and speculative at times.

      (2) The hypotheses for the study are not clearly stated.

      (3) Multiple statistical models are implemented without correcting the alpha value. This leaves the analyses vulnerable to Type I errors.

      (4) It is confusing to understand how many discrete experiments are included in the study as a whole, and how many participants are involved in each experiment.

      (5) The TMS study is significantly underpowered and not robust. Sample size calculations need further explanation (effect sizes appear to be based on behavioural studies?). I would caution an exploratory presentation of these data, and calculate a posteriori the full sample size based on effect sizes observed in the TMS data.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a well-designed and insightful behavioural study examining human adaptation to room acoustics, building on prior work by Brandewie & Zahorik. The psychophysical results are convincing and add incremental but meaningful knowledge to our understanding of reverberation learning. However, I find the transcranial magnetic stimulation (TMS) component to be over-interpreted. The TMS protocol, while interesting, lacks sufficient anatomical specificity and mechanistic explanation to support the strong claims made regarding a unique role of the dorsolateral prefrontal cortex (dlPFC) in this learning process. More cautious interpretation is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, the imprecise targeting of dlPFC (which is not validated), and the lack of knowledge about the timescale of TMS effects in relation to the behavioural task. I recommend revising the manuscript to shift emphasis toward the stronger behavioural findings and to present a more measured and transparent discussion of the TMS results and their limitations.

      Strengths:

      (1) Well-designed acoustical stimuli and psychophysical task.

      (2) Comparisons across room combinations are well conducted.

      (3) The virtual acoustic environment is impressive and applied well here.

      (4) A timely study with interesting behavioural results.

      Weaknesses:

      (1) Lack of hypotheses, particularly for TMS.

      (2) Lack of evidence for targeting TMS in [brain] space and time.

      (3) The most interesting effect of TMS is a null result compared to a weak statistical effect for "meta adaptation"

      Reviewer #4 (Public review):

      Summary:

      Several behavioral experiments and one TMS experiment were performed to examine adaptation to room reverberation for speech intelligibility in noise. This is an important topic that has been extensively studied by several groups over the years. And the study is unique in that it examines one candidate brain area, dlPFC, potentially involved in this learning, and finds that disrupting this area by TMS results in a reduction in the learning. The behavioral conditions are in many ways similar to previous studies. However, they find results that do not match previous results (e.g., performance in anechoic condition is worse than in reverberation), making it difficult to assess the validity of the methods used. One unique aspect of the behavioral experiments is that Ambisonics was used to simulate the spaces, while headphone simulation was mostly used previously. The main behavioral experiment was performed by interleaving 3 different rooms and measuring speech intelligibility as a function of the number of words preceding the target in a given room on a given trial. The findings are that performance improves on the time scale of seconds (as the number of words preceding the target increases), but also on a much larger time scale of tens to hundreds of seconds (corresponding to multiple trials), while for some listeners it is degraded for the first couple of trials. The study also finds that the performance is best in the room that matches the T60 most commonly observed in everyday environments. These are potentially interesting results. However, there are issues with the design of the study and analysis methods that make it difficult to verify the conclusions based on the data.

      Strengths:

      (1) Analysis of the adaptation to reverberation on multiple time scales, for multiple reverberant and anechoic environments, and also considering contextual effects of one environment interleaved with the other two environments.

      (2) TMS experiment showing reduction of some of the learning effects by temporarily disabling the dlPFC.

      Weaknesses:

      While the study examines the adaptation for different carrier lengths, it keeps multiple characteristics (mainly talker voice and location) fixed in addition to reverberation. Therefore, it is possible that the subjects adapt to other aspects of the stimuli, not just to reverberation. A condition in which only reverberation would switch for the target would allow the authors to separate these confounding alternatives. Now, the authors try to address the concerns by indirect evidence/analyses. However, the evidence provided does not appear sufficient.

      The authors use terms that are either not defined or that seem to be defined incorrectly. The main issue then is the results, which are based on analysis of what the authors call d', Hit Rate, and Final Hit rate. First of all, they randomly switch between these measures. Second, it's not clear how they define them, given that their responses are either 4-alternative or 8-alternative forced choice. d', Hit Rate, and False Alarm Rate are defined in Signal detection theory for the detection of the presence of a target. It can be easily extended to a 2-alternative forced choice. But how does one define a Hit, and, in particular, a False Alarm, in a 4/8-alternative? The authors do not state how they did it, and without that, the computation of d' based on HR and FAR is dubious. Also, what the authors call Hit Rate, is presumably the percent correct performance (PCC), but even that is not clear. Then they use FHR and act as if this was the asymptotic value of their HR, even though in many conditions their learning has not ended, and randomly define a variable of +-10 from FHR, which must produce different results depending on whether the asymptote was reached or not. Other examples of usage of strange usage of terms: they talk about "global likelihood learning" (L426) without a definition or a reference, or about "cumulative hit rate" (L1738), where it is not clear to me what "cumulative" means there.

      There are not enough acoustic details about the stimuli. The authors find that reverberant performance is overall better than anechoic in 2 rooms. This goes contrary to previous results. And the authors do not provide enough acoustic details to establish that this is not an artefact of how the stimuli were normalized (e.g., what were the total signal and noise levels at the two ears in the anechoic and reverberant conditions?).

      There are some concerns about the use of statistics. For example, the authors perform two-way ANOVA (L724-728) in which one factor is room, but that factor does not have the same 3 levels across the two levels of the other factor. Also, in some comparisons, they randomly select 11 out of 22 subjects even though appropriate test correct for such imbalances without adding additional randomness of whether the 11 selected subjects happened to be the good or the bad ones.

      Details of the experiments are not sufficiently described in the methods (L194-205) to be able to follow what was done. It should be stated that 1 main experiment was performed using 3 rooms, and that 3 follow-ups were done on a new set of subjects, each with the room swapped.

      We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and insightful comments. We greatly appreciate the time and expertise invested in reviewing our work. The feedback has been invaluable in improving the clarity, rigor, and overall presentation of the manuscript.

      In response to the reviewers’ comments, we have carefully revised the manuscript throughout. The revisions include clarification of the study hypotheses, re-analysis of the TMS data using a mixed ANOVA framework, additional methodological details regarding the TMS procedures and behavioural analyses, expanded justification of the statistical modelling approach, clarification of the room-acoustics paradigm, revision of figures and figure legends, additional discussion of study limitations, and a more balanced interpretation of the TMS findings. We have also substantially revised the Discussion section and improved the overall structure and readability of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is understood that this topical area is necessarily detail-heavy, but if there are ways to streamline the manuscript to more quickly arrive at key results (Figure 3?), the work might have an even greater overall impact.

      We appreciate the reviewer’s feedback and have carefully revised the manuscript to address all comments from all reviewers. However, we have retained the existing order of figures and results to maintain consistency and avoid extensive structural changes that could compromise the clarity and flow of the manuscript.

      (2) Minor point: I believe Equation 1 should be d' = z(H) – z(F).

      Thanks for noticing this, we have fixed the equation.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 115: Towards the end of the introduction, the hypotheses for the current study remain unclear. Please explicitly outline each hypothesis before the Methods section.

      We have now specified the hypotheses tested at the end of the Introduction.

      (2) Line 229: Please state the minimum MEP amplitude criterion used during TMS thresholding (usually 50 µV). Please also reference the EMG hardware and software used to measure MEPs.

      We are grateful to the reviewer for noticing that a few details regarding our TMS procedure were missing in the “Continuous theta-burst stimulation” section of the Methods. We have extensively revised this section to include the required details. We did not, however, use MEP amplitude as a criterion for estimating motor thresholds. For single-pulse TMS-induced motor threshold determination, we used visual observation of the first dorsal interosseous (FDI) muscle twitch (Varnava et al., 2011).

      (3) Line 248: The Coordinate Response Measure corpus contains combinations of callsigns, colours, and numbers. What was the rationale in asking participants to only identify the colour and the number spoken, and not the callsign?

      The rationale for reporting only Color and Number was to ensure an equal number of keyword identifications across phrase lengths—for example, CP0 and CP1 do not contain a callsign. This has now been clarified in the Procedure section of the Methods.

      (4) Lines 327–334: The analyses outlined here are unclear – it appears that there are multiple different statistical tests being conducted on the same outcome variable in this study. Given that the design includes a between-subjects factor (TMS condition) and two within-subjects factors (Rooms, Carrier Length Phrase), could the analyses be simplified by employing a mixed ANOVA as opposed to separate repeated-measures and univariate ANOVAs? If this is, in fact, the analysis that was conducted, please improve the wording. Clarification/correction is also needed surrounding the use of the term “univariate” ANOVA, as this can refer to different statistical tests.

      We appreciate the reviewer pointing this out. We have re-analysed the TMS data using a mixed ANOVA design and reported it as such in the Results section. Although the numerical values have changed, the significant findings and conclusions remain the same.

      (5) Line 340: Whilst the authors identify that an alpha value of 0.05 and Bonferroni corrections were used for statistical inference in two-tailed t-tests, there is no indication of the alpha value used in the interpretation of the ANOVA results. Please include this before the Results section. Furthermore, given that multiple tests are being conducted in this project, the alpha value used to infer statistical significance should be corrected in accordance with the number of hypotheses, to reduce Type I error rates (e.g., 4 hypotheses would result in an alpha inference criterion of α = 0.0125).

      We appreciate the reviewer’s observation. The corrected alpha value was not explicitly reported because all statistical analyses were conducted using IBM SPSS Statistics for Windows, Version 29.0.2.0 (IBM Corp., Armonk, NY; RRID: SCR_002865). In SPSS, Bonferroni corrections are applied by adjusting the p-values rather than the alpha threshold itself. For transparency, we have now clarified in the manuscript the factors included in each statistical comparison to make the tested hypotheses fully explicit.

      (6) Line 346: The sample size calculation used in the current study could be improved. It is unclear why a sample size estimation of ≥18 is used when the actual sample recruited is significantly greater than this (62). Is this because 62 participants were divided across multiple experiments in this paper? Further clarification is needed. The alpha value used in this calculation should also be corrected to account for multiple statistical models.

      We thank the reviewer for noting this. We have clarified that the initial sample size estimation (n = 18) referred to individual ANOVA analyses. In the revised manuscript, we specify in each experimental section (Identity of Sound Environments and Continuous Theta Stimulation) the exact number of participants recruited per experiment, which together sum to a total of 74 participants across all experiments (also clarified in the Participants section of the Materials and Methods).

      (8) Line 438: Many of the statistical tests presented in the Results section have not been outlined in the Methods section or had the rationale explained in the Introduction – this makes the analyses feel confusing and speculative. Explicitly identifying the core hypotheses earlier in the manuscript and clearly stating which hypotheses are confirmatory or exploratory would be an important improvement.

      We appreciate the reviewer noticing this. We have clarified the hypotheses tested in the Introduction (4th and 5th paragraphs) and provided a more detailed account of the statistical analyses in the Materials and Methods (Statistical Analysis section).

      (9) Line 565: It is unclear how the 62 participants recruited in the study were divided across each of the experimental conditions. I would recommend expanding the Participants section in Methods to outline the number of participants involved in each stage of the study.

      Thanks for noticing this discrepancy—this calculation was indeed confusing and incorrect. We ultimately tested a total of 74 participants. We have specified the sample size per experiment in both the “Identity of selected sound environments” and “Continuous theta-burst stimulation” sections of the Methods, and reiterated this in the Results section to prevent confusion.

      (9) Line 792: The Results section as a whole is incredibly complex and lacks structure. This could be condensed significantly – details regarding previous research should be removed from Results, as this should already be outlined in the Introduction as rationale for the current project. It would be clearer to outline each confirmatory and exploratory hypothesis in the Introduction, then identify at each stage in the Results section where it is being tested.

      We appreciate the reviewer’s point. We have rewritten the 4th and 5th paragraphs of the Introduction to outline the confirmatory and exploratory hypotheses. While we have streamlined parts of the Results section to enhance clarity, we have retained contextual information for each analysis to help readers follow the logic of the findings, given the complexity and scope of the study.

      Reviewer #3 (Recommendations for the authors):

      (1) Introduction – Overinterpretation and Hypothesis Clarity. The final paragraph of the Introduction (page 4) discusses the experimental findings and their interpretation, which belongs in the Discussion. This section should instead clearly state hypotheses for both the behavioural and TMS experiments. In particular, the TMS experiment lacks a clear rationale: what mechanism is being tested, and what behavioural outcome is predicted? Please revise this section to focus on the theoretical motivation, clearly defined hypotheses, and expected results, and less on summarizing the results.

      We appreciate this point and have accordingly deleted the final paragraph of the Introduction, as it is already incorporated into the Discussion section. We have stated the hypotheses/rationale more clearly in the Introduction.

      (2) TMS – Mechanism, Timeframe, and Clarity. The manuscript does not adequately explain how TMS produces long-lasting effects relevant to the task, which occur minutes (or possibly longer) after stimulation.

      (a) What is the specific timeframe between stimulation and behavioural testing?

      The timeframe between stimulation and behavioural testing has been clarified in the Methods section, e.g.: “Participants exposed to ‘real’ or ‘sham’ TMS completed the familiarization and behavioural task right after cTBS procedures.”

      (b) What is the evidence that TMS to the prefrontal cortex affects function on this timescale?

      We have added the following paragraph to the Discussion to address the timescale of continuous theta-burst stimulation (cTBS): cTBS, as used in our study, typically induces aftereffects lasting 20–50 minutes (Huang et al., 2005; Wischnewski & Schutter, 2015). While these effects are well established in the motor cortex—with motor-evoked potential changes persisting for up to one hour—recent evidence suggests that similar durations of cortical modulation can also occur in the prefrontal cortex (Taylor et al., 2025). Specifically, studies applying inhibitory rTMS to the dorsolateral prefrontal cortex (dlPFC) during cognitive tasks have demonstrated functional effects lasting up to one hour in healthy participants (Wagner et al., 2006). Furthermore, Tupak et al. (2013) showed that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels, reflecting decreased cortical activity, for at least 45 minutes—the same duration as the experimental task in our study. Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), the fNIRS-measured alterations in local cerebral blood oxygenation provide an indirect but reliable indicator of TMS-induced neural modulation within this timescale.

      (c) Can post-stimulation effects be objectively measured or confirmed?

      Although no objective post-stimulation measures were collected in the present study, we acknowledge this as a limitation. However, previous research has shown that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels—reflecting decreased cortical activity—for at least 45 minutes (Tupak et al., 2013). Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), these findings indicate that fNIRS can serve as an indirect but reliable method for confirming TMS-induced neural modulation. We plan to incorporate such objective measures in future studies.

      We thank the reviewer for raising these important points and have addressed them by updating the Methods and Results sections and adding a paragraph to the Discussion. We would like to clarify, however, that the TMS effects observed in our study are not weak: the statistically significant differences between sham and TMS conditions were accompanied by large effect sizes, indicating that bilateral inhibitory stimulation of the dlPFC produced a robust and consistent effect across participants—specifically, a reduction in performance, reflecting decreased improvement in speech understanding with increasing exposure to the reverberant environment.

      (3) Figure 3A – Anatomical Specificity and Interpretation. Figure 3A implies precise stimulation of dlPFC and its projections to auditory cortex (A1), but the authors cannot actually target dlPFC or its connections specifically with this approach. Rather, the TMS protocol disrupts an undetermined region of PFC, with diffuse downstream effects. This should be clearly acknowledged in the figure legend and main text.

      We appreciate this comment and have acknowledged this point in the figure legend and Discussion, e.g., Figure 3 legend: “Although the TMS protocol was intended to target the dlPFC, it likely affected adjacent prefrontal regions, leading to diffuse downstream effects that may have included modulation of A1.”

      (4) Additionally, the Discussion overstates the evidence for a specific dlPFC → AC role in reverberation learning. The weak and poorly localized TMS effect does not support strong claims about this pathway. Please scale back this interpretation and focus more on the robust psychophysical results, which are the manuscript's stronger contribution.

      We thank the reviewer for this comment. We have revised our interpretation to clarify that the proposed dlPFC–auditory cortex link is speculative, and have added caveats regarding the limited spatial precision of TMS targeting and individual variability in its effects, in the Discussion section “A role for dlPFC in statistical learning of room acoustics.”

      (4a) Line 152: Extra comma after “of”? Also, why are there square brackets around “callsigns” etc.?

      Fixed.

      (4b) Line 156: “RRID:SCR_001622” is unexplained and likely unnecessary—consider removing.

      Removed.

      (4c) Line 159: Why was no ramping applied at the end of the noise? Please clarify.

      Similar to Brandewie & Zahorik (2013), no ramping was applied. This has been clarified in the Methods (“Acoustic Stimuli”).

      (4d) Line 207: Methods do not describe the sham TMS protocol—please add this information.

      Thanks for noticing this. Information on the sham TMS protocol has been added.

      (4e) Line 352: Fix bracket formatting.

      Fixed.

      (4f) Line 382: Sentence is grammatically incorrect—please revise.

      Fixed.

      (4g) Line 394: Unclear use of square brackets—clarify or standardize.

      Fixed.

      (4h) Line 401: It is unclear how interleaving the talker and length ensures a different room each trial. Aren't these variables independent?

      The reviewer is correct: the only variable that was pseudorandomized was room order, to prevent carry-over effects, similar to Brandewie & Zahorik (2013). This has been clarified in the Methods section.

      (4i) Figure 1: Clarify that AI-generated images refer only to the room images, not other components.

      Fixed.

      (4j) Figure 2A: Confirm that “overall” includes all speakers and durations—clarify in legend.

      Fixed.

      (4k) The interesting duration effects in Figure 1D are not discussed in the text and appear before overall room effects in Figure 2A—please reorder and comment on these results.

      Fixed.

      (4l) Supplementary Figure 2: Caption contains a typo (“Lecture Room/Open-Plan Office”).

      Typo has been fixed.

      Also, I recommend adding this result to the main figure set—e.g., include overall d′ for all six talkers in Figure 2 alongside rooms (2A) and lengths (2D).

      We thank the reviewer for the suggestion. Including overall performance for all six talkers in Figure 2 would require substantial restructuring and risk making the figure crowded. We have therefore retained these results in the Supplementary Materials (Supplementary 2 and 3), as originally presented.

      (4m) Line 555: Phrase “to better understand” could be clearer—consider rewording.

      This section has been reworded.

      (4n) Lines 587–592: The lack of main effect of TMS is helpful, but more important is whether interactions between TMS and room/length variables occur. Please report these interactions, as they are central to interpreting the TMS effects.

      We appreciate the reviewer highlighting this. We have reviewed this analysis and reported the interaction Condition x CP length as follows: “A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], and no significant interaction Condition x CP length was observed: [F (3,93) =0.48, p=0.69, ŋp2 = 0.01]; confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population”. 

      (4o) Line 601: Reiterate the timeframe of the TMS-behaviour gap. Is there supporting evidence that TMS can affect behaviour over this duration? Could null effects reflect fading TMS efficacy?

      We appreciate the reviewer pointing this out. We have clarified the timeframe of TMS stimulation in both the Methods and Results, e.g.: “The procedure began with TMS manipulation, and although the behavioural task lasted 45 minutes, the inhibitory effects of TMS extended for at least 60 minutes post-stimulation (Huang et al., 2005; Gamboa et al., 2010; Hoogendam, Ramakers, & Di Lazzaro, 2010; Romero et al., 2022).” However, we cannot dismiss individual differences in the duration of TMS effects, nor differences in efficacy duration between anatomical areas (motor cortex vs. dlPFC). We have noted this in the Discussion section “A role for dlPFC in statistical learning of room acoustics” (Pallant, 2011).

      (4p) Figure 3I: The “meta-adaptation” effect is marginal in both Exp 1 (p = 0.03) and Exp 2 sham (p = 0.04). These should be interpreted cautiously, given their statistical fragility.

      We appreciate this comment. We have now calculated effect sizes for all Wilcoxon signed-rank tests (Pearson’s r) and report them. For the two comparisons noted by the reviewer, the effect sizes are medium (Exp 1) and large (Exp 2). We are therefore confident that, even where the p-values are not extremely low, the statistical differences are reliable.

      (4q) Line 696: Reverberation is described as “common,” but it is nearly universal. Consider rephrasing to reflect this.

      We appreciate this suggestion and have rephrased this line.

      (4r) Line 816: The authors state that TMS reduced overall performance, but the earlier ANOVA (lines 587–592) shows no such effect. Please correct this discrepancy.

      We appreciate the reviewer noticing this. This section has been clarified: the lack of statistical significance at lines 587–592 relates to the comparison between a subset of ‘no-TMS-exposed’ listeners and ‘sham’-TMS-exposed listeners, made only to demonstrate the absence of placebo effects in the sham sample. Following the reviewers’ suggestions, we also re-analysed the data using a one-way ANOVA; this slightly changed the numerical values of the reported main effect but did not change the statistical outcome.

      Reviewer #4 (Recommendations for the authors):

      (1) Lines 201–202: It's not clear what is meant by combination and by carrier length here.

      This section has been rewritten for clarity.

      (2) Line 330: What is meant by “Univariate” here? I think this was a mixed ANOVA, with a betweensubject factor of TMS exposure and the remaining factors within-subject.

      We appreciate this suggestion; we have re-analysed this section to use a mixed-ANOVA design. The numerical results differ, but the statistical outcome remains the same.

      (3) Lines 335–343: This is impossible to follow if one does not understand that there were 3 followup experiments.

      Thank you for highlighting that this section was confusing. We have rewritten it to clarify the following: “Three follow-up experiments were performed (univariate ANOVA) to assess whether speech understanding was affected by room context (i.e., the third room in which Open-Plan Office and Underground Car Park were learnt), with one between-subjects factor: room context (levels: Anechoic Room, Living Room, Lecture Room, and Highly Reflectant Room).” We have also added the following earlier in the Methods: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).”

      (4) Lines 414–416: The review of Tsironis et al. (2024) (doi:10.1177/23312165241273399) does not provide strong evidence that there are multiple scales for adaptation to room reverberation (most adaptation effects stabilize within 1 sec).

      We apologize for this mistake, which arose from an issue with our reference manager. It has been corrected to: Robinson, Harper, & McAlpine (2016), Nature Communications, and Simpson, Harper, Reiss, & McAlpine (2014), Journal of Neuroscience.

      (5) Line 427: The supplementary figure shows that many subjects did not achieve asymptotic performance.

      We appreciate the reviewer pointing this out. The fittings have been extensively reviewed; please see the Methods and Results for the new fitting analysis. Indeed, some participants, although very slowly, keep improving over time without reaching clearly asymptotic behaviour. This section, however, referred specifically to the point at which performance stabilised within ±10% of final performance.

      (6) Line 428: The ±10% statistic is random (as discussed below). And why switch to HR now? And what is its meaning when the HR did not converge by the end of the run?

      We appreciate the reviewer raising the inconsistent use of d′ versus HR. d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis; time-course analysis does not allow us to calculate d′ at each trial or time point, owing to the lack of HR and FA values for single trials. We have clarified this in the Methods: “Given how d′ was calculated for |Color| and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data for each |Color| and |Number| affecting the temporal resolution of any generated curve. We therefore analysed the development of individual and average performance in each acoustic environment using cumulative hit rates, applying a 5-point moving average (~7 seconds) to each trace and plotting performance as a function of mean cumulative exposure time.”

      We have also extensively reviewed our fits, following these steps: (i) we compared single- vs. double-exponential fits across 22 participants (Supplementary Figure 1), which showed that double exponentials provided a better R<sup>2</sup> for the majority of participants; (ii) we re-ran all analyses forcing the fits to the HR endpoint; (iii) characterizing taus was not informative for our sample, given flat-like performance for some participants (Supplementary Figure 1)—in these cases, taus do not aid understanding of how performance stabilizes over time, particularly given the use of two taus; and (iv) cutting initial points differs by participant.

      (7) Line 434: Or that they learned/adapted to other characteristics that were fixed.

      Thank you—we have added “adapted to” in the sentence.

      (8) Supplementary tables often show differences, but the actual values are not shown. Also, the tables and figures randomly switch between d' and HR.

      Supplementary tables are intended only to show additional detail not reported in the main text or figures, to avoid redundancy; means (referred to in the Supplementary tables) are always shown in the main figures. We appreciate the reviewer raising the inconsistent use of d′ versus HR, addressed above, and have clarified this throughout the Methods and Results.

      (9) Lines 436–442: There seem to be a lot of issues with the fitting shown in Supplemental Figure 1 and Figure 2B:

      (a) It does not seem to converge, especially for the green line. So, presumably, the asymptotic value obtained for tau_slow was the upper bound set to 2000 s for many subjects' conditions. But those values are never shown—they should be in Supplemental Figure 1.

      (b) Then the FHR value, derived from that, is completely dependent on what the bound was set to, and is therefore arbitrary. And its value of 10% is also arbitrary. Why do this when tau itself of an exponential model represents the time it takes to reach 67% of the asymptotic value, from which one can derive whatever time it should take to reach the final 10%?

      (c) Even the use of the model specified by Equation 2 seems arbitrary. Average data in Figure 2B do not provide strong evidence for two time scales. If the authors are worried about the instability of the data at the beginning, a simple exponential with a weighted fit that prioritizes the later portions seems sufficient.

      (d) Lines 430–434: This conclusion seems wrong, based only on the arbitrary measure chosen for “global likelihood learning.” Looking at Figure 2B, there is no evidence that the green graph reached any asymptote, while for the yellow and blue it appears to have. The authors should try fitting a simple exponential function to it to show that tau is larger.

      We appreciate the reviewer’s comments and have significantly revised these sections of the Methods and Results. In summary, we fitted the data with single- and double-exponential functions. Double exponentials were fitted to the full time course. Single exponentials were fitted to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) to mitigate initial variability (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments. Compared against the truncated single-exponential fit, the double-exponential model retained a lower mean AIC (−420.7 vs. −399.1) and remained the preferred model for approximately half the subjects. The double-exponential model was preferred not only for its automated nature (requiring no manual truncation) but also for the magnitude of improvement: when the single model was superior, the advantage was marginal (ΔAIC = 9.7 ± 1.4), whereas the double model’s advantage was substantial (ΔAIC = 52.8 ± 8.9). We therefore used the double-exponential fit for further analysis.

      (10) Lines 436–446: How can this analysis be performed if asymptotic performance was not achieved in any of the conditions (nothing has plateaued in Figure 2B)? Also, the 10% FHR measure is dependent on the FHR estimate; correlating two measures based on the same measure is, by definition, expected to be correlated. This result seems to reflect that if one's learning is faster within a fixed number of trials (150), one has more opportunity to reach a higher final PCC even if asymptotic performance is identical.

      To clarify, the variables being correlated are not the FHR and ±10% of the FHR values themselves, but rather the time points at which each participant reached ±10% of their individual FHR during the task. This analysis therefore does not involve two measures derived directly from the same estimate. The timing of reaching ±10% of the FHR reflects the learning-settling trajectory rather than the FHR magnitude, so there is no a priori reason for the two measures to be intrinsically correlated. While asymptotic performance was not reached within 150 trials for some participants, the estimated FHR still provides a consistent individual marker of learning rate, allowing comparison of relative learning dynamics across participants and conditions.

      (11) Lines 448–461: Brandewie & Zahorik (2013) show that a large portion of that improvement is due to tuning to the voice and location of the speaker. Also, in the current study, there are some issues with the anechoic condition (see below).

      We thank the reviewer for this comment. This section refers specifically to results related to carrier phrase length, not to speaker identity (addressed separately below) or location, both of which were fixed in our study and therefore unlikely to account for the observed effects. We address the reviewer’s concerns about the anechoic condition in our responses below.

      (12) Line 491: What were the average trial numbers for the steady and initial trials? Also for the anechoic condition?

      We appreciate the reviewer raising this. We analysed the average trial number at which initial and steady trials occurred across a total of 360 trials: for all 22 participants, initial trials mean = 6 ± 4 and steady trials mean = 37 ± 9; sham TMS: initial trials mean = 5 ± 4, steady trials mean = 34 ± 6; real TMS: initial trials mean = 10 ± 9, steady trials mean = 38 ± 12; anechoic condition: initial trials mean = 6 ± 3, steady trials mean = 36 ± 8; and the 11 randomly selected subjects: initial trials mean = 6 ± 5, steady trials mean = 38 ± 7. This information has been added to the relevant Results sections.

      (13) Line 498: It's still not clear when the anechoic condition was performed. Lines 200–205 talk about combinations in which the Lecture Room was swapped, but it's impossible to follow when and how often that occurred. Given that the anechoic room was not included in the same way as the main three rooms, the conclusion at lines 500–504 is questionable.

      Thank you for noticing this. We have rewritten the relevant section of the Methods (“Identity of sound environments”) as follows: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics—i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).” The anechoic room was therefore explored in the same manner as the main three rooms.

      (14) Lines 508–521: Tuning to the talker's voice/location would not predict that the effect would be different for a different voice.

      We thank the reviewer for this point. If the improvement in speech understanding were due to tuning to a specific talker’s voice or location, we would expect the effect to differ across talkers. However, our analysis across six talkers (three female, three male) showed no significant interaction between talker, carrier phrase length, and room (F(30,630) = 0.73, p = 0.85, ηp<sup>2</sup> = 0.034). Although overall performance differed across talkers (main effect of talker: F(5,105) = 27.19, p < 0.001, ηp<sup>2</sup> = 0.56), these differences did not modulate the carrier phrase effect. We therefore conclude that the improvement in speech understanding with increasing carrier phrase length is consistent across talkers.

      (15) Lines 523–536: Neither of these tests addresses the question directly. That would require switching the talker randomly between the carrier and target (or throughout the sentence).

      We thank the reviewer for this comment. We respectfully disagree that our analyses fail to address the question. While our experiment was not specifically designed to test the effect of switching talkers between the carrier and target segments, we examined whether adaptation to a talker could explain the improvement in performance with increasing carrier phrase length through three complementary analyses: (1) a repeated-measures ANOVA testing for interactions between talker and carrier phrase length (see response above); (2) an analysis of potential carry-over effects across consecutive same-talker trials; and (3) an assessment of talker-learning effects in the absence of reverberation (anechoic condition). As detailed in the Results section “Improvements in performance are explained by exposure to the environment, not talker idiosyncrasies,” none of these analyses revealed evidence that talker identity influenced the observed improvement in speech understanding. We therefore conclude that the performance improvements with increasing carrier phrase length are better explained by adaptation to the acoustic environment than to specific talkers.

      (16) Line 544: Why is FHR used in this measure when d' is used for the standard analysis in Figure 2D?

      As noted above, d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis. Given how d′ was calculated for | Colour | and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data affecting temporal resolution. We therefore used cumulative hit rates, with a 5-point moving average (~7 seconds), plotted against mean cumulative exposure time. This has been clarified in the Methods.

      (17) Also, why is FHR, as opposed to HR (which I assume is really PCC), computed across the whole experiment?

      FHR refers to the Final Cumulative Performance. This naming was used to distinguish it from trial-by-trial Hit Rate used in the time-course analysis. FHR is the final data point of the cumulative hit rate—i.e., after all responses have been accumulated in that listening environment. This has been clarified throughout the manuscript.

      (18) Still worse, it's also not clear when these anechoic trials were measured.

      We appreciate the reviewer noting a lack of clarity here. We performed three follow-up experiments in different, naïve groups of listeners, assessing performance across combinations of three rooms, including Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners), assessed in the same way as Brandewie & Zahorik (2013). The anechoic room was therefore explored in the same manner as the main three rooms; this has been clarified in the Methods and reiterated in the Results.

      (19) And specifically, from Supplemental Table 4, it looks like the improvement was considerable (up to 15%), supporting that the effect is occurring. Also, note that there seems to be something numerically wrong in Supplemental Table 4: the improvement CP0–CP1 is −6, CP1–CP2 is −9.667, and CP2–CP3 is −2.5. Based on this, CP0–CP2 is expected to be −15.667 (which it is), but CP0–CP3 is expected to be −18.167, yet it's stated as −13.167.

      We appreciate the reviewer pointing this out. Our statistical analysis does not match the calculations the reviewer derived from Supplemental Table 4. For transparency, we report below the means for each carrier phrase in the anechoic room, exported directly from SPSS, which are the values reported in the manuscript. We have reviewed this section to improve clarity and have included a link to the raw supplemental data.

      CP0: Mean 40.500, SE 4.548, 95% CI [30.212, 50.788]

      CP1: Mean 46.500, SE 5.296, 95% CI [34.519, 58.481]

      CP2: Mean 56.167, SE 3.777, 95% CI [47.624, 64.710]

      CP3: Mean 53.667, SE 3.966, 95% CI [44.695, 62.638]

      (20) Lines 555–557: This sentence seems grammatically incorrect.

      It has been corrected.

      (21) Lines 559–560: The sentence “a brain region implicated in listening performance in noise (Houtgast & Steeneken, 1973; Knudsen, 1929; Lochner & Burger, 1961)” seems to imply that the cited studies support dlPFC being the brain region implicated in hearing in noise. None of these studies does that.

      Thank you for noticing this—this was an error introduced by our reference manager and has been corrected.

      (22) Lines 586–592: What was the “overall performance” measure—d′, HR, or PCC? Also, what is “univariate” analysis here? A mixed ANOVA with a between-group factor of condition and withingroup factors of room and CP length would be appropriate, and the whole group of 22 subjects should be used for the “no-exposure” group, rather than a random selection of an 11-subject subgroup.

      We appreciate this comment and we have revised this analysis to include a mix ANOVA as suggested by the reviewer. It reads as follows in the Manuscript: “Given the potential placebo effects of a ‘sham’ TMS stimulation, we first tested whether our sample of 11 ‘sham’ TMS participants exhibited similar behavioural performance to the larger sample of 22 participants who had not been exposed to any TMS manipulation. A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population.”

      (23) Lines 600–601: By “univariate ANOVA” is meant one-way ANOVA? And why wasn't it a two-way ANOVA with factors of room and sham/real TMS? More importantly, the 10% of FHR measure is arbitrary and should be replaced by standard fitting, as discussed earlier. Looking at Figure 3B, the black line appears near an asymptote while the red one is still growing toward the end, and that should be reflected in tau.

      The ANOVAs in this and other sections have been revised based on the reviewers’ suggestions; mixed ANOVAs have instead been performed and reported, yielding similar results. The fittings and 10% FHR calculations have also been extensively revised. We now show that double exponentials are better suited to our dataset, and that two-tau parameters are not informative about when performance reaches a stable point during the task.

      (24) Also, why is the exposure time on the x-axis different in Figure 3B from Figure 2B (150 vs 500)? And it would be good to see where the across-room average no-TMS data would lie here (or show the equivalent average in Figure 2B).

      We appreciate the reviewer noticing this mismatch. Figure 2B shows the time course for each environment (150 s of exposure to each), whereas Figure 3B shows all environments collapsed (150 s × 3). This is because, for the 22 listeners without TMS exposure, a Rooms main effect was observed, justifying separate time courses per room; however, for listeners exposed to sham and real TMS, no Rooms × TMS interaction was observed, so separating time courses per room was not statistically justified. As the only significant effect was TMS condition, we grouped the time spent across all environments by TMS condition.

      (25) Lines 606–617: This analysis and Figure 3C have the same issues as described for Figure 2C— asymptotic performance was not achieved for many conditions, so the 10% measure is arbitrary, as is the resulting correlation.

      We appreciate the reviewer raising these fitting issues. This part of the manuscript has been extensively revised, including new analyses and figures, although our results have not changed. Additional detail has been added to the Methods (“Speech Performance Analysis and Timecourse Fittings of Mean Cumulative Hit Rates”), and the following summary has been added to the Results (“Statistical learning of reverberant environments occurs over long and short time courses”): we fitted data with single- and double-exponential functions; double exponentials were fitted to the full time course, and single exponentials to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments, and remained preferred when compared against the truncated single-exponential fit (−420.7 vs. −399.1, preferred for roughly half the subjects). The double-exponential model was preferred for both its automated nature and the magnitude of improvement (marginal ΔAIC = 9.7 ± 1.4 when the single model won, versus substantial ΔAIC = 52.8 ± 8.9 when the double model won). We therefore used the double-exponential fit for further analysis.

      (26) Lines 619–620: Was d′ really calculated using FHR (the final value) and a non-final False Alarm Rate? This would be arbitrary. It is still unclear how HR and FAR are defined here. There is no apparent benefit to switching between HR (Figure 3B/C, presumably overall percent correct, PCC), d′ (D, E, F), and back to HR (H, I).

      We appreciate the reviewer raising this. We have clarified in the Methods (“Speech performance analysis and time-course fittings of mean cumulative hit rates”) how Hit Rate, False Alarm Rate, and d′ were calculated.

      (27) Line 655: Figure 3E should be Figure 3F.

      Corrected.

      (28) Lines 661–670: Why was Number only analyzed for initial trials, while Color was analyzed for both initial and steady trials? Also, the choice of trials 9–10 for “steady” is arbitrary and should be shown somewhere in Figure 3B.

      We analysed performance for |Number| on initial trials (1–2) only, for CP0, because performance for this speech token could only improve if positively influenced by short-term, within-trial accumulation of information (acknowledging that | Colour | precedes |Number|). To determine how much knowledge accumulated over repeated exposures — i.e., metaadaptation—we instead needed to compare performance on a speech token whose improvement could only stem from knowledge gained across previous trials, not within a single carrier phrase. We compared | Colour | performance on CP0 between initial trials (1–2) and later, steady trials (9–10). If this hypothesis is supported, it suggests that | Colour | performance for CP0 benefits from meta-adaptive information conveyed across trials as knowledge of the environment’s global structure accumulates—our proxy for meta-adaptation (Figure 3G). The choice of trials 9–10 as “steady” follows work on animal models of meta-adaptation (Robinson, Harper, & McAlpine, 2016), which described a faster adaptation rate after the eighth presentation of an environment. This has been clarified in the manuscript.

      (29) Lines 724–728: This description is confusing. The main effect of “Lecture Room” vs. “Highly Reflectant” context is that performance is very good in the Lecture Room (green line) and poor in the Highly Reflectant Room (purple). Averaging that with OPO and CP and reporting “mean difference = 20.09” (in what units?) as “overall performance” distracts from the main point. Moreover, how can that be entered into an ANOVA when the room contexts differ (LR+OPO+CP vs. HR+OPO+CP)? That ANOVA seems incorrect; it should only be performed on OPO+CP across the two contexts.

      This section has been revised and re-analysed as suggested. Redundant and unnecessary statistical comparisons were removed, retaining only those that show the effect of context on OPO and CP when comparing the different contexts in which these common environments were learned.

      (30) Lines 740–750: Again, it is not surprising that when Living Room replaces Lecture Room—and performance in Living Room is worse than in Lecture Room—the average of LiR+OPO+CP is lower than LER+OPO+CP, if OPO+CP performance is unchanged. The interesting question is whether anything changed in OPO+CP performance, as suggested for the previous point.

      This section has been revised as suggested by the reviewer.

      (31) Lines 752–771: Again, the same issue—the main effect is that performance in Anechoic trials is worse than in Lecture Room or Living Room trials, which alone explains the group difference. I am also sceptical of the finding that Anechoic performance is worse than reverberant performance, contrary to typical spatial-release-from-masking results, where reverberation degrades performance by adding noise energy at the better ear and reducing binaural benefit through decorrelation. This may be an artefact of how target and noise levels were normalized after convolution with HRTFs/BRIRs (or the use of Ambisonics); no acoustic analysis of the stimuli is provided. At minimum, the total received level at the two ears for target and masker in every environment should be reported. Brandewie & Zahorik (2013), using equivalent anechoic and reverberant conditions, never observed reverberant performance to exceed anechoic, contrary to what is stated here (lines 754–755).

      We appreciate the reviewer raising these points. The statistical analysis in this section has been revised: only the common rooms across the three-room conditions (Open-Plan Office and Car Park) were directly compared. Performance in Living Room and Lecture Room was not statistically different (Results, paragraph 4, “Statistical learning of room acoustics is tuned to universally experienced reverberation times”). However, Anechoic and Lecture Room performance remained significantly different (mean difference = 11.08, t(9) = 2.66, p = 0.013, Cohen’s d = 0.84), as did performance in the common rooms when learned in the context of Lecture Room versus other contexts (mean difference = 9.7, F(1,41) = 13.24, p < 0.001, ηp<sup>2</sup> = 0.26).

      While this setup resembles many masking studies, Brandewie & Zahorik (2013) tested four rooms simultaneously, whereas we tested three-room conditions explicitly designed to test environment-mix adaptation. We observed a synergistic relationship between performance in ‘good’ reverberant rooms (Lecture Room, Living Room) and the common but less favourable rooms (Open-Plan Office, Car Park, with longer RT60): in the absence of a ‘good-reverb anchor,’ performance in the common rooms improved less over time, possibly because participants had less to leverage in anechoic environments.

      We agree that verifying at-ear acoustic levels is critical to ruling out a normalization artefact. Our stimuli were normalized in the 41-channel sound field, not at the listener’s ears: source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses, the 41-channel noise energy was scaled to a target of 70 dB, and the 41-channel speech field was scaled to the target SNR; the 41-channel signals were then rendered to two channels using a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related acoustic effects such as head shadow.

      To verify that this did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (see supplied table, Summary Reverb Data). In the anechoic condition, the noise (positioned to the left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, raising noise level at the right ear to 55–56 dB depending on room, substantially lowering the ear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming this is not a normalization artefact but rather a genuine perceptual spatial release from masking, likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added the at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2, and updated the Discussion to clarify this mechanism.

      (32) Lines 755–759: Describing anechoic spaces as “rare” and as rooms whose “walls are treated” states the facts backwards. Open spaces (e.g., a grass lawn) are largely anechoic fields, and people spend considerable time in such environments. An anechoic room may be artificial, but an anechoic (or near-anechoic) space is very common, and the room is simply an attempt to simulate that within an enclosure.

      We appreciate this point and have rewritten this section to reflect it.

      (33) Lines 808–812: This sentence appears incorrect. It refers to “the ability to correctly report keywords spoken in environments with the more extreme—lower or higher—RIRs,” presumably meaning OPO and CP, but these are not the environments with extreme RIRs; or, if referring to An and HR, those were not “encountered in experimental blocks also containing the moderately reverberant Lecture Room or Living Room.”

      Thank you for noting this. We have rephrased this section to refer only to the extreme high-RIR environments encountered.

      (34) Line 843: Should Fig 3Ai be Fig 3A? Also, in that figure there are arrows between dlPFC and A1, and between A1 and (the cerebellum?)—it's unclear what these represent.

      The arrows were intended to represent feedforward and feedback information flow to lower auditory brain centres. We acknowledge they were confusing and have removed them from the figure.

      (35) Lines 929–949: The authors did not account for listeners tuning to voice and location (as now cited via Best et al.), and their own and Brandewie & Zahorik's data show improvement due to carrier phrase even in the anechoic case (with the inconsistency in Supplemental Table 4 noted earlier). A direct test—switching the environment between carrier phrase and target phrase, as in Brandewie & Zahorik and Vlahou et al.—would be needed to fully attribute the effect to reverberation rather than other factors.

      We thank the reviewer for this detailed comment and agree that directly manipulating the environment between carrier and target phrase would provide the most direct test of environment-specific adaptation. While our study did not implement this manipulation, our data provide converging evidence: (1) listeners showed improvement with longer carrier phrases even in the anechoic condition, consistent with previous reports, but this improvement did not interact with talker identity, carrier phrase length, or room, indicating it is not driven by tuning to specific voices or locations; and (2) regarding Supplemental Table 4, the calculations suggested by the reviewer do not match our statistical analysis—we have reported the SPSS-exported means directly (shown above) and reviewed this section for clarity, including a link to the raw supplemental data. Taken together, while we cannot fully quantify the proportion of adaptation attributable to reverberation versus other factors without the direct environment-switch manipulation, our results indicate that the observed improvements primarily relate to exposure to the environment rather than talker-specific effects.

      (36) Lines 951–953: It is unclear what about “understanding speech in background noise” distinguishes this study from previous studies of adaptation to reverberation, many of which also examined speech in noise (as reviewed in Tsironis et al., 2024). Rather than reviewing pertinent studies on adaptation to reverberation for speech tokens, the authors cite abstract noise-texture studies that are only partially relevant, given the prevalence of speech in everyday listening (lines 955–958).

      We appreciate the reviewer raising this point. Our intention was to refer specifically to statistical learning of implicit environmental acoustic features such as reverberation, rather than to speech-in-noise perception per se. We have revised the text accordingly: “A key feature of our study, which distinguishes it from previous investigations of statistical learning of acoustic features in human listeners, is the use of an ethologically valid listening task—understanding speech in background noise while listeners implicitly learn repeated acoustic features.”

      (37) Lines 978–980: In what way? For environments with large T60, a simpler explanation than “ecological validity” is that there is more late reverberant energy in the target acting as a masker, predictable from DRR.

      This section has been rewritten to clarify that it is the decline in performance at longer RT60 that is reminiscent of the decline observed under rTMS.

      (38) Line 981: What does “the better to understand speech in reverberant background noise” mean?

      This sentence has been revised.

      (39) Lines 987–988: When did “performance decline over the course of an experimental session”? Figures 2B, 3B, and 4A all show performance improving over the session.

      We have rephrased this sentence to refer to a decline in overall performance.

      (40) Lines 990–992: Many previous studies report better adaptation to reverberation for some rooms than others (e.g., Brandewie & Zahorik, 2010; Vlahou et al., 2021), but none have reported decreased performance for an anechoic space relative to a reverberant one. This anomaly should be explained and reconciled with the existing literature before invoking ecological explanations such as “ethologically relevant environments.”

      We thank the reviewer for raising this important point. We agree that verifying the at-ear acoustic levels is critical to ruling out a normalization artefact, particularly given our finding that reverberant performance exceeded anechoic performance. To address this directly: our stimuli were normalized in the 41-channel sound field, not at the listener’s ears. The source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses; the 41-channel noise field was scaled to a target of 70 dB and the 41-channel speech field scaled to the target SNR; the 41-channel signals were then rendered to two channels via a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related effects such as head shadow because normalization preceded binaural rendering.

      To verify that this sound-field normalization did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (Summary Reverb Data table). In the anechoic condition, the noise (positioned left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, increasing right-ear noise level to 55–56 dB depending on room, substantially worsening the atear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming the finding is not a normalization artefact but instead reflects a genuine perceptual spatial release from masking—likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added these at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2 to clarify this mechanism.

      (41) Discussion: Given the questions about the results, the discussion might need to be rewritten to only discuss claims that are actually supported.

      The Discussion section has indeed been extensively revised.

      (42) The hippocampus and other areas have been proposed for statistical learning, and studies also show that disruption of DLPFC can boost statistical learning (https://doi.org/10.1016/j.jml.2020.104144).

      We appreciate the reviewer raising this. We have cited Ambrus et al. (doi:10.1016/j.jml.2020.104144) in the Discussion (line 888) as evidence of opposing effects of dlPFC stimulation on statistical learning, and have further revised lines 891–903 of the Discussion to more clearly describe the known projections and functional interactions between dlPFC, hippocampus, striatum, and basal ganglia that support implicit and statistical learning.

      (43) Line 1739: What is “cumulative” here?

      “Cumulative” has been deleted.

    1. eLife Assessment

      This study presents an important study into the molecular function of AT-HOOK MOTIF NUCLEAR LOCALIZED 15 (AHL15), a member of the AHL protein family, identifying it as a potential regulator of three-dimensional gene-loop organization within transcribed gene bodies. The authors support this claim with compelling genome-wide evidence, integrating AHL15 binding profiles with transcriptional and chromatin accessibility changes, as well as demonstrating overlap with genes known to form loops across transcribed regions. The evidence supporting the claims of the authors is convincing. Collectively, these findings will be of broad interest to biologists seeking to understand the fundamental regulatory mechanisms underlying gene expression.

    2. Reviewer #1 (Public review):

      The study by Luden et al. seeks to elucidate the molecular functions of AHL15, a member of the AT-HOOK MOTIF NUCLEAR LOCALIZED (AHL) protein family, whose overexpression has been shown to extend plant longevity in Arabidopsis. To address this question, the authors conducted genome-wide ChIP-sequencing analyses to identify AHL15 binding sites. They further integrated these data with RNA-sequencing and ATAC-sequencing analyses to compare directly bound AHL15 targets with genes exhibiting altered expression and chromatin accessibility upon ectopic AHL15 overexpression.

      The analyses indicate that AHL15 preferentially associates with regions near transcription start sites (TSS) and transcription end sites (TES). Notably, no clear consensus DNA-binding motif was identified, suggesting that AHL15 binding may be mediated through interactions with other regulatory factors rather than through direct sequence recognition. The authors further show that AHL15 predominantly represses its direct target genes; however, this repression appears to be largely independent of detectable changes in chromatin accessibility.

      In addition to the AHL protein family, the globular H1 domain-containing high-mobility group A (GH1-HMGA) protein family also harbors AT-hook DNA-binding domains. Recent studies have shown that GH1-HMGA proteins repress FLC, a key regulator of flowering time, by interfering with gene-loop formation. The observed enrichment of AHL15 at both TSS and TES regions, therefore, raises the intriguing possibility that AHL15 may also participate in regulating gene-loop architecture. Consistent with this idea, the authors report that several direct AHL15 target genes are known to form gene loops.

      Overall, the conclusions of this study are well supported by the presented data and provide new mechanistic insights into how AHL family proteins may regulate gene expression.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Luden et al. investigates the molecular function and DNA-binding modes of AHL15, a transcription factor with pleiotropic effects on plant development. The results contribute to our understanding of AHL15 function in development specifically and transcriptional regulation in plants more broadly.

      Strengths:

      The authors developed a set of genetic tools for high-resolution profiling of AHL15 DNA binding and provide exploratory analyses of chromatin accessibility changes upon AHL15 overexpression. The generated data (CHiP-Seq, ATAC-Seq and RNA-Seq is a valuable resource for further studies. The data suggest that AHL15 does not operate as a pioneer TF, but is likely involved in gene looping.

      Weaknesses:

      The authors have extended the motif analysis to the top 1,000 shared peaks, addressing part of my previous concern. It did not erase my worries about overclaiming completely, but I think it can be considered as a terminological issue, rather than technical one. Therefore, it could be addressed through minor revision, without further analyses.

      Specifically, the absence of a predominant enriched motif does not establish that AHL15 binds non-specifically or lacks sequence preferences. Low motif prevalence limits the proportion of peaks it could explain but does not exclude a genuine binding preference; a binding motif need not be rare in the genomic background (binding can be supported by other factors). Please revise the interpretation in lines 188-194, and the conclusion in lines 204-206 to state that the present analysis did not identify a strongly enriched motif accounting for a substantial proportion of AHL15-associated regions. Any equivalent claims of non-specific binding elsewhere in the manuscript should be toned down.

      Additional minor points:

      (1) Please provide the exact HOMER background-selection procedure, and definition of the 50-bp search windows.

      (2) Figures 2B-C and the related supplementary figures, add the x axis label.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigated the role of AHL15 in regulation of gene expression using AHL15 overexpression lines. Their results do show that more gene are downregulated when AHL15 is upregulated and its binding is not affecting the chromatin accessibility. Further, they investigated AHL15 binds in regions depleted in histone modifications and other epigenetic signatures. Subsequently, they investigated the presence of AHL15 in the gene chromatin loops. They found overlaps with both upregulated and downregulated genes. The methods are appropriately described, but could be improved to include the analysis of self-looping gene boundaries.

      Strengths:

      Their study clearly showed lack of any specific sequence enrichment in the AHL15 binding sites, other than these being AT-rich, suggesting that AHL proteins do not recognize a specific DNA sequence but are recruited to their AT-rich target sites in another way. The study does suggest significant enrichment of AHL15 binding sites at TSS and TES, and AHL15 sites are depleted of any histone marks. They also identified that AHL15 binding sites overlap with self-looping gene boundaries. The authors have addressed the comments raised in the revised manuscript.

      Comments on revised version.

      The authors have addressed the comments raised, in the revised manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study by Luden et al. seeks to elucidate the molecular functions of AHL15, a member of the AT-HOOK MOTIF NUCLEAR LOCALIZED (AHL) protein family, whose overexpression has been shown to extend plant longevity in Arabidopsis. To address this question, the authors conducted genome-wide ChIP-sequencing analyses to identify AHL15 binding sites. They further integrated these data with RNA-sequencing and ATAC-sequencing analyses to compare directly bound AHL15 targets with genes exhibiting altered expression and chromatin accessibility upon ectopic AHL15 overexpression.

      The analyses indicate that AHL15 preferentially associates with regions near transcription start sites (TSS) and transcription end sites (TES). Notably, no clear consensus DNA-binding motif was identified, suggesting that AHL15 binding may be mediated through interactions with other regulatory factors rather than through direct sequence recognition. The authors further show that AHL15 predominantly represses its direct target genes; however, this repression appears to be largely independent of detectable changes in chromatin accessibility.

      In addition to the AHL protein family, the globular H1 domain-containing high-mobility group A (GH1-HMGA) protein family also harbors AT-hook DNA-binding domains. Recent studies have shown that GH1-HMGA proteins repress FLC, a key regulator of flowering time, by interfering with gene-loop formation. The observed enrichment of AHL15 at both TSS and TES regions, therefore, raises the intriguing possibility that AHL15 may also participate in regulating gene-loop architecture. Consistent with this idea, the authors report that several direct AHL15 target genes are known to form gene loops.

      Overall, the conclusions of this study are well supported by the presented data and provide new mechanistic insights into how AHL family proteins may regulate gene expression.

      However, it is important to note that the genome-wide analyses in this study rely predominantly on ectopic overexpression of AHL15 at developmental stages when the gene is not usually expressed. Moreover, loss-of-function phenotypes for AHL15 have not been reported, leaving unresolved whether AHL15 plays a physiological role in regulating plant longevity under native conditions. It therefore remains possible that longevity control is mediated by other AHL family members rather than by AHL15 itself. In this regard, the manuscript's title would benefit from more accurately reflecting this broader implication.

      The ahl15 loss-of-function phenotype has previously been described in Karami et al., 2020 (Nat. Plants), Rahimi et al., 2022a (New Phyt.), and Rahimi et al., 2022b (Curr. Biol.), showing that ahl15 loss-of-function among others results in accelerated vegetative phase change and flowering, a reduced number of leaves produced by axillary meristems in short day grown plants and reduced secondary growth in the inflorescence stem. The dominant-negative ahl15 delta-G allele, expressing a mutant protein lacking the conserved G motif in the PPC domain, shows these phenotypes more clearly in the heterozygous ahl15 +/- background, and is embryo lethal in the homozygous ahl15 background (Karami et al., 2021, Nature Comm.). In addition, we recently show that leaf senescence is significantly accelerated in the ahl15 loss-of-function mutant (Luden et al., 2025, BioRxiv). These results show that AHL15 is involved in several aspects of ageing in Arabidopsis, and we have adjusted the introduction to discuss these previous findings more explicitlyWe agree with reviewer 1 on the possibility that multiple AHLs could have an effect on longevity, which is partially supported by the delayed flowering time observed in the AHL20, AHL27, or AHL29 overexpression lines (Karami et al., 2020, Street et al., 2008). However, the induction of the AHL15-GR fusion alone by DEX shows a clear delay of developmental phase transitions and the aging process in general, indicating that AHL15 by itself is able to extend longevity as other AHLs are not affected by DEX treatment (proven by the fact that their expression is not significantly changed in our RNA-seq analysis of DEX-treated 35S:AHL15-GR seedlings).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Luden et al. investigates the molecular function and DNA-binding modes of AHL15, a transcription factor with pleiotropic effects on plant development. The results contribute to our understanding of AHL15 function in development, specifically, and transcriptional regulation in plants, more broadly.

      Strengths:

      The authors developed a set of genetic tools for high-resolution profiling of AHL15 DNA binding and provided exploratory analyses of chromatin accessibility changes upon AHL15 overexpression. The generated data (CHiP-Seq, ATAC-Seq and RNA-Seq is a valuable resource for further studies. The data suggest that AHL15 does not operate as a pioneer TF, but is likely involved in gene looping.

      Weaknesses:

      While the overall message is conveyed clearly and convincingly, I see one major issue concerning motif discovery and interpretation. The authors state that because HOMER detected highly enriched motifs at frequencies below 1%, they conclude that "a true DNA binding motif would be present in a large portion of the AHL15 peaks (targets) and would be rare in other regions of the genome (background)."

      I agree that the frequency below 1% is unexpectedly low; however, this more likely reflects problems in data preprocessing or motif discovery rather than intrinsic biological properties of the transcriptional factor that possesses a DNA-binding domain and is known to bind AT_rich motifs. As it is, Figure 2 cannot serve as a main figure in the manuscript: it rather suggests that the generated CHiP-Seq peakset is dominated by noise (or motif discovery was done improperly) than that AHL15 binds nonspecifically.

      Since key methodological details on the HOMER workflow are missing in the M&M section, it is not possible to determine what went wrong. Looking at other results, i.e. the reasonably structured peak distribution around TSS/TTS and consistent overlap of the peaks between the replicas, I assume that the motif discovery step was done improperly.

      Therefore, I recommend redoing the motif analysis, for example, by restricting the search to the top-ranked peaks (e.g. TOP1000) and by using an appropriate background set (HOMER can generate good backgrounds, but it was not documented in the manuscript how the authors did it). If HOMER remains unsuccessful, the authors should consider complementary methods such as STREME or MEME, similar to the approach used for GH1-HMGA (https://pmc.ncbi.nlm.nih.gov/). If the peakset is of good quality, I would expect the analysis to identify an AT-rich motif with a frequency substantially higher than 1%-more likely in the range of at least 30%. If such a motif is detected, it should be reported clearly, ideally with positional enrichment information relative to TSS or TTS. It would also be informative to compare the recovered motif with known GH1-HMGA motifs.

      If de novo motif discovery remains inconclusive, the authors should, at a minimum, assess enrichment of known AHL binding motifs using available PWMs (e.g. from JASPAR). As it stands, the claim that "our ChIP-seq data show that AHL15 binds to AT-rich DNA throughout the Arabidopsis genome with limited sequence specificity (Figure 2A, Figure S2-S4)" is not convincingly supported.

      Another point concerns the authors' hypothesis regarding the role of AHL15 in gene looping. While I like this hypothesis and it is good to discuss it in the discussion section, the data presented are not sufficient to support the claim, stated in the abstract, that AHL15 "regulates 3D genome organization," as such a conclusion would require additional, dedicated experiments.

      The motifs discovered by HOMER are ranked by their enrichment over background, of which the highest-scoring motifs are very rare in the AHL15-bound targets, but even rarer in the background, which is why they score highly on the percent enrichment score. As expected by reviewer 2, we identified AT-rich motifs that were present in a larger percentage of AHL15 targets (found in 3-18% of targets, depending on the motif, see for example motif #5 in figure S4A), which can be seen at the right tail of the histograms shown in figures 2B-C and figures S2-S4 B-C. However, these motifs were also common in the background and were therefore not considered as significantly enriched in the AHL15-bound regions, with a target:background ratio of <2. As most of these motifs were flagged by HOMER as possible false-positives, and to limit the size of the (supplemental) figures, we did not show each of the motifs identified by HOMER in table form, but the full tables of de novo motifs identified by HOMER, including possible false-positive results are included in Additional file 3.

      Although the identification of AT-rich motifs shows that AHL15 (and very likely most other AHL proteins as well) binds AT-rich regions, it does not sufficiently explain the binding of AHL15 to its target genes, as these motifs are found at almost equal frequencies in non-AHL15-bound regions. In addition, a sequence found at this frequency in the genomic background is, in our view, too unspecific to be considered as a transcription factor binding site. Based on this, we concluded that AHL15 lacks a specific binding motif that can define the genes it binds.

      We have updated the methods section to include more details on the HOMER analysis and have also run the analysis in the top1000 shared peaks as suggested by reviewer 2 for both AHL15 and AHL29 ChIP-seq peaks, which showed that unlike in AHL29, a clear AT-rich motif cannot be found for AHL15 (Additional file 1: Figure S5).

      Reviewer #3 (Public review):

      Summary:

      This study investigated the role of AHL15 in the regulation of gene expression using AHL15 overexpression lines. Their results do show that more genes are downregulated when AHL15 is upregulated, and its binding does not affect the chromatin accessibility. Further, they investigated AHL15 binds in regions depleted in histone modifications and other epigenetic signatures. Subsequently, they investigated the presence of AHL15 in the gene chromatin loops. They found overlaps with both upregulated and downregulated genes. The methods are appropriately described, but could be improved to include the analysis of self-looping gene boundaries.

      Strengths:

      Their study clearly showed a lack of any specific sequence enrichment in the AHL15 binding sites, other than these being AT-rich, suggesting that AHL proteins do not recognize a specific DNA sequence but are recruited to their AT-rich target sites in another way. The study does suggest significant enrichment of AHL15 binding sites at TSS and TES, and AHL15 sites are depleted of any histone marks. They also identified that AHL15 binding sites overlap with self-looping gene boundaries.

      Weaknesses:

      The claim that AHL15 acts as a repressor and genes regulated by it are downregulated needs to be investigated based on AHL15 binding sites, to show enrichment/ depletion of AHL15 binding sites in overexpressing genes and repressed genes. The authors should provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title. Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression may be helpful to understand the significance of the study. Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful. It is not clear how the overlap of AHL15 peaks with self-looping genes has been carried out.

      A metagenome plot of AHL15 binding around genes that are differentially expressed upon DEX treatment can be found in Figure 3F. This analysis shows that AHL15 binding near differentially expressed genes is more pronounced compared to all AHL15-bound genes, and that AHL15 binding near the TSS is especially enriched for upregulated genes.

      As also suggested by reviewer 2, we ran a motif enrichment analysis on the differentially expressed genes that are bound by AHL15 to see if any motifs are enriched compared to the background and overrepresented in the AHL15-bound genes. Again, this did not reveal an AT-rich motif nor a motif that was conserved between up- and downregulated AHL15-bound genes (Additional file 1: Figure S6).

      Plant longevity in 35S:AHL15-GR Arabidopsis plants treated with DEX has been reported previously.. DEX treatment extended vegetative development after flowering resulting in polycarpy (Karami et al., 2020, Nature Plants), enhanced secondary growth resulting in woody stems (Rahimi et al., 2022, Current Biol.) and recently we showed that it delays leaf senescence in Arabidopsis (Luden et al., 2025, bioRxiv). All these observations have now been incorporated in the results section where the p35S::AHL15-GR plants are first presented. In addition, we show that 35S:AHL15-GR plants treated a single time with DEX at 10 days after germination show a significantly delayed flowering time in figure 4C-D of this manuscript.

      The enrichment of AHL15 ChIP-seq peaks in self-looping genes will be analyzed as suggested and compared to a random set of genes as a control, and the methods section will be updated to clarify how the analyses on self-looping genes were carried out.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors need to correct Line 28 on page 10:

      By comparing the AHL15 ChIP-seq data with 1766 previously identified self-looping genes -> By comparing the AHL15 ChIP-seq data with 1792 previously identified self-looping genes (Liu et al., 2016)

      Note: 1792 genes with self-loops were first reported by Liu et al. (2016)

      Liu, C., Wang, C.M., Wang, G., Becker, C., Zaidem, M., and Weigel, D. (2016). Genome-wide analysis of chromatin packing in at single-gene resolution. Genome Res 26, 1057-1068.

      These numbers will be corrected in the text.

      Reviewer #2 (Recommendations for the authors):

      (1) The newly generated datasets have been deposited only as raw sequencing reads. For reusability and reproducibility, the authors should also provide processed data accompanied by detailed metadata.

      We will upload the processed data and corresponding metadata to GEO.

      (2) Figures 4B and 5E are of low quality; can they be improved?

      We have submitted the original high-quality images to the publisher, which should resolve the issue.

      (3) Supplementary Table 5 lacks description (columns do not have names).

      This has been fixed.

      Reviewer #3 (Recommendations for the authors):

      Suggestions:

      (1) Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful.

      This analysis will be performed and included in the revised manuscript as Additional file 1: Figure S6.

      (2) Investigate AHL15 binding sites to show enrichment/ depletion of AHL15 binding sites in overexpressing genes than repressed genes.

      This analysis has been done, please see figure 3F.

      (3) Provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title.

      For the effect of AHL15-GR induction by DEX on vegetative phase change and flowering time, please see figure 4C-D and Rahimi et al., (2022, New Phytologist). For other phenotypic changes induced by DEX treatment of 35S:AHL15-GR plants, please see Karami et al. (2020; Nature Plants), Rahimi et al., (2022, Current Biology) and Luden et al., (2025; BioRxiv). Text has been added to the results section where the 35S:AHL15-GR line is first introduced to refer to these previous publications.

      (4) Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression, may be helpful to understand the significance of the study.

      This analysis has been done and included in the revised manuscript as Additional file 1: Table S1.

      (5) Describe how the overlap of AHL15 peaks with self-looping genes has been carried out.

      The methods section has been updated with detailed information on this analysis.

    1. eLife Assessment

      This important study reveals the roles of two lytic transglycosylases in the progression of spore formation in the research model species Myxococcus xanthus. Solid evidence is provided for the roles of these two enzymes in spore formation and for interplay between their functions and peptidoglycan synthetic systems. These findings may have broader implications for studies of peptidoglycan metabolism across a range of species.

    2. Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to all points of the previous reviews.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to most points of the previous review.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted in the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. The conclusions concerning the PG degradation during sporulation needs to be clarified, as described below. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation". This needs to be reflected also in title of the paper.

      We have addressed the reviewer’s concerns and updated the text. 

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 100-125: I am still concerned about clarity in the description of the muropeptides from spores. The claim is that 90% of the recovered muropeptides are anhydro LTG products. If LTGs had truly cleaved so many of the glycosidic bonds such that 90% of the muramic acid was now in the anhydro form, then all of the PG in the spores would be in small fragments (dimers and trimers containing 4-6 sugars and 8-12 amino acids), and these would be soluble and lost during purification. Whatever the form of the PG, it must not be easily soluble, so it must be either larger than that or bound to something else. The fact that the muropeptides are solubilized by muramidase indicates that there are some NAM-NAG bonds remaining, and cleavage of these should release some non-anhydro muropeptides. The spore muropeptide chromatograms have a large, late "mound" of UV-absorbing material that was released by muramidase digestion, and two of the identified anhydro products are present in this mound. It is not clear how these two anhydro products were quantified within this mound and what other (presumably) muropeptide species might be present in this mound. Some explanation of how these two muropeptides were quantified is needed.

      We thank the reviewer for this careful point. As the reviewer notes, our purification procedure recovers only sedimentable PG, and any PG fragments solubilized by LTG activity would be lost during the washing steps and therefore not represented in our analysis. We have now explicitly acknowledged this limitation and discussed its implications in the revised text:

      “Consequently, any PG fragments solubilized by LTG activity during sporulation would be lost at this stage, and the muropeptides we detect derive from the more highly crosslinked material that survives the procedure.

      “Despite the overall decline in identified muropeptides, anhydro-muropeptides a minor component of vegetative PG were enriched in both spore types (Figure 1).”

      I feel that the authors need to address some of this uncertainty in the results and discussion. They might be able to say that anhydro-muropeptides represent 90% of the "identified muropeptides" but need to acknowledge that there is a great deal of unidentified material released by the muramidase digestion. The data might also indicate that the LTG activity solubilizes much of the PG, which is lost, and results in recovery of only highly cross-linked muropeptides that might survive the PG purification process.

      We agree that two of the identified anhydro species elute within a broad, unresolved region of the chromatogram that accounts for a substantial fraction of the muramidase-released material. Because reliable assignment and quantification of individual components within this region is not possible, we concluded that expressing anhydro-muropeptides as a fraction of the total identified muropeptides could be misleading. We have therefore removed the quantitative statement and now describe the enrichment of anhydromuropeptides in spores relative to vegetative cells qualitatively:

      “The identified muropeptides, however, represent only a part of the material released by muramidase from spores: a substantial, as a late-eluting portion of the chromatogram could not be assigned, and its composition remains unknown.”

      This would not eliminate the conclusion that "their abundance in spores indicates that certain LTGs must play essential roles (perhaps change to "might play important roles") in M. xanthus sporulation", which leads to the remaining studies in the paper.

      Following the reviewer's suggestion, we have softened the conclusion of this section to state that LTGs "may play important roles" in sporulation.

      (2) The authors have changed the statement about a "new mechanism of sporulation" at the beginning of the discussion, but this language is still in the paper title. Something more like "Programmed peptidoglycan degradation plays an important role in Myxococccus sporulation"

      Following the reviewer's suggestion, we changed the title to “A novel mechanism for morphological change during bacterial sporulation based on programmed peptidoglycan degradation”.

    1. eLife Assessment

      This valuable study provides convincing evidence for deficits in aversive taste learning and taste coding in a mouse model of autism spectrum disorders. Specifically, the authors found that Shank3 knockout mice exhibit behavioral deficits in learning and extinction of conditioned taste aversion, and calcium imaging of the gustatory cortex identified impaired neuronal responses to taste stimuli. This paper will likely be of interest to researchers studying how learning and sensory processes are affected by genetic causes of autism spectrum disorders.

    2. Reviewer #1 (Public review):

      Summary:

      The study from Wu and Turrigiano investigates how disruption of taste coding in a mouse model of autism spectrum disorders (ASDs) affects aversive learning in the context of a conditioned taste aversion (CTA) paradigm. The experiments combine 2photon calcium imaging of neurons in the gustatory portion of the anterior insular cortex (i.e., gustatory cortex) with behavioral training and testing. The authors rely on Shank3 knockout mice as a model for ASDs. The authors found that Shank3 mice learn CTA more slowly and extinguish the memory more rapidly than control subjects. Calcium imaging identified impairments in taste evoked activity associated with memory encoding and extinction. During memory encoding, the authors found less suppressed neuronal activity and increased correlated variability in Shank3 mice compared to control. During extinction, they observed a faster loss of taste selectivity and degradation of taste discriminability in mutants compared to controls.

      Strengths:

      This is a well-written manuscript that presents interesting findings. The results on the learning and extinction deficits in Shank3 mice are of particular interest. Analyses of neural activity are well conducted and provide important information on the type of impaired cortical activity that may correlate with behavioral deficits.

      Weaknesses:

      The authors did an excellent job addressing the weaknesses highlighted in my first assessment.

    3. Reviewer #2 (Public review):

      Summary

      Wu and Turrigiano investigate how Shank3 loss affects experience-dependent changes in sensory representations during conditioned taste aversion learning and extinction. Using longitudinal two-photon calcium imaging in the gustatory cortex, the authors show that Shank3 knockout mice acquire taste aversion more slowly but, after additional conditioning, reach an aversion comparable to wild-type mice; this learned aversion then extinguishes more rapidly. At the neural level, knockout mice exhibit reduced stimulus-evoked suppression and increased correlated variability; while learning and extinction are accompanied by changes in the reliability, selectivity and population-level discriminability of taste representations. The revised manuscript more clearly distinguishes baseline genotype-dependent differences in cortical activity from learning-associated changes and appropriately frames the relationship between neural activity and behaviour as associative rather than causal.

      Strengths

      A major strength of the study is the combination of longitudinal cellular-resolution imaging with a behavioural paradigm that allows cortical population activity to be followed across acquisition, retrieval and extinction. This provides a rich description of how sensory representations evolve as learned value changes, and how these dynamics differ following Shank3 deletion. The observation that knockout mice eventually acquire a robust aversion but subsequently extinguish it more rapidly is particularly useful because it separates impaired acquisition from subsequent instability of the learned association.

      The revised manuscript has substantially addressed several concerns raised in the original review. Importantly, the authors now show that reduced stimulus-evoked suppression is already evident during pre-learning habituation and that increased coactivity therefore appears to reflect, at least in part, a pre-existing network property rather than a consequence of learning. They additionally report that the level of coactivity at the beginning of conditioning correlates with subsequent behavioural acquisition, providing a useful link between individual variation in cortical activity and learning performance. This analysis strengthens the association between cortical network state and behaviour without establishing causality.

      Another useful addition concerns the potential contribution of licking behaviour to taste decoding. Because sampling behaviour necessarily differs as animals acquire an aversion, separating sensory representations from movement-related activity is difficult in this paradigm. The authors now perform decoding during the ten-second post-sampling period and find above-chance decoding after licking has ceased, making it less likely that differences in licking alone explain the principal population-decoding results.

      The manuscript is also clearer in its anatomical and conceptual terminology. The recordings are now appropriately described as being from gustatory cortex rather than implying coverage of the broader anterior insular cortex, and the conclusions have been restricted primarily to conditioned taste aversion rather than generalised to cognitive flexibility more broadly. The latter is particularly important because whether these findings generalise to reversal learning, probabilistic learning, or other forms of adaptive behaviour remains unknown.

      Weaknesses

      The principal remaining limitation is mechanistic. The experiments establish a robust association between Shank3 deletion, altered cortical activity and altered learning dynamics, but they do not establish the causal relationships among these observations. The new correlation between early coactivity and subsequent learning is informative, but manipulating the relevant network property would ultimately be required to determine whether increased correlated variability contributes directly to slower acquisition. The authors now acknowledge this limitation and have appropriately removed language implying causality.

      Similarly, the cellular or circuit origin of the altered correlated variability remains unresolved. Reduced inhibition is an interesting potential explanation, but the authors do not directly measure interneuron function or inhibitory transmission here. Consequently, the proposed relationship between Shank3 loss, altered inhibition, increased correlated activity, and impaired sensory encoding should remain a hypothesis emerging from the results rather than a demonstrated mechanism.

      A further limitation is the absence of a full conditioned-stimulus-only Shank3 knockout control group. The newly analysed habituation recordings partly address this issue by demonstrating reduced suppression and a tendency toward increased coactivity before learning, and the cross-session decoding analysis suggests that naïve knockout cortical populations can nevertheless distinguish water from saccharin. These analyses considerably improve interpretation of the existing experiment, although they are not equivalent to longitudinal comparison with a knockout control group undergoing the complete protocol without aversive conditioning.

      Finally, the interpretation of population activity in terms of taste identity, learned value and their interaction remains necessarily limited by the task design. The results clearly demonstrate experience-dependent changes in population discriminability, but because taste identity, learned value, and sampling behaviour covary during conditioned taste aversion and extinction, the present experiments cannot fully separate the precise information represented by these population changes.

      Overall assessment

      The revision has addressed the major interpretational concerns raised in the previous review and has strengthened the manuscript through several useful additional analyses. In particular, distinguishing pre-existing cortical abnormalities from learning-associated changes, relating early coactivity to subsequent behaviour, controlling more carefully for licking-related activity, and removing causal language better align the conclusions with the evidence. The study therefore provides solid evidence for altered experience-dependent sensory coding and learning dynamics following Shank3 loss, while the mechanisms connecting these phenomena remain an important question for future work.

    4. Reviewer #3 (Public review):

      In this study Wu & Turrigiano investigate an ethologically relevant form of associative learning (conditioned taste aversion-CTA) and its extinction in the Shank3 KO mouse model of ASD. They also examine the underlying circuits in anterior insular cortex (AIC) simultaneously, using two-photon calcium imaging through GRIN lens. They report that Shank3 KO mice learn CTA slower and suggest that this is mediated by a reduction in tastant-stimulus activity suppression of AIC neurons and reduced signal-to-noise ratio due to increased noise correlations in AIC neurons. Interestingly, once Shank3 KO mice do acquire CTA, they extinguish the aversive memory more rapidly than wild-type. This accelerated extinction is accompanied by a faster loss of neuronal and population-level taste selectivity and coding in the AIC compared to WT mice.

      This is an important study that uses in vivo methods to assess circuit dysfunction in a mouse model of ASD, related to sensory perception valence (in this case taste). The study is well executed, the data are of high quality, and the analyses procedures are detailed. Furthermore, the behavioural paradigm is well thought, particularly the approach for assessing extinction through repeated retrieval sessions (T1-T5), which effectively tests discrimination between saccharin and water rather than relying solely on lick counts or total consumption as a measure of extinction. Finally, the statistical tests used are appropriate and justified.

      Comments on revised version.

      The authors have addressed all comments satisfactorily.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The experiments rely on three groups: CS-only WT, CTA WT, and CTA KO. Can the authors provide a rationale for not having a CS-only KO group?

      We did not include the CS-only (KO) group because longitudinal in vivo recordings in behaving animals are technically demanding, and our primary goal was to follow and compare the dynamics across learning in WT and Shank3 KO mice. However, we recorded an additional habituation water session one day before CST1, which allows us to address (1) the decoding performance in naïve KO animals, and (2) whether higher correlated noise is already present in KO animals before learning. To this end, we first trained and tested the classifier on cross-registered data from the habituation (HAB) water session and the CST1 saccharin session across the CSonly (WT), CTA (WT), and CTA (KO) groups. Because this comparison is made across sessions, it is not exactly comparable to discrimination within a session, but it does show that GC responses in naïve KO animals can discriminate water from saccharin at baseline. These data are now included in Figure 6 – figure supplement 2 and are described in lines 329-338.

      Interestingly, we also observed a significant reduction in the amplitude of suppressed responses during the habituation water session in KO compared to WT animals, as well as a trend toward higher stimulus-evoked neuronal coactivity (data included in Figure 2 – figure supplement 2). This suggests that reduced suppression is already present in GC of KO animals before learning, potentially contributing to slower CTA acquisition (now described in lines 209-215).

      (2) The authors design an effective behavioral paradigm comparing consumption of water and saccharin and tracking extinction (Figure 3). This paradigm shows differences in licking across distinct behavioral conditions. For instance, during T1, licking to water strongly differs from licking to saccharin for both WT and KO. During T2, licking to water strongly differs from licking to saccharin only for WT (much less for KO), and licking to saccharin in WT differs from that in KO. These differences in taste sampling across conditions could contribute to some of the effects on neural activity and discriminability reported in Figures 5 and 6. That is, sucrose and water trials may be highly discriminable because in one case the mouse licks and in the other it does not (or licks much less). The author may want to address this issue.

      This is an important point. As noted by the reviewer, active licking can modulate neuronal activity in the gustatory cortex independent of taste identity (Neese et al., 2022). Because our paradigm required animals to voluntarily sample tastants of different valences, motivated differences in licking are inherently tied to taste value, making it difficult to fully disentangle taste-evoked responses from lick-related activity.

      However, taste exposure is known to induce prolonged neural responses that persist beyond the sampling phase (Juen et al., 2024). In our recordings, we included a 10second post-sampling epoch. We thus trained and tested classifiers using calcium traces during this post-sampling period; in particular, we divided the 10-second duration into five 2-second bins, matching the length of the tastant delivery phase, and analyzed the classifier built within each bin. We found that in both WT and KO animals, decoding performance was consistently above chance throughout the postdelivery period (now added to Figure 6 – figure supplement 1, lines 321-329), suggesting that decoding accuracies during sampling likely reflect taste rather than licking.

      (3) Are there any omission trials following CTA? If so, they should be quantified and reported. How are the omission trials treated with regard to the analyses?

      On the day following each CST session, animals underwent a water-only session (i.e., saccharin was omitted) to minimize context–malaise association. During these sessions, animals resumed licking both in the total lick counts and in the number of trials they engaged in, to levels comparable to pre-conditioning behavior. We did not observe significant differences between the WT and KO groups during these omission sessions. This point has been mentioned in the Methods section of the revised manuscript (lines 544-547).

      (4) The authors describe the extinction paradigm as "alternative choice". In decision-making, alternative choice paradigms typically require 2 lateral spouts to report decisions following the sampling from a central spout. To avoid confusion, the authors may want to define their paradigm as alternative sampling.

      We have revised this terminology to “alternative sampling” to avoid confusion with the classical alternative-choice paradigms.

      (5) Figure 4 reports that CTA increases the proportion of neurons that consistently respond to saccharin and water across days. While the saccharin result could be an effect of aversive learning, it is less clear why the phenomenon would generalize to water as well. Can the authors provide an explanation?

      Water and saccharin activated an overlapping population of neurons in GC. When we further quantified their tuning properties in the lifetime plots (Figure 4), we found that neurons responsive to both stimuli showed the most stable responsiveness across days, compared to neurons that responded only to saccharin or only to water (Author response image 1). Because the water-responsive and saccharin-responsive groups in Figure 4 both include this subset of dual-responsive neurons, this likely explains why both plots show increased reliability. This effect on reliability of single-cell responses is thus distinct from changes in the ability to discriminate between tastants at the population level (Fig. 6).

      Author response image 1.

      GC neurons responding to both water and saccharin are more stable during CTA extinction. Lifetime plot showing significant responses of the same neurons responding to only water (blue), only saccharin (magenta), and to both saccharin and water (gold) across test sessions (T1-5) in the CTA (WT) group.

      (6) The recordings are performed in the part of the anterior insular cortex that is typically defined as "gustatory cortex" (GC). Given the functional heterogeneity of the anterior insular cortex (AIC) and given that the authors do not sample all of the anteroposterior extent of AIC, I would suggest being more explicit about their positioning in GC. Also, some citations (e.g., Gogolla et al, 2014) refer to the posterior insular cortex, which is considered more inherently multimodal than GC. GC multimodality is typically associative in nature, as only a few neurons respond to sound and light in naïve animals.

      Our stereotaxic coordinates targeted the conventional gustatory region within AIC (see revised manuscript Methods section, lines 489-490). We have revised the terminology throughout the manuscript to more explicitly reflect this anatomical positioning.

      (7) It would be useful to add summary figures showing the extent of viral spread as well as GRIN lens placement.

      Revised Manuscript Figure 1B shows a representative example of confirmed GRIN lens placement and the viral spread of GCaMP. In most cases, GCaMP expression is confined to GC, with minimal spread to the piriform cortex and along the injection track. Depth and GCaMP expression in GC were further validated during two-photon imaging.

      (8) I encourage the authors to add Ns every time percentages are reported. How many neurons have been recorded in each condition? Can the authors provide the average number of neurons recorded per session and per animal?

      We now included these numbers in the revised manuscript (lines 157-158, 162-163, 253, 268, 271-272).

      (9) It looks like some animals learned more than others (Figure 1E or Figure 3C). Is it possible to compare neural activity across animals that showed different degrees of learning?

      We thank the reviewer for this suggestion – we now show a significant correlation between the magnitude of CTA and the coactivity metric in Figure 1 Figure supplement 3; we elaborate in our Response to Reviewer #3 Public Review 1.

      Reviewer #2 (Public review):

      (1) Causality: The paper infers that increased correlated variability causes learning deficits, but no causal tests (e.g., optogenetic modulation of inhibition or interneuron rescue) are presented to confirm this.

      Although we now provide data showing that correlated variability prior to learning is significantly correlated with the magnitude of CTA (see Response to Reviewer #1 Public Review 1above), we agree that we cannot infer causality without additional manipulations. While it might be possible to manipulate correlated variability by targeting inhibition within GC, optogenetic and chemogenetic manipulations of inhibition are likely to impact behavior through multiple mechanisms; for example, enhancing PV-interneuron activity in visual cortex profoundly impairs vision-dependent learning (Bissen et al. 2026, Leman et al. 2025). Thus, testing this would require finding a paradigm that specifically restores synchronization to WT levels without over-inhibiting the network, which is beyond the scope of the current study. We have rewritten the manuscript throughout to remove the inference of causality, and instead describe these two findings as being “associated” (see e.g. lines 95, 219-220, 379-382).

      (2) Behavioural scope: The study focuses exclusively on taste aversion; generalisation to other flexible learning paradigms (e.g., reversal or probabilistic tasks) is not addressed.

      Our study is focused on conditioned taste aversion (CTA) acquisition and extinction, which provides a well-established model for examining the formation and updating of aversive associative memories. We agree that cognitive flexibility encompasses a broad range of behavioral paradigms, and in the revised manuscript have sought to confined our conclusions to CTA. Whether the mechanisms identified here extend to other forms of flexible learning, such as reversal or probabilistic learning, or even to other sensory-stimulus-guided behavior, will require future investigation.

      (3) Mechanistic insights: While providing interesting findings of altered sensory perception and extinction of learning-related signals in AIC, it offered nearly no mechanistic insights. This makes the interpretation, especially on how generalisable these findings are, difficult. Also, different reported findings are "potentially" connected, but the exact relation between increased correlated variability and faster loss of taste selectivity cannot be assessed.

      In a new analysis we find that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3). This new piece of data (now added to revised manuscript lines 215218) provides a link between these two findings, and suggests that baseline coactivity levels in GC can influence the speed of CTA learning. We agree that this association does not imply causation, and have taken pains to avoid stating this.

      Reviewer #3 (Public review):

      (1) The authors don't make a causal link between the behaviour and AIC neurophysiology, both the percentage of suppressed cells and the coactivity measurements. For the % of suppressed cells, it seems that both WT and KO cells are suppressed in the transition between CST1 and CST2 (Figure 1L), yet only the WT mice exhibit CTA (at least by CST2). For the taste-elicited coactivity measure, it seems that there is an increase in coactivity from CST1 to CST2 in WT (Figure 2C - blue, although not statistically tested?), but persistently higher coactivity in KO. Is this change of coactivity in WT important for the expression of CTA? Plotting behavioral performance (from Figure 1G) against coactivity (from Figure 2C) for each animal would be informative.

      This is a good suggestion (also made by the other reviewers), and we now show that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3).

      (2) Shank3 KO cells already show an increase in baseline coactivity (Figure 2- figure supplement 1), and the authors never examine CS-only responses in the KO group, therefore making it difficult to determine whether elevated coactivity and noise correlations reflect a generalized AIC abnormality in Shank3 KOs (perhaps through impaired PV-mediated inhibition in insular cortex - Gogolla et al, 2014) that is not directly responsible/related to CTA?

      We agree that the increased coactivity in KO animals likely reflects a general network defect in AIC before CTA learning. This baseline change is unrelated to the taste stimuli, because (1) the coactivity is already elevated before the animals receive the taste stimuli (that is, a baseline abnormality), (2) the neuronal responses to taste delivery during CST1 is indicative of CS-only responses, as it occurs before the injection of LiCl, when the associative learning process is initiated, and (3) a trend toward higher coactivity is already present in the habituation water session before CST1 (Figure 2 and Figure 2 – figure supplement 2).

      We have clarified our description of these findings to avoid claiming that the increased coactivity “causes” poor learning performance (lines 379-382).

      (3) How do the authors interpret the large range of lick ratios (Figure 1G) for WT (almost bi-modal distribution)? Is there a within-subject correlation with any of the neurophysiological measurements to suggest a relationship between AIC neurophysiology and behavioural expression of CTA?

      See response to Point 1 above.

      (4) Indeed, CTA appears to be successfully achieved for Shank3 KO mice delayed by 1 day, as the level of saccharin aversion during the first retrieval session (T1) is comparable between Shank3 KO and WTs. In this context, not extending the first part of the paradigm to include CST3 seems to be a missed opportunity. Doing so would have allowed for within-cell and within-subject comparison of taste-elicited pairwise correlation across the learning and to investigate the neural mechanism of delayed extinction in KOs more effectively.

      We did not include a third CST session because when we analyzed the lick counts, KO animals already formed robust CTA after CST2 that was indistinguishable from WT animals. This suggests that the faster loss of CTA memory during extinction is due to a faster extinction process, rather than a weaker CTA memory from the outset. Adding a third CST could potentially lead to a memory that is harder to extinguish. Whether Shank3 KO mice would exhibit faster loss of memory in this scenario is an open question that would be interesting to explore in a future study.

      (5) How to interpret Figure 5F: Absolute discriminability is lower for T5 for CTA WT and CTA KO compared to CS-only? Why would AIC neurons have less information on taste identity by the end of extinction than during the unconditioned (CS-only) condition? And if that is the case, how is decoding accuracy in Figure 6C higher in T5 for CTA WT vs CS-only?

      We appreciate the reviewer's confusion about the discrepancy between our single-cell and population-level discriminability results. We speculate that in the CS-only state, individual AIC neuronal responses mostly reflect taste identity. However, after learning (in the CTA group), these neurons develop “mixed selectivity” (Tye et al., 2024), encoding not only identity but also the learned valence and extinction history. The lower single-cell discriminability after extinction (T5) in Figure 5F suggests that, although taste identity may remain constant, the learned history (e.g., "this taste used to be dangerous, but now it's safe") has shifted. This mixing of information makes each cell a weaker discriminator on its own.

      However, the higher population decoding accuracy in Figure 6C demonstrates that the entire population of neurons can work together more effectively. The learning process could reorganize the neural ensemble in such ways that our support vector classifier (SVC) is able to identify and combine the relevant signals within the population, even when the valence of taste stimuli has changed, to better decode stimuli and outperform the non-learned state. This suggests that the brain shifted to a more robust, population-based coding strategy for complex, learned information, which is resistant to changes in selectivity at the single-cell level. The finding that population coding is robust to single-neuron variability has also been reported in other cortical regions (Montijin et al., 2016).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Mechanistic experiments: Consider inhibitory neuron-specific imaging or manipulation (e.g., optogenetic enhancement of interneuron activity) to test whether restoring inhibition rescues learning flexibility.

      We have addressed the limitation and potential issues for manipulating cortical inhibition in Response to Reviewer #2 Public review 1.

      (2) Clarify limitations: Explicitly acknowledge the correlational nature of neuralbehavioural relationships in the Discussion.

      We have removed language that implies a causal relationship throughout, and have emphasized the correlational nature of our findings in the Discussion section of our revised manuscript (lines 379-382)

      (3) Enhance clarity: Simplify some dense methodological sections and expand figure legends to guide interdisciplinary readers.

      We have adjusted the Methods section and figure legends as needed for better readability.

      Individual Comments for Authors:

      (1) L83-90: Confusingly written, not easy to understand for someone not knowing the paradigm in detail.

      - What are the different stages? Memory encoding? Leaning? Extinction

      - More reliable in taste responsiveness - what does that mean?

      We have emphasized the behavior stages where each finding was observed in the revised manuscript (lines 82-91)

      (2) L112: Not sure if these references support the "crucial", since they do not seem to be causal.

      We have reworded this for accuracy (line 111).

      (3) Figure 1: panels h and i in the heat maps, it looks like that in the KO animal, activity is more suppressed from CST1 to 2?

      Panels j, l, m, and Figure 2: Neuronal suppression is already higher in CST1; therefore, there is no CTA effect but a general "perceptual" issue in the Shank3 model. The only effect seems to be a potential reduction in activation in CST2 in KO animals.

      This point has been discussed in Reviewer #3 Public Review 2.

      (4) Clarify in text. Especially with the sentence in the next paragraph, it might be confusing: "We wondered what other features of AIC activity during CTA acquisition might differ between WT and Shank3 KO mice."

      We have rewritten this in the revised manuscript (lines 169-170).

      (5) Clarify which are CTA-dependent and which are general (e.g., if writing suppression during CTA acquisition, it implies that it is related. But these changes were present before CTA.

      We have clarified this in the revised manuscript (lines 209-215).

      (6) Figure 4: Mainly shows a CTA-related increase in reliability in their taste responsiveness. This is not addressed anywhere else in the document and is not taken up in the discussion. How could it be related to the other findings, and what is its relevance? Please elaborate (e.g., in the discussion) or potentially remove?

      We measured response reliability, as stabilization of stimulus-evoked responses has been reported in other sensory cortices across different learning tasks. Yet, it remained unclear whether CTA learning would induce similar changes in AIC. We took advantage of our longitudinal recording to address this question and believe that this piece of evidence will contribute to the research community that studies taste and learning in general. In addition, what is striking to us is that while the taste selectivity is degraded faster in KO animals, their response reliability is largely preserved. This suggests that these two sensory stimulus-related neuronal properties may involve distinct cellular and/or circuit mechanisms.

      (7) Figure 5: Problematic to compare T5 between both groups, since T5 is lower than T4 in WT (against the trend) and T4 is an outlier in KO. e.g., if compared at T5, completely different results? Or why is there significance between T1 and T2 but not between T1 and T4 in KO? Could the authors address this point?

      In Figure 5B, the slightly lower average for WT animals at T5 was driven by a single outlier, and there was no statistically significant difference between T4 and T5 (corrected post hoc t-test, WT, T4 vs. T5, p = 0.4097). Therefore, it does not contradict the trend toward an overall increase in nonselective neurons during CTA extinction. For KO animals, the lower average at T4 than T5 (corrected post hoc ttest, KO, T4 vs. T5, p = 0.0082) was intriguing, and one possible explanation is that neurons in the KO group might undergo more dynamic and variable changes in their responsiveness during CTA extinction, fluctuating before finally stabilizing.

      Comparing T5 instead of T4 thus ensures that neuronal responsiveness is stabilized and reflects an “extinct” CTA memory more truly.

      General Comments:

      (1) While changes in SNR were observed in Shank3 models, the mechanism underlying decreased correlated variability has not been reported to date. Since decreased variability is usually associated with improved SNR ratio, it might be worth highlighting the distinction between "signal" and "noise" as separated in your analyses to make it more understandable for the reader.

      We have described in the Results section what signal and noise correlations indicated and how they were separated in our analyses in both the Results and Methods section of the revised manuscript (lines 185-193, lines 673-681).

      (2) What is the origin of the increased correlated variability?

      We have discussed that reduced cortical feedback inhibition could be a potential source of increased correlated variability in the Discussion section of our revised manuscript (lines 370-375).

      (3) Is the variability generally increased between trials (bigger fluctuations between trials for each neuron), or is the variability of each neuron similar, but they are just more correlated (more synced)?

      Our pilot analysis did not detect any evident changes in the response variability for each neuron across trials; thus, we think that in KO animals, neuronal responsivity becomes more correlated and synchronized.

      Reviewer #3 (Recommendations for the authors):

      (1) Point in line 422-424: Rephrase the closing statement of the discussion as you have shown that mutant mice are actually able to update their behaviour (in fact faster) when the valence of the sensory input changes.

      The “reduced ability to update behavior when the valence of a sensory input changes” refers to the finding that KO animals learned CTA more slowly; i.e., they were unable to timely adjust their behavior after malaise. We have rephrased this for clarity (line 448-449)

      (2) The Figure 6 legend does not correspond to panels D and E in the figure. Νο I, J in figure.

      We have fixed this mismatch in the revised manuscript.

      Minor concerns:

      (1) Cue/lick/taste-responding neurons greatly overlap and are not exclusively selective (Figure 1- figure supplement 2). Is there a genotype difference for the % of selective neurons (i.e., ones that only respond during cut/lick/taste) or the % of overlap?

      When we quantified the stimulus responsivity in KO animals, we also identified neurons that were activated by cues, licks, or tastes. Their respective percentages and overlap did not differ significantly from those in the WT group, indicating that the modality of KO neurons across different sensorimotor cues is not compromised in the KO condition (Author response image 2).

      Author response image 2.

      Neurons in WT and Shank3 KO animals show comparable responsiveness to sensorimotor stimuli during conditioning. (A) Percentage of neurons activated by the cue (left), lick movement (middle), and the tastant (right) in the CTA (KO) group (B) during the first conditioning session (CST1). (B) Venn diagram showing the overlaps among cue-, lick-, and tastant responsive neurons in (Figure 1 - figure supplement 2 C) and (A).

      (2) For Figure 1: The authors could also express consumption as a % of consumed (trial-averaged licks) over the number of trials. It is mentioned that mice undergo daily training sessions consisting of 'approximately 30 trials' (line 114). This can give an indication of how strong the learning is between cta1 and cta2 and how strong the genotype difference is.

      We are not sure if dividing trial-averaged licks over the number of trials would provide additional information, as the trial-average lick is already normalized to the number of trials.

      (3) Figure 4: Why is there a different number of neurons in C vs G?

      The figures B, C, D showed neurons that were activated by saccharin, and the figures F, G, H showed neurons that were activated by water. In all experimental groups, the numbers of neurons responsive to saccharin and water were different (i.e., B vs. F, C vs. G, D vs. H). The exact numbers were included in the corresponding figure legends in the revised manuscript.

      (4) Figure 5B: The grey background box is moved to the left.

      We kept the current figure format, as it effectively presents the mean, fitted mean, error bars, and individual animal data.

      (5) In line 142: (1-2), (2-3), (3-4), the numbers in parentheses are confusing.

      We have relabeled this as epoch 1-2, epoch 2-3, and epoch 3-4 in both text and figures for clarity (lines 145-146).

      (6) Line 188: Do the authors mean noise correlations?

      Rosenbaum et al. and Khoury et al. indeed measured correlated variability (noise correlation) in their study. On the other hand, Rothschild et al. did not specifically separate the noise from signal activities, which more likely reflect the coactivity measured in our case. We have rewritten this for accuracy (line 196).

      (7) Where mentioning in the CS-only group, please explicitly state the CS-only WT group.

      We have relabeled this throughout our revised manuscript.

      (8) In lines 273-274: if the comparison is the reduction in discriminability being faster for the KO animals that had CTA, the correct comparison should be CSonly KO vs CTA KO.

      We think that the better comparison to test how fast taste discriminability is reduced would be to perform post-hoc tests comparing T1 vs T2 within genotypes. We did not see significant changes between T1 and T2 in either genotype, which was reported in the figure legends of the reviewed preprint (lines 1140-1141).

    1. eLife Assessment

      This important study uses diffusion magnetic resonance imaging to non-invasively map the white matter fibres connecting the zona incerta and cortex in humans. The authors present compelling evidence to indicate that these connections are organized along a rostro-caudal axis. The findings will be of interest to researchers interested in neuroanatomy and cortico-subcortical connectivity.

    2. Reviewer #1 (Public review):

      Summary:

      This is a study which used 7T diffusion MRI in subjects from a Human Connectome Project dataset to characterize the zona incerta, an area of gray matter whose involvement has been demonstrated in a broad range of behavioral and physiologic functions. The authors employ tractography to model white matter tracts that involve connections with the ZI and use clustering techniques to segment the ZI into distinct subregions based on similar patterns of connectivity. The authors report a rostral-caudal organization of the ZI's streamlines where rostrally-projecting tracts are rostrally-positioned in the ZI and caudally-projecting tracts are caudally-positioned in the ZI.

      Strengths:

      The paper presents robust findings that demonstrate subregions of the human ZI that appear to be structurally distinct using a combination of spectral clustering and diffusion map embedding methods. The results of this work can contribute to our understanding of the anatomy and structural connectivity of the ZI, allowing us to further explore its role as a neuromodulatory target for various neurological disorders.

      Weaknesses:

      There should be further discussion of the clustering methods employed and why they are appropriate for the pertinent data. Additionally, the limitations of analyzing solely the cortical connections of the zona incerta should be addressed, as anatomical studies of the ZI have shown significant involvement of the ZI in tracts projecting to deep brain regions.

      Comments on the latest version:

      I reviewed the file and am more than satisfied with the authors responses and edits.

    3. Reviewer #2 (Public review):

      Summary:

      Haast et al. investigated the organization of the zona incerta (ZI) in the human brain based on its structural connectivity to the neocortex. They found that the ZI is organized according to a primary rostro-caudal gradient, where the rostral ZI is more strongly connected to the prefrontal cortex and the caudal ZI to sensorimotor cortex. They also found that the central region of the ZI is differently connected to neocortex compared with the rostral and caudal regions and could be important as a deep brain stimulation target for the treatment of essential tremor.

      Strengths:

      I think the overall quality of this work is great, and the results are presented in a very clear and organized manner. I particularly appreciate the effort that the authors put into validating the results using 7T and 3T data, as well as test-retest data.

      Weaknesses:

      The initial version of the manuscript left me with a couple of minor concerns that the authors addressed in the revised version. These were related to the clinical relevance of the work (now addressed with the lower-resolution data), and to the initial emphasis on a dorso-ventral gradient that the data did not strongly support.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study that used 7T diffusion MRI in subjects from a Human Connectome Project dataset to characterize the zona incerta, an area of gray matter whose involvement has been demonstrated in a broad range of behavioral and physiologic functions. The authors employ tractography to model white matter tracts that involve connections with the ZI and use clustering techniques to segment the ZI into distinct subregions based on similar patterns of connectivity. The authors report a rostral-caudal organization of the ZI's streamlines where rostrally-projecting tracts are rostrally-positioned in the ZI and caudally-projecting tracts are caudally-positioned in the ZI.

      Strengths:

      The paper presents robust findings that demonstrate subregions of the human ZI that appear to be structurally distinct using a combination of spectral clustering and diffusion map embedding methods. The results of this work can contribute to our understanding of the anatomy and structural connectivity of the ZI, allowing us to further explore its role as a neuromodulatory target for various neurological disorders.

      Weaknesses:

      There should be further discussion of the clustering methods employed and why they are appropriate for the pertinent data. Additionally, the limitations of analyzing solely the cortical connections of the zona incerta should be addressed, as anatomical studies of the ZI have shown significant involvement of the ZI in tracts projecting to deep brain regions.

      We are grateful to the reviewer for recognizing the strengths of our study, as well as for providing constructive suggestions to further strengthen the manuscript.

      In response to the reviewer’s feedback, we have expanded our discussion of the clustering methods employed, including the rationale for using spectral clustering in combination with diffusion map embedding, and clarified why this approach is well-suited to connectivity-based parcellation of the ZI.

      Additionally, we have expanded the Discussion to address the limitations of focusing exclusively on cortical connections. As the reviewer correctly notes, anatomical studies have demonstrated that the ZI has extensive connections with deep brain regions, and our approach therefore represents only a partial view of its connectivity. We have previously demonstrated the feasibility of reconstructing subcortical pathways using in vivo diffusion MRI (Kai et al., NeuroImage, 2022), providing a foundation for extending the present framework beyond cortical connectivity. However, as iterated below in our specific response to reviewer 1, we believe this deserves a separate thorough investigation. Nonetheless, we now explicitly discuss this limitation in the revised Discussion and outline directions for future work incorporating subcortical connectivity analyses.

      Reviewer #2 (Public review):

      Summary:

      Haast et al. investigated the organization of the zona incerta (ZI) in the human brain based on its structural connectivity to the neocortex. They found that the ZI is organized according to a primary rostro-caudal gradient, where the rostral ZI is more strongly connected to the prefrontal cortex and the caudal ZI to the sensorimotor cortex. They also found that the central region of the ZI is differently connected to the neocortex compared with the rostral and caudal regions, and could be important as a deep brain stimulation target for the treatment of essential tremors.

      Strengths:

      I think the overall quality of this work is great, and the results are presented in a very clear and organized manner. I particularly appreciate the effort that the authors put into validating the results using 7T and 3T data, as well as test-retest data.

      Weaknesses:

      That being said, I was left with a couple of concerns after reading the paper.

      - Although the authors discussed animal evidence for a dorsal-ventral organization of the ZI, I thought that the evidence they presented for it in this paper was not so convincing. In Figure S5, the second gradient (G2) shows a clear dorsoventral pattern, but this pattern seems to primarily separate the ZI and H fields rather than show an internal topology of the ZI. This is more likely the case given that there are two bands (superior and inferior) of high G2 values surrounding a single band (middle) of low G2 values. The evidence for the rostrocaudal gradient, on the other hand, is quite convincing.

      - HCP data is still too advanced for clinical translation. Although 3T is becoming more and more prevalent for presurgical planning, the HCP 3T dataset is acquired with a voxel size of 1.25mm, which is a far higher resolution than the typical clinical scan. It would be very useful for clinical readers to see what individual subject replicability looks like if the data were acquired at the more typical voxel size of 2mm. This could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for their positive evaluation of our work and for highlighting the clarity of the results, and our validation efforts across 7T, 3T, and test-retest datasets.

      Regarding the reviewer’s concern about the evidence for a dorsal-ventral organization, we agree that the rostro-caudal gradient is more prominent and convincing in our data, while the dorsal-ventral pattern is less robust. As the reviewer points out, the second gradient (G2) in Figure S5 may primarily reflect differences between the ZI and surrounding H fields, rather than a clear internal subdivision within the ZI itself. We have revised the Discussion to clarify this interpretation, emphasizing that our evidence for a dorsal-ventral organization is more tentative and requires further validation, particularly in light of prior animal literature.

      We also appreciate the reviewer’s important point regarding clinical translation. Indeed, the HCP datasets, both at 7T and 3T, use acquisition parameters (e.g., 1.25 mm voxel size at 3T) that exceed those of typical clinical scans. We therefore assessed the replicability of our findings in data acquired at more clinically representative resolutions (i.e., 2 mm voxel size at 3T). Details concerning this analysis are outlined in our response to reviewer 2 below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) When using "spectral clustering" to segment the ZI per its structural connectivity, it is unclear how k=6 clusters were chosen. Moreover based on previous reports of rodent ZI cytoarchitecture into rostral, dorsal, ventral, and caudal regions there is disagreement with the topographic organization of six clusters presented in this manuscript. Moreover, "diffusion map embedding" was not described or cited. This technique of dimensionality reduction and how it was applied to the data should be specifically described.

      We thank the reviewer for this thoughtful comment. We agree that selecting the optimal number of clusters in data-driven approaches such as spectral clustering is inherently challenging in the absence of a definitive ground truth. To address this, we computed alternative cluster solutions across a range of k values (k=2-8), which are presented in Supplementary Figure 2B. Our decision to focus on k=6 was guided by prior cytoarchitectonic descriptions of the rodent ZI by Romanowski et al., 1985, who delineated six distinct sectors (pars rostropolaris, pars dorsalis, pars ventralis, pars magnocellularis, pars retropolaris, and pars caudalis). Thus, our approach aligns the data-driven clustering with established anatomical subdivisions. Additionally, we found that k=6 provided a meaningful level of granularity for probing location-dependent neuromodulatory effects within the ZI, as discussed in the revised Discussion section (‘Discrete subregions of the zona incerta using spectral clustering’, second paragraph).

      Finally, we acknowledge the lack of clarity concerning “diffusion map embedding” in the original submission. We have now more explicitly mentioned diffusion map embedding in the “Connectivity gradients” paragraph in the Methods section, which includes the relevant references as well as description on how it was applied to the data.

      (2) "Strong correlations with cognitive terms and cortical hierarchies were particularly evident for clusters 3-5 (Figure 7a-b), which have the highest number of connecting streamlines (Figure 3c). Cluster 5, located near the central sulcus, is significantly linked to movement related cognitive processing (Pearson's r = 0.623, p < 0.005) and CogPC1 (Pearson's r = 0.593, p < 0.005). Cluster 4, situated more anteriorly, overlaps with regions involved in working memory (Pearson's r = 0.494, p < 0.001) with a high level of expression of serotonin (5-HT1B) receptors (Pearson's r = 0.413, p < 0.005) (Beliveau et al., 2017). Cluster 3, located further towards the frontal pole, is associated with mood (Pearson's r = 0.473, p < 0.005) and impulsivity (Pearson's r = 0.473, p < 0.001), and the sensorimotor association axis (Pearson's r = 0.583, p < 0.005) (Sydnor et al., 2021). Clusters 1, 2, and 6, characterized by the least number of connecting streamlines (Figure 3c), were relatively weakly associated (i.e., low Spearman's coefficient and/or within the spatial autocorrelation range) with cognitive terms or cortical hierarchies."

      The validity of identifying correlations with the spectral clusters and functional connectivity in tasks related to "keywords", and reporting similarities between these data and the regions most strongly connected via tractography, is a bit questionable. These results are based on studies that may not report all relevant findings. There is a bias towards regions that are more commonly studied with fMRI. Moreover, claims can be made about assigning functions specific to a brain region for almost any structure/function relationship (with some exceptions).

      We agree that correlations between connectivity-defined ZI clusters and functional annotations derived from external datasets should be interpreted with caution, given potential biases in the available literature (e.g., overrepresentation of well-studied cortical regions in fMRI meta-analyses) and the inherent risk of over-assigning functions to structural subdivisions. Our intention was not to make definitive claims about the functional specialization of individual ZI subregions, but rather to provide an exploratory framework for situating the ZI within broader cortical hierarchies and functional domains. These analyses are intended to generate hypotheses and to offer preliminary insight into how connectivity-based subdivisions of the ZI may relate to cognition and behavior.

      We agree with the reviewer that future studies should be specifically designed to address these questions more directly, for example, by combining connectivity-informed parcellations of the ZI with task-based or resting-state fMRI in the same subjects. Such targeted approaches will be necessary to rigorously establish the integration of the ZI within the brain’s functional organization.

      (3) The ZI's connections to many subcortical structures have also been reported in rodents and non-human primates. Moreover, the authors describe the efficacy of DBS of the caudal ZI in alleviating symptoms in patients with essential tremor, which indicates modulation of the dentato-rubro-thalamic tract fibers that project to subcortical structures such as the VIM thalamus, red nucleus, and cerebellum. The atlas in the study was characterized per the ZI's cortical connections only. These concerns should be addressed in the discussion.

      We agree with the reviewer that incorporating subcortical connections is essential for a comprehensive understanding of ZI connectivity. Building on our prior work demonstrating the feasibility of subcortical tractography (Kai et al., NeuroImage, 2022), we propose a systematic investigation of in vivo subcortico-incertal tractography as a critical next step. We believe that a dedicated investigation is required to systematically evaluate subcorticoincertal tractography, optimize reconstruction of key pathways (e.g., the dentato-rubrothalamic tract), and determine how these subcortical connections contribute to the topographic organization of the human ZI. We now explicitly discuss these considerations and identify them as an important direction for future work. Such (currently ongoing) work will not only clarify how subcortical inputs shape the topography of the ZI, but will also enable targeted optimization of tractography parameters to maximize the reliable reconstruction of specific but key pathways, including the dentato-rubro-thalamic tract. We believe this line of investigation will be important in advancing both the anatomical characterization and translational relevance of the ZI.

      Reviewer #2 (Recommendations for the authors):

      (1) Re: data quality compared to the clinic, this could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for this valuable suggestion. While this was suggested as potential future work, we felt adding this analysis would strengthen the manuscript. In response, we repeated our analyses using diffusion MRI data that more closely approximates a clinical acquisition with a lower spatial resolution (i.e., 2 mm vs. 1.25 mm isotropic) and number of diffusion-encoding directions (i.e., 130 vs. 270, Kasa et al., NeuroImage Clin., 2022). We have included these analyses in the revised manuscript.

      Reassuringly, the principal rostro-caudal gradient of cortico-incertal connectivity was preserved, demonstrating that the dominant organizational feature of the ZI is robust even under clinically representative acquisition conditions. However, finer-grained parcellations were less consistent with the original HCP analyses. In particular, cluster solutions with larger numbers of clusters (k > 3) became increasingly variable, indicating that differentiation of subtle connectivity-defined subregions benefits from the higher spatial and angular resolution afforded by research-grade diffusion MRI.

      We believe these findings provide a more nuanced assessment of the translational potential of our approach. They suggest that the large-scale topographic organization of the ZI can be recovered using clinically realistic diffusion MRI, while also highlighting the current limitations of routine clinical acquisitions for resolving finer anatomical subdivisions.

      We have incorporated these results and their implications into the revised manuscript.

      (2) Figure 6 legend labels: (c) and (d) should be (b) and (c).

      Thank you for highlighting this discrepancy. We have corrected as proposed.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have responded critically to all issues raised in the initial round of reviews.]

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular.

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    4. Reviewer #3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful comments on the manuscript. In response to their suggestions, we have:

      Improved hardware calibration flexibility and documentation (Rev 1).

      Clarified the optical specifications of the system, including axial resolution and working distance (Rev 1).

      Updated Figure 2 and Figure S1 (Revs 1 and 2).

      Corrected typographical errors and clarified terminology throughout (Revs 1 and 2). In addition, we have a new Zapit release (v1.0.4, which includes release notes), that contains many improvements and bug-fixes including suggestions from reviewers. 

      Public Reviews:

      Reviewer 1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      We thank the reviewer for their assessment of the manuscript, particularly the reference to finding the balance between modularity and an integrated package.

      (1) Command signals

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      We thank the reviewer for this suggestion. Zapit uses a linear calibration by default, which works well for high-quality diode lasers with built-in power control. For users with EOMs or AOMs we have now implemented a feature that allows creation of the appropritate sigmoid calibration curve. This is in Zapit version 1.0.4 and the process for generating the calibration curve is documented on our GitBook doc site. and the commit containing most of the changes is here. We also include a third order polynomial fit, which we hope will help users of some cheaper lasers where the control function has non-linearities. The appropriate non-linear fit is chosen automatically. We describe this in the legend and main text associated with Fig. 7F.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      The number of grid lines and their spacing are set via GUI options. The initial galvo-toimage mapping uses a field-centred affine guess based on the user’s "scanners.voltsPerPixel" setting, and the setup process is described in the user guide. The software ships with a suitable default gain value, which is unlikely to require modification. Invert flags for X and Y are also provided. Once beam locations are recorded, a similarity transform is fitted for residual offset, rotation, and scale. The latest Zapit release includes bug-fixes associated with the centering of the initial calibration point grid in the field of view.

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      We kept the mapping fixed for simplicity in both the build instructions and the code. Zapit will require a dedicated DAQ so we do not anticipate re-wiring is a hurdle. Nonetheles, the code is open-source and such a change is possible: the settings file would need to be augmented and the functions that write the analog output waveforms modified. 

      (2) Laser and optics

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      This suggestion mirrors our own thought process, but we opted not to add a ray-tracing rendering to Figure 1, as doing so comprehensively would require illustrating additional optical principles that would detract from accessibility. However, building on the reviewers recommendation, we now provide references explaining the underlying scanning principles for interested readers (Schottdorf et al. 2025, referenced in Fig. 2). The relationship between scanner angle and beam position is also shown diagrammatically in Figure 7.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      This is an excellent point. With our specifications (0.8 mm beam diameter, 473 nm wavelength, 200 mm focal length objective), the effective NA is approximately 0.002, yielding a Rayleigh range of approximately 37.6 mm. The beam must therefore travel nearly 4 cm from focus before the point-spread function doubles in width, making the system highly insensitive to skull curvature. We have added this calculation and noted its practical advantage in the revised manuscript. (Section 2.2).

      What is the working distance?

      The Plossl scan/objective lens is housed at the end of the lens tube, giving a working distance of approximately 20 cm for the 200 mm focal length objective. We have added this information to Figure 2.

      Reviewer 2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank the reviewer for their positive assessment of the manuscript, and their vision for Zapit as an important community resource.

      Reviewer 3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

      We thank the reviewer for their detailed assessment of Zapit.

      Concerns & comments:

      (1) While the authors argue that it offers the best utility to affordability trade-off - faster than motorized drivers and require much less power than DMDs (100X) and less expensive/easier to use compared to SLMs, in the current form, the manuscript does not clearly list the limitations of the approach. At such, in my opinion, the authors should include side by side comparisons (perhaps as a table). For example, clear statements should be included with respect to comparisons in lateral (x-y), axial (z) spatial resolution, as well as temporal sequential aspect of Zappit and other photo-stimulation techniques involving DMDs or SLMs.

      Section 3.2 (Comparison to other approaches) compares the scanner-based approach to related techniques. Whilst this is brief, we believe it is adequate because the resolution is ultimately limited by tissue scattering. Indeed, we demonstrate that the radius of neural inhibition ~10 times larger at 2 mW than the lateral PSF (Figure 8D). The size of a DMD pixel on the brain will likely also be smaller than the excitation area, but it does depend on the imaged size of the DMD on the brain. Since that can vary from system to system, a comparison of even theoretical resolution is not straightforward. In terms of spatial patterning, DMDs and SLMs allow for arbitrary shapes to be created on the brain and we point this out in section 3.2. 

      (2) Is power really a limitation in terms of the laser sources? Or is this a disadvantage mainly because using less power has beneficial effects on the tissue health? It may be useful to provide metrics of comparisons along these lines between Zappit and DMD-based approaches.

      The reviewer highlights an important distinction. Too much laser power can cause phototoxicity, and can make neurons more excitable due to heating. It also results in a larger region of stimulation and increases the chance of off-target effects. However, when activating multiple sites with a galvo-based system like Zapit, dwell time goes down ~linearly as the number of “simultaneous” stimulation sites increases. Therefore, even if the same power at the sample is maintained (and the same risk of phototoxicity), peak power must go up to provide the same average power at each site. For example, stimulating 20 points at 40 Hz requires 10 times the peak power compared to stimulating 2 points at 40 Hz.

      The same is true of DMD-based approaches. For example, the Mightex recommend a 1 to 4 W laser to run their Polygon DMD-based photostimulation system over an area the size of the mouse dorsal cortex (personal communication); Kauvar, et al. 2020 used a 5 W laser to cover an area 7 mm across using a Polygon system. The cost of such a laser and the Polygon alone likely exceeds 60,000 USD. 

      Other than prices and logistics, the laser powers needed at the sample and their duty cycles are essentially the same across approaches and so there are no meaningful comparisons we can provide in this domain. However, based on the reviewer’s comments, we now clarify the effects of heating from high laser powers on neural excitability in the discussion (Section 1.1).

      (3) Arbitrary-scanning vs random scanning may be more appropriate to describe to strategy.

      We agree with the confusion surrounding “random-scanning” and have chosen the phrase “laser-scanning” rather than “random access”, which is the term used in the original pre-print. 

      Reviewer 1 (Recommendations for the authors):

      Consider noting that most other galvo scanner models can probably be used.

      We have added: "Other scanners could also be substituted, as can other lasers, lens combinations, and laser wavelengths etc.” (Section 4.3).

      The basic version of Zapit requires MATLAB (although alternatives are possible and guidance/code is provided), which is not unreasonable but may limit adoption.

      We acknowledge this point. MATLAB is widely available in academic neuroscience laboratories, and we provide a Python-based interface, but a complete conversion is beyond the scope of this manuscript. We hope that the open-source code will be adapted to other languages by the community as needed.

      Where laser power is mentioned (e.g., Discussion: "we recommend using 1-2 mW time-averaged light power..."), it is not always clear if this is at the laser or in the specimen plane; clarifications would be helpful.

      Thank you, we have clarified throughout that reported laser powers refer to measurements at the specimen plane.

      Abstract: "causally manipulating" - the "causal" part is redundant and can be dropped, or replaced with "transient" (more relevant).

      Changed to "transiently".

      Section 2.2: "The modular and open-source nature of Zapit means that exciting new configurations" - strike "exciting".

      Corrected.

      Figure 7D - y-axis label text is clipped.

      Corrected.

      Reviewer 2 (Recommendations for the authors):

      (1) In Figure S1 - Our paper (Heindorf et al.) is incorrectly listed as not making code available - the "camera controlled laser stimulation" software we used is part of Iris2p (that is freely available on SourceForge and linked as such in the paper). Granted, it's not user-friendly or easy to find in the large Iris2p software package - but it is technically "available".

      We apologise for this oversight. We have updated the Figure S1 legend to note that “available” code in this context refers to a dedicated and documented standalone package. We have also added links to both Pinto and the Keller-lab software packages.

      (2) In Figure 3: "and THE beam goes"

      Corrected.

      (3) In Figure 4: "the useR will be prompted"

      Corrected.

    1. eLife Assessment

      In this study, Yuan and colleagues carry out transcriptomic and epigenomic experiments to study open chromatin regions and transcripts that change upon larval settlement in the sponge Amphimedon. The authors present compelling evidence to show that sponge larvae prepare for receiving an environmental cue (sunset) by extensively modifying their chromatin accessibility. The study represents a fundamental advance in understanding the fine genetic control of larval settlement and has significance beyond the immediate field of sponge larval biology.

    2. Reviewer #1 (Public review):

      Summary:

      Yuan and colleagues present a thorough study of gene activation before and during metamorphosis in sponge larvae, combining in depth analyses of staged transcriptomes and chromatin accessibility profiling (ATACseq). Amongst several very interesting findings, the study reveals that the acquisition of settlement competence, which arises in response to decreasing light at sunset, is characterized by changes in chromatin accessibility that anticipate strong transcriptional shifts occurring as metamorphosis starts. Another notable finding is a set transcription factors amongst the genes strongly up-regulated at the onset of metamorphosis. In addition, larvae exposed to constant light, a condition that stalls metamorphosis, were found to activate metabolic pathways that are not normally expressed in swimming larva, Together the findings provide a rare level of understanding into how environmental conditions can promote deployment of alternative developmental programs in planktonic larvae.

      Strengths:

      This is a comprehensive and rigorous study of a phenomenon of wide interest. It will inspire researchers working on other species to look for similar, environmentally-driven "anticipatory" epigenetic mechanisms. It also provides a wealth of detailed information on genes, notably transcription factors, that are candidates for involvement in regulating specific metamorphosis transitions - and beyond. The data presented here are thus undoubtedly a rich and valuable resource.

      Weaknesses:

      It is not always straightforward to connect the conclusion statements in the text to the figures, or to grasp the aims and logic of the workflows.

    3. Reviewer #2 (Public review):

      Summary:

      It is demonstrated that sponge larvae prepare for receiving the environmental cue (sunset) by extensively modifying their chromatin accessibility in the absence of large gene expression changes - "anticipatory" chromatin remodeling. This program can be offset by modifying the cue (making light constant), leading to a novel molecular state. These are the most important novel findings. Then, around metamorphosis there is dramatic regulation of transcription factors and their targets. While this result is hardly surprising, it is nice to see it demonstrated thoroughly in a phylogenetically fundamental animal study system.

      Strengths:

      This is a top-notch study of a key life cycle transition in an organism of great phylogenetic importance, involving concurrent gene expression and chromatic accessibility profiling (to the best of my knowledge, this has never been done in non-bilaterians and likely anywhere outside Vertebrata). The result is highly non-trivial. There is also an additional experiment modifying the key environmental cue (constant light), adding additional insight.

      Weaknesses:

      The paper presents a great amount of material, which somewhat obscures results related to the major novel finding ("anticipatory" chromatin remodeling). Among them is the fact that there is no significant statistical association between regions opened during remodeling and genes subsequently regulated during metamorphosis. This lack of direct association suggests that there may be more to the story than just priming the genome for metamorphosis.

    4. Reviewer #3 (Public review):

      Summary:

      In their manuscript, Huifang Yan and colleagues perform RNA-seq (CEL-seq) and ATAC-seq experiments to profile the transcriptome and chromatin accessibility of sponge larvae across larval competence, settlement and early postlarval development. Amphimedon, the sponge species that they use, is amenable to lab experiments and can therefore be a convenient model for experimenting with this otherwise difficult to assay ecological parameters and cues. They had previously observed that light conditions (diminished light) at sunset are Sucritical for larvae to enter a pre-settlement stage and prime them for settlement and metamorphosis. In this paper they report that these conditions induce a gain of accessibility in many genes, including transcription factors, and that altering these conditions by providing continuous light at sunset affects this reprogramming event.

      Strengths:

      The above is a very interesting observation, one that the authors speculate that could have a broader significance and be a theme in many more larvae. I agree with the authors that this is an important finding and I think that the paper will be interesting for the broad readership of eLife. If this is the case, the authors open up a new theme of chromatin regulation, extensively studied in mammalian contexts, but severely understudied in pretty much every other context.

      Weaknesses:

      I think however that their paper often reports the data in a difficult to follow way, and that other sorts of analyses would have made the results more accessible for the broad readership.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Yuan and colleagues present a thorough study of gene activation before and during metamorphosis in sponge larvae, combining in-depth analyses of staged transcriptomes and chromatin accessibility profiling (ATACseq). Amongst several very interesting findings, the study reveals that the acquisition of settlement competence, which arises in response to decreasing light at sunset, is characterized by changes in chromatin accessibility that anticipate strong transcriptional shifts occurring as metamorphosis starts. Another notable finding is a set of transcription factors amongst the genes strongly up-regulated at the onset of metamorphosis. In addition, larvae exposed to constant light, a condition that stalls metamorphosis, were found to activate metabolic pathways that are not normally expressed in swimming larvae. Together, the findings provide a rare level of understanding into how environmental conditions can promote deployment of alternative developmental programs in planktonic larvae.

      Strengths:

      This is a very comprehensive, well-documented and rigorous study of a phenomenon of wide interest. It will inspire researchers working on other species to look for similar, environmentally-driven "anticipatory" epigenetic mechanisms. It also provides a wealth of detailed information on genes, notably transcription factors, that are candidates for involvement in regulating specific metamorphosis transitions - and beyond. The data presented here are thus undoubtedly a rich and valuable resource.

      We thank reviewer #1 for finding our study on sponge metamorphosis interesting and compelling, and that it is likely to inform future studies on gene regulation and activity in environmentally regulated developmental processes and metamorphosis.

      Weaknesses:

      I see no significant weaknesses; however, the documentation of the data is very compressed, with all the findings contained in 4 multi-panel figures with succinct legends. It is not always straightforward to connect the conclusion statements in the text to the figures. Although the relevant data is available in supplementary files, I would appreciate more help in navigating the data to assess the support for key conclusions, if possible, illustrating each text conclusion explicitly in the main figures.

      Thank you sincerely for these suggestions on how to better present the results. We agree that the figures and associated legends are succinct. To rectify this in the hope of improving clarity and accessibility, we have (i) created two new figures by splitting our original four figures into six, and (ii) expanded figure legends to provide more explanatory details. We also made minor additions to the main Results text to more fully explain some results (see also reviewer #3’s comments).

      Specifically, we:

      (1) Removed panel K (heat map of TF expression) from original Fig. 1 and created a new figure (new Fig. 2) that focuses solely on TF expression and emphasises the extraordinarily high expression of many TFs. We also moved into the new Fig. 2 a panel from the original SFig. 1 that documents the larval cell types that express these most highly expressed abundant TF transcripts. This new figure should provide the reader with a clearer perspective on high TF expression in the larval competence and the initiation of metamorphosis.

      (2) Removed panel F from the original Fig. 4 (now Fig. 5) to create a new, expanded Figure 6 that presents a stand-alone summary of the main findings of this work; that is, environmental regulation of competence and early metamorphosis. This allowed us to (i) incorporate the constant light experiment into the summary figure, and (ii) provide a more detail explanation in the legend.

      Reviewer #2 (Public review):

      Summary:

      It is demonstrated that sponge larvae prepare for receiving the environmental cue (sunset) by extensively modifying their chromatin accessibility in the vicinity of genes that are going to be regulated during metamorphosis, in the absence of large gene expression changes. This program can be offset by modifying the cue (making light constant), leading to a novel molecular state.

      Strengths:

      This is a top-notch study of a key lifecycle transition in an organism of great phylogenetic importance, involving concurrent gene expression and chromatic accessibility profiling (to the best of my knowledge, this has never been done in non-bilaterians and likely anywhere outside Vertebrata). The result is highly non-trivial. There is also an additional experiment modifying the key environmental cue (constant light), adding additional insight.

      We thank reviewer #2 for their efforts and for appreciating the approaches we employed to understand environmental induction of sponge metamorphosis. In addition to the phylogenetic importance of sponges, their pelagobenthic life cycle is likely shared with disparate bilaterians (but not with vertebrates and other chordates, whose metamorphoses are probably derived).

      Weaknesses:

      I have only a couple of suggestions.

      (1) Not all new pre-emptively opened OCR regions are associated with genes that are going to be regulated during metamorphosis. Is their association with such genes statistically significant? (Fisher's exact test?)

      Thank you for raising this helpful point. In following your suggestion to statistically test this, we determined that a Fisher’s Exact Test was not appropriate because that test is generally used only for small samples or tables with expected counts below 5; in our data, all four expected cell frequencies are well above 5 (minimum = 228.6) and N = 25,149. Thus we instead tested for an association between chromatin accessibility and differential gene expression using the more appropriate Pearson's chi-squared test. We found no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.18), and have added these details into the Results (lines 344-46) as follows: “Consistent with this interpretation, 62% of all genes that are differentially expressed in 1 hps postlarvae (3032) have proximal chromatin regions already accessible in competent larvae (Supplementary Tables 3 and 9), although statistically we find no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.184).”

      (2) Re: extended discussion on possible reasons for activation of specific transcription factor families. I feel it is not terribly useful since it is hardly more than guesswork. The authors should consider condensing this part to better emphasize the major (and most unexpected) large-scale regulation patterns.

      We agree with this appraisal and have modified the beginning of Discussion to highlight the large-scale and rapid changes of overall gene expression. This emphasises the regulatory processes – TF expression and chromatin state changes – that must be in place to allow such transcriptional changes to occur. We feel this addition enhances the focus on TF activation and regulation at competence and early metamorphosis, especially given the scale and level of change, with most of TFs being expressed at very high levels (i.e. top 5% of all gene expressed). As outlined above, we created a new figure focussed on highly expressed TFs (new Fig. 2) to hopefully further highlight this phenomenon. It would be of great interest to know if this is conserved amongst animals with a pelagobenthic life cycle and rapid metamorphosis.

      (3) Re: enrichment analysis based on significant genes (Figure 1H): Even though it is a common practice, there is nuance: as we all know very well, many genes pass a significance threshold not because they are highly differentially regulated (i.e., show large fold-change), but because they are more abundantly expressed overall and so the statistical power for them is greater. A good example is ribosomes - before we realized what was happening, they would show up as enriched in almost every experiment of ours, which was not very useful since their fold-change was quite trivial. I see the authors have ribosome enrichment too, and I suspect there are a few more functional groups that made it because they tend to express highly on average. Ideally, we want to see what is enriched among highly regulated genes, not among abundantly expressed genes. Because of this we moved to compute enrichment based only on fold-change, using the GO_MWU package (https://github.com/z0on/GO_MWU). I suggest authors give it a shot, to see if the enrichment results become more interpretable. GO_MWU is also very powerful to analyze enrichment in WGCNA modules, in case the authors want to try that.

      Thank you for this interesting insight and advice. We applied the GO_MWU package to our gene expression dataset. Overall, these new results corroborated the original KEGG enrichment analysis, largely identifying GO biological processes, cellular components and molecular functions consistent with the previously identified KEGG molecular and cellular processes operating at larval competence and 1 hps. These include genomic regulatory processes underlying transcriptional changes and morphogenetic processes that occur in the first hour of metamorphosis, and which are also highlighted in a recent BioRxiv paper (https://doi.org/10.64898/2026.04.23.719999) from our group.

      We have added (i) results from the GO MVU analyses to Supp. Fig. 1 and Supp. Table 4, (ii) the following statement to Fig. 1 legend: “GO-MWU analysis of upregulated genes reveal stage-specific enrichments largely consistent with the KEGG analysis (Supplementary Fig. 1 and Supplementary Table 4).”, and (iii) a brief description of this approach into the Methods.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, Huifang Yan and colleagues perform RNA-seq (CEL-seq) and ATAC-seq experiments to profile the transcriptome and chromatin accessibility of sponge larvae across larval competence, settlement and early postlarval development. Amphimedon, the sponge species that they use, is amenable to lab experiments and can therefore be a convenient model for experimenting with this otherwise difficult to assay ecological parameters and cues. They had previously observed that light conditions (diminished light) at sunset are critical for larvae to enter a pre-settlement stage and prime them for settlement and metamorphosis. In this paper, they report that these conditions induce a gain of accessibility in many genes, including transcription factors, and that altering these conditions by providing continuous light at sunset affects this reprogramming event.

      Strengths:

      The above is a very interesting observation, one that the authors speculate could have a broader significance and be a theme in many more larvae. I agree with the authors that this is an important finding, and I think that the paper will be interesting for a broad readership. If this is the case, the authors open up a new theme of chromatin regulation, extensively studied in mammalian contexts, but severely understudied in pretty much every other context.

      We thank reviewer #3 for their positive assessment of our findings, pointing out the novelty of this research and its broad relevance.

      Weaknesses:

      I think, however, that their paper often reports the data in a difficult-to-follow way, and that other sorts of analyses would have made the results more accessible for a broad readership. Here, I present some suggestions that the authors might want to take into account to improve their results.

      Reviewer #1 also commented on how the results were difficult to follow. Based on your and their comments, we have reworked parts of this section, and the figures and legends. Details of these changes are listed above and can be viewed in the new version with track changes on.

      We note that no further specific suggestions were visible to us in your review.

    1. eLife assessment

      This study presents a reaction-coupled molecular simulation framework to investigate how enzymatic post-translational modifications regulate biomolecular condensates and where reactions preferentially occur within them. These findings are important, in particular, to the emergence of spatially heterogeneous reaction fluxes and enhanced activity at condensate interfaces. The work is timely and has implications beyond the immediate subfield, especially for understanding condensates as biochemical reaction centers. The strength of evidence is solid overall: the computational framework is well motivated and broadly supports the main conclusions, although some aspects of the model implementation, parameter choices, and controls would benefit from further clarification and validation.

    2. Reviewer #1 (Public review):

      Summary:

      A well-presented computational work on how post-translational modifications take place from a thermodynamic and mechanistic point of view.

      Strengths:

      A model capable of recapitulating complex phenomena to simulate, such as phosphorylation.

      Weaknesses:

      The methodology relies on multiple user-defined parameters that alter the setup. Below are the specific concerns:

      (1) The abstract reads: 'First, reactions that weaken favorable interactions are thermodynamically suppressed within condensates. As a consequence, regulation of condensate solubility is most efficient when PTMs tune interactions to values close to the solubility threshold.' This phrasing is confusing. PTMs that weaken interactions will, in principle, shift the solubility to higher values, and therefore closer to the thermodynamic conditions that allow for condensate formation, but if those transitions are most efficient when they tune interactions to values close to the solubility threshold, they will not be suppressed? The authors have to make this statement clearer to readers, which is particularly important in the abstract of the manuscript

      (2) The computational framework used by the authors is reasonable given the coarse-grained nature of the model required to study this problem. However, there are multiple user-defined variables that require further validation in order to make their choices justifiable:

      a) NN and NK have the same interaction epsilon value, as well as K-K. This would effectively make kinase form condensates on their own if the right stoichiometry was imposed. This should be re-evaluated with a reduced K-K interaction to validate whether the conclusions remain invariant with this assumption.

      b) N, P and K beads also have the same molecular diameter. This should be justified, for instance, with solvent available surface area calculations to determine the excluded volume for the different species, or by citing other works that support this approach.

      c) The phosphorylation reaction takes place when 2 particles are found within a 1.5sigma distance; however, this value choice is not justified. The authors should prove how variations to this choice affect their conclusions.

      (3) Figure 2a is informative although not entirely intuitive to follow. It can be concluded, as the authors mention, that the capacity to form condensates decreases as phosphorylation is favored. Therefore, as the text also says, there is a higher fraction of P particles as lambdaP increases; however, the fraction NP/NT appears to decrease based on the color scale. According to the figure caption, this is meant to represent the fraction of P particles over the total, but this should not exceed 1. This should be clarified and better explained in a revised version. Moreover, the information related to this, shown in Figure S4a, is very informative, and I advise the authors to include it as part of the main set of figures, as it will potentially help many readers to follow the manuscript better. Moreover, in Figure S4a, some lines appear to be disconnected; this probably comes from trajectory merging; the authors must check this.

      (4) 'At the interface, N and K concentrations remain relatively high, but scaffold proteins experience fewer stabilizing interactions, lowering the energetic cost of phosphorylation. This leads to enhanced reaction activity specifically at the boundary between phases.' This statement perfectly explains why the density profiles of K and N do not match the phosphorylation probability curve. This probability is determined by the energetic impact of the reaction and the probability of encountering each other in space, but also on the short timescale diffusion: N and K proteins have greater access to more microstates at the interface and can access them faster, while having enough density to encounter each other. It would be interesting for the authors to prove or invalidate this argument. At the very least, it should be mentioned.

      (5) Characterizing the real impact of condensate interfaces in real size condensates is an interesting approach, nonetheless this paragraph lacks most of the necessary details to be robust and obtain any reliable conclusion out of it in its current form:

      a) There are several CALVADOS parametrizations, which one do the authors use? The force field must be cited.

      b) R is not well defined in the caption.

      c) In the rendered images, periodic boundary conditions appear not to be implemented; this should be clarified

      d) In the Intermolecular energy profiles, it seems that FUS-LC has no condensate bulk.

      e)How is the interface width calculated?

      f) Why do the authors choose the energy profile and not density? Or other observables such as the radius of gyration.

      g) The interface width will depend on the temperature, and how distant this temperature is from the critical temperature for phase separation. Currently, this information is lacking.

      h) Variations in the temperature, quantity used to define the interface, should be addressed in order to make the interfacial importance claim robust.

      (6) 'Since the size of typical cellular condensates rarely exceeds the 2 μm diameter'. This statement should be supported by multiple references.

      (7) Figure S2 should be improved: The use of 'weak' or 'strong' labels is subjective and it is unclear which parameters are being used. Moreover, 'exp reference' is not described or cited. Furthermore, this plot shows what appear to be sketched curves. The calculation of the coexistence densities to construct this phase diagram is trivial for the system studied here; the authors should provide direct estimates.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Thermodynamic principles of enzymatic regulation in biomolecular condensates from reaction-coupled molecular modeling" reports results of a very coarse-grained molecular dynamics (MD) simulation of three types of particles that undergo phase separation and chemical reactions. The study focuses on a molecule that can exist in two states (phosphorylated vs. unphosphorylated), whereas the third species is an enzyme that catalyzes the reaction in one direction of the transition. This topic of chemically active multicomponent mixtures is timely, and its connection to biomolecular condensates is well motivated in the introduction. Using their minimal setup, the authors find that driven reactions affect the composition of the dilute and dense region, and can even suppress phase separation completely. Moreover, the reaction fluxes are heterogeneous and exhibit a pronounced peak at the interface. The authors then interpret this acceleration of reactions in light of the role of condensates as reaction centers and conclude that the effect they report is relevant in cells.

      Strengths:

      A strength of the manuscript is the setup of a minimal system to understand the complex roles of enzymatically controlled, active reactions in condensates. This is a subtle topic since the physics of phase separation, implying non-ideal systems, affects the reactions, which can then no longer be described in a dilute approximation. The authors tried to take special care to ensure thermodynamic consistency, which is key to describing the interplay of the two effects accurately. The authors also perform relevant numerical tests, e.g., by determining the effect of reactions on the phase diagram, and measuring densities and reaction fluxes carefully. Moreover, they attempt to interpret the obtained quantities with intuitive pictures, although this is not always convincing.

      Weaknesses:

      The main weakness of the work is its presentation: I could not follow the detailed setup of the model since important details (such as the concrete interaction potentials and particularly the implementation of the reactions) are not described thoroughly. In particular, the implementation of the actively driven reaction is mysterious, and a proper negative control without activity is missing. Such a control is crucial since it would allow testing whether the implementation of the code ensures local detailed balance and thus thermodynamic consistency. Furthermore, such a passive null model would surely help in establishing the effects of activity, which is currently unclear. Since I don't understand the detailed setup, I cannot judge whether the major result, namely that reactions are accelerated at interfaces, is correct. It might very well be true, but I'm unable to judge this based on the current presentation. I detail my criticism in separate points below, and I hope that addressing these points helps the authors to improve their manuscript:

      (1) I find the model setup unclear, which makes it difficult for me to gauge the correctness of the results. I think the details of the model need to be explained in more detail, both in the main text and the SI. There are three different aspects that I find lacking:

      1a) The setup of the interactions in the model is unclear. First, only single interaction parameters \epsilon_i are specified, but the simulation likely needs to specify pairwise interactions. Table S2 is not particularly helpful since it uses a different notation (Is this switch of notation necessary?). Second, the statement that k_BT = 0.75 is confusing since k_BT should have units of energy. Third, the authors mention "a truncated and shifted LJ potential", but it is unclear whether this is the potential they use or not. Given this lack of details, I would not be able to immediately repeat the simulation, even without reactions.

      1b) In the main text (page 6), it is unclear whether an active or passive system is studied. The subsequent results suggest that the system is active (and thus has sustained energy fluxes), but this needs to be explained in detail. In any case, I would like to see an explicit activity parameter, so that a passive control can be added. For instance, I would expect that a passive system remains the same when reaction rates are changed, but the energetics are kept constant. In any case, since it is not explained how activity enters the system, the subsequent results are unclear to me.

      1c) It is unclear how thermodynamic consistency (i.e., local detailed balance) is ensured. For instance, how is the formation of a scaffold-enzyme complex performed (page 6)? The SI is also light on details. For instance, it is unclear whether the transitions in Equations (1-8) refer to individual particles (in this case, how are bimolecular reactions implemented?) or refer to densities or even overall particle counts. It is also suspicious that the reactions are given for the "dilute limit", whereas the main text clearly discusses condensed regions. It is also unclear how ΔU is calculated and what it means. Generally, I would expect a detailed discussion of detailed balance conditions in the SI, if not even in the main text. Here, it might help to clearly separate thermodynamic aspects (involving detailed balance) from kinetic considerations.

      (2) If the authors study an active system, reporting averaged concentrations and partition coefficients might be misleading. Generally, active systems develop gradients in the dilute and dense regions (Figure 1), so any averaged measurement (such as a density) will depend on the size of the region. It is thus necessary to clearly define the measurements in the main text (so readers know how to interpret the plots), and the discussion needs to be much more careful, particularly when comparing to phase diagrams, which typically discuss thermodynamically large systems.

      (3) The strongly increased fluxes at the interface are a bit suspicious, particularly since they are hardly visible in passive systems (Figure S6). The enhancement might originate from large reaction rates, leading to a short reaction-diffusion length scale (although I would then expect balanced reactions in the phases). Alternatively, they might originate from the microscopic details, such as the cut-off length of the kinase-interaction range. In any case, it would be important to establish a negative control using a passive setup and (based on this) explain the observed flux increase clearly.

      (4) I am not convinced by the statement about the bias of reactions in the dense and dilute phase, e.g., on page 12. I would find it extremely helpful to start with a passive system (without energy input), where I would expect balanced chemical potentials, so that all reactions are balanced at all points in the system. Already in this case, there might be accelerated reactions at the interface (due to kinetic details), but the two reactions need to be balanced to have overall homogeneous chemical potentials. Starting from such a base state, one could then discuss how activity biases reactions in a certain direction. While the current explanation correctly mentions that a transition can be hindered by energetics, it fails to account for the different abundances. In a passive system, these two aspects are perfectly balanced, leading to a linking of reaction rate constants and partition coefficients (e.g., see https://arxiv.org/abs/2202.13646). These aspects need to be discussed much more cleanly (and they are intimately linked to the setup of the model, which is currently unclear; see point 1).

      (5) I think the title is misleading. The authors do not establish new "Thermodynamic principles". I am also not sure what they mean by "reaction-coupled molecular modeling". Finally, the 
"enzymatic regulation" is hardly discussed in the main text, which instead seems to emphasize the accelerated reactions at the interface.

    1. eLife Assessment

      This study addresses a recent discovery by others that electroconvulsive therapy (ECT) generates seizure activity and spreading depolarization (SD), reflected in large increases in calcium, which can be observed through imaging fluctuations in neuronal calcium. Revisions addressed some concerns of reviewers but not the main issues. Therefore, the consensus was that the work is useful but the evidence showing that SD, rather than seizures, confers the neuroplastic and other therapeutic effects of ECT is incomplete.

    2. Reviewer #1 (Public review):

      The work shows that the ECS-induced calcium waves recapitulate several hallmarks of CSD, including hemodynamic alterations, distinct propagation patterns, and elevated Fos expression.

      Update of the weakness section.

      (1) I still have concerns about the use of Fos staining as the exclusive marker of neuroplasticity. Of course, Fos drives many forms of plasticity. But in addition to the role, Fos upregulation also reflects a recent history of elevated neuronal activity. This has been reported in numerous papers, including those cited in the article (e.g. Mahringer et al., 2019; Tyssowski et al., 2018). If the authors claim that their main finding is that ECS-induced CSD is necessary to drive Fos, then the relevant background should be presented in the Introduction.

      (2) The authors did not sufficiently address the concern regarding the bilateral suppression of EEG following unilateral calcium waves. Unilateral CSD is known to depress EEG only in the affected (ipsilateral) cortex. I suspect that the bilateral EEG suppression can be driven by the bilateral seizure (phase III oscillations) that is always coupled with the ECS-induced calcium waves.

      (3) Using the term 'phase III oscillations' instead of 'seizures' is confusing because the word "oscillations" covers a wide range of brain oscillations - from normal to pathological ones. First, the authors base the terminology change on the absence of "the ictal spikes characteristic of epileptic seizures..." during the phase. However, anesthesia can suppress ictal spiking and the phase III activity can represent an aborted seizure induced under anesthesia. Second, the authors cite the study by Brumback and Staton (1982) to justify their use of the term "phase III oscillations". However, the cited work describes the phase III as a seizure with "spike/polispike-wave activity".

      (4) Cortical SD cannot invade the hippocampus in the in vivo brain, although SD can occur in the hippocampus in response to generalized seizures. Invasion of CSD suggests its non-synaptic propagation via the contiguous gray matter. The following studies showed that CSD cannot invade the hippocampus non-synaptically: Fifkova E. Spreading EEG depression in the neo-, paleo- and archicortical structures of the brain of the rat. 1964 (PMID: 14138725); Yoshida et al. Identification of the extent of cortical spreading depression propagation by Npas4 mRNA expression. 2015 (doi: 10.1016/j.neures.2015.04.003). Concerning the papers cited in the discussion ("In mice, the ECS driven CSD likely invades hippocampus" (Mitlasóczki et al., 2025)), and response (Bahari et al., 2020; Bonaccini Calia et al., 2022), hippocampal SD was triggered by focal seizures, independent of CSD, in the experiments.

    3. Reviewer #3 (Public review):

      In this revised version of the manuscript, the authors have addressed several of the Reviewer's prior questions and comments. However, the main elements of the critique remain largely intact in the absence of additional experiments.

      Major:

      (1) The authors and Reviewer disagree as to the extent to which the main findings replicate prior work vs. represent true novelty. In their response, the authors state, "This appears incorrect. It was known that direct cortical stimulation (as was done in the Rosenthal paper) can drive a calcium event that resembles a CSD. It was also known that CSD can drive Fos expression. What was not known is that the Fos expression driven by ECS is fully explained by the calcium event (putative CSD)." First, what appears incorrect is the author's statement that Rosenthal et al. use direct cortical stimulation. So far as the Reviewer can discern, Rosenthal et al. in fact did not use intracranial or direct cortical stimulation, as stated repeatedly in the rebuttal. Instead, Rosenthal et al. use stimulating electrodes that are attached to the skull (i.e., non-penetrating), as clearly shown in Figure 1 of that paper. Describing this as "intracranial stimulation" and equating it with direct cortical stimulation used to induce CSD in Leão (1944), is incorrect and misleading.

      (2) The authors ask for guidance in improving the clarity of their messaging around this and cite the Abstract as clearly stating what they view as the novelty here. Yet, the key section of the Abstract reads: "We show that CSD drives increased expression of the immediate early gene Fos, a key marker of neuronal plasticity, and is associated with factors that predict positive ECT therapeutic outcome. Our results suggest that the therapeutic efficacy of ECT may be mediated by CSD. This challenges the seizure-centric model and implies that CSD, a currently unmonitored neurophysiological event, may serve as a more relevant biomarker for predicting and optimizing therapeutic outcomes of ECT." The Reviewer finds this quite misleading, as it suggests (at least to this Reviewer) that the authors view themselves as showing, for the first time, that ECT generates CSD and that this may be the therapeutic mechanism of ECT, while not clearly defining what is new. Instead, one could perhaps more accurately write: "Prior work has shown that ECT generates CSD and that CSD is associated with increases in the immediate early gene Fos, a key marker of neuronal plasticity. We provide data suggesting that ECT-associated CSD in fact drives this increased Fos expression. Our results imply the hypothesis that the therapeutic effect of ECT may be mediated by CSD-induced Fos expression and could serve as a biomarker of ECT efficacy."

      (3) In claiming novelty, the authors further assert that Rosenthal et al. provided "no confirmation beyond the observation of the calcium event that this is indeed a CSD." This also appears to be incorrect or misleading. Rosenthal et al. show simultaneous calcium and intrinsic optical imaging of hemoglobin dynamics associated with the propagating event. Rosenthal et al. also show DC-coupled electrophysiology demonstrating the slow potential shift characteristic of CSD using auricular ECS (the very stimulation model that the authors argue distinguishes their work from Rosenthal et al.). Hence, speculation that the "cortical seizure Rosenthal and colleagues find is a methodological artifact" is unsupported and should be removed.

      (4) In the rebuttal, the authors clarify the novelty of their manuscript as "that ECS drives CSD dependent immediately early gene expression (Fos)." The Reviewer agrees with this, yet as summarized in the bullet point, finds this novelty to be relatively limited and not particularly surprising, given that Fos is robustly induced in association with CSD in various contexts. This novelty could be markedly enhanced by linking increased Fos expression mechanistically to the therapeutic efficacy of ECT. Instead, any relevance of the CSD-dependent Fos induction is unclear, and statements to that effect, while interesting, are hypothetical, yet follow logically from Rosenthal et al. Any association between CSD occurrence and clinical outcome in human patients could establish predictive value but would also not prove causality, as suggested. Also unclear is how Fos might be used as a biomarker in humans, and how this might be valuable as a biomarker beyond CSD itself (which can actually be recorded), is unclear. The authors should state these limitations.

      Minor:

      (1) Regarding the microprism preparation, the added three-week postoperative interval is helpful toward alleviating this concern. Still the absence of spontaneous CSD does not address the fact that cortical incision and the CSD likely elicited during implantation could alter subsequent susceptibility to experimentally evoked CSD. This should simply be acknowledged as a limitation.

      (2) The Reviewer recognizes that the authors have added Tables S3 and S4 detailing stimulation parameters. However, the response that properties of the traveling calcium event do not depend on stimulation charge does not resolve the concern regarding analyses in which stimulation parameters are used to predict whether CSD occurs. Manually selecting stimulation parameters rather than selecting them randomly or systematically may confound these analyses, an issue which could simply be acknowledged.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1) The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy.

      This is correct, our claim is that ECS drives spreading depression dependent IEG expression in cortex. We are not claiming this is the only consequence of ECS. The extracortical response to ECS could contribute to the treatment, and we discuss this in the hippocampus specifically (see Discussion). In current clinical practice, cortical surface EEG is used as a biomarker for predicting treatment outcome. Given that the CSD is necessary to drive Fos expression in cortex, we think it is warranted to speculate that monitoring CSD during ECT is worth exploring as a potentially superior biomarker of treatment outcome. Especially given that this is readily feasible (Discussion).

      Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      The idea that Fos is a marker of neuronal activity is outdated. The primary correlate of Fos expression is neuronal plasticity and learning-related circuit modifications, rather than just neuronal activity – see e.g. (Bolhuis et al., 2001; de Hoz et al., 2018; Fleischmann et al., 2003; Kimpo and Doupe, 1997; Mahringer et al., 2022, 2019; Nakadate et al., 2012; Roy et al., 2016; Ryan et al., 2015; Tanaka et al., 2018; Tyssowski et al., 2018; Watanabe et al., 1996; Yassin et al., 2010). We have now referenced some of those articles in the manuscript (Introduction). But more importantly, our claim is that the CSD is necessary to drive Fos. The fact that CSD can occur in a single hemisphere allows, within a brain, to control for direct ECS response and phase III oscillations. In unilateral CSD, contralateral hemispheres displayed Fos levels in the cortex that were similar to sham mice. Investigating the exact plasticity pathway that is triggered by the CSD is an interesting question we are pursuing in follow-up work, but does not influence our conclusions here.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      The reviewer is correct in that postictal suppression is currently one of the best correlates of positive clinical outcome used in the clinic. However, note that EEG-based correlates are generally weak predictors of treatment outcomes: even recently identified correlates are unreliable between patients cohorts and account for less than 60% of the clinical outcome (Francis-Taylor et al., 2020; Scangos et al., 2019). See also our response to comment (2) of reviewer 3 on this topic.

      The primary problem in the interpretation of postictal suppression of EEG activity is that it is unclear what the source of the EEG signal is in this case. During phase III oscillation following the CSD, cortex is silent and all of the EEG is likely driven by thalamic input to cortex. Note, source localization in EEG does not help as the source of the signal is likely the thalamic axons in cortex (Huels et al., 2023) (or their postsynaptic, and in this case subthreshold, effects). Whatever the source of the cortical surface EEG may be, we know that following CSD it cannot be cortical activity, hence, whatever is driving postictal suppression is also not cortical. If we had to speculate, postictal suppression is likely an exhaustion of thalamus in attempting to drive cortex without getting any excitatory feedback. During an epileptic seizure, thalamus oscillates at delta frequencies and cortex responds at each peak of the oscillation, generating ictal spikes on the EEG (Meeren et al., 2002; Polack et al., 2009). The fact that ictal spikes are missing in an “optimal” phase III oscillation is consistent with our findings that the CSD completely silences cortex for minutes. Thus, during an ECS-driven phase III oscillation, cortex does not respond to thalamic input and at some point thalamus runs out of energy to maintain the oscillation - it is likely the combination of thalamus running out of energy and a silent cortex that gives rise to postictal suppression. Note we do find an asymmetry for the mid-oscillation amplitude (Figure 6A), which in this interpretation can be explained by the thalamo-cortical circuit attempting to compensate for the absence of cortical feedback by increasing its drive to the CSD-affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      This is a misunderstanding. We claim (and show) that the CSD is necessary to drive immediate early gene expression following ECS. Based on this we speculate that the CSD “may serve as a more relevant biomarker for predicting and optimizing therapeutic outcomes of ECT” than EEG biomarkers.

      We do not “state [that there is] a beneficial role of CSD” – we actually address the possibility that SD may not be therapeutically beneficial (Discussion). The suggestion of investigating behavioral consequences of CSD in mice is interesting, but likely not the most relevant follow-up work. This is for two reasons:

      (1) Very different from humans, most mouse behavior does not depend on cortex (Kawai et al., 2015; Pandey et al., 2026). Thus, any potential behavioral consequences of ECS in mice are difficult to interpret in the context of clinical relevance. We don’t think there is any known biomarker based on mouse behavior that has any direct translational value to psychiatric treatments. The reviewer may disagree with this assessment but just consider that no mouse behavior assay has ever been instrumental to the development of a novel psychiatric treatment.

      (2) The more direct – and simpler approach – is to directly test whether CSD correlates with treatment outcomes in patients. Our primary aim with this paper is to inspire exactly this. Note, we are currently pursuing this as well, of course. Monitoring CSDs in humans during ECT is likely possible with fNIRS or other hemodynamic-based measurements (Discussion).

      However, to reiterate – the claim of the manuscript is exactly as stated in the title. Thus, demonstrating the clinical relevance of the CSD is outside of the scope of the current manuscript, but will be trivial to prove by measuring treatment outcomes while routinely monitoring patients for CSD.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

      Reviewer #1 (Recommendations for the authors):

      The results of sham tests are shown only for the immunohistochemical data. However, results from control experiments should be provided for the EEG and calcium data.

      Sham data are included in Author response image 1. We are not sure how they support our conclusions and have left them out of the manuscript. A sham stimulation just characterizes ongoing activity under anesthesia. The more direct comparison is the difference between pre- and post-ECS activity. Baseline widefield calcium imaging data is already shown in Figure 1, 2, 6; baseline mouse EEG data is already shown in Figure 1 and 6. Baseline EEG in patients is not available in the present dataset.

      Author response image 1.

      Sham ECS does not result in phase III oscillations or a CSD. (A) Representative spectrograms (top) and raw traces (bottom) of concurrent EEG (left) and widefield calcium imaging (right) during a sham ECS session. (B) Population raster plot (top) and two example activity traces of neurons (bottom) during a sham ECS two-photon recording.

      Moreover, baseline (pre-ECS) calcium and EEG activity should be shown in Figures 1 and 6 to correctly assess the changes induced by stimulation.

      We are not sure we understand as pre-ECS data are already shown in Figures 1 and 6. We assume the reviewer may mean “more baseline data should be shown”? We have extended the time scale of the relevant panels in Figure 1, 2 and 6 to include 30 s instead of 10 s of pre-ECS data.

      Using the term 'oscillations/phase III oscillations' instead of 'seizures' throughout the text is confusing because the word covers a wide range of brain oscillations - from normal to pathological ones.

      This is indeed confusing – but the confusion arises from the often imprecise usage of the term “seizure” in the ECT literature. The EEG response to ECS is not equivalent to that of an epileptic seizure. The description of a “phase III oscillation” is a more precise description of the EEG signature. It was introduced by (Brumback and Staton, 1982) and is not our terminology. We dedicate an entire paragraph in the introduction to this distinction. We are not sure how to make this clearer in the current manuscript. Continuing to describe the EEG response to ECS as “seizure” is inaccurate, and mechanistically misleading, especially given that cortex is silent during the phase III oscillation following the CSD (i.e. cortex cannot be “seizing” during this phase of the EEG response, see our point above on postictal EEG suppression).

      The presentation of clinical data is scarce and unclear. e.g., the authors claim that the oscillation frequency is similar in mice and humans (lines 206-208). However, in Figure 1F, oscillations during the early post-ECS phase have twice the frequency (6-7 Hz) in the patient EEG recording (6-7 oscillations per 1 second) compared to the mouse EEG (3 Hz, i.e., 6 oscillations per 2 seconds during 73-75 s). It seems a bit odd because Figure 1H shows that the frequency does not exceed 5 Hz, even in patients.

      Please excuse, this is our mistake. The x-axis label in the human data of Figure 1F was incorrect. It should have read 1 to 3 s, not 1 to 2 s. Window sizes were of course matched between mice and human data and are all 2 s. The mistake is now corrected. The frequencies in patients and mice are compared in Figure 1H and are not different.

      A part of the discussion (lines 627-635) is based on factual inaccuracy: cortical SD cannot invade the hippocampus in the in vivo brain (only in slices), although SD can occur in the hippocampus in response to generalized seizures.

      It would be helpful if the reviewer would back up this claim with references. Short of this we are left to speculate - we suspect the reviewer may be referring to earlier work in the rat cortex, that claimed that a cortical SD can only invade hippocampus if glia has been impaired (Largo et al., 1997). This is now an outdated model: there are more recent reports of CSD invading the hippocampus (Bahari et al., 2020; Bonaccini Calia et al., 2022). If the reviewer has specific concerns with any of these papers, we would be happy to discuss in more detail, but the reviewers’ claim seems unfounded here.

      It is reasonable to expect that bilateral stimulation produces bilateral calcium/SD waves. In Rosenthal's experiments, this situation was most common. In the present study, bilateral ECS triggers mostly unilateral calcium waves. Do you have any idea what the reasons for the result are? As stimulation parameters (polarity, intensity, and frequency) have been shown to control the occurrence of SD waves, their unilateral pattern suggests non-uniform stimulation conditions. I am curious whether the uni- or bilateral pattern of calcium/SD waves depends on the ECS parameters.

      The main difference between our patient and mouse data is the polarity of stimulation. For patient data, the polarity of the current alternates with every pulse, while in our mouse data the stimulation was always right unipolar. This is discussed when we mention the asymmetry of SD in our data (Results, Methods) - we suspect the reviewer may have missed this.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      (1) The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      This is correct, our work confirms that the calcium event following ECS is a CSD. The main claim of our paper is that ECS drives immediate early gene expression in a CSD-dependent manner.

      However, note that the Rosenthal paper concluded that the calcium event is a CSD while providing only little evidence for that claim – mind you, we agree with their interpretation, but provide more evidence for the conclusion. We are happy to discuss the limitations of the Rosenthal paper as highlighted below more prominently in the manuscript, if the reviewer thinks this would be helpful, but we think that is likely not necessary.

      Briefly, in the data presented in (Rosenthal et al., 2025), the only support for the calcium response being a CSD is the speed of propagation and the DC shift reported in extracellular recordings. We add to this by showing that the spread of the calcium event follows the pattern expected by a CSD through cortical layers (Figure 4), results in heterogeneous returns to baseline calcium levels (Figure 4), causes vasoconstriction (Figure S5), travels at the speed expected of a CSD regardless of stimulation parameters (Figure S3) and causes Fos expression (Figure 5).

      Most importantly, the ECS as used by Rosenthal and colleagues is not a mouse model of ECT, in the sense that is not a scalp electrical stimulation. The stimulation method is fundamentally different between our two articles: (Rosenthal et al., 2025) implanted stimulating electrodes directly in the mice’s dorsal cranium. Direct cortical stimulation is well known to be able to cause CSD (Leao, 1944). However, it is unclear whether direct cortical stimulation is a useful model for ECT. We suspect that Rosenthal and colleagues were led to believe that auricular stimulation does not work because they saw no evidence of a cortical seizure in calcium recordings following stimulation. (See discussion on the confusion of phase III oscillation and seizure in the ECT literature.) We suspect that EEG recordings following direct cortical stimulation would reveal a very different EEG pattern from that observed in patients. This highlights the importance and novelty of the comparison of mouse and human EEG we present in Figure 1. Note, in our auricular stimulation preparation we do not observe any seizure-like activity in cortex that lasts beyond the stimulation (compare Rosenthal’s Figure 1E vs Figure 1F here). Given the EEG similarity we show between mice and patients, we suspect the cortical seizure Rosenthal and colleagues find is a methodological artifact. This is puzzling to us, as Rosenthal and colleagues do briefly mention a single mouse example with an extracellular electrophysiology recording compatible with CSD following auricular stimulation (Supplementary Figure 2).

      Thus, not only is it necessary to add evidence to the interpretation that the ECS-driven calcium event is indeed a CSD, but also that it can be triggered by a stimulation method that successfully replicates the known EEG response of human patients.

      (2) The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSDCa2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

      We have added additional references to the CSD-Fos literature in the discussion. Regarding the role of Fos as a marker of plasticity rather than activity, we discuss this point in our reply to comment (1) of reviewer 1. Concerning the expression of other genes, we already referenced the TRKB/BDNF pathways (Discussion). We have now added references on RNA-seq following CSD, which show that Fos is one of the most differentially expressed genes following CSD (Dell’Orco et al., 2023). Regarding the use of Fos knockout/knockdown lines, please see our reply to the reviewer’s comment (4).

      Reviewer #2 (Recommendations for the authors):

      (3) The Results and Discussion sections should be revised to better reflect the prior discovery of CSD following ECT in rodents and initial evidence in humans, as discussed in the first point in the Weaknesses section above.

      The prior discovery of the fact that ECS can drive CSD is first mentioned in the third sentence of the abstract “However, this view is challenged by the recent finding that electroconvulsive stimulation (ECS) can trigger a cortical spreading depression (CSD).” (The abstract has no references, but the introduction should make it clear what is meant).

      There is an entire paragraph of the introduction discussing the Rosenthal results.

      The first time we discuss our results (end of the Introduction) we say: “Consistent with previous work (Rosenthal et al., 2025), we observed a slow travelling calcium event that appeared to be a CSD.” The Rosenthal paper is cited in 8 times in total throughout the manuscript.

      We are unsure what the reviewer is asking us to do here. As mentioned above, if any revision is warranted regarding that reference, it should be to clarify that the Rosenthal paper used direct intracranial stimulation rather than ECS, and did not fully confirm the calcium event as a CSD – but they should be credited for finding the first preliminary evidence for CSD following ECS in patients.

      (4) To support the authors' statement that "CSD is the primary driver of plasticity is based on its role in driving Fos expression", additional experiments with Fos knockdown or knockout (or alternative interventions) are needed.

      Our main claim is that CSD causes Fos expression in the cortex following ECS. The argument that Fos can be used as a marker of plasticity follows from the literature, not from the experiments done here – this would require a form of functional plasticity measurement, which is outside the scope of this paper (and might not be the most relevant direction, see our reply to comment (3) of reviewer 1). The only observation that a Fos genetic manipulation would give us is the lack or reduction of Fos expression following CSD, which would be orthogonal to the points we make here.

      (5) It would be helpful to use a more specific term than "Ca2+ event" in Figure 2D and throughout the related Results section. It is assumed that this is the large propagating Ca2+ event attributed to CSD, but the terminology is important, as all the other events in the recording (including during Pre-ECT and the Direct ECT period) are also Ca2+ events.

      We have added a clear definition of what we mean on first usage (Results). The reason to call it a calcium event, and not a CSD, is that we did not want to jump to conclusions. We do think that the event is a CSD, but conclusive proof of that is still lacking (in both our work and that of Rosenthal). It is conceivable that the event propagates via a different mechanism than a CSD.

      (6) The authors should comment on differences among rodent models of ECT stimulation, especially with direct and ear clip methods, as discussed in the context of translational value (Theilmann et al., 2014).

      The Theilmann 2014 paper compares ECS delivered via auricular stimulation and intracranial electrodes. They compare the two stimulation methods, but use stimulation currents, total charges, and stimulation duration that were not matched. In the case of stimulation current, those used in auricular stimulation are approximately 8 fold higher than what they use for intracranial stimulation. Moreover, the stimulations were performed in awake rats. This is scientifically - and ethically, even for 2014 - questionable in light of the fact that this is aimed at developing a model for ECT, which is always done under general anesthesia. They conclude that cortical stimulation has less adverse effects and is more effective in reducing immobility in a forced swim test. The confound in the interpretation of these results is that all stimulation was performed in awake animals. The reason this is no longer done in humans is that it is extremely painful. Direct cortical stimulation is likely much less painful (for the same reason TMS is less painful). This would explain their findings of increased adverse effects with auricular electrodes. Given the differences in stimulation parameters used, the difficulty of calibrating equivalent doses of auricular and intracortical stimulations, as well as the small effect sizes reported, we don’t think the results allow for any solid conclusions as to which method is more effective in reducing immobility in a forced swim test. However, even if one would assume the intracortical ECS is more ‘effective’, this is hardly relevant, as we are interested in using mouse ECS as a model for human ECT. One could speculate that intracranial ECS might also exhibit higher clinical benefit than surface ECS in patients, but that is not the scope of our research, and probably not clinically relevant. Finally, while the authors do include EEG recordings in the rat – these were not compared to patient recordings, and from visual inspection do not resemble patient EEG recordings that we have seen. We would argue that the best rodent model of ECT stimulation is the one that triggers neuronal activity most similar to that observed in patients. We have added a brief discussion of these points to the corresponding Methods section.

      Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      This appears incorrect. It was known that direct cortical stimulation (as was done in the Rosenthal paper) can drive a calcium event that resembles a CSD. It was also known that CSD can drive Fos expression. What was not known is that the Fos expression driven by ECS is fully explained by the calcium event (putative CSD). This is particularly relevant as most people still erroneously assume Fos is a marker or neuronal activity, and prior work has come to the conclusion that ECT does not drive, but likely downregulates Fos expression (Calais et al., 2013; Morinobu et al., 1995; Park et al., 2014; Winston et al., 1990). We have added a more prominent discussion section on this point.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      The importance of Figure 1 is to demonstrate that ECS delivered with auricular electrodes in mice causes an EEG signature that is very similar to that seen in patients. We do not claim we are the first to describe characteristics of phase III oscillations in either patients or mice (we have added the Murakami reference to the manuscript). The reply to comments (1) and (6) of reviewer 2 highlights why this comparison is so important – it was not done in the Rosenthal paper the reviewer mentions below, and it is not clear whether the intracortical stimulation used there even drives a comparable EEG response (given the calcium activity shown, we suspect the answer is no). To the best of our knowledge this direct comparison is novel – but again the key novelty of our work we highlight is the one described in the title.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      It has indeed been demonstrated that ECS delivered using intracranial electrical stimulation can trigger CSD-like events (e.g. Rosenthal et al.). However, the fact that localized intracranial electrical stimulation can trigger a CSD has been shown quite a while ago already (see e.g. (Leao, 1944)). This is not the case for surface stimulation the way it is done in ECT and the way we do it. Nevertheless, note we give full credit to the Rosenthal work for making this connection. Our main contribution – as highlighted by the title – is showing that the Fos expression driven by ECS is fully explained by the CSD.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1.

      If the reviewer has references for this claim, we would be happy to add to the manuscript – we are not aware of any such work. As far as we are aware, this is still an area under active investigation – calcium signals correlate (locally) strongly with shank recordings (Wei et al., 2020), but how this translates into an EEG signal is speculative.

      It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      The reviewer may be jumping to conclusions here. It is correct that this has been shown for a CSD. The more important question (and the reason we did this experiment) is whether the calcium event triggered by ECS is indeed a CSD. We try to be careful to describe it as a calcium event in the results (mind you the calcium event in the Rosenthal is very likely a CSD as their intracortical stimulation (‘ECS’) is likely equivalent to the electrical stimulations used in the discovery of the CSD). We then perform a series of comparisons to see whether the calcium event has the known characteristics of a CSD – and we conclude everything we test is consistent with it being a CSD. Note, once again, that we do not claim novelty in any of this – the primary novelty is the link between ECS, Fos and CSD.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      That is correct. The question however is how much of the Fos expression is explained by CSD. Prior work that has looked at Fos expression in response to ECS has found that ECS results in a slight reduction of Fos expression (Calais et al., 2013; Park et al., 2014). We suspect this is the result of not triggering a CSD. We do not claim to have discovered that CSD induces Fos expression. The novel contribution, which is the main claim of the paper, is that in the context of ECS the entirety of the cortical Fos expression can be explained by the CSD. This links the relative contributions of multiple components of the ECS response to a known marker of neuronal plasticity.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      There is probably a misunderstanding here. We argue and show that there is no cortical seizure following ECS – that is why we refer to the EEG responses as phase III oscillations (characteristic of a silent cortex). And there is little prior evidence that the calcium events triggered by ECS are indeed a CSD (we think this is likely the case, but demonstrating this conclusively will require further work). If the reviewer has references for this, we would be happy to discuss. Note, the Rosenthal et al. paper just assumes (probably correctly) that they are a CSD, but does not demonstrate this. We don’t fully demonstrate this either, we just provide additional evidence. But once again, the novelty is in the title of the manuscript, and we do not claim any other novelty to the best of our reading of our manuscript. If the reviewer has a particularly misleading passage in mind, we are happy to rephrase.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      We are not sure what the reviewer means by “cFOS expression is nonspecific to plasticity”. Does the reviewer mean Fos has other roles beside driving neuronal plasticity? That is very likely correct, but it is unclear how that is relevant. Fos is a key driver of a number of neuronal plasticity pathways (Chaudhuri et al., 2000; Cohen and Greenberg, 2008; Cruz et al., 2015; Durchdewald et al., 2009). Which exact pathways are driven by ECS is an interesting question that we are currently pursuing in follow-up work. Given what we know about Fos expression, it is probably well within reasonable bounds to conclude that Fos increases result in neuronal plasticity. Fos expression directly drives network plasticity (Yap et al., 2021): "our findings indicate that Fos expression has an instructive role in orchestrating persistent circuit modifications”. Likewise, Fos induction is strictly required for experience-dependent representational plasticity during learning (de Hoz et al., 2018): “locally blocking c-Fos expression caused […] decreased cortical experience-dependent plasticity, without affecting baseline excitability or basic auditory processing.”. Calcium activity explains about 15% of the variance of Fos expression (Mahringer et al., 2022). This is likely driven by the correlation between activity and plasticity, not by a direct necessity for Fos expression to maintain neuronal activity. This is consistent with the finding that Fos as a transcription factor does not function to maintain spiking activity; it is part of the gene-regulatory machinery that converts patterned synaptic input into lasting plastic change. Fos is induced by NMDA/Ca2+, ERK, and CREB/Elk signaling rather than by firing alone (Fields et al., 1997; Xia et al., 1996), and those same pathways are required for the transcriptional program that stabilizes long-term potentiation and other durable synaptic modifications (Davis et al., 2000). As an AP-1 transcription factor, Fos drives downstream gene expression, so its appearance is better read as entry into a plasticity-related nuclear program than as a measure of ongoing excitability (Minatohara et al., 2015; Morgan and Curran, 1991; Sheng and Greenberg, 1990). That interpretation is consistent with our work showing that early Fos expression preferentially marks neurons that later undergo the strongest learning-related functional changes (Mahringer et al., 2019) and tracks functional reorganization during learning rather than simple recent activation (Mahringer et al., 2019).

      The link between CSD and therapeutic effect is a speculation we make in the manuscript, not a conclusion. We are of course in the process of performing follow-up work to test whether CSD in patients is a better predictor of treatment outcome than EEG based metrics – no experiment we can do in mice will be able to test the hypothesis that CSD is the mediator of the therapeutic benefit of ECT.

      Regarding the power of EEG metrics to predict therapeutic outcomes, we fully agree with the reviewer. The literature on the reliability of EEG metrics computed from phase III oscillation data is controversial and noisy – this is something we establish in the introduction to motivate our research into other biological processes that could explain how ECT works. Note however, this is certainly not a fringe view – see e.g. comment (2) of reviewer 1: “Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy“ We think the reason for this is that a CSD is necessary for therapeutic benefit, but only has minor effects on the phase III oscillation – note this is a hypothesis based on our results that is trivial to test in patients (which are in currently investigating).

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      Whenever possible, we use mice for multiple experiments in the lab. This is done to reduce the total number of mice used for experiments for ethical reasons. For the mice in question, ChrimsonR was injected in the retrosplenial cortex (Methods), which was originally done to stimulate locally the axons projecting in the imaging area (primary visual cortex here). Thus the labelling is very sparse, and perfectly compatible with two-photon GCaMP6f imaging. Fluorescence emission is far too weak to activate ChrimsonR.

      Qualitatively, one can mentally compare the light power we use to activate optogenetic tools, which tends to be blindingly bright (one shouldn’t look into the optogenetic stimulation laser), with the fluorescence emission from two-photon imaging of calcium indicators, which tends to be barely visible by eye.

      Quantitatively, one can estimate this as follows: At 510nm emission, each photon carries an energy of about . Assuming a neuron that strongly expresses GCaMP6f under two-photon excitation would emit a very high 10<sup>7</sup>photons per second (Har-Gil et al., 2018), its total emission power would be P = N<sub>photons</sub> * E ≈ 4pW. Assuming this is spread over the surface of the cell (sphere with 10 µm diameter), this translates to . This is several orders of magnitude below the value of 1 mW/mm<sup>2</sup> irradiance required to activate ChrimsonR modestly at peak absorption, which is 80nm away from GCaMP6f emission (Klapoetke et al., 2014). Note that this would be true even when ChrimsonR is injected at the site of imaging (Vasilevskaya and Keller, 2026).

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      We suspect that the question is driven by a misunderstanding. While it is likely that the implantation triggers a CSD (and likely so does a standard two-photon window implantation), the implantation surgery and the experiments/imaging are separated by at least 3 weeks (Methods). We have never observed spontaneous CSDs in the days and weeks following an implantation surgery.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      We have added additional details as requested by the reviewer (Methods).

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      We have added two tables (Table S3 and Table S4) that displays the stimulation parameters used for each figure, as well as the distribution of parameters for mice and patients. More importantly, the properties of the travelling calcium event do not depend on the stimulation charge (Figure S3), which removes this confounding factor and supports the idea that the calcium event is a CSD.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      That is correct – only the cortical expression of Fos depends on the cortical CSD (see our reply to comment (1) of reviewer 1). Quantification of the whole-brain Fos expression following ECS is outside the scope of the manuscript, but it is something we are currently pursuing for separate publication. We have now reworked the section describing the Fos expression to make it clear that we are only talking about cortical expression of Fos (Results).

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      That was indeed our speculation based on ongoing work on TMS in the lab – but it was poorly phrased and unnecessary. We have rephrased.

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

      It is speculative indeed – but the speculation is not unfounded and has been made previously. The primary evidence is correlative in that schizophrenia is characterized by a reduction in coupling between thalamus and frontal areas of cortex – see e.g. (Giraldo-Chica and Woodward, 2017; Vinogradov et al., 2023). We have argued in previous work that combining this with computational models of psychosis (Sterzer et al., 2018), it is not unreasonable to speculate that a reduction in the coupling between thalamus and frontal areas of cortex could explain psychosis (Keller and Sterzer, 2024). The value of that speculation here is that it forms a testable hypothesis for the mechanism of action of ECT. Nevertheless, we have rephrased the statement slightly to make it clearer that this is still speculation.

      Reviewer #3 (Recommendations for the authors):

      The authors should design/execute an experiment(s) to attempt to prove causality between CSD, calcium influx, cFOS expression, and therapeutic effect. At a minimum, the authors need to revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field.

      Regarding novelty – see discussion above.

      Regarding therapeutic effects – we are in the process of testing this in patients. Using CSD measurements during ECT to test whether CSD is a better predictor of treatment outcome. I suspect that this will, however, require many more labs to come to firm conclusions. Our aim here is to inspire these experiments. That is why we speculate about clinical relevance in the abstract and the discussion.

      References

      Bahari, F., Ssentongo, P., Liu, J., Kimbugwe, J., Curay, C., Schiff, S.J., Gluckman, B.J., 2020. Seizure-associated spreading depression is a major feature of ictal events in two animal models of chronic epilepsy. https://doi.org/10.1101/455519

      Bolhuis, J.J., Hetebrij, E., Den Boer-Visser, A.M., De Groot, J.H., Zijlstra, G.G.O., 2001. Localized immediate early gene expression related to the strength of song learning in socially reared zebra finches. European Journal of Neuroscience 13, 2165–2170. https://doi.org/10.1046/j.0953816x.2001.01588.x

      Bonaccini Calia, A., Masvidal-Codina, E., Smith, T.M., Schäfer, N., Rathore, D., Rodríguez-Lucas, E., Illa, X., De la Cruz, J.M., Del Corro, E., Prats-Alfonso, E., Viana, D., Bousquet, J., Hébert, C., Martínez-Aguilar, J., Sperling, J.R., Drummond, M., Halder, A., Dodd, A., Barr, K., Savage, S., Fornell, J., Sort, J., Guger, C., Villa, R., Kostarelos, K., Wykes, R.C., Guimerà-Brunet, A., Garrido, J.A., 2022. Full-bandwidth electrophysiology of seizures and epileptiform activity enabled by flexible graphene microtransistor depth neural probes. Nat. Nanotechnol. 17, 301–309. https://doi.org/10.1038/s41565-021-01041-9

      Brumback, R.A., Staton, R.D., 1982. The Electroencephalographic Pattern during Electroconvulsive Therapy. Clinical Electroencephalography 13, 148–153. https://doi.org/10.1177/155005948201300306

      Calais, J.B., Valvassori, S.S., Resende, W.R., Feier, G., Athié, M.C.P., Ribeiro, S., Gattaz, W.F., Quevedo, J., Ojopi, E.B., 2013. Long-term decrease in immediate early gene expression after electroconvulsive seizures. J Neural Transm (Vienna) 120, 259–266. https://doi.org/10.1007/s00702-012-0861-4

      Chaudhuri, A., Zangenehpour, S., Rahbar-Dehgan, F., Ye, F., 2000. Molecular maps of neural activity and quiescence. Acta Neurobiol Exp (Wars) 60, 403–410. https://doi.org/10.55782/ane-2000-1359

      Cohen, S., Greenberg, M.E., 2008. Communication between the synapse and the nucleus in neuronal development, plasticity, and disease. Annu Rev Cell Dev Biol 24, 183–209. https://doi.org/10.1146/annurev.cellbio.24.110707.175235

      Cruz, F.C., Javier Rubio, F., Hope, B.T., 2015. Using c-fos to study neuronal ensembles in corticostriatal circuitry of addiction. Brain Res 1628, 157–173. https://doi.org/10.1016/j.brainres.2014.11.005

      Davis, S., Vanhoutte, P., Pages, C., Caboche, J., Laroche, S., 2000. The MAPK/ERK cascade targets both Elk-1 and cAMP response element-binding protein to control long-term potentiation-dependent gene expression in the dentate gyrus in vivo. J Neurosci 20, 4563–4572. https://doi.org/10.1523/JNEUROSCI.20-12-04563.2000

      de Hoz, L., Gierej, D., Lioudyno, V., Jaworski, J., Blazejczyk, M., Cruces-Solís, H., Beroun, A., Lebitko, T., Nikolaev, T., Knapska, E., Nelken, I., Kaczmarek, L., 2018. Blocking c-Fos Expression Reveals the Role of Auditory Cortex Plasticity in Sound Frequency Discrimination Learning. Cereb Cortex 28, 1645–1655. https://doi.org/10.1093/cercor/bhx060

      Dell’Orco, M., Weisend, J.E., Perrone-Bizzozero, N.I., Carlson, A.P., Morton, R.A., Linsenbardt, D.N., Shuttleworth, C.W., 2023. Repetitive spreading depolarization induces gene expression changes related to synaptic plasticity and neuroprotective pathways. Front. Cell. Neurosci. 17. https://doi.org/10.3389/fncel.2023.1292661

      Durchdewald, M., Angel, P., Hess, J., 2009. The transcription factor Fos: a Janus-type regulator in health and disease. Histol Histopathol 24, 1451–1461. https://doi.org/10.14670/HH-24.1451

      Fields, R.D., Eshete, F., Stevens, B., Itoh, K., 1997. Action potential-dependent regulation of gene expression: temporal specificity in ca2+, cAMP-responsive element-binding proteins, and mitogen-activated protein kinase signaling. J Neurosci 17, 7252–7266. https://doi.org/10.1523/JNEUROSCI.17-19-07252.1997

      Fleischmann, A., Hvalby, O., Jensen, V., Strekalova, T., Zacher, C., Layer, L.E., Kvello, A., Reschke, M., Spanagel, R., Sprengel, R., Wagner, E.F., Gass, P., 2003. Impaired Long-Term Memory and NR2AType NMDA Receptor-Dependent Synaptic Plasticity in Mice Lacking c-Fos in the CNS. J. Neurosci. 23, 9116–9122. https://doi.org/10.1523/JNEUROSCI.23-27-09116.2003

      Francis-Taylor, R., Ophel, G., Martin, D., Loo, C., 2020. The ictal EEG in ECT: A systematic review of the relationships between ictal features, ECT technique, seizure threshold and outcomes. Brain Stimulation 13, 1644–1654. https://doi.org/10.1016/j.brs.2020.09.009

      Giraldo-Chica, M., Woodward, N.D., 2017. Review of thalamocortical resting-state fMRI studies in schizophrenia. Schizophr Res 180, 58–63. https://doi.org/10.1016/j.schres.2016.08.005

      Har-Gil, H., Golgher, L., Israel, S., Kain, D., Cheshnovsky, O., Parnas, M., Blinder, P., 2018. PySight: plug and play photon counting for fast continuous volumetric intravital microscopy. Optica 5, 1104. https://doi.org/10.1364/OPTICA.5.001104

      Huels, E.R., Kafashan, M., Hickman, L.B., Ching, S., Lin, N., Lenze, E.J., Farber, N.B., Avidan, M.S., Hogan, R.E., Palanca, B.J.A., 2023. Central-positive complexes in ECT-induced seizures: Possible evidence for thalamocortical mechanisms. Clin Neurophysiol 146, 77–86. https://doi.org/10.1016/j.clinph.2022.11.015

      Kawai, R., Markman, T., Poddar, R., Ko, R., Fantana, A.L., Dhawale, A.K., Kampff, A.R., Ölveczky, B.P., 2015. Motor cortex is required for learning but not for executing a motor skill. Neuron 86, 800– 812. https://doi.org/10.1016/j.neuron.2015.03.024

      Keller, G.B., Sterzer, P., 2024. Predictive Processing: A Circuit Approach to Psychosis. Annu Rev Neurosci 47, 85–101. https://doi.org/10.1146/annurev-neuro-100223-121214

      Kimpo, R.R., Doupe, A.J., 1997. FOS Is Induced by Singing in Distinct Neuronal Populations in a Motor Network. Neuron 18, 315–325. https://doi.org/10.1016/S0896-6273(00)80271-8

      Klapoetke, N.C., Murata, Y., Kim, S.S., Pulver, S.R., Birdsey-Benson, A., Cho, Y.K., Morimoto, T.K., Chuong, A.S., Carpenter, E.J., Tian, Z., Wang, J., Xie, Y., Yan, Z., Zhang, Y., Chow, B.Y., Surek, B., Melkonian, M., Jayaraman, V., Constantine-Paton, M., Wong, G.K.-S., Boyden, E.S., 2014. Independent optical excitation of distinct neural populations. Nat Methods 11, 338–346. https://doi.org/10.1038/nmeth.2836

      Largo, C., Ibarz, J.M., Herreras, O., 1997. Effects of the Gliotoxin Fluorocitrate on Spreading Depression and Glial Membrane Potential in Rat Brain In Situ. Journal of Neurophysiology 78, 295–307. https://doi.org/10.1152/jn.1997.78.1.295

      Leao, A.A.P., 1944. Spreading depression of activity in the cerebral cortex. Journal of Neurophysiology 7, 359–390. https://doi.org/10.1152/jn.1944.7.6.359

      Mahringer, D., Petersen, A.V., Fiser, A., Okuno, H., Bito, H., Perrier, J.-F., Keller, G.B., 2019. Expression of c-Fos and Arc in hippocampal region CA1 marks neurons that exhibit learning-related activity changes. https://doi.org/10.1101/644526

      Mahringer, D., Zmarz, P., Okuno, H., Bito, H., Keller, G.B., 2022. Functional correlates of immediate early gene expression in mouse visual cortex. Peer Community Journal 2. https://doi.org/10.24072/pcjournal.156

      Meeren, H.K.M., Pijn, J.P.M., Luijtelaar, E.L.J.M.V., Coenen, A.M.L., Silva, F.H.L. da, 2002. Cortical Focus Drives Widespread Corticothalamic Networks during Spontaneous Absence Seizures in Rats. J. Neurosci. 22, 1480–1495. https://doi.org/10.1523/JNEUROSCI.22-04-01480.2002

      Minatohara, K., Akiyoshi, M., Okuno, H., 2015. Role of Immediate-Early Genes in Synaptic Plasticity and Neuronal Ensembles Underlying the Memory Trace. Front Mol Neurosci 8, 78. https://doi.org/10.3389/fnmol.2015.00078

      Morgan, J.I., Curran, T., 1991. Stimulus-transcription coupling in the nervous system: involvement of the inducible proto-oncogenes fos and jun. Annu Rev Neurosci 14, 421–451. https://doi.org/10.1146/annurev.ne.14.030191.002225

      Morinobu, S., Nibuya, M., Duman, R.S., 1995. Chronic antidepressant treatment down-regulates the induction of c-fos mRNA in response to acute stress in rat frontal cortex. Neuropsychopharmacology 12, 221–228. https://doi.org/10.1016/0893-133X(94)00067-A

      Nakadate, K., Imamura, K., Watanabe, Y., 2012. Effects of monocular deprivation on the spatial pattern of visually induced expression of c-Fos protein. Neuroscience 202, 17–28. https://doi.org/10.1016/j.neuroscience.2011.12.004

      Pandey, A., Kang, S., Pacchiarini, N., Wyszynska, H., Masseri, Z., O’Neill, J., Honey, R.C., Fox, K., 2026. Secondary Somatosensory Cortex Is Required for Learning but Not Execution of a Tactile Discrimination. Eur J Neurosci 63, e70390. https://doi.org/10.1111/ejn.70390

      Park, H.G., Yu, H.S., Park, S., Ahn, Y.M., Kim, Y.S., Kim, S.H., 2014. Repeated treatment with electroconvulsive seizures induces HDAC2 expression and down-regulation of NMDA receptorrelated genes through histone deacetylation in the rat frontal cortex. Int J Neuropsychopharmacol 17, 1487–1500. https://doi.org/10.1017/S1461145714000248

      Polack, P.-O., Mahon, S., Chavez, M., Charpier, S., 2009. Inactivation of the Somatosensory Cortex Prevents Paroxysmal Oscillations in Cortical and Related Thalamic Neurons in a Genetic Model of Absence Epilepsy. Cereb Cortex 19, 2078–2091. https://doi.org/10.1093/cercor/bhn237

      Rosenthal, Z.P., Majeski, J.B., Somarowthu, A., Quinn, D.K., Lindquist, B.E., Putt, M.E., Karaj, A., Favilla, C.G., Baker, W.B., Hosseini, G., Rodriguez, J.P., Cristancho, M.A., Sheline, Y.I., William Shuttleworth, C., Abbott, C.C., Yodh, A.G., Goldberg, E.M., 2025. Electroconvulsive therapy generates a postictal wave of spreading depolarization in mice and humans. Nat Commun 16, 4619. https://doi.org/10.1038/s41467-025-59900-1

      Roy, D.S., Arons, A., Mitchell, T.I., Pignatelli, M., Ryan, T.J., Tonegawa, S., 2016. Memory retrieval by activating engram cells in mouse models of early Alzheimer’s disease. Nature 531, 508–512. https://doi.org/10.1038/nature17172

      Ryan, T.J., Roy, D.S., Pignatelli, M., Arons, A., Tonegawa, S., 2015. Memory. Engram cells retain memory under retrograde amnesia. Science 348, 1007–1013. https://doi.org/10.1126/science.aaa5542

      Scangos, K.W., Weiner, R.D., Coffey, C.E., Krystal, A.D., 2019. An electrophysiological biomarker that may predict treatment response to ECT. J ECT 35, 95–102. https://doi.org/10.1097/YCT.0000000000000557

      Sheng, M., Greenberg, M.E., 1990. The regulation and function of c-fos and other immediate early genes in the nervous system. Neuron 4, 477–485. https://doi.org/10.1016/0896-6273(90)90106-p

      Sterzer, P., Adams, R.A., Fletcher, P., Frith, C., Lawrie, S.M., Muckli, L., Petrovic, P., Uhlhaas, P., Voss, M., Corlett, P.R., 2018. The Predictive Coding Account of Psychosis. Biol Psychiatry 84, 634–643. https://doi.org/10.1016/j.biopsych.2018.05.015

      Tanaka, K.Z., He, H., Tomar, A., Niisato, K., Huang, A.J.Y., McHugh, T.J., 2018. The hippocampal engram maps experience but not place. Science 361, 392–397. https://doi.org/10.1126/science.aat5397

      Tyssowski, K.M., DeStefino, N.R., Cho, J.-H., Dunn, C.J., Poston, R.G., Carty, C.E., Jones, R.D., Chang, S.M., Romeo, P., Wurzelmann, M.K., Ward, J.M., Andermann, M.L., Saha, R.N., Dudek, S.M., Gray, J.M., 2018. Different Neuronal Activity Patterns Induce Different Gene Expression Programs. Neuron 98, 530-546.e11. https://doi.org/10.1016/j.neuron.2018.04.001

      Vasilevskaya, A., Keller, G.B., 2026. A functional influence based circuit motif that constrains the set of plausible algorithms of cortical function. eLife 15. https://doi.org/10.7554/eLife.110827.1

      Vinogradov, S., Chafee, M.V., Lee, E., Morishita, H., 2023. Psychosis spectrum illnesses as disorders of prefrontal critical period plasticity. Neuropsychopharmacology 48, 168–185. https://doi.org/10.1038/s41386-022-01451-w

      Watanabe, Y., Johnson, R.S., Butler, L.S., Binder, D.K., Spiegelman, B.M., Papaioannou, V.E., McNamara, J.O., 1996. Null Mutation of c-fos Impairs Structural and Functional Plasticities in the Kindling Model of Epilepsy. J Neurosci 16, 3827–3836. https://doi.org/10.1523/JNEUROSCI.16-1203827.1996

      Wei, Z., Lin, B.-J., Chen, T.-W., Daie, K., Svoboda, K., Druckmann, S., 2020. A comparison of neuronal population dynamics measured with calcium imaging and electrophysiology. PLoS Comput Biol 16, e1008198. https://doi.org/10.1371/journal.pcbi.1008198

      Winston, S.M., Hayward, M.D., Nestler, E.J., Duman, R.S., 1990. Chronic electroconvulsive seizures down-regulate expression of the immediate-early genes c-fos and c-jun in rat cerebral cortex. J Neurochem 54, 1920–1925. https://doi.org/10.1111/j.1471-4159.1990.tb04892.x

      Xia, Z., Dudek, H., Miranti, C.K., Greenberg, M.E., 1996. Calcium influx via the NMDA receptor induces immediate early gene transcription by a MAP kinase/ERK-dependent mechanism. J Neurosci 16, 5425–5436. https://doi.org/10.1523/JNEUROSCI.16-17-05425.1996

      Yap, E.-L., Pettit, N.L., Davis, C.P., Nagy, M.A., Harmin, D.A., Golden, E., Dagliyan, O., Lin, C., Rudolph, S., Sharma, N., Griffith, E.C., Harvey, C.D., Greenberg, M.E., 2021. Bidirectional perisomatic inhibitory plasticity of a Fos neuronal network. Nature 590, 115–121. https://doi.org/10.1038/s41586-020-3031-0

      Yassin, L., Benedetti, B.L., Jouhanneau, J.-S., Wen, J.A., Poulet, J.F.A., Barth, A.L., 2010. An embedded subnetwork of highly active neurons in the neocortex. Neuron 68, 1043–1050. https://doi.org/10.1016/j.neuron.2010.11.029

    1. eLife Assessment

      This study presents a valuable contribution to comparative cognitive neuroscience by directly mapping functional homologues of the human multiple-demand network in macaques using a matched spatial maze task. However, the evidence is incomplete due to methodological asymmetries in the task design that appear to affect the results in ways that are not fully analyzed and discussed. The work will be of interest to researchers studying the evolution of cognitive control and cross-species neuroimaging.

    2. Reviewer #1 (Public review):

      Summary:

      The "multiple-demand" (MD) system is a well-known finding of human brain imaging and is thought to play a central role in cognitive control. To directly compare the MD system in humans and monkeys, Mione et al. used functional magnetic resonance imaging to measure whole brain activation in a multi-step saccadic maze task. In humans, the authors found a distributed pattern of brain activity close match to the canonical MD network and extending to adjacent regions of dorsal attention and other networks. While there was good correspondence between monkey and human data, differences were also notable in lateral frontal cortex, dorsal parietal cortex, and sensorimotor cortex.

      Strengths:

      Though previous data hint at a corresponding network in the macaque, there has been no direct comparison to human data. This study provides a direct cross-species comparison with whole-brain data of fMRI, and the findings suggest an extended and strongly interconnected brain network recruited by increased cognitive challenge.

      Weaknesses:

      In previous human imaging, the MD system is defined by overlapping activation for many kinds of cognitive demand. In the present work, however, the authors used just a single task. Although there is some overlap between putative monkey MD network and canonical MD network identified in human imaging, it should be cautious to link current findings to MD system based on limited task events.

    3. Reviewer #2 (Public review):

      Summary:

      Mione et al. aim to resolve a long-standing question in comparative neuroscience: whether the macaque brain contains a functional analogue to the distributed human multiple-demand (MD) network. To address this, the authors employ a direct cross-species fMRI comparison using a multi-step saccadic maze task in humans and a simplified two-step version in macaques. By contrasting goal-directed navigation against a control condition that requires similar motor responses but no strategic planning, the study isolates the neural signatures of cognitive control across species.

      Strengths:

      The most compelling aspect of this work is its methodological alignment. Previous attempts to compare these systems often relied on comparisons of human BOLD signals and macaque single-unit recordings. By running parallel fMRI protocols, the authors establish a shared measurement basis that allows for a more direct comparison. The resulting activation maps demonstrate conserved network topology in dorsolateral and dorsomedial frontal cortex and insula.

      Weaknesses:

      However, there is concerning inter-individual variability in the macaque data, along with a notable design asymmetry between the human and monkey tasks. In the human experiment, 2-, 4-, and 6-step trials were mixed within the same blocks, whereas macaques performed only 2-step problems. This design difference likely places human participants in a state of sustained proactive cognitive control (Braver, 2012), as they must remain prepared for more demanding trials at any moment. This elevated baseline arousal may potentially inflate MD network activation even during the simpler 2-step trials in humans, making direct comparisons with the macaque data difficult.

      The authors acknowledge this concern and note that similarities between species survive despite this procedural difference. However, their own individual-level data raise doubts about how robust these "similarities" really are. A defining feature of the MD system is consistent recruitment across individuals performing the same task. In this study, however, the two monkeys show markedly different activation maps in parietal cortex. The conjunction map (Supplementary Figure 6) does not support the claim that lateral and medial parietal regions are commonly activated across both animals at any meaningful statistical threshold, even when using a permissive threshold of |t| > 1.5, corresponding to approximately p ≈ 0.13 - 0.14 (two-tailed).

      This suggests that the group-level parietal activations are largely driven by a single animal, raising questions about whether these regions can reliably be considered as part of the monkey MD network. Notably, in Premereur et al. (2018), the only other direct macaque fMRI study of a cognitively demanding task, their conjunction map between the two animals also failed to show common parietal activation during task switching. Taken together, these findings suggest that either the monkey MD network is not as similar to its human counterpart as claimed, or the monkey version of the maze task was simply not sufficiently challenging to fully engage parietal MD nodes.

      Because of the design asymmetry described above, it remains difficult to determine which of these possibilities is correct (a genuine species difference versus insufficient task demands in macaques). The authors' current claims treat the group-level activations, which appear largely driven by one animal, as reliable evidence for homology, but without consistent individual-level support, such conclusions remain tentative.

      In summary, the authors demonstrate functional correspondence between human and macaque cognitive control networks in dorsolateral and dorsomedial frontal regions. However, assertions that this extends to lateral and medial parietal cortex are not consistently supported at the individual level. While the overall conclusions are plausible, they would be significantly strengthened by demonstrating consistent activation across both animals in all claimed MD regions, ideally with a properly thresholded conjunction map.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The "multiple-demand" (MD) system is a well-known finding of human brain imaging and is thought to play a central role in cognitive control. To directly compare the MD system in humans and monkeys, Mione et al. used functional magnetic resonance imaging to measure whole-brain activation in a multi-step saccadic maze task. In humans, the authors found a distributed pattern of brain activity close match to the canonical MD network and extends to adjacent regions of dorsal attention and other networks. While there was good correspondence between monkey and human data, differences were also notable in the lateral frontal cortex, the dorsal parietal cortex, and the sensorimotor cortex.

      Strengths:

      Though previous data hint at a corresponding network in the macaque, there has been no direct comparison to human data. This study provides a direct cross-species comparison with whole-brain data from fMRI, and the findings suggest an extended and strongly interconnected brain network recruited by increased cognitive challenge.

      Weaknesses:

      In previous human imaging, the MD system is defined by overlapping activation for many kinds of cognitive demands. In the present work, however, the authors used just a single task. Although there is some overlap between the putative monkey MD network and the canonical MD network identified in human imaging, there should be caution in linking current findings to the MD system based on limited task events.

      In the Discussion, we acknowledge the limitation of using a single task. With this one task, however, the canonical MD network is clearly shown in our human data. Accompanying activation, especially of the canonical dorsal attention network, likely reflects the specific spatial demands of the maze task. In the monkey data, much of this dorsal attention activity is not seen (e.g. superior and medial parietal), likely reflecting limited power. Instead, there is distributed overlap with the previous limited monkey studies than have contrasted higher with lower cognitive demand. Though we agree that task-specific activations may contribute to our results, these arguments suggest that, in large part, our method does successfully identify a distributed set of multiple-demand regions. At the same time we acknowledge the desirability of further work to examine a wider range of task demands.

      Reviewer #1 (Recommendations for the authors):

      (1) Though the whole-brain data obtained by fMRI can provide a direct comparison between species, a single cognitive task might be insufficient to link the findings with the MD system. A cognitively challenging task likely activates multiple regions of the MD system; however, it may also recruit some task-specific regions, which do not belong to the canonical MD network. Furthermore, this is probably the reason that the dorsal attention network showed the strongest activation rather than the core MD for the current visuospatial maze task.

      In the Discussion, we acknowledge the limitation of using a single task (p. 15-16), and the likely contribution of task-specific activations to our data, especially involving the dorsal attention network (p. 15).

      (2) Ideally, a meta-analysis recruiting more fMRI studies on humans and monkeys when they perform various similar cognitive tasks may strengthen the evidence that there is a comparable MD system between species.

      Many human meta-analyses, of course, show the common MD system. For monkeys, however, at least to our knowledge, there are insufficient studies for a similar meta-analysis. Instead we discuss overlaps between the current activation findings and two previous studies of respectively antisaccades and task switching (p. 15), suggesting that, in monkey as in human, there is convergence for different kinds of demand.

      (3) Different from human subjects, monkeys usually require substantial training before fMRI scanning. More details about the training of the two monkeys should be given. In addition, training may reshape cognitive task activations. This potential impact should also be discussed.

      We now address this point in the Discussion (p. 18). As we note, similar results for the two species apparently survive even large differences in protocol. A training summary has been added to Methods (p. 27).

      (4) According to the description in the text, the two monkeys have obvious differences in cognitive task activations. It is necessary to show individual-level brain activations as well.

      Individual results and a conjunction map are shown in Supplementary Figure 6 (see accompanying text on p. 13). As expected, the conjunction map had substantial similarity to the findings from the two animals combined. Similarities and differences between animals are addressed in the Discussion (p. 17).

      (5) Based on the current research content, the title seems to be too general.

      For the reasons given above (see point 2), we think our title is reasonable.

      Reviewer #2 (Public review):

      Summary:

      Mione et al. aim to resolve a long-standing question in comparative neuroscience: whether the macaque brain contains a functional analogue to the distributed human multiple demand (MD) network. To address this, the authors employ a direct cross-species fMRI comparison using a multi-step saccadic maze task in humans and a simplified two-step version in macaques. By contrasting goal-directed navigation against a control condition that requires similar motor responses but no strategic planning, the study isolates the neural signatures of cognitive control across species.

      Strengths:

      The most compelling aspect of this work is its methodological alignment. Previous attempts to compare these systems often relied on comparisons of human BOLD signals and macaque single-unit recordings. By running parallel fMRI protocols, the authors establish a shared measurement basis that allows for a more direct comparison. The resulting activation maps clearly demonstrate conserved network topology across dorsomedial frontal, lateral, and medial parietal, and insula cortices. Combining these results with recent research on functional and structural connectivity further supports the idea that these networks evolved across species and provides a helpful starting point for future comparative studies. The findings will be highly useful for researchers investigating the evolutionary origins of domain-general cognitive control, as well as for neuroimaging methodologists developing cross-species alignment pipelines.

      Weaknesses:

      However, there are several differences in how the two groups were studied that make it harder to compare the results precisely. The human task mixed 2-, 4-, and 6-step trials within the same experimental blocks, whereas macaques performed only 2-step trials. This design difference likely places human participants in a state of sustained proactive cognitive control (Braver, 2012), as they must remain prepared for highly demanding trials at any moment. This elevated baseline arousal may artificially inflate MD network activation during the simpler 2-step trials in humans, making direct magnitude comparisons with the macaque data difficult.

      This is a reasonable concern, which we note in the Discussion. Crucially, as we point out, similarities between species appear to survive this and other differences in procedure.

      Additionally, the general linear model combined correct and error trials into a single regressor. Given that macaques exhibited substantially higher error rates, this approach risks diluting task-specific planning signals with activity related to error monitoring and reward prediction errors. The preprocessing pipeline also applied a 4 mm full-width half-maximum smoothing kernel to macaque data acquired at 1.5 mm resolution. Relative to the smaller size of the macaque brain, this kernel is quite large and likely blurs fine-grained topographical distinctions. This may partly explain why the macaque lateral frontal cortex shows a single dorsal activation patch rather than multiple discrete patches seen in humans.

      These are also reasonable concerns. To address them, we have added supplementary analyses using (a) only correct trials, and (b) a smaller smoothing kernel (see Supplementary Figure 3 and 4). In both cases, there is some loss of power, but otherwise similar results.

      Furthermore, there is concerning inter-individual variability in the macaque data. Normally, a functional network like the MD system is identified by consistent activation across all individuals. In this study, however, the two monkeys show substantially different activation maps and behavioral patterns. This lack of consistency renders the group-level results questionable, as it is unclear whether the group-level map represents a unified biological system or merely an average of disparate individual maps.

      Results for individual animals have now been supplemented with a conjunction map (Supplementary Figure 6). As expected, the conjunction map is similar to the map obtained by pooling data from the two animals.

      Finally, the subcortical activations shown in Figure 7 require more precise anatomical localization to confidently distinguish cerebellar nodes from adjacent brainstem structures.

      The slices shown in Figure 7 have been amended to better show the detail of activation outside cerebral cortex, in particular in cerebellum.

      The authors demonstrate a broad functional correspondence between human and macaque cognitive control networks, moving the field beyond speculative homology. The data suggest that an extended, interconnected network is recruited by cognitive challenge in both species; however, the strength of this claim is limited by the inter-individual variability and methodological constraints noted above. Assertions of precise topological equivalence should therefore be tempered. The absence of ventrolateral prefrontal and strong dorsal parietal activations in the macaque group analysis may reflect genuine biological differences, but could also stem from limited statistical power, excessive smoothing, or task design asymmetries. While the overall conclusions are plausible, they would be significantly strengthened by a more explicit discussion of these limitations and additional analytical clarifications regarding individual-level consistency.

      Indeed, as we note in the Discussion, there is a strong possibility that ventrolateral frontal and dorsal parietal activation are missing in our monkey data because of limited statistical power. Further work would be needed to address these possible limitations.

      Reviewer #2 (Recommendations for the authors):

      (1) Please discuss how the mixed-difficulty block design in humans may induce a state of proactive cognitive control that elevates baseline MD activation compared to the fixed 2step macaque condition. A supplementary analysis comparing early versus late block 2-step trials in humans would help clarify whether activation magnitudes reflect sustained task-set maintenance or transient trial demands.

      We now note the potential importance of this in the Discussion (p. 18), and point out that similarities between species appear to survive this difference in procedure.

      (2) Please clarify the rationale for combining correct and error trials in a single regressor.

      Given the higher macaque error rates, please provide a supplementary GLM restricted to correct trials only for the monkey data. This will demonstrate whether the core activation topography remains consistent when error-related signals are excluded.

      A new analysis addresses this point (Supplementary Figure 3). As we note, restricting analysis to correctly-completed problems somewhat reduces power but leaves major features of the results intact.

      (3) Please justify the use of a 4 mm FWHM smoothing kernel for macaque data, given cortical thickness and brain size differences. If feasible, re-run the analysis with a smaller kernel or surface-based smoothing to assess whether finer topographical distinctions emerge in lateral frontal and parietal cortices.

      This new analysis has also been run (Supplementary Figure 4), again with reduced power but major features of the results intact. We also refer to previous work indicating choice of either 3 mm or 4 mm smoothing for macaque fMRI (p. 13), approximately matching the smoothing needed to align electrophysiological and fMRI maps (Issa et al., 2013, J.Neurosci.).

      (4) Please detail how the human '2-step problems only' analysis in Supplementary Figure 1 was specified in the GLM. Explicitly state whether 4- and 6-step trials were modeled as separate regressors of no interest to prevent hemodynamic bleed-over from contaminating the 2-step beta weights.

      Indeed, 4- and 6-step trials were removed using regressors of no interest, as now specified in Methods (p. 24).

      (5) Please provide higher-resolution slices or probabilistic atlas overlays for the cerebellar and subcortical activations in Figure 7. This will help clearly distinguish cerebellar hemispheres from adjacent brainstem structures and ensure anatomical labeling is accurate.

      To address this question, additional slices have been added to a revised Figure 7, in particular adding detail to cerebellar activation.

      (6) Please address the substantial inter-individual variability in the macaque data, particularly regarding Monkey B's performance. Even after excluding five poor-performing sessions, Monkey B only reached about 65% accuracy on the first step of the maze task. This suggests that the animal was largely guessing rather than following a strategic, goal directed plan. The notably higher accuracy on step 2 (>80%) could simply be a selection bias artifact, as the task terminates immediately following step 1 errors, meaning step 2 trials only occur when step 1 was already correct. Including an animal that likely did not fully grasp the overarching task rule in a cohort of N=2 raises serious concerns about signal dilution and the reliability of the group-level maps. Please explicitly justify why Monkey B's data were retained despite these performance concerns, or consider excluding this animal and acquiring data from a third, behaviorally reliable subject. At minimum, report a conjunction map showing regions strictly active in both animals alongside the combined analysis, and present individual subject maps more prominently to improve transparency.

      Though we appreciate this concern, we do not think the data indicate that monkey B failed to use the maze goal to constrain choices. As shown in Figure 3, the great majority of errors were timing errors (mostly not waiting for go signal). Excluding these, for step 1, of choices directed to one of the two available alternatives, 88% were correct (Figure 4, compare “correct” with “wrong open location”).

      As noted above, we have added a conjunction map (Supplementary Figure 6) to our previous presentation of individual data. We note (p. 13) that, as expected, this map strongly resembles results from the combined-animal analysis.

    1. eLife Assessment

      This valuable descriptive study describes the expression of a developmentally relevant transcription factor in the adult Tribolium brain. The evidence supporting the claims is convincing and based on a very detailed and rigorous analysis of light microscopy data, which, however, lacks single-cell resolution. This neuroanatomical study is of interest to the field of insect neural development and neuroscience.

    2. Reviewer #1 (Public review):

      [Editors' note: The reviewing editor has assessed the revisions. The authors have addressed the previous minor concerns of the reviewers, added more details on the generation of the brainbow constructs and have made the image stacks available on public repositories.]

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

    3. Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a non-standard laboratory organism.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

      Weaknesses:

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects.

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single-cell labelling”, which admittedly limits precision.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function.

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an effect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately.

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question but a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative.

      Indeed, we do not reach single-cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content.

      We also think that combining our transgenic line with dopamine expression was sufficient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism.

      Thank you for this encouraging comment.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      Comments:

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect:

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/ or motor learning in motor neurons https://pubmed.ncbi.nlm.nih.gov/38779314/ or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest.

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal thresholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal-directed navigation, respectively."

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors claim that several dop-positive cell types resemble cell types described in Drosophila - to be able to better compare the two, it would be great to have the supplement and main figure pictures in one figure.

      We have added Fig. 11 from the main text to suppl. Fig. 3 for comparison

      (2) In Figure 11 and the corresponding supplement, it would be great to have some landmarks to better understand the expression patterns.

      In the legend, we have now referred to Figs. 3, 4 and 5 for depictions of these neurons within the neuropil reconstructions

      (3) For a better understanding of neurotransmitter expression, is it possible to figure out if dopamine is rather coexpressed with Glut- or ChaT-positive neurons?

      Very interesting idea. Unfortunately, the first author of the study has graduated and left the lab, such that we are unable to add this piece of information.

    1. eLife Assessment

      This valuable study uses technically challenging long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The study provides solid evidence that ORNs have a circadian pattern of spontaneous firing that is Orco-dependent although Orco transcript abundance does not itself show circadian rhythmicity, and that cAMP can modulate Orco-dependent activity. Together with the computational work, the study proposes the provocative hypothesis that an Orco-centered post-translational feedback loop generates the circadian rhythm of spontaneous firing. Nevertheless, this mechanistic interpretation will require direct testing in future work.

    2. Joint Public Review:

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

      We thank the referees and the editors for their efforts and their positive feedback. As requested, we clarified throughout our manuscript that our proposal of an Orco-centered post-transcriptionally controlled feedback loop in the plasma membrane, which we term PTFL clock, is a novel hypothesis that needs to be tested in future experiments.

    1. eLife Assessment

      This important study combines a five-year field experiment with a meta-analysis to quantify effects of grazing on ecosystem CO₂ fluxes as modified by wetness. The evidence supporting the conclusions is convincing, but several aspects of the methods and discussion would benefit from clarification to improve reproducibility and interpretability.

    2. Reviewer #3 (Public review):

      Combining a five-year field experiment with a global meta-analysis, Wu et al. investigate how grazing intensity influences ecosystem carbon dioxide (CO₂) fluxes in grasslands and how these effects are regulated by environmental conditions such as grazing duration, wetness index, and soil temperature and moisture responses.

      The authors show that the response of net ecosystem productivity (NEP) to light grazing shifts from negative to positive along a wetness gradient, whereas heavy grazing consistently suppresses NEP across wetness conditions. Importantly, this pattern is supported by both the field experiment and the meta-analysis, suggesting that may help maintain moderate levels of grazing can potentially enhance both plant productivity and carbon sequestration under favorable moisture conditions.

      The integration of experimental data with a global synthesis is a particular strength of the study, allowing the authors to evaluate grazing impacts across both temporal variability (precipitation fluctuations in the field experiment) and spatial variability (wetness gradients across global grasslands). Overall, the conclusions are well supported by the data.

      Overall, the principal conclusions are generally supported by the reported results, and the study provides useful evidence that the effects of grazing on grassland carbon cycling depend on both grazing intensity and environmental context. The comparison between field and synthesis results is potentially valuable for understanding why grazing effects vary among grassland systems. However, some aspects of the meta-analysis remain insufficiently documented. In particular, the study-selection numbers presented in the new PRISMA diagram require clarification, and the manuscript does not clearly explain how individual response ratios were weighted when estimating the pooled effect sizes. Resolving these reporting and methodological issues would improve the reproducibility and interpretation of the synthesis.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study combines a five-year field experiment with a meta-analysis to quantify the effects of grazing on ecosystem CO<sub>2</sub> fluxes as modified by wetness. The results of this study are potentially valuable, but the Methods description is incomplete and compromises reproducibility and interpretability. A major caveat of this study is that the assessment of CO<sub>2</sub> fluxes in time and space is not complete.

      We appreciate the comments. The incomplete assessment of CO<sub>2</sub> fluxes in time and space has been added in the Limitations and implications for future study Section in the revised manuscript. Method description has been revised as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      Lines 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      Lines 468-470, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that was not monitored.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study integrates long-term (7th-11th year) ecosystem CO<sub>2</sub> flux measurements from a continuously grazed typical steppe in Inner Mongolia with a global meta-analysis of grazing experiments (585 observations) to systematically investigate how the wetness index modulates the effects of grazing intensity on net ecosystem productivity (NEP) in grasslands.

      Key findings include:

      (1) Heavy grazing significantly reduced gross primary productivity (GPP) and ecosystem respiration (ER) in the typical steppe but did not significantly affect NEP.

      (2) Globally, grazing significantly reduced GPP, ER, and NEP, and a higher wetness index and aboveground biomass (AGB) enhance the positive response of NEP to grazing.

      (3) Under light and moderate grazing, the response of NEP to grazing was positively correlated with the wetness index, a relationship potentially mediated by a higher plant relative growth rate (RGR) in wetter years.

      This study holds considerable practical significance. In the context of global change, comprehending the regulatory function of water is of great importance for the adaptive management of grasslands.

      Strengths:

      (1) The study cleverly combines long-term in situ observations (revealing temporal dynamics and potential mechanisms) with a global meta-analysis (testing the generality of patterns), forming a complete evidence chain from "process understanding" to "pattern verification".

      (2) Focusing on the 7th to 11th years of grazing treatments avoids the common "initial disturbance effects" observed in short-term grazing experiments and truly captures the steady-state response of the ecosystem after it has reached a new equilibrium.

      (3) By analyzing plant relative growth rate (RGR) and aboveground biomass (AGB), this study provides mechanistic clues regarding how the wetness index modulates grazing effects. The finding that plant compensatory growth is enhanced in wetter years is a crucial pathway explaining the variation in the NEP response.

      (4) The meta-analysis systematically integrates published literature with a substantial sample size (585 observations) and broad geographical coverage (Figure 1B).

      (5) The finding that light grazing promotes carbon sinks in wet years, while heavy grazing reduces productivity even in wet years, has direct implications for adaptive grassland management under climate change.

      Thank you for the comments. We would like to express our gratitude for your positive and insightful evaluation of this manuscript.

      Weaknesses:

      Although the paper does have strengths in principle, there are weaknesses and areas for improvement in the paper. In particular:

      (1) Incomplete mechanistic chain: While RGR and AGB data offer valuable clues for mechanistic interpretation, the causal chain from "increased wetness → higher RGR → maintained NEP" remains incomplete. Key intermediate processes, such as soil moisture dynamics, nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, are absent, making the mechanistic explanation somewhat speculative.

      Thank you for the comments. The data of the soil moisture dynamics, composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain. However, the structural equation modeling could not be set up if C<sub>3</sub> and C<sub>4</sub> plant communities were included (Fig. S8). Leaf photosynthetic physiological parameters were not collected in this study. We have made following changes in the revised manuscript:

      Lines 339-341, page 11: “In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Lines 472-474, page 15: “Furthermore, key intermediate processes, such as soil nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, should also be investigated in the future.”

      (2) Inadequate exploration of heterogeneity: The global meta-analysis reveals significant effect heterogeneity. Currently, only a few moderators, such as wetness index, precipitation, and temperature, are analyzed. Potential moderators, including grazing history, grassland type (typical steppe/alpine meadow/savanna), livestock type (cattle/sheep/mixed), and soil type, are not adequately explored.

      Thank you for the comments. Grazing duration was presented in Fig. S13. The heterogeneity analysis of grassland types and livestock types was supplemented as Fig. S11 and Fig. S12 in the Supplementary documents. We have also made following changes in the revised manuscript:

      Lines 367-370, page 12: “Grazing decreased GPP, ER and NEP in desert and temperate grassland, but had no significant effect on GPP, ER and NEP in alpine grassland and savanna (Fig. S11). In addition, cattle and sheep grazing decreased GPP and NEP, but livestock mixed grazing did not show significant effect on GPP, ER and NEP (Fig. S12).”

      (3) Integration of long-term experiment and meta-analysis could be tighter.

      Thank you for the comments. Integration of long-term experiment and meta-analysis has been revised in the revised manuscript as follows:

      Lines 449-463, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term grazing field experiments revealed the mechanisms underlying the effects of grazing on ecosystem CO<sub>2</sub> fluxes. Global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analyses consistently revealed that wetness was an important factor in regulating the ecosystem CO<sub>2</sub> fluxes in response to grazing. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP (Morgan et al., 2016; Owensby et al., 2006). Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness modulated the effects of grazing on NEP in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      Reviewer #2 (Public review):

      Overgrazing in grasslands is a widespread, enormous problem, and its consequences need more research attention.

      However, I have some concerns about the methods in this manuscript.

      (1) The plots were grazed from May to September. For the CO<sub>2</sub> measurements, it states growing season, and that data was collected on sunny days from 9 to 11 am. However, annual values are reported, and no details on how these were calculated. With no measurements outside of 9 to 11 am, and only on non-sunny days, and only during the growing season, I question how accurate the annual numbers are.

      Thank you for the comments. The ecosystem CO<sub>2</sub> fluxes were measured during growing season; the description of annual CO<sub>2</sub> have been revised to “growing-season” CO<sub>2</sub>. The ecosystem CO<sub>2</sub> fluxes were measured on sunny days during the growing season (May-September) between 9:00 and 11:00 a.m. It can represent typical daytime carbon flux dynamics during the growing season (Fan et al., 2011; Li et al., 2010). In addition, previous studies showed that the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. represented the average ecosystem CO<sub>2</sub> fluxes of the day (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), so we use the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. to calculate the total ecosystem CO<sub>2</sub> fluxes of the day. We have also revised it in the limitation for future study section. We have made following changes in the revised manuscript:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      Lines 468-472, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that were not monitored. Multiple time-point sampling should be adopted; the sampling frequency and points should also be increased in the future field study.”

      (2) This study uses one chamber, 0.5 by 0.5 m, for a total area of 0.25 m<sup>2</sup>. This is pretty small, and with no replication, I would like to see more detail on how these are placed, especially when one of the dominant species is a bunchgrass. Sampling on or off a bunchgrass will give very different readings, and neither will be representative of the entire plot. The same for the soil temp and water measurements, there is no detail on how many replicates are in each plot, and working randomly does not work when there are spatial patterns caused by the bunchgrasses in the plots.

      Thank you for the comments. Two replicates of the chambers in each plot were set, and their specific sites were labeled in Fig. S2. The study area was located in a typical steppe dominated by Stipa grandis and Leymus chinensis. Therefore, we selected quadrats mainly containing these two species. We have made following changes in the revised manuscript:

      Line 186-187, page 6: “In each plot, the chambers of two replicates were measured.”

      (3) MAP and MAT are important in this study, and no mention is given if this data is from this site or a neighboring site, and if so, what the distance to this location is.

      Thank you for the comments. The data of MAP and MAT were provided in the revised manuscript as follows:

      Lines 193-195, page 7: “The precipitation and air temperature were collected from the Xilinhot Meteorological Observation Station, which was near the experimental site in this field experiment.”

      (4) Vegetation was sampled in five locations in each plot. I note no detail on the belowground biomass, how many reps, what area, or which depth? This essential data is missing. The same goes for the RGR, and here the authors talk about grazing events. Whereas earlier in section 2.2, it implies continuous grazing every day during the growing season. RGR needs more details, such as how many times, its replication, etc.

      Thank you for the comments. The detailed method for how to collect belowground biomass and how to calculate the RGR has been supplemented in the revised manuscript as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      (5) Hypothesis 2 is vague: "act as key factors". A more precise hypothesis would be better, otherwise it just leads to p-hacking.

      Thank you for the comments. The misleading description has been deleted.

      (6) More generally, neither hypothesis is really based on the introduction, as the authors report mixed results in the literature. Hence, the more specific hypothesis 1 reads like it is HARKED, based on the results that the authors found. This is not a good research practice. See the following papers on this topic:

      Murphy & Aquinis. 2019. HARKing: How bad can cherry-picking and question trolling produce bias in published results? Journal of Business and Psychology 34:1-17

      Bishop. 2019 Rein in the four horses of irreproducibility. Nature 568: 435

      Fraser et al. 2018. Questionable research practices in ecology and evolution. Plos One 13:7

      Parker et al. 2016. Transparency in ecology and evolution: real problems, real solutions. Trends in ecology and evolution 31:9

      Thank you for the suggestion. The hypotheses have been deleted in the revised manuscript to avoid HARKED mistakes.

      (7) Figure 2 mentioned n=3, which is correct. However, the variance around the means in the figures is tiny, which raises questions about the replication that is actually used here. I would like to see the entire statistics tables, including DF and sample sizes, included as an appendix, so that the reader can evaluate this much more.

      Thank you for the comments. In this study, the data presented in Fig.2 were expressed as mean± standard error (SE):

      In addition, we have provided the complete statistical table in Supplementary Table S2 according to the suggestion of the reviewer.

      (8) Figure 3, now only low, medium, and high grazing are reported as changes from the control. This can be misleading as the reader can't see how the control varies along the various gradients. I suggest including a figure with the raw data for each as an appendix. In addition, the sample size in 3d, e, f is much higher, and it looks to me like the authors used both the reps and years together. This is not good practice. In addition, full statistics tables should be included in the appendix.

      Thank you for the comments. The relationships between ecosystem CO<sub>2</sub> fluxes and wetness index have been supplemented in Fig. S6. We have also provided the complete statistical table in Supplementary Table S3.

      (9) In Table S1, since there are already a number of recent meta-analyses on this topic, I think there needs to be a stronger justification for this one.

      Thank you for the comments. The results of our five-year field monitoring experiments showed that carbon fluxes and their components were significantly influenced by the wetness index. Therefore, we explored whether such effects also occur in grazing experiments at the global scale. The results demonstrated that similar patterns indeed exist worldwide, indicating that our meta-analysis is meaningful and necessary. In addition, previous studies have addressed related topics, net ecosystem productivity has mostly been treated as an auxiliary variable rather than the primary focus of investigation (Jiang et al., 2020; Shi et al., 2022; Zhang et al., 2022; Zhou et al., 2019). In previous studies, meta-analysis literatures on grazing and NEP were limited and not adequate. Our meta-analysis is more comprehensive and reflects the reliability of the results. We have made following changes in the revised manuscript:

      Lines 491-492, page 16: “In previous studies, meta-analysis literatures of effects of grazing intensities on NEP were limited and not adequate.”

      Lines 512-514, page 16: “In summary, the meta-analysis of this study presents the first comprehensive assessment of how annual wetness index affects the response of ecosystem CO<sub>2</sub> fluxes to grazing across global grasslands.”

      (10) Regarding the grazing-induced CO<sub>2</sub> fluxes, these are only based on the plants and do not incorporate the animal CO<sub>2</sub> flux, nor the animal litter CO<sub>2</sub> flux, as they were penned at night outside the plots. Thus, the grazing impact is inflated in the data reported here and does not really represent GPP, NET, or RE. This needs to be written about in the discussion section.

      Thank you for the comments. The objective of this study was to evaluate the carbon exchange processes from vegetation and soil under grazing disturbance. Thus, the grazing effects reported in this study were the vegetation-soil ecosystem scale rather than the net carbon balance of the entire grazing system. Consequently, our conclusions regarding the effects of grazing on grassland ecosystem carbon exchange remain reliable and ecologically meaningful.

      We have made following changes in the manuscript:

      Lines 483-490, page 16: “There are also limitations in evaluating the effects of grazing on grassland-livestock ecosystem CO<sub>2</sub> fluxes in this study. Specifically, carbon emissions derived from animal respiration were not included in this study. Therefore, the results of the field experiment and meta-analysis in this study should not be interpreted as a complete carbon budget assessment of the grassland-livestock ecosystem. Our primary objective was to evaluate the carbon exchange processes between vegetation and soil under grazing disturbance. Thus, the grazing effects reported in the field experiment and meta-analysis of this study were responses at the vegetation- soil ecosystem scale rather than the net carbon balance of the entire grazing system.”

      Reviewer #3 (Public review):

      Combining a five-year field experiment with a global meta-analysis, Wu et al. investigate how grazing intensity influences ecosystem carbon dioxide (CO<sub>2</sub>) fluxes in grasslands and how these effects are regulated by environmental conditions such as grazing duration, wetness index, and soil temperature and moisture responses.

      The authors show that the response of net ecosystem productivity (NEP) to light grazing shifts from negative to positive along a wetness gradient, whereas heavy grazing consistently suppresses NEP across wetness conditions. Importantly, this pattern is supported by both the field experiment and the meta-analysis, suggesting that moderate levels of grazing can potentially enhance both plant productivity and carbon sequestration under favorable moisture conditions.

      The integration of experimental data with a global synthesis is a particular strength of the study, allowing the authors to evaluate grazing impacts across both temporal variability (precipitation fluctuations in the field experiment) and spatial variability (wetness gradients across global grasslands). Overall, the conclusions are well supported by the data.

      However, several aspects of data acquisition, analysis, and presentation could be clarified to further strengthen the reproducibility and interpretation of the results:

      (1) A PRISMA-style flow diagram would be helpful for the meta-analysis to clearly illustrate the study selection process and facilitate interpretation of the dataset.

      Thank you for the comments. We have added the PRISMA-style flow diagram in the revised supplementary materials. See Fig. S9.

      (2) Grazing intensity requires a clearer definition and, where possible, standardization between the field experiment and the studies included in the meta-analysis. Key parameters such as the number of stock per area, days per rotation or per year, and total years of grazing should be clearly defined. In addition, the criteria used to classify grazing intensity into LG, MG, and HG in the meta-analysis should be explicitly described.

      Thank you for the comments. In our meta-analysis, the classifications of grazing intensity into light, moderate, and heavy grazing were provided in previous studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the USDA criteria (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf).

      The criteria was that: light grazing: approximately equal to a maximum of 40% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; moderate grazing: approximately equal to a maximum of 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; Heavy grazing: greater than 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season (November 15).

      We have made following changes in the manuscript:

      Lines 271-277, page 9 “(e) The classifications of grazing intensity (light, moderate, and heavy grazing) were primarily based on the definitions provided in the original studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the modified Grazing Intensity Classes proposed by the USDA (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf) (Yin et al., 2023).”

      (3) In several sections of the manuscript, it is difficult to distinguish whether phrases such as "this study" or "our study" refer specifically to the field experiment or to the overall study, including both the experiment and the meta-analysis. Clearer wording distinguishing these components would improve readability.

      Thanks for the suggestion. The specific distinctions have been revised to clarify whether they refer to the field experiment, meta-analysis or the comprehensive conclusions of the results from both field experiment and meta-analysis.

      Lines 50-53, page 2: “Overall, the meta-analysis and field experiment jointly provide global perspectives on the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity and improve our knowledge of the factors influencing the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity.”

      Lines 191-193, page 6: “The annual precipitation and mean annual air temperature (MAT) for the field study area from 2019 to 2023 were obtained from the China Meteorological Data Service Centre (http://data.cma.cn/).”

      (4) The discussion attributes the non-significant annual NEP response to intra-annual precipitation variability, with grazing enhancing NEP under wet conditions but suppressing it under dry conditions. Another potential explanation may be that grazing affects GPP and ecosystem respiration (ER) at similar magnitudes (i.e., RR(GPP) ≈ RR(ER)), resulting in limited net changes in NEP.

      Thank you for the comments. Another potential explanation has been revised in the manuscript as follows:

      Lines 385-389, page 13: “Grazing decreased the responses of ecosystem CO<sub>2</sub> fluxes and plant biomass in global grasslands, but only NEP was not significantly affected by grazing in our field experiment (Fig. 4). The lack of a significant response in NEP may be because the site in this field study was managed for year-round continuous low-intensity grazing (Liang et al., 2021). In addition, grazing affected GPP and ER at similar magnitudes, resulting in limited net changes in NEP.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see from the above reviews and the specific recommendations below, the first issue you should resolve is a full description of methodological detail such that readers can, in principle, repeat your study. They have to know how you placed the chamber to have a representative measure of vegetation (averaging between high and low biomass patches). They also must know how to calculate grazing intensity and assign the values to the three classes. They want to see your raw data and statistical analyses (including formulae, replication, and degrees of freedom). Consider non-linear relationships of grazing and wetness. Try to better integrate the results from the experiment and the meta-analysis. Finally, make sure that you develop hypotheses from the prior knowledge presented in the introduction. For instance, instead of stating that NEP responses would shift from negative to positive with increasing wetness index, simply hypothesize that wetness mitigates the negative effects of grazing on CO<sub>2</sub> fluxes, even if you find that this is not true under heavy grazing.

      Thank you for the comments. The description of methodological details has been supplemented in the revised manuscript. The results from the experiment and the meta-analysis have been integrated.

      Reviewer #1 (Recommendations for the authors):

      General suggestions:

      (1) The Introduction section and the assumptions should be rewritten and improved. The logicality of the introduction should be revised to better prioritize and contextualize the research problem. The reader is lost since the links between assumptions and previous knowledge are not clear.

      Thanks for the suggestion. The Introduction section and the assumptions have been rewritten as follows:

      Lines 135-147, page 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable? In this study, we investigated the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes by combining a long-term (7-11 years) field experiment conducted in a typical steppe and a meta-analysis of global grasslands. The objectives of this study were to: (i) investigate the effects of grazing intensity with annual wetness fluctuations on ecosystem CO<sub>2</sub> fluxes (GPP, ER and NEP) covering the 7th to 11th years of a continuous grazing experiment in a typical steppe as well as the meta-analysis in global grasslands; (ii) explore how environmental factors (particularly wetness index, soil moisture and temperature, and grazing intensity) regulate the effects of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (2) In the "Materials and methods" section, you need to provide the reason why you chose the "wetness index" in this study, rather than other drought indices (such as Standardized Precipitation Evapotranspiration Index (SPEI), Aridity Index (AI)).

      Thank you for the comments. Although the standardized precipitation evapotranspiration index (SPEI) would be a more appropriate indicator for this study, its calculation requires relatively long and continuous climate data series, which were difficult to obtain in our global meta-analysis. This limitation was particularly important because our study also included a meta-analysis, for which complete climatic datasets were often unavailable from the collected literature. In contrast, the Aridity Index (AI) cannot adequately reflect interannual variability. We have supplemented this limitation in the revised as follows:

      Lines 482-483, page 15: “Furthermore, more drought or wetness indices should be investigated in future studies of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (3) For the results and discussions, the results are currently presented in parallel (long-term experiment first, then meta-analysis). It is recommended to add a dedicated integration paragraph in the discussion.

      Thank you for the comments. The dedicated integration paragraph has been supplemented in the revised manuscript as follows:

      Lines 450-464, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term experiments revealed the mechanisms underlying the effects of grazing on plant characteristics and ecosystem CO<sub>2</sub> fluxes under control conditions. And global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analysis consistently revealed that wetness was an important factor in regulating the effect of ecosystem CO<sub>2</sub> fluxes. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP. Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness determined whether grazing promoted or inhibited the carbon sink function in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      (4) Make sure that the whole manuscript has undergone professional proofreading, and check it carefully to avoid language mistakes.

      Thank you for the comments. The manuscript has been revised by professional proofreading to avoid language mistakes.

      Specific suggestions:

      (1) Lines 126-131: It is recommended to explicitly state three levels of research questions: (i) How does grazing intensity affect ecosystem CO<sub>2</sub> fluxes? (ii) Does the wetness index modulate this effect? (iii) Are these relationships globally generalizable? This will provide a clear logical thread for the paper.

      Thanks for the suggestion. We have made following changes in the revised manuscript:

      Lines 135-139, pages 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable?”

      (2) Lines 181: Briefly justify the choice of the De Martonne wetness index (Equation 1) in the introduction (why this index over other aridity indices).

      Thank you for the comments. The De Martonne wetness index is easier to obtain compared to other drought index, as it only requires annual average temperature and precipitation data for calculation, facilitating the statistics of global meta-analysis. The justification of the De Martonne wetness index has been revised in the manuscript:

      Lines 120-125, pages 4-5: “Annual precipitation is one of the climatic parameters, while the wetness index (WI) serves as a more integrative climatic indicator that incorporates both precipitation and temperature, thereby reflecting the overall water surplus or deficit (Song et al., 2019). A higher wetness index (WI > 30) indicates sufficient water availability for plant growth, whereas a lower wetness index (WI ≤ 30) suggests the water availability may be limited (De Martonne, 1926).”

      (3) Lines 173-174: Please specify the exact timing of flux measurements (e.g., "measured three times per month between 9:00 and 11:00 AM" is already stated, but add "on sunny and calm days" to ensure consistent conditions).

      Thanks for the suggestion. We have supplemented the exact timing of flux measurements in the revised manuscript as follows:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> fluxes represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      (4) The RGR and AGB data provide important clues for explaining the moderating role of the wetness index, but the mechanistic chain can be further refined. It is recommended to use the structural equation model.

      Thanks for the comments. The structural equation model has been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Suggested addition: If data on soil moisture, soil nutrients (e.g., ammonium, nitrate), leaf photosynthetic parameters (e.g., maximum photosynthetic rate, stomatal conductance), or community composition are available, please incorporate them into the analysis to test a more complete mechanistic pathway.

      Thanks for the comments. The data of the composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      (5) The current meta - analysis has established the moderating role of the wetness index, but there is room for further exploration: Test for non-linearity. Could the relationship between the wetness index and the grazing effect size (lnRR of NEP) be non-linear in the global data? Consider fitting models that include a quadratic term for the wetness index or using generalized additive models (GAMs). If a threshold is identified, report the threshold estimate and its confidence interval and discuss its management implications.

      Thank you for the comments. Following the reviewer's suggestion, we further examined whether there is a nonlinear relationship between the wetness index (WI) and the magnitude of grazing effect (NEP_RR). We compared a linear mixed-effects model including a first-order term for WI with a quadratic mixed-effects model, and set Study ID as a random effect in both models. The results indicated that the AIC value of the quadratic model (256.56) was higher than that of the linear model (241.48), suggesting that adding the second-order term did not improve the model fitting precision. Based on this, we retained the simpler linear model in the revised manuscript.

      Author response table 1.

      The comparison of linear mixed-effects model and quadratic mixed-effects model

      Note: if the AIC value was lower, the model fitting accuracy was higher.

      (6) Section 4.1: This section is quite long. Consider splitting it into 2-3 paragraphs, discussing: (i) overall grazing effects on CO2 fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) differential effects of grazing intensities.

      Thanks for the suggestion. Section 4.1 has been split into four paragraphs according to the suggestion of the reviewer. (i) overall grazing effects on ecosystem CO <sub>2</sub> fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) integrated discussion of field study and meta-analysis.

      (7) Integrating the long-term experiment and meta-analysis. Does the effect size observed in the long - term experiment (e.g., an 85.83% increase in NEP under LG in the wettest year) align with the average effect size from the global meta - analysis under similar conditions? If not, what are the potential reasons? (e.g., specificity of the typical steppe, methodological differences between chamber and eddy covariance measurements) Is the mechanism identified in the long - term experiment (e.g., increased RGR) likely to be common globally? What are the joint management implications from both parts of the study? Are there contexts where caution is needed in extrapolating the findings? (e.g., alpine meadows might be more sensitive to grazing).

      Thank you for the comments. The long-term experiment and meta-analysis have been integrated in the revised manuscript. The results of field experiment showed that light grazing increased NEP by 85.83%. However, the response of light grazing was not significant in meta-analysis. Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions. The different results in effect size and statistical significance were mainly because the long-term experiment was conducted in a typical grassland ecosystem, which may possess a strong compensatory growth capacity under moderate water conditions. Therefore, light grazing can enhance NEP by increasing plant photosynthetic rate, promoting new leaf growth, and improving community resource utilization efficiency. However, the global meta-analysis integrated different grassland types, climatic conditions and grazing durations. Consequently, the average effect of global meta-analysis may be diluted by the high heterogeneity among ecosystems. The data to calculate RGR were not available in the original studies of the meta-analysis, so we could not include RGR in the meta-analysis. Therefore, it is still unclear if the mechanism of increased RGR could be common globally. The joint management implications have been revised as follows:

      Lines 393-400, page 13: “Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions (Fig. 3B, 6B and 7H). Light grazing usually stimulates leaf regrowth following defoliation, and these new leaves often are more physiologically active than the older leaves that contribute much of leaf area in ungrazed treatment (Polley et al., 2008), which likely imply a stronger leaf photosynthesis and C sink (Reich et al., 2007). Considering factors such as different grassland types (Fig. S11), livestock grazing modes (Fig. S12), climatic conditions, and grazing durations, the future grazing studies on NEP should focus more on ANPP and RGR, and extrapolate the results cautiously.”

      (8) Figure 8: The conceptual diagram is clear and effectively summarizes the main findings. Briefly explain the meaning of the arrows in the caption.

      Thank you for the suggestion. We have revised Fig. 8 accordingly:

      “Schematic summary of the effects of grazing intensity on ecosystem carbon dioxide (CO<sub>2</sub>) fluxes and biomass in global grasslands in the meta-analysis. The blue and black arrows indicate negative effects on grazing and grazing intensities, respectively. Asterisks (<sup>*</sup>) indicate significant effects on variables at P < 0.05. GPP, gross primary productivity; ER, ecosystem respiration; NEP, net ecosystem productivity; LG, light grazing; MG, moderate grazing; HG, heavy grazing.

      (9) Please check all references for consistency with eLife style. Some entries currently have inconsistent formatting (e.g., some include issue numbers, while others do not; page number formatting varies).

      Thank you for the comments. We have checked the references one by one to revise them consistent with eLife style.

      Reviewer #3 (Recommendations for the authors):

      (1) Method citation:

      (a) Please cite the original method references rather than studies that applied the methods. For example, the original publication introducing the wetness index (WI) is: De Martonne, E. Une nouvelle fonction climatologique: l'indice d'aridité. La Météorologie 2, 449-458 (1926).

      (b) Please also cite the R packages used in the analysis. One straightforward approach is the function citation() in R.

      Thank you for the comments. We have cited the references related to the original methods and cited the R packages used in the analysis.

      (2) Several expressions would benefit from clarification:

      (a) Line 182: What has been "referred to as soil moisture"?

      Thank you for the comments. Soil moisture refers to the volumetric water content of the soil, which has been clarified in the revised manuscript as follows:

      Lines 199-200, pages 7: “Soil moisture was the soil volumetric water content.”

      (b) Line 184-185: Do you mean "soil temperature and soil moisture were measured simultaneously with ecosystem CO<sub>2</sub> flux measurements."?

      Thank you for the comments. Yes, we simultaneously measured soil temperature and soil moisture using temperature and moisture probes while measuring ecosystem CO <sub>2</sub> flux measurements. We have made following changes in the revised manuscript as follows:

      Lines 198-199, page 7: “Soil temperature and moisture at a depth of 0-10 cm were measured simultaneously with ecosystem CO <sub>2</sub> flux measurements, using the probes of the LI-8100 system.”

      (c) Line 200: Please clarify what is meant by "the caged plots"?

      Thank you for the comments. The misleading words have been deleted in the revised manuscript.

      (d) Line 242-244: Do you mean "data from non-grazed treatments were excluded when additional treatments were present"?

      Thank you for the comments. Yes, data from non-grazed treatments were excluded when additional treatments were present. We have made following changes in the revised manuscript:

      Lines 268-270, page 9: “(c) Data from non-grazed treatments were excluded when additional treatments (e.g., fertilization, experimental warming, or precipitation manipulation) were present.”

      (e) Line 263: Should this refer to RR<sub>++</sub> instead of RR? Please clarify how RR<sub>++</sub> (or lnRR++) was calculated from the reported response ratios and study weights.

      Thank you for the comments. RR is the dependent variable of the model, representing the response ratio for each observation. RR<sub>++</sub> typically refers to the pooled effect size obtained after all RRs are weighted and subjected to a mixed-effects model, which primarily corresponds to the estimated value of the model intercept β<sub>0</sub>. We have made following changes in the revised manuscript:

      Lines 293-303, page 10: “A linear mixed-effects model, with ‘study’ included as a random factor, was employed to estimate the weighted response ratio (RR<sub>++</sub>) across studies or within a specific group, fitting with restricted maximum likelihood using the ‘lmer’ function in the ‘lme4’ package (Feng et al., 2023).

      where β<sub>0</sub> is the coefficient, π<sub>study</sub> denotes the random effect associated with ‘study’ (accounting for autocorrelation among observations from the same study), and ɛ corresponds to the residual sampling error. We checked the normality of the model residuals using the ‘check_normality’ function in the ‘performance’ package. When the assumption of normality was violated, bootstrapping with 999 iterations was performed using the ‘boot’ package to derive the 95% confidence interval (CI) for each RR<sub>++</sub> (Chen et al., 2021).”

      (3) Line 222: Please list all environmental predictors considered in the analysis.

      Thank you for the comments. The environmental predictors considered in the analysis have been listed in the revised manuscript:

      Lines: 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      (4) Line 343: The non-significant response of NEP appears only under MG, while both LG and HG significantly decrease NEP (Fig. 5). It may be helpful to discuss this grazing-intensity-dependent response more explicitly.

      Thank you for the comments. The reason why the non-significant response of NEP appeared only under MG was that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP. This was consistent with the hypothesis of moderate disturbance, which suggested that ecosystem functions remain stable under moderate levels of disturbance. We have made following changes in the revised manuscript:

      Lines 390-393, page 13: “Similarly, the reason why the non-significant response of NEP appeared only under MG may be that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP.”

      (5) Line 397: The phrase "global-scale study" typically refers to experiments conducted worldwide. "Studies across global grasslands" may be more precise here.

      Thank you for the comments. We have made following changes in the revised manuscript:

      Lines: 475-477, page 15: “Our meta-analysis has limitations due to the relatively small sample size, which stems from the scarcity of studies across global grasslands exploring the effects of grazing intensity on ecosystem CO<sub>2</sub> fluxes.”

      (6) Line 403: ...have large amounts of grassland for grazing, "but rarely investigated".

      Thank you for the comments. We have made following changes according to the suggestion of the reviewer in the revised manuscript as follows:

      Lines 480-482, pages 15-16: “In addition, future studies could conduct more experiments of ecosystem CO<sub>2</sub> fluxes in South America, Africa and Oceania, which have large amounts of grasslands for grazing, but were rarely investigated.”

      (7) Please remember to cite Figure 8 in the text.

      Thank you for the comments. Figure 8 has been cited in the manuscript as follows:

      Lines 411-414, page 13: “Moderate grazing decreased GPP and NEP under lower WI, but the responses of GPP and NEP to moderate grazing were similar under higher WI in global grasslands (Figs. 6 and 8), indicating that higher wetness offset the response of GPP and NEP to moderate grazing.”

      (8) Figure S1:

      (a) The orange line in panel A appears to represent monthly temperature rather than mean annual temperature.

      (b) Same for the precipitation, the data shown here should be monthly values rather than annual means.

      (c) Consider using a color different from orange for the wetness index in panel C, unless the variable shown is temperature instead.

      Thank you for the comments. We have revised Fig. S1 according to the suggestion of the reviewer as follows:

      Supplementary page 5: “Fig. S1 The monthly total precipitation and average air temperature (A), annual precipitation (B) and wetness index (C) in the study area of the grazing intensity experiment in the typical steppe from 2019 to 2023.”

      (9) Figure S2B: Grazing intensity labels appear in Chinese in the figure. These should be translated into English, and the information on grazing intensity should be provided in the legend or the figure here, as well as in the Method section.

      Thank you for the comments. We have replaced the figures included the Chinese labels. See Fig. S2.

      (10) Figure S6B: LG significantly affects the relationships between wetness index and ER (P<0.05), but a regression line is missing from the panel.

      Thank you for the comments. We have added the regression line between wetness index and ER in Supplementary Fig. S6.

      (11) Figure S7A: "Mean" annual precipitation

      Thank you for the comments. The “Mean” has been revised in Figure S10A.

      References

      Fan, Y., Zhang, X., Wang, J., & Shi, P. (2011). Effect of solar radiation on net ecosystem CO2 exchange of alpine meadow on the Tibetan Plateau.Journal of Geographical Sciences, 21(4), 666-676.

      Jiang, Z., Hu, Z., Lai, D., Han, D., Wang, M., Liu, M., Zhang, M., & Guo, M. (2020). Light grazing facilitates carbon accumulation in subsoil in Chinese grasslands: A meta-analysis. Global Change Biology, 26(12), 7186–7197.

      Li, X., Fu, H., Guo, D., Li, X., & Wan, C. (2010). Partitioning soil respiration and assessing the carbon balance in a Setaria italica (L.) Beauv. Cropland on the Loess Plateau, Northern China. Soil Biology and Biochemistry, 42(2), 337-346.

      Niu, S., Wu, M., Han, Y., Xia, J., Li, L., & Wan, S. (2008). Water‐mediated responses of ecosystem carbon fluxes to climatic change in a temperate steppe. New Phytologist, 177(1), 209-219.

      Rong, Y., Johnson, D. A., Wang, Z., & Zhu, L. (2017). Grazing effects on ecosystem CO2 fluxes regulated by interannual climate fluctuation in a temperate grassland steppe in northern China. Agriculture, Ecosystems & Environment, 237, 194-202.

      Shi, R., Su, P., Zhou, Z., Yang, J., & Ding, X. (2022). Comparison of eddy covariance and automatic chamber‐based methods for measuring carbon flux. Agronomy Journal, 114.

      Wan, L., Liu, G., & Su, X. (2025). Global meta-analysis reveals different grazing management strategies change greenhouse gas emissions and global warming potential in grasslands. Geography and Sustainability, 6(3), 100251.

      Yin, M., Gao, X., Kuang, W., & Tenuta, M. (2023). Soil N<sub>2</sub>O emissions and functional genes in response to grazing grassland with livestock: A meta-analysis. Geoderma, 436, 116538.

      Yu, H., Wang, X., Wu, Y., Wang, C., Yan, R., Xu, D., & Xin, X. (2025). Light grazing tends to enhance ecosystem carbon sequestration and resource use efficiency in a meadow steppe of northern China. Agricultural and Forest Meteorology, 372, 110690.

      Zhang, R., Tian, D., Chen, H. Y. H., Seabloom, E. W., Han, G., Wang, S., Yu, G., Li, Z., & Niu, S. (2022). Biodiversity alleviates the decrease of grassland multifunctionality under grazing disturbance: A global meta-analysis. Global Ecology and Biogeography, 31(1), 155–167.

      Zhou, G., Luo, Q., Chen, Y., Hu, J., He, M., Gao, J., Zhou, L., Liu, H., & Zhou, X. (2019). Interactive effects of grazing and global change factors on soil and ecosystem respiration in grassland ecosystems: A global synthesis. Journal of Applied Ecology, 56(8), 2007–2019.

    1. eLife Assessment

      This important study provides convincing evidence supporting the existence of transposable element (TE)-gene chimeric transcripts in the Drosophila brain and helps reconcile previously conflicting findings. The authors demonstrate that differences in computational approaches, experimental design, and Drosophila lines can substantially influence the detection of these transcripts. The work highlights the importance of standardized approaches for studying TE-derived transcripts, although their broader biological significance remains to be established.

    2. Reviewer #1 (Public review):

      Summary:

      Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to accurately detect TE insertions from RNA-seq. In this short response, Choucri and Treiber clearly show that differences in the tools used between their study and that of Azad et al. likely explain the contrasting results, along with RT-PCR failure to design primers that match the chimeric transcript and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.

      Strengths:

      The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and to compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. The methods are clear and thorough. The discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.

      The biological function, if any, of such chimeric transcripts remains to be determined by further analysis, including chimeric transcripts with low to high overall contribution to gene expression (Figure 1B).

    3. Reviewer #3 (Public review):

      Summary:

      This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.

      Strengths:

      The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.

      Weaknesses:

      The authors should mention that combining PCR-amplified cDNA generation with short-read sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct long-read ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv) . Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that, contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to detect TE insertions from RNAseq accurately. In this short response, Choucri and Treiber clearly demonstrate that differences in the tools used between their study and that of Azad et al. likely account for the contrasting results, along with RT-PCR failure in designing primers that would match the chimeric transcript, and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.

      Strengths:

      The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. Finally, the discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.

      Weaknesses:

      I think it is necessary to add more detail to this article, for instance, the differences between TEchim and Tidal could be laid out more precisely.

      We thank the reviewer for this helpful suggestion and agree that a more explicit comparison improves the clarity of the manuscript. Briefly, TIDAL and TEChim are designed to answer different questions. TIDAL is an insertion/deletion caller that was primarily built to determine whether a TE insertion exists at a genomic locus and how it is distributed across strains or populations. It was not designed to resolve splice junctions between exons and TEs. TEChim, by contrast, is purpose-built to detect breakpoint-spanning reads that span exon junctions and putative splice sites within a TE. This difference in design has several concrete consequences in the algorithms used:

      (1) TIDAL clusters reads that support a candidate breakpoint within a window of twice the sequencing read length (e.g. 300nt for 150nt reads). This works well for calling genomic insertions, but it can be too restrictive for a splice junction, which may fall at a variable position within the gene and TE. TEChim does not rely on fixed-window clustering, but instead maps all reads, groups them to individual nt positions within the genome, and then filters for events that were detected in more than one biological replicate.

      (2) TEChim reconstructs long in-silico reads from overlapping paired-end reads, using FLASH, prior to alignment. In-silico paired-end reads are analysed, and full-read merged fragments provide single-nucleotide resolution for a breakpoint. This increases accuracy and confidence in split reads.

      (3) TEChim explicitly intersects candidate breakpoints with annotated exon/intron/UTR. Structures and canonical splice donor- and acceptor sites.

      We have now added a concise summary of the underlying principles of TEChim in the Methods section, and added details to the comparison of TIDAL with TEChim in the discussion. We hope that this addition makes it clearer why the two pipelines produce different results from the same input data.

      Regarding the roo example, one of the caveats of this family, along with others, is the presence of simple repeats. It would be important to show that the simple repeats are not interfering with the read mapping.

      We thank the reviewer for raising this important point. We agree that simple repeats can complicate read mapping and that this requires careful consideration. We have now added to the results the exact locations and lengths of the three known and annotated repeat regions within roo (Domínguez, 2021) and show that the breakpoints we report map more than 4kb away from these regions. In addition, the splice junctions we identify (at positions 5190 and 5462) recur at the same position relative to the roo consensus sequence across multiple genomic insertions, and biological replicates. If these calls were artifacts of copyspecific simple repeats that interfered with read mapping, then we would expect breakpoints to vary between with each insertions local sequence, rather than converge on the same breakpoint. We therefore consider it unlikely that simple repeats account for the observed splice junctions. We have expanded our discussion to make this reasoning clearer.

      Regarding the experiments, if we are looking for a standardized protocol, then we should have a detailed material and methods section, with every experiment, replicate, and PCR temperature clearly defined.

      We thank the reviewer for this suggestion and agree that a more detailed description of the experimental procedures will improve the reproducibility of the study. We have substantially expanded the Materials and Methods section, including the number of biological replicates, primer information, PCR conditions and other methodological details relevant for reproducing the experiments. In addition, we have written a detailed manual for the updated version of TEChim that we used here.

      Finally, and in my opinion, more importantly, the use of RT negative controls on the RT PCRs, along with DNA PCRs to show insertion presence, is mandatory for testing the presence of chimeric genes. Of course, water negative PCR controls are also needed, and unfortunately, absent from Figure 3.

      We thank the reviewer for this helpful suggestion. We have repeated the RNA extraction on 3 new samples, and this time also included minus-RT aliquots for each of the three biological replicates. We have run our PCR for the chimeric transcript between Beadex and opus on all these samples, and in addition on a water control. All these new results are now shown in Figure 3, and confirm our previous conclusions.

      We now also provide results from our DNA testing for the opus insertion in Beadex. We conduct these at regular intervals in our lab to ensure the insertion remains stable in our stock, and mentioned the results in the original version of this manuscript, but we agree that it is important to also show the raw data of this in this study. We use primers at the up- and downstream end of the opus insertion. The downstream pair gave a single band at the predicted size, which we confirmed by Sanger sequencing. This data is now presented in Figure 3 – Figure supplement 1. The PCR around the upstream end resulted in the expected band of 888bp, and two additional bands. Sanger sequencing of the 888bp band produced signals for both Beadex and opus, but we did not get a reliable signal across the precise breakpoint (see Author response image 1). We think this might partly be due to a tandem repeat at the beginning of the opus LTR, which may interfere with Sanger sequencing. Taken together, our data provides strong evidence that the opus insertion is present in our flies.

      Author response image 1.

      Sanger sequencing results of the 888bp band: Two segments of the raw Sanger sequencing trace are shown; the intervening, unmapped section is omitted. The left segment (grey) shows clean, high-confidence signal matching the Bx locus. The right segment (pink) shows the signal falling to near baseline within the opus LTR, so base calls in this region are log-confidence, and the trace does not resolve the breakpoint. Numbers above the sequence indicate position within the raw sequencing read.

      Reviewer #2 (Public review):

      Summary:

      This study by Choucri and Treiber aims to directly address a recent critique regarding the role of transposable elements (TEs) in diversifying the neural transcriptome of Drosophila. The authors seek to demonstrate that TEs are not merely genomic "noise" but are frequently and reliably "exonized" into brain-specific mRNA. By introducing an upgraded computational pipeline, TEChim, and conducting precise experimental validations, the authors set out to show that TE-mediated splicing represents a genuine biological phenomenon that expands the molecular repertoire of the nervous system.

      Strengths:

      The study's primary strength lies in its rigorous technical "forensic" analysis of previous failed replication attempts. The authors convincingly demonstrate that the lack of signal in the opposing study stemmed from a fundamental methodological mismatch: the software used by the critics (TIDAL) is logically incapable of detecting splice sites located within TE sequences. Importantly, the authors complement this computational clarification with definitive experimental evidence through an effective "experimental rescue." By employing correctly designed primers and matching the genetic backgrounds of the fly strains, thereby accounting for genomic polymorphisms, they successfully validated all seven loci that were previously reported as undetectable. This dual-pronged strategy, addressing both algorithmic bias and experimental design, establishes a more robust technical benchmark for the detection and validation of TE-derived exons in neural tissues.

      Weaknesses:

      While the technical rebuttal is highly convincing, the scope of the study remains primarily defensive. As a response to a prior critique, the work focuses on establishing the existence and detectability of chimeric TE-derived transcripts rather than exploring their broader functional consequences. As a result, there is limited new insight into how these TEmodified isoforms influence neural circuit function or organismal behavior.

      We agree with the reviewer that the primary focus of this study is to establish a robust protocol for the detection and validation of TE-derived chimeric transcripts, and to resolve discrepancies raised by Azad et al. We believe this provides an important technical and conceptual framework for future studies investigating the functional impact of TE-driven genetic variation.

      In addition, the detection and validation of these events remain technically demanding, requiring deep sequencing and specialized bioinformatic expertise, which may limit broader adoption by laboratories without dedicated computational resources.

      We agree that the robust detection and validation of TE-derived chimeric transcripts is technically demanding, requiring both high-quality sequencing data and specialised computational analyses. However, we believe that these methodological challenges are justified by the biological insights that can be gained. By providing an updated computational approach together with experimental validation, we hope to facilitate further research into this phenomenon.

      Reviewer #3 (Public review):

      Summary:

      This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.

      Strengths:

      The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.

      Weaknesses:

      The authors should mention that combining PCR-amplified cDNA generation with shortread sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct longread ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv). Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.

      We thank the reviewer for this excellent suggestion and agree that long-read RNA sequencing represents a powerful approach to characterise chimeric transcripts, because it can capture full-length transcripts and reduce ambiguity of mapping short reads onto repetitive sequences. We have now expanded the Discussion to highlight the advantages of these technologies and to cite the suggested studies. At present, long-read RNA sequencing remains challenging for Drosophila brain samples, because the total amount of input RNA is limited. As a consequence, amplification is usually required, which itself can introduce artefacts, which we showed previously (Treiber and Waddell, 2017). While we agree that long-read sequencing will be an important approach for future studies, we believe that the combination of computational analysis and targeted experimental validation presented here provides robust evidence for the existence of chimeric TE-gene transcripts.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      To maximize the impact of the work, the authors should consider adding a technical guide that outlines the implementation of the updated TEChim pipeline and provides a standardized protocol for designing chimeric RT-PCR primers. This section should include a clear software workflow and a precise primer design strategy, emphasizing junction spanning probes and the necessity of genomic confirmation, to establish these methods as the definitive technical standard for studying transposable elements in the nervous system.

      We thank the reviewer for this constructive suggestion and want to be upfront about what can and cannot be standardized here. Because TE sequences are repetitive, the specific primer pair that works best for a given locus cannot necessarily be predicted in advance, and some empirical testing of candidate primer pairs is unavoidable. What we can standardise, and now describe explicitly in the Methods section, is, firstly the use of TEChim to identify candidate splicing events, and secondly, the logic behind how we choose primers. Several candidate primer pairs are tested for a given genomic locus, and all resulting bands are confirmed using Sanger sequencing. Together with an expanded TEChim manual on GitHub, we believe this study gives other research groups a clear and reproducible starting point, while remaining honest that, as with most repeat-adjacent primer design, some locus-specific optimization remains necessary.

    1. eLife Assessment

      This valuable study demonstrates that a multi-step differentiation program in bacteria combining a bistable switch with two quorum-sensing systems is capable of generating autonomous and self-organized spatial patterns. The evidence for the core engineering system using fluorescent reporters support patterning across several conditions is convincing and has significant implications on the process of cell differentiation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data and to make interesting predictions (which require future validation).<br /> The data presented are convincing and support the claims, the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FP-expressing cells in the population) when changing the initial ratio of green:blue FP-expressing cells.

      Comments on revised version:

      The authors' replies to my comments were very satisfactory. I think the paper has been strengthened.

    3. Reviewer #2 (Public review):

      The revised manuscript by Boni et al. is substantially improved, and in my view the authors have responded constructively to the concerns raised during the first review. Most importantly, they now explicitly distinguish the bistable green/blue states generated by the toggle switch from the reversible red and yellow quorum-sensing outputs and acknowledge that the latter do not constitute irreversible differentiated states. This clarification improves the conceptual accuracy of the work.

      The other revisions are also well justified. The authors clarify that Fig. 2d is derived quantitatively from flow-cytometry data and explain how the cross-sections in Fig. 2e were obtained; they introduce the nullcline interpretation of the toggle-switch landscape and distinguish stochastic single-cell fate from predictable population-level proportions. They also provide more information on pLux engineering, report the functional forms and parameters of fitted curves, quantify the different HSL induction ranges, explain the heuristic treatment of entry into stationary phase, and clarify image selection and sender-receiver distance calculations.

      The principal strength remains the systematic engineering of a sequential multi-circuit programme in a single bacterial genetic circuitry (spread over two plasmids). The work combines bistable symmetry breaking, LuxI/LuxR-mediated communication and an orthogonal CinI/CinR layer to generate spatially organised colony-level behaviours. The DBTL engineering cycle is also convincing: substantial cross-talk in the initially tested Lux/Las quorum-sensing combination was experimentally identified and addressed by switching to the more orthogonal Lux/Cin architecture, while subsequent promoter leakiness was identified and mitigated by replacing pLux with pLuxLac.

      Overall, I consider the revisions sufficient to address all the important concerns.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggle-switch-based symmetry breaking with quorum-sensing interactions to generate colony-scale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      Comments on revised version:

      The authors have satisfactorily addressed my previously raised issues.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data.

      The data presented are convincing and support the claims; the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FPexpressing cells in the population) when changing the initial ratio of green:blue FPexpressing cells.

      We thank Reviewer #1 for the Summary and for highlighting the strengths of our work.

      Weaknesses:

      (1) No spatial pattern emerges. There are some isolated colonies that turn on the downstream FPs, but I do not see a pattern, really. Nonetheless, some colonies do differentiate (i.e. they turn on additional FPs).

      The pattern does not arise within single colonies, but when looking at groups of colonies: a green sender is surrounded by a circle of blue-red receivers, which is in turn surrounded by blue-only receivers. This organization can be seen as a collective bullseye pattern (green centre, red annulus, blue background). When multiple green colonies are present in the plate, each gives rise to its own bullseye pattern. We have now clarified this detail in the text:

      Lines 244-245 – “This motif is reminiscent of a bullseye pattern emerging not within a single colony, but in groups of colonies”

      (2) The mathematical model appears somewhat superfluous. While it can clearly reproduce the data, it is not used to make interesting predictions, changing parameters (and not initial conditions) that guide further experimental implementations.

      It is true that we presented the mathematical model primarily as being capable of recapitulating our experimental results. However, the model helped us identify the ideal set of initial conditions (i.e. inducers concentrations) for Figure 6, and to better understand the dynamics of 3O-C6-HSL and 3O-C14-HSL diffusion in our system (Supplementary Movie 6). While changing parameters would be theoretically possible (e.g. diffusion coefficients, protein production rates...) we have limited exploration in this direction, as fine-tuning a single molecular parameter, all other things being equal, is experimentally challenging.

      Since several comments highlighted this weakness, we have now employed our mathematical model to predict patterns theoretically achievable with the sequential differentiation program under substantially different experimental conditions. Please find a detailed answer to this point below (see Reviewer #1 (Recommendations for the authors), point 8).

      Future directions:

      The utility of this differentiation process (e.g. in metabolic engineering or for the study of biofilm formation and antibiotic resistance) will become clearer once the FPs are substituted with functional proteins that exert an effect on the cells.

      We agree with Reviewer #1 that, for the purposes of an application, we would have to replace the fluorescent proteins with functional proteins. In the last paragraph of the discussion, we outline several potential applications of our differentiation system, including division of labor, biocomputation, biosensing, and engineered living materials. However, expanding the system in this direction is beyond the scope of this manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) This sentence is confusing to me: ‘When cells were pre-cultured with 100 nM aTc and 1 mM IPTG, the resulting colonies were almost entirely homogeneously green and blue, respectively.’ It reads as if both inducers were added in the same culture. I would rewrite this as: ‘When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.’

      We adapted the text as suggested:

      Lines 127-129 – “When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.”

      (2) It would be good to explain why the follow-up experiments are conducted with inducers at the beginning of the differentiation process, given that, in the absence of inducers, there is spontaneous symmetry breaking as noted by the authors: ‘When cells were pre-cultured in absence of inducers, both green and blue colonies grew in the plate (Figure 2b).’ Is it to obtain consistent results across biological replicates?

      We used defined inducer concentrations to control the ratio of senders: receivers. In the absence of inducers, the population consisted of approximately equal proportions of senders and receivers. Under these conditions, nearly all blue colonies were close to a green colony and exhibited high expression of the red reporter. To enable spatial patterning, we therefore required a population strongly biased towards the receiver state. Accordingly, we used low concentrations of IPTG to shift the population to this state (as explained in lines 235-241 of the revised version). For testing and characterization of the system, we spotted senders alongside receivers (e.g. in Fig 7C and some supplementary figures). In these experiments, pre-culturing with high concentrations of aTc or IPTG ensured that the entire population was in the desired state. We have now revised the text to clarify that the choice of inducer concentrations ensured specific population ratios and the reproducibility of the differentiation assay:

      Lines 235-241 – “We therefore selected a condition that would consistently generate a population strongly biased towards the receiver state, with only a few sparse senders to produce the diffusible signal. Having characterized the TS differentiation landscape (Figure 2), we decided to pre-culture cells starting from the green state in presence of 9, 12 and 18 µM IPTG, then plated at a cell density of 500-2000 colonies per plate, in absence of any positional information. This resulted in few sparse green senders densely surrounded by blue receivers, in a highly reproducible manner.”

      (3) In the legend to Figure 4b, please indicate what ‘experimental points’ are (single colonies?)

      We adapted the figure caption as suggested:

      Figure 4b – “Hexagons, circles and triangles represent the average fluorescence intensity of single colonies.”

      (4) This sentence was confusing to me: ‘The time-lapses supported our hypothesis that the differentiation key processes (colony growth, HSL production and detection, expression of the red reporter) happened simultaneously.’ I thought the whole idea behind the work was to have a step-wise differentiation process and that HSL production had to happen first, then its detection and finally the downstream activation of the red FP expression.

      We acknowledge that this sentence was poorly phrased and may have caused confusion. Our multi-step differentiation program indeed operates in a sequential, stepwise manner at the single-cell level: a blue receiver cell can only transition to the red state after detecting 3O-C6-HSL, and red colonies can only transition to the yellow state after detecting 3O-C14-HSL. At the colony level, one might therefore expect colonies to first appear green or blue and only later acquire red fluorescence. However, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period. Experiments at the single-cell level might make the step-wise progression blue - red - yellow more obvious than colony-level time-lapses. We have revised the text to clarify this point:

      Lines 273-280 – “We therefore collected time-lapse videos of the plate assay, first inducing receiver cells with pure 3O-C6-HSL, then spotting sender cells alongside receivers, and finally performing the full differentiation assay (Supplementary Movies 1, 2, 3, 4). Whilst one might expect colonies to first appear green or blue and only later acquire red fluorescence, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period.”

      (5) The authors write: ‘The combination of spatial and temporal information allowed us to develop a mathematical model that could qualitatively recapitulate the patterning properties of the system’. It is strangely put in my opinion. One is ‘allowed’ to develop a mathematical model under other circumstances, too. I suggest rewriting. For example: ‘We developed a mathematical model combining spatial and temporal information to quantitatively ...’.

      We adapted the text as suggested.

      (6) In Figure 6b, the scale bar is missing

      Thank you for noticing this, we have now added the scale bar.

      (7) Figure 7c: Why is the green colony huge compared with the others? The formation of such huge colonies seems to occur sometimes, but not always (it is seen also in Supplementary Figure 12a and in Supplementary Figure 13, but nowhere else; by the way, in Supplementary Figure 12a there is a spatial pattern, with a clearly visible red ring in the otherwise black colony!).

      The huge green colonies occur only in the experiments where 1 µl of green senders were inoculated at specific locations. In most of the differentiation assays, each colony arises from a single cell at unpredictable locations. For Figures 7c, S13a (previously 12a) and S14 (previously S13) we wanted a single green sender colony in the centre of the plate surrounded by blue receivers. To achieve this, as we briefly explained in the methods, we decided to spot 1 µl of culture at OD=1, therefore these green colonies originate from roughly 8×10<sup>5</sup> cells, which justifies their larger size. We have now clarified this detail in the corresponding figure captions:

      Figures 7c, S13a, S14 – “The sender colonies in this figure do not originate from one single cell, but from 1 µl of a culture of sender cells at OD=1, which explains their larger size (see Methods).”

      Concerning the red rings in figure S13a (previously S12a), we hypothesise they are the effect of the leakiness of the pLux promoter. Observing expression of mCherry in the green sender colony was actually the main indication that the sequential differentiation circuit was not working as planned: if mCherry was being expressed in the sender state, then probably cinI was also being expressed, therefore the green sender colony was undesirably producing 3O-C14-HSL. We therefore replaced pLux with pLuxLac, resulting in tighter control over the mCherry-cinI operon in the green sender state. The improvement associated with this modification can be appreciated in Figure 7c, where the green sender colonies do not display any red fluorescence.

      Regarding the specific pattern of this red signal, i.e. a ring, as opposed to an homogeneous signal in the entire colony, we have different hypotheses, but we have not investigated it thoroughly. Since the colony does not originate from a single cell but from a 1 µl inoculum, the effect might partially be ascribed to a ‘coffee ring stain’ phenomenon: upon absorption of the droplet, the outer ring could feature higher cell density compared to the centre, leading to higher red signal. In other cases, we have sometimes seen ring patterns arising in homogeneous colonies due to growth and temperature effects.

      (8) It would be nice to see the model predictions being used to create a system with different properties than the actual one. Can the authors modify parameters such as diffusion of the quorum-sensing molecule or the gene expression response to it, and see how the output would change? Changing the type of quorum-sensing molecule experimentally is doable, as is the addition of, for instance, a delay module in the gene expression module. Nonetheless, I am aware that this would require some time to do, but it would, in my opinion, strengthen the paper.

      We thank Reviewer #1 for the valuable suggestions. We have now employed the mathematical model to test the pattern resulting from various diffusion coefficients combinations and added Supplementary Figure S15 and a paragraph in the main text. Choosing diffusible signals with different diffusion coefficients would indeed allow to generate more complex patterns, including concentric rings where the expression of the red and yellow reporters does not overlap. While we agree that experimentally validating this point would be of considerable interest, we believe that changing diffusion rates is challenging and beyond the scope of a typical three-month revision, as it would most likely require substantial optimization and fine-tuning.

      Concerning the delay module, experimentally it could be implemented by introducing an intermediate step with a transcription factor modulating the expression of the CinI synthase. While we agree it would technically be feasible, we did not implement it due to time constraints. We would expect the effect of a delay module to be the following: production of C14-HSL would be delayed, therefore the front wave of C14-HSL would be consistently lagging behind the front wave of C6-HSL. The distance between the two fronts would be proportional to the delay introduced by the module. We hypothesise this would generate patterns with small yellow circles inside larger red circles. Our simulations already capture these dynamics by varying diffusion coefficients, and given the absence of experimental data to constrain parameters, we decided not to include a hypothetical delay module in our mathematical model.

      See Supplementary Figure S15

      Lines 378-385 – “Our model suggests that varying the diffusion coefficients of the two diffusible signals could generate more complex patterns, particularly concentric rings where the expression of red and yellow reporters does not overlap (Supplementary Figure S15). Interestingly, the largest region of mCitrine reporter expression is predicted when the diffusion of the first signal is slow, while that of the second is fast. This suggests that sufficient local accumulation of the first signal is required to reach the threshold for production of the second signal, which can then spread further away to produce an outer blue-yellow region.”

      Reviewer #2 (Public review):

      In this manuscript, the authors implement a three-step genetic programme in E. coli that converts an initially homogeneous population into spatially structured sender, receiver, and ‘matured’ receiver colonies on agar without externally supplied positional information. They combine a TetR/LacI toggle switch for symmetry breaking, LuxI/LuxR quorum sensing for a paracrine signalling step, and CinI/CinR for an autocrine signalling-like maturation step, and complement the experiments with a mathematical model that qualitatively reproduces pattern formation over a range of initial conditions.

      While the article has many strengths such as a clear conceptual framing using Waddington landscapes, a modular and carefully optimised circuit design, thorough experimental characterisation of the toggle and quorum-sensing modules, integration of spatial modelling with experiments, and generally clear writing and figures, I think it will benefit the article to clarify the definition and stability of ‘differentiated’ states, clarify several quantitative and modelling aspects, better explain how fitted curves and promoter engineering were done, and improve some figure design and wording to avoid ambiguity.

      We thank Reviewer #2 for the summary and for highlighting the strengths of our work. Please find below our detailed responses addressing the weaknesses and gaps you identified.

      Detailed comments below:

      (1) P5-8 / and more generally: A major concern is that producing a reporter output is not, by itself, differentiation. For a state to be credibly called ‘differentiated’, it should be stable (self-maintained) over relevant timescales, ideally in the absence of the inducing context. As written, the manuscript sometimes seems to equate cell type with reporter expression. I strongly suggest adding a short subsection explicitly defining state versus output, and for each claimed state, stating whether it is stable/bistable or unstable/reversible, with evidence. Concretely, the authors should enumerate: a) Togglederived sender versus receiver: stable? under what conditions (inducer ranges, hysteresis window)? b) Paracrine-induced ‘red’ receivers: is this a stable differentiated state, or a context-dependent induction requiring proximity to senders? c) ‘Mature’ (yellow) state: does it persist after removal from the spatial signal field? If not, it should be described as an induced output programme rather than a mature lineage state.

      At present, later sections (and the ‘maturation’ language) risk over-stating what is demonstrated.

      We acknowledge Reviewer #2’s concern regarding the stability and irreversibility of our states and agree this is a limitation in our work. The hysteresis property of the toggle switch is well documented in previous studies (Litcofsky et. al., 2012; Barbier et. al., 2020), so we did not formally re-quantify it. However, the fact that pre-culturing cells without inducer yielded predictable green: blue ratios and very few colonies exhibiting both states indicates that the toggle switch is both irreversible and stable under our experimental conditions. For the second and third steps, the expression of the fluorescent proteins is maintained throughout the duration of the experiment and for several hours to days thereafter. It is indeed established that cells in stationary phase have protein half-lives on the scale of tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025). However, we agree that re-streaking these cells in the absence of sender cells would eventually lead to the cessation of fluorescent protein expression, making the quorum sensing-mediated differentiation steps reversible.

      You can find further details on this point in our reply to Reviewer #3 (Public review), point 1.

      We have now adjusted the wording throughout the manuscript, and in particular we modified the abstract to highlight that our system ‘mimics’ cell differentiation. We have also added sections to explain how red and yellow are reversible cellular programs and not stably differentiated cellular states:

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ONOFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      (2) Figure 2d: It is unclear whether this panel is intended to be qualitative (schematic/ illustrative) or generated from quantitative data. The legend should explicitly state the origin (e.g., representative image, averaged data, simulation output, schematic) and, if quantitative, what was measured, how many replicates, and how the visualisation was constructed.

      We acknowledge the potential confusion generated from this image. While its purpose is to visually illustrate the toggle switch differentiation landscape in 3D, the figure is derived from quantitative data, specifically from the same flow cytometry data used for Figure 2c. Basically, it is the density plot of the (GFP, mCerulean) events recorded for a single replicate (initial state: mixed, induced with 0.003 mM IPTG), but the density is represented with a third dimension (depth) instead of colour intensity (as we did for Figure 2e). We have now adjusted the figure caption to clarify the figure purpose and how it was generated:

      Figure 2d – “Representative 3D energy landscape: valleys represent the green and blue stable states, red line represents the separatrix. The landscape was generated from quantitative flow cytometry data from a single replicate in c (initial state: mixed, induced with 0.003 mM IPTG). We measured GFP and mCerulean fluorescence intensity from 50,000 single cells (see Methods). The depth of each (x,y) point in the landscape corresponds to the number of recorded cells with a specific (GFP, mCerulean) intensity. This specific condition was chosen for illustrative reasons, to show the landscape in an almost symmetrical condition.”

      (3) Figure 2e: The cross-sectional line is described as meant to be comparable, yet the leftmost plot appears to have a different slope from the others. The authors should explain whether this reflects a different scaling/normalisation, a different underlying dataset/condition, or simply a plotting artefact. If these are fitted trends, report the fit function (see also the comment on fitted lines below).

      The observation is correct, the cross-sectional lines do not have the same slope. They are obtained by connecting the local minima of the green and blue population, i.e. the two points with highest density in the density plots. As the mean fluorescence intensity of the two population is not the same across conditions, the resulting lines have different slopes. We chose this strategy instead of using fixed-slope cross-sectional lines to capture the maximum depth of each valley. We have now adjusted the figure caption to clarify how the cross-sectional lines were generated:

      Figure 2e – “Flow cytometry density plots of 5 representative conditions from c (initial state: mixed, inducer concentration indicated below each plot). For each density plot, we computed the coordinates of the local minima in the green and blue populations, then generated an orthogonal plane crossing them. Black lines represent the projections of these planes on the (GFP, mCerulean) plane. [...]”

      (4) Around P7-8: (saddle/separatrix description): When describing the saddle or separatrix between the two valleys, it would be helpful to briefly connect this more directly to a quantitative dynamical-systems perspective: for instance, the intersection of nullclines and how nullcline geometry changes under IPTG/aTc induction. This will make the landscape picture more complete for readers familiar with the original genetic toggle switch work (Garder et al., 2000).

      We thank Reviewer #2 for the suggestion. We have now included the concept of nullclines in our description of the differentiation landscape in the main text:

      Lines 153-159 – “Mathematically, the profile of the toggle switch landscape corresponds to the number of intersections between the nullclines (the curve where the derivative of a given species over time equals zero), which can be one or three (Gardner et. al., 2000). If they intersect only once, there is only one minimum on the potential curve, which corresponds to a single stable state. If they intersect three times, there are two minima and one maximum, which corresponds to two stable states, and the maximum represents the separatrix.”

      Lines 168-171 – “From a quantitative dynamical-systems perspective, the addition of aTc or IPTG modified the geometric shapes of the nullclines, shifting the three solutions. If the inducer concentration is large enough, the nullclines intersect only once, producing a single stable steady state (Gardner et. al., 2000).”

      (5) P9, lines 157-159: The current phrasing (‘in absence of noise, the system would be fully deterministic... in living cells, however, stochastic bursts... change the trajectory’) risks conflating predicting population-level percentages with predicting colony-level trajectories. It would help to clearly separate (i) the ability to predict the overall fraction of ON/OFF (green/blue) colonies from inducer conditions (which is largely deterministic at the population level) from (ii) the intrinsically stochastic choice of state made by any given founder cell and its colony.

      We acknowledge that this wording has generated confusion, but we are not completely sure we understood the suggestion made by Reviewer #2. If we understood correctly, the concern is that single-cell trajectories and population-level ratios are substantially different and should not be confused, and should be investigated and modelled differently.

      To clarify, our initial sentence referred to the toggle switch differentiation landscape, not to the likelihood of a cell or a colony to be in a given state. In absence of noise, the TS landscape is deterministic: one of the two states is always the stronger attractor, for the same initial conditions cells would fall 100% of the times in that valley. In living cells, however, there are sources of noise, which make it possible for two cells starting from the same initial conditions to end up in two different valleys. The stochastic bursts in gene expression can ‘push cells back up’ and beyond the separatrix. Very large noise would allow cells to end up in the green and blue valley regardless of the TS previous state or the inducer concentration.

      The differentiation induced by the toggle switch when the system is initiated close enough to the separatrix is stochastic. The probability of observing a certain ratio green: blue at the population level reflects exactly the likelihood of a single cell falling into one of the two stable states. Our mathematical model does not take into account molecular details (synthesis and degradation/dilution of proteins, affinity of transcription factors for their cognate promoter...) to mimic the stochastic choice at the cell level. Instead, it estimates the TS differentiation landscape from the population-level percentages based on thousands of single cells.

      To avoid confusion, we have removed the initial sentence and replaced it with an explanation of the link between individual cell trajectories and the resulting population-level ratios:

      Lines 159-164 – “When the system is close to the separatrix, cells can progress towards both the green and the blue destiny. The differentiation trajectory of a single cell is largely stochastic and is influenced by bursts of gene expression. Once the TS is locked in one state, the progeny of that cell maintains a memory of that state, hence a colony has the same state as the founder cell. At the population level, the ratio of green and blue colonies reflects the probability of each cell to fall into the green or the blue valley.”

      (6) P11, lines 193-195 (promoter engineering): The main text currently only refers to screening variants and choosing pLux76; I suggest briefly stating in the main text (not only in the supplement) what was changed (for example, promoter box variants, core promoter strength modifications) and what design criteria were used (reduced leakiness, increased dynamic range).

      We adapted the text as suggested:

      Lines 206-211 – we carried out a screening of pLux promoter variants, testing combinations of Lux boxes (G1 [iGem part BBa K1216007] and pLux76 (Grant et. al., 2016) and promoters with different strengths (100% and 54%), in order to identify variants with minimal leakiness and a high fold-change. We identified pLux76 (Grant et. al., 2016) as the regulatory region with the highest fold-change and sufficiently low leakiness among our candidates (Supplementary Figure S4)”

      (7) Use of fitted lines (Figures 2, 4, 5, 7): Wherever fitted curves are overlaid on data, the authors should indicate in the figure legend the explicit form of the fit as well as the fit equation/ parameters. As a reader, it is difficult to interpret what is empirical smoothing versus what is a mechanistic functional form.

      In most cases, with the exception of Figure 2c, the fitted curves have an illustrative purpose and result from empirical smoothing of the data points. Unless otherwise stated, the resulting parameters were not implemented in our mathematical model. For Figure 2c, instead, the experimental data is fitted with a mechanistic functional equation and the calculated parameters were implemented in our mathematical model.

      We have now added the equations and parameters of the lines fitting the experimental points, either directly in the figure caption or in Supplementary Information Tables (Tables VII to XII). In the latter case, the exact table is referenced in the corresponding figure caption.

      (8) P13, lines 232-235: The comparison between induction directly with C6-HSL and induction from sender colonies is qualitative (‘significantly smaller range’). The authors should provide distances (for example, in mm) for the induction range in each case and, if possible, approximate total HSL amounts or concentrations, so that the reader can appreciate the magnitude of the difference.

      We have now calculated the induction range as the distance where half-maximal induction is observed, which allowed us to compare the induction range across conditions. We adapted the text accordingly. We also provide an estimate of the amount of 3O-C6-HSL produced by a green sender colony, based on simulations of our mathematical model, but we highlight that the two conditions are substantially different and care should be used when comparing them: in one case, a fixed amount of C6-HSL is present from t=0 and simply diffuses outwards; in the second case, a source continuously produces C6-HSL that progresses as a wavefront, therefore the concentration profiles and diffusion dynamics are different.

      Lines 220-226 – “The induction range obtained with 100 picomoles of pure 3O-C6-HSL was approximately 6-8 mm (50% of maximal induction at 3.74 mm). The induction range around sender colonies was significantly smaller (50% of maximal induction at 1.8 mm, Figure 4b). Mathematical simulations yielded a similar slope when using 0.25 picomoles of 3O-C6-HSL (data not shown), even though care should be used when comparing diffusion of a fixed amount of inducer with continuous production from a growing source.”

      (9) P13, lines 259-262: The authors model the transition to the stationary phase via a monotonically decreasing sigmoid in time for biosynthetic capacity. What is the rationale or literature basis for this approach to model entry into the stationary phase? The authors should cite prior work and clarify why this form is appropriate here, versus alternatives (nutrient diffusion limitation, logistic growth with resource depletion, etc.).

      In most mathematical models involving pattern formation, entry of cells in stationary phase is simply treated as a step-wise function (Zwietering et al., 1990): the entire population is active (activity = 1) until it suddenly becomes inactive (activity = 0). However, experimental evidence suggests that the decay in activity is better represented by a smooth decreasing function (Gefen et al., 2014). While several frameworks and mathematical equations exist to describe loss of cell viability (for example under scenarios of heat inactivation (Mafart et al., 2002; Van Boekel et al., 2002), or nutrient limitation), we could not find previous works modeling the loss of metabolic activity upon entry in stationary phase. Again, experimental evidence suggest that switches in metabolic state are rather heterogeneous and occur stochastically at the single cell level (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to implement a simple heuristic approximation that captured well our experimental data. We have now added a paragraph in the section Mathematical modelling - Bacterial activity, supported by the appropriate references, to justify our reasoning:

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – In our system, production of diffusible molecules and fluorescent reporters happens on a timescale of several hours, therefore we decided to take into account the entry of cells in stationary phase. Traditionally, bacterial activity has been modelled as a step-wise inactivation function (Zwietering et al., 1990). However, experimental evidence has shown that bacteria support a low constant rate of protein expression even while growth-arrested, suggesting a low decay in bacterial activity (Gefen et al., 2014). Even when considering cell death upon heat inactivation, a Weibull frequency distribution model is preferred over a step-wise viability function (Mafart et al., 2002; Van Boekel et al., 2002). While we could not find mathematical descriptions specifically for entry in stationary phase, several single-cell metabolism studies support gradual, asynchronous state transitions (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to describe the loss of activity function in our system via a monotonically decreasing sigmoid function, that captured well our experimental data.”

      (10) Figure 6c: Are the areas of the plate shown in each column the same field of view across conditions/time, or are these simply representative regions selected per condition (possibly from different plates)? The caption/legend should clarify whether these are matched locations and how images were chosen.

      They are indeed representative regions per each condition. Each condition is a different plate, as the differentiation assays requires plating the culture homogeneously on a fresh plate without inducers. The plates used for imaging were the same used for the quantification in Figure 6b, four images per plate were collected, one representative image was chosen for Figure 6c. We have now adapted the figure caption to clarify how images were collected and chosen:

      Figure 6c – “Representative microscopy images of the spatial patterns generated by cells harbouring the 2-step system, pre-cultured with different inducer concentrations (indicated at the bottom of each column). Each column displays one representative image (from four locations imaged) of the seven plates in b. Rows (from top to bottom): GFP channel, CFP channel, mCherry channel, composite image.”

      (11) Figure 7a: The combination of solid, dashed, and dash-dot arrows/lines is visually hard to read. I suggest replacing the dash-dot line with a fully dotted line or using different colours (if consistent with journal style) to improve readability.

      Thank you for noticing this. We have now replaced the dash-dot line with a simple dotted line in Figures 3c, 7a, S12a (previously S11a) and S13b (previously S12b) to improve readability.

      (12) Figure 7e and similar analyses: The authors should explain in the Methods and/or captions how ‘distance from sender colonies’ is computed when multiple senders exist. Is the distance always measured to the nearest sender, and how are cases handled where a receiver is in the overlapping influence of several senders? This clarification is important for interpreting the fitted curves.

      The calculation of the distance in 4b, 5b and 7e was performed only in cases where sender colonies were sufficiently sparse (i.e., more than 1 cm away) to assume each receiver was under the influence of a single sender. When collecting the microscopy images, we carefully avoided to image fields of view with green senders just outside the edges of the images. Images with multiple senders (e.g., 7c) would require a non-trivial calculation to compute the relative contribution of each sender to the red intensity of each receiver. We have now clarified this detail in the Methods:

      Lines 558-560 – “When collecting microscopy images with senders and receivers, we carefully avoided to image fields of view with green senders just outside the edges of the images.”

      Lines 590-593 – “Calculation of the distance between receiver and sender colonies was performed only for images collected from plates with few sparse (i.e., more than 1 cm away) senders, ensuring that each receiver only sensed the 3O-C6-HSL produced by a single sender colony.”

      Reviewer #2 (Recommendations for the authors):

      (1) P8, ‘upon transformation’: This phrasing is ambiguous and can be misread as referring to DNA transformation. I recommend changing ‘upon transformation’ to ‘after transition’ (or similar) to avoid confusion.

      The phrasing indeed refers to the DNA transformation of circuits into the cells. The two plasmids were co-transformed into MG1655, cells recovered for approximately 1 h, then the bacteria were plated on solid medium in absence of inducers. The resulting colonies, which we refer to as ‘upon transformation’, were a mixture of green and blue (both expressed at low intensity, as highlighted in Supplementary Figure 2). We only used these colonies for Figure 2c, middle row. For all other experiments, we selected colonies that had been pre-differentiated in one of the two states via chemical inducers, and showed strong expression of the respective reporter (Supplementary Figure 2). We have now adapted the figure caption to make this detail more explicit:

      Figure 2c – “For the initial state green and blue, cells were taken from colonies that were homogeneously green or blue, while for the initial state mixed, cells were taken from a colony obtained immediately after transformation of the circuit plasmids into cells.”

      (2) More generally, consider disambiguating terminology around ‘cell type’, ‘state’, and ‘output’, since the current wording occasionally implies stable fate commitment where the data (as presented) may instead support reversible, context-driven induction.

      We went carefully through the text and adapted it appropriately, highlighting which states are stable and which are reversible. See our detailed reply to your point 1 (Public review) above.

      Reviewer #3 (Public review):

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggleswitch-based symmetry breaking with quorum-sensing interactions to generate colonyscale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      However, I think the paper needs to tone down the extent to which the system demonstrates multi-step differentiation or morphogenesis, which is not critical for making the paper valuable. Only the first step of their circuit design (Figure 1), the toggle switch, generates stable alternative states. The latter steps are mainly signal-dependent reporter activation states layered on top of the blue receiver state, rather than true fate transitions. The authors explicitly state that red expression is added without replacing the blue identity, and they also acknowledge that red cells lose their identity upon restreaking unless they remain near sender cells. That substantially weakens the differentiation analogy and makes the Waddington framing too strong.

      We acknowledge that our sequential program does not recapitulate all the features of a differentiation trajectory, and in particular we do recognize the only two stable states are the identities associated with the toggle switch, while red and yellow are outputs indicating activation of a new molecular program. A complete differentiation program would require irreversible cell fate determination, as we mention in the discussion. For the same reason, in our illustrative Waddington landscape (Figure 1) we only represent two valleys, or minima, corresponding to the green and blue states, while the activation of the red and yellow programs does not result in further valleys. We modified the text at various points to clarify this important distinction, and the fact that our system mimics some key steps happening during multicellular differentiation, without claiming that expression of a fluorescent reporter is a stably differentiated state. Notably, we modified the abstract to highlight that our system ‘mimics’ cell differentiation.

      We believe the differentiation analogy and the Waddington framing remain valuable in our work, considering the physical constraints and timescale of our system. While it is true HSL-induced molecular program would eventually deactivate in absence of signal, this would require a significantly longer amount of time than the one we use to observe patterns. Also, upon storage of Petri dishes at 4 °C, the pattern (including red and yellow) remains stable for a few weeks. Finally, our main purpose was to investigate the capacity of a sequential program to generate autonomous spatial patterns, without the need for human intervention. The removal of a colony from the local signal field, e.g. to re-streak it on a fresh plate, represents a strong disturbance of the pattern, not dissimilar from early developmental biology experiments where transplant of tissue portions would sometimes result in fate reprogramming.

      You can find further details on this point in our reply to Reviewer #2 (Public review), point 1.

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ON-OFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      A related concern is that the 3rd step does not introduce a new spatial organizing rule. The authors show that the second signal remains confined to cells already receiving the first signal, and explicitly conclude that it functions only as an autocrine cue rather than a second paracrine layer. As a result, the 3-step system seems more like an added local readout or maturation layer. Overall, the main 2-step outcome is sparse green sender colonies surrounded by red-expressing blue receivers, with distant receivers remaining blue. That is a valid engineered pattern, but it is still a local, threshold-response circuit architecture.

      The comment is correct, we indeed refer to the 3rd step as ‘maturation’ in the manuscript, as the newly emerged red population activates a new molecular program, which cannot be activated in blue receivers that have not been exposed to C6-HSL. While we had initially expected that production of a second diffusible signal could result in signal propagation, therefore generating an outer yellow ring surrounding the red ring, our experiments showed C14-HSL only acted locally, and our mathematical simulations confirmed that the front wave of the second signal was always lagging behind the first one. We have now employed our mathematical model to explore which conditions would support a new patterning rule, and added Supplementary Figure S15. If the diffusion coefficients of C6-HSL and C14-HSL were 10-100 times smaller and 2-5 times larger respectively, a yellow-only area could appear around the red area.

      Regarding the last sentence, we are unsure about the concern raised and which alternatives Reviewer #3 would recommend. We fully agree our genetic program leverages a local, threshold-response circuit architecture, it is indeed a reaction-diffusion system based on quorum sensing signalling. Many natural patterning and morphogenesis systems do rely on local, rather than global, interactions, yet they are capable of forming complex and hierarchical structures.

      The autonomy claim should be toned down and stated more precisely. The plate patterning occurs without externally imposed spatial gradients, which is a strength. However, by design, the overall system behavior depends strongly on pre-culture inducer conditions that set the sender:receiver ratio, and this externally imposed history is central to the final pattern. This property is tied to how the circuit is designed where steps 2 and 3 largely respond to symmetry breaking introduced in step 1, which is dependent on both history and initialization on the plate. In particular, currently the pattern formation process is quite variable (e.g. figure 5), depending on how different colonies flip the toggle switch, and consequently, how many become senders and how many become receivers. It would have been fascinating if they could also demonstrate the differentiation within individual colonies, leading to intra-colony patterns. This aspect should at least be discussed.

      We would like to clarify that, by ‘autonomous’ we refer to the reaction-diffusion system that starts functioning upon the seeding of the cells on the plate. All previous steps (including the cell culturing, dilution and plating) correspond to setting the initial conditions. We agree with the description of Reviewer #3 concerning the features of our system, but we argue that autonomy and dependence on the initial conditions are two different properties. We claim our system is both autonomous and sensitive to initial conditions. Upon initialization of the system on the plate, there is no further intervention, or nudging (e.g. time-dependent light stimuli, addition or removal of inducers at specific times...), therefore the system is autonomous. It is also dependent on the initial conditions, which are set homogeneously for all the cells (no positional information, each cell experiences the exact same conditions). We believe this is a strength of our system, which allows to generate a variety of diverse yet reproducible spatial patterns. Independency from initial conditions would consistently generate the same pattern, regardless of the initial inducer concentration, which might be interesting for certain applications (e.g., ensuring a fixed population ratio, provide robustness to environmental variability...), but it is not universally superior.

      In order to clarify that we only refer to the reaction-diffusion system on the plate as the autonomous component, we have now added a sentence in the main text:

      Lines 132-134 – “In our differentiation assay, cell culturing, dilution and plating correspond to setting the initial conditions for the system. Upon initialisation on the plate, no further intervention was performed, therefore the system evolved in an autonomous fashion.”

      Concerning the intra-colony patterns, that was admittedly our initial interest. We explored conditions that would consistently generate colonies with sectors, for example by inducing cells in one of the two states and then growing them on agar supplemented with the opposite inducer. However, we immediately observed that, due to the close proximity of sender and receiver bacteria, and the rapid diffusion relative to the timescale of gene expression, all blue sectors were also red. This outcome effectively eliminated the distance-dependent nature of the quorum sensing response. In order to take full advantage of the diffusible system, we focused instead on well-separated homogeneous colonies. We have now added Supplementary Figure 5, the corresponding figure caption, and we briefly discuss in the main text the occurrence of intra-colony patterns and why we did not investigate them further:

      See Supplementary Figure S5

      Lines 228-231 – “We then tested the potential of the 2-step differentiation system to generate self-organized spatial patterns. We observed rare motifs arising within the sporadic colonies showing both green and blue sectors. Due to the close proximity of the sender and receiver bacteria within a colony, the blue sectors always showed strong red signal (Supplementary Figure S5).”

      The mathematical model is useful in guiding both the characterization of parts, modules and the overall system. However, the claims around its quantitative predictive power should also be made narrower. The simulations are built from multiple fitted and partly hand-tuned components, including toggle-switch response curves, colony-growth rules, diffusion, reporter-response functions, and activity decline. This supports a calibrated qualitative reconstruction of the observed patterns, but not a strong predictive or mechanistic validation.

      We accepted the suggestion and replaced ‘predicted’ with ‘recapitulated’ or ‘simulated’ at various locations in the main text:

      Lines 105-107 – “Throughout the work, experimental results were used to develop a mathematical model, providing us with insights that guided further experimental efforts.”

      Lines 292-293 – “We developed a mathematical model combining spatial and temporal information to qualitatively recapitulate the patterning properties of the system.”

      Lines 316-318 – “We therefore employed the mathematical model to simulate the patterns generated from the 2-step differentiation system for different initial blue: green ratios. The model suggested a variety of outcomes [...]”

      Other specific points:

      (1) Given the topic of the work, the authors should cite closely relevant studies in programming pattern formation, including: Cao et al, Cell 2016 Collective space-sensing coordinates pattern scaling in engineered bacteria Rajasekaran et al, Cell 2024 A programmable reaction-diffusion system for spatiotemporal cell signaling circuit design Lu et al, BioRxiv 2024 Discovery of interpretable patterning rules by integrating mechanistic modeling and deep learning

      We have added these references to the introduction and discussion.

      (2) The model assumes identical diffusion coefficients for C6-HSL and C14-HSL despite their substantially different molecular sizes and hydrophobicities. This assumption could distort kinetic lag with differential diffusion in explaining the autocrine confinement of the third step. Its impact should at least be explored in the simulations.

      The comment is very appropriate. To the best of our knowledge, no exact values have been published for C6-HSL and C14-HSL diffusion in water, let alone for diffusion in agar. However, estimates in water range between 3 ·10<sup>−6</sup> cm/s and 5 · 10<sup>−6</sup> cm/s at 25°C, depending on the chain length (i.e. a factor of 1.7 at most). In response to your concern, we have now computationally explored the effect of using signals with different diffusion coefficients and their impact on the resulting spatial pattern (Supplementary Figure S15). Varying the diffusion coefficient of the second diffusible signal does not substantially change the pattern: lower diffusion leads to the yellow region being slightly more concentrated around the green sender and more intense. Higher diffusion has the opposite effect, with yellow areas being wider but less intense. Conversely, changing the diffusion coefficient of the first signal leads to significant changes in the final pattern, suggesting that the kinetics of C6-HSL accumulation and dispersal have a strong influence on the production of C14-HSL and therefore of the yellow signal.

      (3) The mCherry response parameters change significantly between the 2-step and 3-step systems. The authors acknowledged this change but did not provide a clear explanation.

      We believe that the elements contributing to the different mCherry response in the 2-step and 3-step systems are: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression. We have now added a paragraph in the section Mathematical modelling Bacterial activity that lists those differences. We have also added Supplementary Figure S20, containing the experimental data that was fitted to obtain the parameters listed in the second row of Table II. Finally, we highlight that for the implementation of our mathematical model we were only interested in the response function to varying 3O-C6-HSL concentrations and not in the exact values of individual parameters.

      Supplementary Material, section A. “Mathematical modelling, subsection 3. mCherry production and 30-C6-HSL sensing – The parameters fitted for mCherry production in the 2-step and 3-step systems are quite different. We hypothesize that the following elements contribute to the difference: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression.”

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – For the implementation of our mathematical model, we were not interested in analysing individual parameters values, but rather in the resulting response function to varying 3O-C6-HSL concentrations. Further analysis would be needed to determine whether these parameters are statistically significant, but this is beyond the scope of this present manuscript.”

      (4) The 3-step system is evaluated at only a single condition with no simulation comparison, in contrast to the systematic 11-condition validation of the 2-step system.

      The rationale for not repeating the 11-condition assay with the 3-step system is that the patterns would be substantially the same as Figure 6c (red colonies would also be yellow). Furthermore, the Nikon SMZ25 stereo microscope we used to take images with a large field of view (ideal for the differentiation assay in Figure 6c) would not allow us to discriminate well between green and yellow colonies, making the interpretation of the patterns difficult. However, following up on your comment, we have now used our mathematical model to produce a representative image corresponding to Figure 6d, that we included as Supplementary Figure S16, exploring the effect of varying the green: blue ratio with the 3-step system.

      Reviewer #3 (Recommendations for the authors):

      Please see above. My comments are largely about improving the rigor and clarity of the writing, particularly those related to conceptual claims, as well as some modeling analysis to strengthen their conclusions.

      We thank Reviewer #3 for the suggestions. See our detailed replies to your points above.

    1. eLife Assessment

      This important study employs a closed-loop, theta-phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats and reports that disrupting theta-timescale coordination impairs performance of challenging aspects of spatial behaviors, while sparing hippocampal replay and spatial coding in hippocampal place cells. Technically rigorous experiments were performed, and solid evidence is provided to support the claims. The findings are expected to advance theoretical understanding of learning and memory operations and to provide practical implications for the application of similar optogenetic approaches.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Joshi and colleagues demonstrates that the precise theta-phase timing of spikes is causal for CA1 hippocampal theta sequences during locomotion on a linear track and is necessary for learning the cognitively demanding outbound component of a hippocampus-dependent alternation task (W-maze), independently of replay during immobility. To reach these conclusions, the authors developed a theta-phase-specific, closed-loop manipulation that used optogenetic activation of medial septal parvalbumin (PV) interneurons at the ascending phase of theta during locomotion. This protocol preserved immobility periods, allowing a clean and elegant dissociation from SWR-associated replay.

      The manuscript is well written and was a pleasure to read. The work described if of high quality and introduces several notable advances to the field:

      a) It extends prior studies that manipulated theta oscillations by examining precise temporal structure (specifically theta sequences) rather than only LFP features.

      b) The closed-loop manipulation enabled dissociation between deficits in theta sequences during a behavioural task and SWR-associated replay activity.

      c) As controls, the authors included rats with suboptimal viral transduction or optic-fibre placement, and, within subjects, both stimulation-on (stim-on) and stimulation-off (stim-off) trials. Notably, sequence disruption persisted into stim-off periods within the same session.

      Overall, this is a strong manuscript that will provide valuable insights to the field.

      After revision, the manuscript has been substantially strengthened. The authors did incorporate the vast majority of the reviewer's comments and have expanded the discussion of prior medial septal manipulations, clarified the rationale for their theta-sequence analyses, added analyses of SWR and replay in the rest/sleep box as well as provided additional methodological and histological validation.

      The new rest-box analysis is a great addition and directly addresses my request to distinguish aSWRs from events during longer off-track rest periods (rSWR). The data support the narrower conclusion that no large group difference was detected in rest-box ripple rate or duration.

      The new observation (in response to reviewer #2, point3.2) that on the W-track theta power does not fully recover during stimulation-off periods does change the interpretation of the results. It means that these epochs are then not a physiologically recovered control condition. Therefore, the persistent disruption of theta sequences during the middle block cannot, alone, demonstrate that disrupting sequences during the earliest experience produced a lasting plasticity-related effect. It could also reflect a lingering network effect of the stimulation that persists after laser delivery has stopped. In my opinion the discussion should present at least these two alternatives: the disruption of early experience-dependent plasticity, as well as the incomplete physiological recovery from the preceding stimulation. The linear-track recovery data is helpful, but it does not guarantee the same mechanisms/effects will be present on the novel Wmaze (versus the familiar linear track).

    3. Reviewer #2 (Public review):

      Summary:

      The authors of this study developed a closed-loop optogenetic stimulation system with high temporal precision in rats to examine the effect of medial septum (MS) stimulation on the disruption of hippocampal activity at both behavioral and compressed time scales. They found that this manipulation preserved hippocampus single-cell-level spatial coding but affected theta sequences and performance during a spatial alternation task. The performance deficits were observed during the more cognitively demanding component of the task and even persisted after the stimulation was turned off. However, the effects of this disruption were confined to locomotor periods and did not impact waking rest replay, even during the early phase of stimulation-on. Their conclusion is consistent with previous findings from the Pastalkova lab, where MS disruption (using different methods) affected theta sequences and task performance but spared replay (Wang et al., 2015; Wang et al., 2016). However, it differs from a recent study in which optogenetic disruption of EC inputs during running affected both theta sequences and replay (Liu et al., 2023).

      Strengths:

      The experiments were well designed and controlled, and the results were generally well presented.

      Comments on revised version.

      The authors of this study addressed all my concerns, some of them successfully. The stimulation disrupted theta oscillations, making quantification of theta sequences problematic. Within the constraints of their experimental design, the authors tried their best to address my concerns. Therefore, I am satisfied with the current version and express no further comments.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study employs a closed-loop, theta-phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats and reports that disrupting theta-timescale coordination impairs performance of challenging aspects of spatial behaviors, while sparing hippocampal replay and spatial coding in hippocampal place cells. The findings are expected to advance theoretical understanding of learning and memory operations and to provide practical implications for the application of similar optogenetic approaches. The experiments were viewed as technically rigorous, but the strength of evidence provided in the current version of the manuscript was viewed as incomplete, mostly due to limited analyses and the descriptions of some of the experimental protocols.

      We thank all reviewers for their overall assessment, thoughtful comments, and suggestions. We have now addressed each of the reviewers’ comments in detail and updated the manuscript on bioRxiv (URL: https://www.biorxiv.org/content/10.1101/2025.09.15.675587v2). In addition, we have shared the raw data, intermediate analysis files, and the complete repository to facilitate replication of the analysis and figures.

      Code repo: github.com/LorenFrankLab/ms_stim_analysis

      Data repo: dandiarchive.org/dandiset/001634

      Docker containers (see GitHub repo for use instructions):

      - Database: https://hub.docker.com/r/samuelbray32/spyglass-db-ms_stim_analysis

      - Python notebooks: https://hub.docker.com/r/samuelbray32/spyglass-hub-ms_stim_analysis

      (1) Novelty and contrast with earlier manipulations:

      We now explicitly contextualize our results with prior pharmacological (Wang et al., 2016; Wang et al., 2015; Koenig et al., 2011; Brandon et al., 2014), systemic (Robbe & Buzsaki 2009; Petersen and Buzsáki 2020), and behavioral (Drieu et al., 2018) manipulations that also assessed some of the physiological features we evaluated. This contrast helps us highlight both the insights and the discrepancies observed in the prior approaches. We also more clearly explain the novelty and importance of our specific approach for temporally and physiologically precise manipulation. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (2) Additional analysis on SWRs during rest:

      Since submitting the manuscript, we have conducted additional analysis on the rate and length of SWRs in the rest box (rSWRs) and found that the rate and length are also indistinguishable between targeted and control animals (effect of manipulation between control and targeted animals; rSWR rate: p=0.45; rSWR length: p=0.94, mixed-effects model). We also find evidence for sequential neural representations (“significant replay”) in the rest box when the encoding was performed in the behavioral arena. Example trajectories and full analysis per animal are shown in the new supplementary figure (Figure S6). These results are consistent with our observations on aSWR rate, length, and content in the behavioral arena. Additionally, based on the reviewer’s recommendation, we have evaluated the fraction of ripples with continuous trajectories during the rest box before the W-Track experience and in the subsequent sleep epochs after the first exposure. We find that with experience on the track, the proportion of continuous replays increases on average in both control and transfected animals, and both groups of animals show an overlapping range of continuous trajectory lengths.

      (3) Theta sequence measurement in the absence of theta:

      We now explicitly explain why our manipulation makes it more appropriate to measure sequential hippocampal representations during locomotion (i.e., theta sequences) without using theta oscillation or an epoch-averaged, relatively large sliding window as a reference. The key insight here is that our manipulation suppresses theta and thus makes it difficult or impossible to accurately identify theta phase. We explain that while theta-phase-based approaches were used in prior work; these prior analyses may have confounded the absence of hippocampal theta sequences during locomotion by the inability to detect theta oscillatory phase reliably. We show that our method of using clusterless Bayesian decoding, in which we estimate the decoded position at every 2ms timestep, is indeed able to capture endogenous hippocampal sequences even without imposing any requirements of aligning to theta oscillations, thus providing an unbiased estimate of the rhythmicity of hippocampal spatial representations.

      (4) Additional analysis on place cell stability and tuning:

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Joshi and colleagues demonstrates that the precise theta-phase timing of spikes is causal for CA1 hippocampal theta sequences during locomotion on a linear track and is necessary for learning the cognitively demanding outbound component of a hippocampus-dependent alternation task (W-maze), independently of replay during immobility. To reach these conclusions, the authors developed a theta-phase-specific, closed-loop manipulation that used optogenetic activation of medial septal parvalbumin (PV) interneurons at the ascending phase of theta during locomotion. This protocol preserved immobility periods, allowing a clean and elegant dissociation from SWR-associated replay.

      The manuscript is well written and was a pleasure to read. The work described is of high quality and introduces several notable advances to the field:

      (a) It extends prior studies that manipulated theta oscillations by examining precise temporal structure (specifically theta sequences) rather than only LFP features.

      (b) The closed-loop manipulation enabled dissociation between deficits in theta sequences during a behavioural task and SWR-associated replay activity.

      (c) As controls, the authors included rats with suboptimal viral transduction or optic-fibre placement, and, within subjects, both stimulation-on (stim-on) and stimulation-off (stim-off) trials. Notably, sequence disruption persisted into stim-off periods within the same session.

      Overall, this is a strong manuscript that will provide valuable insights to the field. I have only minor comments:

      (1) As the authors note, it is striking that both behavioural performance and spike patterns are altered during stim-off trials. They propose that "disruption of theta sequences during the initial experience in an environment is sufficient to have lasting effects," implying that rapid, experience-dependent plasticity is driven by sequential firing. Does this imply that if rats were previously trained on the task, subsequent stim-on and stim-off trials would yield different outcomes, with stim-off trials showing improved performance and intact theta sequences? For example, if the sequence of one-third stim-on, one-third stim-off, one-third stim-on were inverted to off-on-off, would theta sequences be expected to emerge, disappear, and potentially re-emerge? While I am not asking for additional experiments, I think the discussion could be extended in this aspect.

      Alternatively, could the number of stim-off trials (one third of the total) be insufficient to support learning/induce plasticity? In the controls, ~50-100 trials appear necessary to achieve high performance.

      We think it is likely that pretraining would result in a different outcome, although we did not test this possibility. We have modified the discussion to address this point:

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to recover. This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (2) In line with the point above, the authors characterise the behavioural changes induced by MS optogenetic stimulation specifically as a "learning deficit," as rats failed to improve across 300 trials in an initially novel environment (W-maze). While they present this as complementary to prior demonstrations of impaired performance on previously learned tasks (Zutshi et al., 2018; Quirk et al., 2021; Etter et al., 2023; Petersen et al., 2020), an alternative interpretation is a working-memory deficit. This would produce the same behavioural pattern, with reference memory (the less cognitively demanding trials) remaining intact despite stimulation and concomitant changes in theta sequences. This interpretation would also be consistent with work in certain disease models, where reduced synaptic plasticity and working-memory deficits co-occur with preserved place coding despite impaired theta sequences (e.g., Viana da Silva et al., 2024; Donahue et al., 2025).

      We agree that traditionally deficits in alternation tasks have been termed “working memory” but we also note that this may confuse some readers, as memory in these tasks does not engage persistent activity throughout delays in areas like the prefrontal cortex.

      (3) It was not immediately clear whether SWR-associated activity was derived from the interleaved ~15-min rest sessions in a rest box, or from periods of immobility or reward consumption in the maze (aSWR, as in Jadhav et al 2012). Regardless, it would be informative to compare aSWR events within the maze to rest-box SWRs that may occur during more prolonged slow-wave episodes (even if not full sleep). This contrasts with Liu et al. (2024), who analyzed replay during ~1.5-h sleep sessions.

      We thank the reviewer for this comment and suggestion. We will now explicitly mention in the manuscript that we have measured awake sharp wave ripples (aSWRs) on the track during immobility periods. In addition, in line with this and another reviewer’s recommendation, we have included analyses on the proportion of rest SWRs (rSWRs) between control and targeted animals in Supplementary Figure 6, replicating our findings during aSWRs. However, we note that the differences between Liu et al.’s (2024) study and ours. While they waited and analyzed replay during 1.5 hours of sleep sessions, in our study, and in the 1-day w-track learning protocol, sleep sessions are typically shorter (15-20 minutes). That said, control animals in our tasks with an intact hippocampus (previous studies) and intact theta sequences (our study control animals) can learn the task in one day, so a longer replay period is not necessary to learn the task.

      Reviewer #2 (Public review):

      Summary:

      The authors of this study developed a closed-loop optogenetic stimulation system with high temporal precision in rats to examine the effect of medial septum (MS) stimulation on the disruption of hippocampal activity at both behavioral and compressed time scales. They found that this manipulation preserved hippocampus single-cell-level spatial coding but affected theta sequences and performance during a spatial alternation task. The performance deficits were observed during the more cognitively demanding component of the task and even persisted after the stimulation was turned off. However, the effects of this disruption were confined to locomotor periods and did not impact waking rest replay, even during the early phase of stimulation-on. Their conclusion is consistent with previous findings from the Pastalkova lab, where MS disruption (using different methods) affected theta sequences and task performance but spared replay (Wang et al., 2015; Wang et al., 2016). However, it differs from a recent study in which optogenetic disruption of EC inputs during running affected both theta sequences and replay (Liu et al., 2023).

      Strengths:

      The experiments were well designed and controlled, and the results were generally well presented.

      Weaknesses:

      Major concerns are primarily technical but also conceptual. To further increase the impact of this study by contrasting findings from different disruptions, it is necessary to better align the analysis and detection methods.

      We thank the reviewer for their assessment and critical questions. We have addressed each of the comments below. As we note in our responses, our inclusion criteria were based on our analysis approach, in which we aimed to measure the impact of our manipulation where possible for each animal, and ideally at the level of every 20-minute run epoch. This is a strength of our experimental approach, and we will explicitly explain that in a next version of the manuscript.

      Major concerns:

      (1) To show that MS disruption does not affect spatial tuning, the authors computed the KL divergence of tuning curves between stimulation-on and stimulation-off conditions. I have two main questions about this analysis:

      (1.1) The authors seem to impose stringent inclusion criteria requiring a large number of spikes and a strong concentration of tuning curves. These criteria may have selected strongly spatially tuned cells, which are typically more stable and potentially less vulnerable to perturbations. Based on the Figure 2 caption, it seems that fewer than 10% of cells were included in the KL divergence analysis, which is lower than the usual proportion of place cells reported in the literature. What is the rationale for using such strict inclusion criteria? What happens to the cells that are not as strongly tuned but are still identified as significant place cells?

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      (1.2) The KL divergence was computed between stimulation-on and stimulation-off conditions within the same animal group. However, the authors also showed that MS stimulation had lasting effects on theta sequences and performance even during stimulation-off periods. Would that lasting effect also influence spatial tuning? Based on these questions, the authors should perform additional analyses that directly measure spatial tuning quality and compare results across control and experimental groups - for example, spatial information of spikes (Skaggs et al., 1996), tuning stability, field length, and decoding error during running.

      To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). Mean decoding error between animals depends on the recording quality. We have reported these values around stimulus times at the choice point in the previous version of the manuscript (Sup. Figure 3G). We also find that the distribution of place field coverage is not statistically different within and across animals (stimulation on versus stimulation off: p = 0.17, control versus targeted: p = 0.9, interaction of stimulation and targeting: p = 0.8, linear mixed effects model).

      (2) The authors compared their results with those from Liu et al. (2023) and proposed that the different outcomes could be explained by different sites of disruption. However, the detection and quantification methods for theta sequences and replay differ substantially between the two studies, emphasizing different aspects of the phenomenon. I am not suggesting that either method is superior, but providing additional analyses using aligned detection methods would better support the authors' interpretations and benefit the field by enabling clearer comparisons across studies. In the current analysis, the power spectrum of the decoded ahead/behind distance only indicates that there is a rhythmic pattern, without specifying the decoding features at different theta phases. Moreover, the continuous non-local representations during ripples could include stationary representations of a location or zigzag representations that do not exhibit a linear sequential trace. Given that, the authors should show averaged decoding results corrected by the animal's actual position within theta cycles and compute a quadrant ratio. For replay analysis, they could use a linear fit (as in Liu et al., 2023) and report the proportion of significant replay events.

      In Liu et al., 2023 study theta sequences were quantified by explicitly segmenting theta cycles into phase quadrants and evaluating the structure of the decoded representations within these phase-defined windows. This approach can be applied in studies that have a stable theta that can provide a temporal reference frame and where there is not enough spatial coverage in the spikes to identify the extent of ahead/behind representations. This analysis is hence not ideal to detect theta sequences in our data, as it is prone to errors due to our disruption of theta oscillatory activity itself. To account for this, we have used a Bayesian clusterless decoding approach in which we can measure the structure of hippocampal sequential representations even without imposing any restrictions on their temporal order. This method allows us to reliably capture theta sequences during locomotion in control animals (Fig. 4B) and their disruption in targeted animals (Fig. 4F).

      Since our experimental paradigm suppresses theta oscillations themselves, we used a clusterless decoding approach (as in Joshi et al., 2023) to obtain an unbiased estimate of rhythmicity for hippocampal spatial representations. Briefly, we estimated the peak of the posterior at every 2ms time step and computed the distance between that value and the actual position of the animal (decode-to-animal distance). We confirmed that, as expected, in control animals, we could detect “theta sequences” as in prior studies without explicitly requiring theta oscillatory cycle windows.

      We have now evaluated the distributions of SWRs that are labeled as continuous and have a trajectory displacement > 10 cm and consistently find continuous replays during both aSWRs and rSWRs (Fig. 5, Fig. S6). We also find that these distributions overlap between control and targeted animals (n=1216 targeted, n=3212 control, p=0.06).

      Author response image 1.

      (3) The finding that theta sequences and performance were impaired even during stimulation-off periods is particularly interesting and warrants deeper exploration. In the Discussion, the authors claim that this may arise from "the rapid plasticity engaged during early learning." However, this explanation does not fully account for the observation. Previous studies have shown that theta sequences can develop very rapidly (Feng et al., Foster lab, 2015; Zhou et al., Dragoi lab, 2025). If the authors hypothesize that rapid plasticity during early stimulation-on disrupts the theta sequence, then the plasticity window must also be short and terminate during the subsequent stimulation-off period. Otherwise, why can't animals redevelop theta sequences during stimulation-off? The authors should conduct additional analyses during the stimulation-off periods of the W-maze task. For example:

      (3.1) What is the spike-theta phase relationship? Do the phases return to normal or remain altered as during stimulation-on?

      We thank the reviewer for this question. We have now looked at theta power on the W Track and find that theta power does not fully recover on the W Track even during stimulation-off periods. We have included this in the results. See Figure S4.

      (3.2) Is there a significant place-field remapping from stimulation-on to stimulation-off? (Supplementary Figure 3F includes only a small subset of cells; what if population vector correlations are computed across all cells, or Bayesian decoding of stimulation-on spikes is performed using stimulation-off tuning curves?)

      We have addressed this question by computing the correlation between the peak of the place fields between stimulation-on and stimulation-off conditions and find that the distributions are largely overlapping between control and targeted animals on both the linear (Spearman correlation of place field peaks between stim-on and stim-off intervals, control vs targeted, p=0.68, t-test, n=6 control, n=4 targeted epochs) and wtrack (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.91, t-test, n=36 control, n=18 targeted epochs).

      (3.3) The authors should also discuss why the stimulation-off epochs were not sufficient to support learning, and if the stimulation-off place cell sequences could have supported replay.

      We do not find the aSWR-associated replay to be impacted as a result of our manipulation. We have not conducted a specific experiment to test the impact of longer stimulation-off periods on the formation of place cell sequences, but in response to this and another question from Reviewer 1, we have added the following speculation in the discussion.

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (4) Citations and/or discussion of key studies relevant to the current work are missing: Wang et al. in Pastalkova lab 2015-2016 studies for disruption of theta sequence (but not place cell sequence) disrupting learning but not replay, Drieu et al. in Zugaro lab 2018 study on disruption of theta sequence affecting sleep replay, Farooq and Dragoi 2019 for association between a lack of theta sequence and presence of waking rest replay during postnatal development, etc. The authors should discuss what the conceptually new findings in the current study are, given the findings of the previous literature above.

      We thank the reviewer for this question. We have substantially modified the introduction to include this prior work and highlight that our manipulation enabled us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (5) The assessment of theta sequence is not state-of-the-art:

      (5.1) Detecting the peak of cross-correlograms between neurons (CCG) relates to behavioral timescale CCG, not the theta sequence one; for the theta sequence, the closest to zero local peak should be used instead.

      Here we think we failed to explain our analyses clearly, as we did exactly that analysis. The cross-correlation peak in Fig. 2D is the peak within the theta timescale (+/- 100ms lag), not a slow behavior timescale (e.g. +/- 1s lag). In the revised version of the manuscript, we have improved the explanation so that this confusion does not arise.

      (5.2) How were other methods of detecting theta sequences performing on the stimulation-on/stimulation-off data: Bayesian decoding, firing sequences?

      In the absence of detectable theta oscillations, it is inappropriate to use the commonly used metric of detecting theta sequences. That method also assumes that the theta oscillatory cycle is the correct temporal “reference” for hippocampal theta sequences. To our knowledge, there is no direct evidence for this. Thus, in our manuscript, we have used two approaches to identify hippocampal spatial representations during stimulation-on and stimulation-off periods:

      (1) Clusterless Bayesian decoding approach

      (2) Pairwise correlations between neurons

      Analyses using these approaches provide an unbiased method to detect hippocampal spatial-temporal sequences during locomotion without using theta oscillations as a reference. Indeed, in control animals, we recover the endogenous theta timescale correlation and sequence structure during locomotion. Using the same approach in targeted animals reveals that even though we can measure hippocampal spatial sequential representations in targeted animals, the timing between them is altered.

      (5.3) How was phase precession during stimulation-on/stimulation-off?

      We cannot do this analysis for the theta manipulation condition since there is an unreliable phase estimate in the absence of theta on W Track. Based on the reviewers’ comments, we have now evaluated the theta phase precession on the linear track 10Hz stimulation condition. Phase precession could be observed even when evaluated against the entrained LFP. Further, at a population level, we observed that autocorrelograms followed the entrained 10Hz LFP in the stimulation-on condition compared to the stimulation-off condition (n=41 neurons, solid lines are medians, shaded areas 25/75 percentiles). We have added these results in Supplementary Figure 5.

      (6) It would be important to calculate additional variables in the replay part of the study to compare the quality of replay across the 2 groups:

      (6.1) Proportion of significant replay events out of the detected multiunit events.

      We assume the reviewer is recommending the analysis of significant replays as defined by performing a linear fit on the trajectory. We recognize that there is a fundamental diversity in the replay architecture, as has been shown using clustered and clusterless decoding approaches, and that assuming that only the replays with a linear fit are significant might bias us toward those trajectories. In aSWRs, we have shown that the replay events that are labeled “continuous” are equivalent in proportion between control and targeted animals. In the current version of the manuscript, for rSWRs, we similarly computed the proportion of replay events with a continuous trajectory that traverses at least 10cm on the w-track and found no differences between control and targeted groups. We also found the proportion of continuous rSWRs to increase in targeted animals, similar to control animals. We have included these results in the new Figure S6.

      (6.3) The average extent of trajectory depicted by the significant replay events in the targeted compared to the control, stimulation-on/stimulation-off.

      Here, we show the distributions of continuous replay trajectories (> 10 cm) between control and targeted animals, showing overlapping distributions (p=0.06). See Author response image 1.

      Reviewer #3 (Public review):

      Joshi et al. present an elegant and technically rigorous study examining how the temporal structure of hippocampal spiking during locomotion contributes to spatial learning. Using a closed-loop, theta phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats, the authors demonstrate that disrupting theta-timescale coordination impairs performance on the cognitively demanding component outbound trajectory of a spatial alternation task, while sparing hippocampal replay, place coding, and the simpler inbound learning. The work aims to dissociate the role of theta-associated temporal organization during navigation from sharp-wave ripple-associated replay during subsequent rest periods, providing a mechanistic link between theta sequences and learning. The findings have important implications for models of septo-hippocampal coordination and the functional segregation between online (theta) and offline (SWR) network states. That said, there are a few conceptual and methodological issues that need to be addressed.

      We thank the reviewer for the thoughtful comments and suggestions to strengthen our work. In a revised manuscript, we explicitly address prior publications where manipulations (either pharmacological or behavioral) have shown a dissociation between temporal sequence formation, place coding, and replay in the introduction and highlight the specific insights obtained from using our approach. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      One concern is the overall novelty of this work; the dissociation between online temporal sequence and offline replay events following memory deficits has previously been shown by Wang et al., 2016 elife. While the authors discuss Lui et al., 2023, which demonstrates MEC activation of inhibitory neurons at gamma frequencies during locomotion disrupts theta sequences, subsequent replay and learning (line 65-66), they do not reference Wang et al., 2016 who performed a very similar study with MS pharmacological inactivation, and report large decreases in theta power, attenuated theta frequencies together with behavioural deficits but SWR replay persisted. Given strong similarities in the manipulation and findings, this study should be discussed.

      We agree that this important study should be cited in the introduction, and have now included that in the revised introduction. Importantly in Wang et al., 2016 elife paper, they showed that while replay can exist after periods where theta sequences are disrupted using muscimol inactivation in the medial septum, the rate of replay was higher than in control, leaving open the possibility that the increased replays might contribute to consolidation of non-task memories. We also note that the manipulation was applied for hours, making it impossible to attribute a precise physiological basis for poor behavioral performance.

      Along the same lines, it should be noted that Brandon et al. (2014, Neuron) demonstrated that hippocampal place codes can still form in novel environments despite MS inactivation and loss of theta, indicating that spatial representations can emerge without intact septal drive. Referencing this study would strengthen the discussion of how temporal coordination, rather than spatial coding per se, underlies the learning deficits observed here.

      Thank you. We have referenced the study appropriately in the updated manuscript.

      Our findings, and previous dissociations between precise timescale and place field properties (Petersen and Buzsáki 2020; Liu et al. 2023; Wang, et al., 2016, Brandon et al., 2014) further suggest that different circuits with different time constants are responsible for processing spatial and temporal information in the hippocampal circuit. Spatial/contextual information may arrive from regions with slower timescales (such as the cortex), making them less susceptible to sub-second brief disruptions, while precisely timed inputs from the medial septum coordinate the tightly controlled timing offsets between hippocampal neurons. We hypothesize that learning requires the intersection of these two streams of information in the hippocampal network, and is impaired by the disorganization of the precise temporal templates in which internal plans can be matched to external inputs.

      The conclusion that disrupting "theta microstructure" impairs learning relies on the assumption that the observed behavioral deficits arise from altered temporal coding from within hippocampal CA1 only. However, optogenetic modulation of medial septal PV neurons influences multiple downstream regions (entorhinal cortex, retrosplenial cortex) via widespread GABAergic projections. While the authors do touch on this, their discussion should expand to include the network-level consequences of entorhinal grid-cell disruption and how this could affect temporal coding both online and offline.

      We agree and have expanded our previous discussion to include that possibility.

      Modified in discussion:

      “Understanding precisely why this temporal organization is critical will require more distributed measurements. Notably, MS targets include multiple cortical and subcortical targets (Joshi 2017), and our manipulation may have disrupted precise spike timing throughout these regions. Key amongst these regions include the pre- and para-subiculum, retrosplenial area and the entorhinal cortex, which also receive dense PV projection in addition to the CA3 and DG (Joshi et. al., 2017; Viney et. al., 2018; Salib et al., 2020). Disrupting the spatial code or spike-timing in these regions may contribute to the disruption of sequential activity we have observed. However, we do note that the first response of the stimulation to spiking activity in CA1 is consistent with a strong disinhibitory input to CA3, with spike latencies less than 20 milliseconds. Additionally, monitoring regions beyond the temporal cortex would be informative given the broad coordination between hippocampal theta and other systems (Joshi et al. 2023; Eichenbaum 2017; Buño and Velluti 1977; Berg, Whitmer, and Kleinfeld 2006; Ledberg and Robbe 2011).”

      The finding that replay content, rate, and duration are unchanged is critical to the paper's claim of dissociation. However, the analysis is restricted to immobility on the track. Given evidence for distinct awake vs. sleep replay, confirming that off-track rest and post-session sleep replays are similarly unaffected would confirm the conclusions of the paper. If these data are unavailable, the limitation should be acknowledged explicitly. Moreover, statistical power for detecting subtle differences in replay organization or spatial bias should be added to the supplement (n of events per animal, variability across sessions).

      We thank the reviewer for this suggestion. We have now explicitly evaluated replay properties during rest and indeed confirm that replay rate, length, and content in this manipulation are indistinguishable between targeted and control animals. We have now added statistical power to these claims by reporting the number of events per animal and variability across sessions. As we mentioned above, where possible, we have attempted to replicate each analysis per session (20min session), and all analyses are replicable for each targeted and control animal. Our internal controls are the power of our approach, as there can be significant animal-to-animal variability.

      The exact protocol for optogenetic stimulation is a bit confusing. For the task, the first and final third (66%) of trials were disrupted and were only stimulated when away from the reward well and only when the animal was moving. What proportion of time within "stimulated" trials remained unstimulated? Why were only 66% of trials stimulated?

      We developed a stimulation protocol that included an interval in which stimulation was not applied (stimulation-off) periods to assess the impact of the stimulation at the level of each animal and epoch. This analytical approach gives us the power to study the impact of the stimulation and our experimental approach to the spiking patterns and neural activity observed. We will modify the explanation in the methods. Based on the reviewer’s comment, we have now calculated the proportion of time within the epoch where the laser was on. The laser is ON for ~10% of the total time in the epoch. We implemented a closed-loop algorithm that had both the spatial location of the animal and the theta phase of a reference electrode as online inputs. The trigger was applied when three conditions were met: a spatial inclusion criterion, a speed criterion, and theta phase criteria.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1D & G: should include the frequency band of the filtered trace in the figure caption.

      We have now included the frequency band of the filtered trace (5-11Hz) in the figure caption.

      (2) The observation that hippocampal cells respond to stimulation within ~5-10 ms (Figure 2C), while theta power decreases significantly after 200 ms (Figure 1G), is interesting. Do the authors have any hypothesis explaining this discrepancy?

      Theta is likely best understood as the result of a complex feedback loop between the medial septum and various structures in the hippocampal formation. This would make it robust to individual perturbations. We now discuss this in the text.

      Modified in results:

      “Note, that while we can measure the response to hippocampal cells within 5-10ms, theta power decreases gradually over a period of 200 milliseconds. This is consistent with the view that theta oscillatory activity that can be measured in the hippocampus is a result of a multi-region feedback loop that involves various cortical and subcortical networks, a feature that may make it robust to individual perturbations.”

      (3) Figure 2A: The caption states that "in control rats, endogenous theta sequences are apparent." This is unclear. The purple region marks the ascending phase, but theta sequences are typically defined across peak-to-peak cycles. The shading may distract from this; consider adding dashed lines to indicate theta peaks.

      We thank the reviewers for this suggestion, and we have modified the visualization for this figure per the reviewers’ suggestion.

      (4) Figure 2B (top panel): How are the cells ordered? It appears they are sorted by peak firing time before stimulation (<0). This could misleadingly suggest a sequence before stimulation but not after. Some cells have peak firing in the 0-40 ms range-what is the intended interpretation of this ordering?

      We have modified the spiking by the spiking order after stimulation onset.

      (5) Line 239: Typo - "as for the linear track" should be "as for the W track."

      Modified.

      (6) Lines 279-282: The statement "This suggests a surprising level of preserved representational movement across frequencies ranging from 6 to 12 Hz" is unclear. What does "preserved representational movement" mean? A simpler explanation for the matching of power-spectrum peaks to stimulation frequency could be that stimulation increases firing rates, biasing decoding toward locations with higher mean firing, producing rhythmic fluctuations at the stimulation frequency under Poisson decoding assumptions.

      We have now added examples of the decode to animal distance in the different stimulation conditions to supplement this statement. We will also acknowledge that the change in firing rate might contribute to the observation.

      (7) Line 313: The phrase "while sparing learning on the interleaved trials where a less cognitively demanding choice was required" may be confusing. Although the authors refer to inbound runs, readers may interpret "interleaved trials" as stimulation-off trials.

      We agree and have modified this phrasing.

      (8) The claim that "representations of locations more distant from the animal were preserved during theta disruption" (line 350) is unclear-where is this shown?

      Fig S3G: Distribution of max decode to animal distance within theta cycles

      Consistent with the pairwise analyses (Supplementary Figure 3B-D), the ∼8 Hz peak in the power spectrum of the ahead/behind distance was significantly larger in control animals than in targeted animals across all conditions (Figure 4I; pooled comparison p’s < 10<sup>−4</sup>, hierarchical bootstrap p’s < 0.05). At the same time, the maximal extent of locations represented ahead and/or behind the animal near the choice point, where choices must be made on outbound and inbound trials, did not differ between control and targeted animals (Supplementary Figure 3G). Thus, our findings indicate that the precise timing of non-local representations during theta was disrupted, but the spatial extent was not.

      (9) The title of Figure S1 appears twice.

      Thank you. We have edited it.

      Reviewer #3 (Recommendations for the authors):

      (1) It should be noted that there are several referencing errors that should be addressed. Please check the following:

      (a) Lines 77-81 - a few incorrect references.

      (b) Line 155 - Disruption of hippocampal theta has consistently shown to preserve spatial properties of place cells: Brandon et al., 2014, Koenig et al., 2011.

      Thank you. We have corrected these references.

      (2) Figure 1C- Did the authors stain for colocalization with virus and PV in the MS?

      Yes, our viral construct has eYFP expressed together with channelrhodopsin, and this rat line and viral construct have been previously standardized (Yu et al., 2018, Lepperod et al., 2021). In addition, we conducted three standardization experiments and visually inspected the overlap between eYFP-positive cells and parvalbumin-expressing neurons (86/86 YFP-expressing neurons tested positive for PV). An example of the overlap is now included in Supplementary figure 7.

      (3) Figure 1 - Shows an impressive reduction of theta power. Can Figure 1G be extended to show theta recovery immediately following stimulation?

      Our stimulation protocol restricted the stimulation to periods that were within the spatial and speed inclusion criteria. The period immediately following the stimulation on every trial is reward delivery, during which the animal has already slowed down, and we do not expect high theta power. Based on the reviewer’s suggestion, we have inspected the first few trials after the stimulus is turned off in the linear track and have confirmed that theta power immediately recovers. We have added these additional figures in Sup. Figure S4H.

      Author response image 2.

      (4) Do authors have examples of the same trajectory where temporal coding is intact in baseline and disrupted during stimulation? Does an intact theta sequence ever develop in the target animals?

      Yes, target animals do exhibit intact spatial sequences during the stimulation-off periods on the linear track and w-track (example in a targeted animal during the stimulation-off condition on linear track below). As we show in Figure 4E and Supplementary Figure S4G, on the w-track, while spatial sequences exist, each cycle’s duration is not consistent across the behavioral experience. Thus, on average, power spectrum of the ahead-behind distance is not rhythmic at 8Hz.

      (5) Have the authors computed spike-phase relationships or shown phase-position to evaluate phase precession in individual cells in both the 10 Hz stim vs the phase-specific?

      As discussed above, due to unreliable phase estimates during theta suppression, we have chosen to base our analysis of sequential structure on temporal cross-correlation and decoding analysis. However, in response to the reviewer’s question, we have evaluated phase precession under 10Hz stimulation. Consistent with our overall results for the ahead-behind distance, we find that individual cells phase-precess with the newly entrained theta. Additionally, at a population level, we are able to visualize a clear shift in peak spiking frequency.

      However, we agree that these results do not completely rule out the contribution of additional spikes, and in a revised version of the manuscript, we have included that possibility.

      (6) Did authors perform other types of stimulations that either drove the dominant frequency out of theta range (gamma) or completely desynchronize the system (by stimulating along all different times of theta to perform a phase-specific "scramble")?

      We did not attempt a gamma or scrambling stimulation condition.

      (7) There is overall inconsistency in the formatting of references throughout the text.

      We apologize for these errors and have rectified them in the updated version.

    1. eLife Assessment

      This important study uses biochemical and single-molecule approaches to characterize how yeast Cdc13 assembles on single-stranded telomeric DNA and contributes to telomere protection. The evidence supporting the proposed model for Cdc13 assembly and its role at telomeres is solid. The findings provide insight into how Cdc13 contributes to the maintenance and protection of chromosome ends. This work will be of interest to researchers in telomere biology, DNA replication, biochemistry, and single-molecule biophysics.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how the Saccharomyces cerevisiae telomere-binding protein Cdc13 assembles on a 12-nucleotide single-stranded telomeric DNA substrate. Using complementary smFRET and CoSMoS measurements, together with dimerization- and DNA-binding-defective mutants, mass photometry, photobleaching analysis, and kinetic modeling, the authors assign state II to a DNA-bound Cdc13 monomer and state III to a stable complex containing two Cdc13 molecules. They propose that the stable DNA-bound dimer forms predominantly through sequential recruitment of two monomers, while direct binding of a preformed dimer represents a less frequent pathway. This work addresses an important mechanistic question in telomere biology because the pathway of Cdc13 assembly may influence telomere recognition, end protection, and recruitment of telomere-maintenance factors.

      Strengths:

      Overall, the manuscript is well written, and the combination of two complementary single-molecule approaches is a strength. The central model is interesting and potentially important.

      Weaknesses:

      Several issues require clarification or additional analysis.

      Major comments:

      (1) The mass photometry results are central to the mechanistic model and should be presented more prominently.

      The conclusion that Cdc13 binds DNA predominantly through sequential monomer recruitment depends strongly on the oligomeric state of Cdc13 at the concentrations used in the single-molecule experiments. The mass photometry results currently provide the principal direct evidence that Cdc13 is predominantly monomeric at 5 nM but exists as a mixture of monomers and dimers at 20 nM. These results should therefore be included in a main figure rather than only in the Supplementary Information. The authors should also provide a complete description of the mass photometry in the Methods section.

      Related to lines 191-192, the authors should discuss the estimated cellular or nuclear concentration and abundance of Cdc13 and compare these values with the experimental concentrations at which states II and III are populated. Because the relevant quantity may be the effective local concentration at a telomere rather than the average nuclear concentration, this distinction should also be acknowledged. Such a discussion is needed to establish under what physiological conditions sequential monomer loading versus binding of a preassembled dimer would be expected.

      (2) 75% labeling efficiency must be clearly defined and incorporated into both the stoichiometric and kinetic analyses.

      In line 251, the authors state that the labeling efficiency of DY-649P1-Cdc13 is 75%, but it is not clear how this value was measured. The authors should state whether 75% refers to the efficiency of the sortase reaction, the fraction of labeled molecules in the final purified preparation, or a value inferred from the plateau in Figure 3B. If it was inferred from the binding plateau, the plateau below 100% could also arise from inactive or inaccessible DNA molecules, incomplete colocalization detection, inactive protein, or an effect of the fluorophore on binding. An independent measurement, such as absorbance-based determination of the dye-to-protein ratio, quantitative gel analysis, or intact-mass analysis, would be preferable.

      Incomplete labeling has direct consequences for the interpretation of Figure 3. With a labeling probability of 0.75, a true Cdc13 dimer would contain zero, one, or two fluorophores. Thus, even among detectable dimers, 40% would appear as one-step photobleaching events. A one-step event therefore cannot automatically be equated with a monomer without correcting for labeling efficiency. The authors should quantitatively account for incomplete labeling when inferring the relative monomer and dimer populations from the photobleaching data.

      The same issue is even more important for the CoSMoS kinetic model. The authors should incorporate labeling efficiency into the observation model, or at minimum perform simulations or a sensitivity analysis demonstrating that the inferred transition rates and state assignments are robust to 75% labeling.

      (3) Apparent state I→III transitions do not by themselves demonstrate direct binding of a preformed Cdc13 dimer.

      At lines 326-330, the authors interpret state I→III transitions as direct binding of a solution dimer and conclude that sequential binding is approximately eightfold faster than direct dimer binding. This interpretation is not yet sufficiently established. An observed I→III transition demonstrates only that no state II intermediate was resolved; it does not distinguish true binding of a preformed dimer from sequential binding in which the state II lifetime is shorter than the temporal resolution of the experiment.

      This concern is particularly important because the CoSMoS experiments were conducted at only 0.3-1.25 nM Cdc13, whereas mass photometry indicates that Cdc13 is predominantly monomeric even at 5 nM. In addition, the smFRET signals were averaged over a sliding window of ten 50-ms frames, which could obscure short-lived intermediate states.

      The authors should show representative raw traces containing apparent I→III transitions in a supplementary figure and quantify the shortest state II dwell time that could be detected under the acquisition, smoothing, and HMM procedures used.

      (4) The interpretation of the WT-Cdc13/Cdc13^R635C mixture requires further clarification.

      For Figures 2G-H and lines 209-218, the reduced state III population in the mixture of 2.5 nM WT Cdc13 and 2.5 nM Cdc1^R635C is interpreted as evidence for a solution monomer-dimer equilibrium and formation of a nonfunctional WT-mutant heterodimer. However, at least two nonexclusive explanations should be considered:

      a) Formation of WT-mutant heterodimers in solution could reduce the concentration of free WT monomers and WT homodimers available to form state III.<br /> b) A WT-mutant heterodimer, or recruitment of Cdc13^R635C to a DNA-bound WT molecule through protein-protein interactions, could produce a DNA-bound complex that cannot adopt the state III conformation because only one subunit has an intact DNA-binding interface.

      The authors should discuss these possibilities explicitly and clarify expected FRET states.

      For direct visual comparison, Figure 2G should include the FRET histograms for 2.5 nM WT Cdc13 alone and 5 nM WT Cdc13 alone, in addition to the WT-mutant mixture. The concentrations of both the initially loaded WT Cdc13 and the WT or mutant protein added during the chase experiment in Figure 2I should also be stated in the main text and figure legend.

      (5) The physical basis of the different FRET values for states II and III should be explained earlier and more carefully.

      The assignment of state II and state III to one and two bound Cdc13 molecules is supported by the combined smFRET and CoSMoS results. However, the manuscript should explain earlier why the addition of a second Cdc13 molecule is expected to produce a further decrease in FRET. Because the fluorophores are attached to the DNA, the different FRET values imply a change in the distance, orientation, or local photophysical environment of the DNA-linked dyes when the second Cdc13 binds. CoSMoS establishes a change in protein stoichiometry, but it does not by itself establish that the DNA has undergone a particular conformational change.

      The discussion at lines 404-423 suggests that the second Cdc13 induces a rearrangement of the first Cdc13-DNA complex. This is a reasonable hypothesis, but it should be presented as an inference rather than as a demonstrated DNA conformational transition. References 44-46 describe different RPA binding modes and rearrangements of protein-DNA contacts; they do not directly demonstrate the specific DNA conformational change proposed here. The authors should either provide more direct support or revise the discussion accordingly. The distance estimates should also be described cautiously because they assume that dye orientation and photophysical properties are unchanged between states.

      The rationale for the internally positioned Cy3 constructs in lines 147-149 should also be explained more clearly. Why does it demonstrate that Cdc13 cannot bind duplex DNA?

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript presents an interesting and potentially important single-molecule study of Cdc13 assembly on telomeric ssDNA. The experimental observations are intriguing, particularly the identification of distinct FRET states associated with different Cdc13 occupancies. However, I have substantial concerns about whether the current data support the central mechanistic conclusion as strongly as the authors claim. In particular, the manuscript does not yet clearly distinguish between the observation that two Cdc13 molecules are associated with DNA and the stronger mechanistic claim that Cdc13 loads sequentially as monomers and subsequently dimerizes on DNA.

      Major concerns

      (1) The central mechanistic conclusion is stronger than the evidence

      The manuscript's principal model is that Cdc13 proceeds through the pathway monomer → DNA-bound monomer → recruitment of a second monomer → stable DNA-bound dimer. However, the experiments establish this sequence only indirectly.

      The authors show that FRET state II is associated with one Cdc13 molecule. FRET state III is associated with two Cdc13 molecules. Cdc13-DM produces state II but not state III. WT Cdc13 can undergo I→II→III transitions. Direct I→III transitions also occur.

      These observations are consistent with sequential binding, but they do not uniquely demonstrate that the second Cdc13 molecule first binds as a monomer and subsequently undergoes dimerization on DNA. In particular, the direct I→III events indicate that a preassembled dimer or another cooperative pathway can also contribute.

      I therefore recommend substantially tempering statements such as "Cdc13 initially loads onto telomeres as a monomer." A more defensible formulation would be: "The data support a kinetically favored sequential pathway involving an initial Cdc13 binding event followed by recruitment of a second Cdc13 molecule." The authors could still present sequential loading as the preferred model, but the language should clearly distinguish a kinetically supported pathway from a uniquely established molecular mechanism.

      (2) The assignment of FRET states II and III to exactly one and two Cdc13 molecules requires stronger validation

      The interpretation of states II and III as one- and two-Cdc13 states is central to the entire mechanistic model. Because labeling efficiency, incomplete labeling, photophysics, and heterogeneous molecular populations can all affect the observed distributions, the authors should provide a quantitative probabilistic model. Specifically, the authors should calculate the expected distribution of one- and two-labeled-Cdc13 species given the experimentally determined labeling efficiency and compare these expectations with the observed FRET-state populations. This analysis would provide an important independent validation of the proposed stoichiometric assignments.

      (3) The TG12 substrate raises an important stoichiometric and conformational question

      The manuscript argues that a single Cdc13 molecule binds approximately 11 nt, while two Cdc13 molecules can bind a 12-nt ssDNA substrate. This raises an immediate mechanistic question: if one Cdc13 occupies approximately 11 nt, how can two Cdc13 molecules simultaneously associate with only 12 nt of ssDNA? This issue should be addressed experimentally rather than only structurally or schematically. A particularly informative experiment would be to systematically vary ssDNA length. The probability and kinetics of state III formation could then be quantified as a function of substrate length. Such an experiment would determine whether formation of the two-Cdc13 state genuinely requires additional DNA and whether the two proteins occupy overlapping or distinct regions of the substrate.

      (4) The "salt-resistant" interpretation is overstated

      The manuscript repeatedly describes state III as "salt-resistant" and uses this observation to support the existence of a highly stable physiological complex. However, the reported experiment examines only a relatively modest range of NaCl concentrations (50, 75, and 100 mM).

      I recommend either expanding the salt-dependence analysis substantially or using more quantitative language. For example, the authors could report the fraction and lifetime of state III as a function of salt concentration and define explicitly what they mean by "salt-resistant." The current data support persistence under the tested conditions, but they do not by themselves establish exceptional physiological stability.

      (5) The mass-photometry experiment does not establish the physiological solution equilibrium

      The mass-photometry data are useful for demonstrating that Cdc13 can exist in monomeric and dimeric forms, but the current experiment does not establish the equilibrium between these species under physiologically relevant conditions.

      A concentration series would be valuable. The observed monomer/dimer populations should be fit to an explicit equilibrium model to obtain an apparent dimerization constant and assess how strongly the equilibrium depends on Cdc13 concentration. This would also help connect the solution behavior to the single-molecule observations and determine whether the observed DNA-bound dimer could plausibly arise from a pre-existing solution dimer.

      (6) The WT + R635C mixing experiment is overinterpreted

      The interpretation of the WT + R635C experiment is currently complicated and somewhat speculative. The experiment is potentially informative, but the conclusions appear stronger than what can be directly inferred from the data. The authors should clearly distinguish between the observations directly supported by the mixing experiment and the mechanistic interpretation proposed from them. In particular, the experiment does not necessarily establish the precise sequence of DNA binding and dimerization events.

      (7) The requirement for DNA-binding activity in both Cdc13 molecules is not fully established

      The manuscript concludes that both Cdc13 molecules must possess DNA-binding activity. However, the R635C experiment does not clearly distinguish between:<br /> two independently DNA-bound Cdc13 molecules; and one Cdc13 molecule directly bound to DNA plus a second molecule whose DNA-binding surface is required for allosteric stabilization of the dimer.

      This distinction is mechanistically important. Additional experiments using DNA-binding-defective mutants in defined heterodimeric configurations would help determine whether both molecules directly contact DNA or whether DNA binding by one molecule promotes recruitment/stabilization of the second through protein-protein interactions.

      (8) Stronger controls are needed for FRET-state assignment

      The FRET states are treated as discrete molecular states, but alternative explanations for heterogeneous FRET populations should be considered more explicitly.

      Important controls would include: concentration-dependent FRET measurements in the absence of DNA binding; fluorescence controls to determine whether the observed states could arise from dye-protein interactions; Cdc13 mutants with altered DNA-binding specificity; alternative dye positions; demonstration that the major FRET states are reproduced with independent labeling configurations. The duplex-positioned Cy3 controls, which show little FRET change, are useful. However, they do not completely exclude the possibility that protein-induced changes in DNA conformation contribute to the observed FRET states. Independent labeling geometries would substantially strengthen the assignment.

      (9) The physiological relevance of the 12-nt substrate requires better justification

      The authors use TG12 as their primary substrate and state that telomeres contain approximately 12-14 nt of ssDNA during most of the cell cycle. This rationale requires greater biological context. Telomere length and the extent of the exposed G-rich strand are dynamic and heterogeneous, and Cdc13 has established functions throughout telomere replication. The authors should explain more carefully why TG12 is biologically representative and how the proposed mechanism is expected to behave on substantially longer telomeric substrates. The TG25 experiment is useful, but at present it functions primarily as a stoichiometric observation. A systematic substrate-length analysis, as suggested above, would turn this observation into a mechanistic test.

      (10) The relationship to existing structural studies requires deeper discussion

      The manuscript presents sequential loading as a novel mechanism, but existing structural and biochemical studies of Cdc13 dimerization and Cdc13-DNA architecture are essential for interpreting these observations. The authors should explicitly reconcile their proposed model with the existing structural literature. In particular, they should address:

      Does the known Cdc13 dimerization interface permit simultaneous DNA binding by both subunits?

      Is the dimerization interface compatible with the proposed DNA-bound state II?

      Could DNA binding alter the dimerization interface?

      Are the two Cdc13 molecules predicted to bind overlapping or distinct portions of the telomeric sequence?

      Can the structural models accommodate the apparent stoichiometry on a TG12 substrate?

      Without this reconciliation, the proposed sequential-loading mechanism remains somewhat disconnected from the established structural framework.

      Other important issues:

      (11) The Kd comparisons are confusing

      The manuscript should include a table summarizing the different Kd values and explicitly explaining why they differ. In particular, describing 3.7 nM as the Kd for the first monomeric binding step while reporting an apparent Kd of 1.2 nM requires careful kinetic and statistical justification. The authors should distinguish clearly among microscopic Kd values, apparent Kd values, and parameters inferred from kinetic models.

      (12) The direct-dimer pathway deserves greater attention

      The occurrence of direct I→III transitions is mechanistically important and should not be treated primarily as an exception to the sequential pathway. A more balanced conclusion would be:

      "Both pathways contribute to formation of the final Cdc13-DNA complex, with the sequential pathway being kinetically favored under the experimental conditions." This interpretation appears better aligned with the data and would still constitute a strong mechanistic conclusion.

      (13) The "kinetic proofreading" interpretation is currently speculative

      The Discussion proposes that the first Cdc13 monomer provides a kinetic proofreading step. This is an interesting hypothesis, but it is not directly demonstrated by the current experiments.

      I recommend changing this to language such as "a kinetic proofreading-like mechanism may be possible" unless the authors can provide direct evidence that the first binding event selectively promotes productive complex formation or rejects nonproductive substrates.

      (14) Biological-function claims should be clearly separated from the in vitro findings

      The manuscript frequently connects the stable state III complex with telomere protection, telomere length regulation, CST formation, Est1 recruitment, Pol α recruitment, and prevention of DNA degradation. None of these functions are directly tested in the present study.

      The experiments establish a biochemical/single-molecule mechanism in vitro. They do not establish that state III is the functional protective species in vivo.

      This distinction should therefore be maintained throughout the Abstract, Discussion, Key Points, and concluding statements. The authors can appropriately discuss these possibilities as<br /> implications or hypotheses, but should avoid presenting them as demonstrated functions of state III.

    1. eLife Assessment

      This valuable proof-of-concept study tests and validates a novel magnetic resonance-based method - proton-observed proton-edited MRS - for tracking glucose metabolism in the brain beyond simple substrate uptake, without the need for specialist hardware typically restricted to research settings. The approach offers a promising way to assess excitation/inhibition balance by detecting labelled glucose metabolites. The pre-clinical data in mouse models provide convincing support for the authors' hypothesis. The evidence in the human brain is incomplete: the results would require further optimisation of the experimental set-up and validation in a larger group of volunteers. This work will be of broad interest to researchers studying psychiatric, neurodevelopmental, and related brain disorders.

    2. Reviewer #1 (Public review):

      Excitation/inhibition (E/I) balance between excitatory (glutamate) and inhibitory (GABA) neurotransmission is being increasingly studied using magnetic resonance spectroscopy (MRS), for example in autism spectrum disorder, schizophrenia and attention deficit/hyperactivity disorder. These are typically measured using standard single-voxel MRS methods (eg PRESS, sLASER) to measure glutamate/glutamine or "Glx" and spectral editing methods (eg MEGAPRESS, MEGA-sLASER) techniques to measure GABA. Such methods only give a measure of the total MR-visible metabolite concentration in the voxel. That is, they don't distinguish between glutamate/GABA involved in neurotransmission or in other metabolic processes.

      Cherix et al present a method for measuring glucose metabolism, with the potential to be used on a standard clinical MRI scanner. This proof-of-concept study focused on measuring proton signals from glucose metabolites, including lactate, glutamate, GABA & Glx. The method works by administering 13C universally labelled glucose (where all six carbon atoms are substituted with 13C). When the glucose is metabolised, 13C label is incorporated into specific positions within its metabolites. Protons attached to 13C don't produce a signal in 1H-MRS in a subsequent MEGA-sLASER scan, leading to a drop in the signal as the labelled metabolite concentration builds up. At the same time, "satellite resonances" appear for protons coupled to 13C, which increase as the labelled metabolite concentration builds up. Metabolite concentrations were inferred using a simple dynamic model.

      The main strength of the method is that it enables dynamic metabolic information that would typically only be available with multi-nuclear MRS capability (13C or 2H) to be achievable using standard preclinical or high field (>= 7T) human MRI systems, with widely available spectral-editing acquisitions.

      The results in the mouse spectra seem very convincing for lactate and GABA/Glx. For the human scans, changes in lactate weren't detectable, which is not surprising given how little lactate appears in the normal brain. In the discussion, the authors argue the method can potentially be used in a standard, 3T clinical scanner. It may be too soon to conclude that, as it's not yet clear there would be sufficient SNR in spectra at that field strength. Additionally, the heteronuclear coupling constants are quite high. The authors recognise that this may complicate detection of satellite resonances due to signal dephasing. Another potential complication at 3T (or 2.9T) is the potential for the satellite resonances to come close to the GABA peak at 3 ppm. More accurate measurements of coupling constants will allow that to be determined. Another potential limitation of the method is macromolecule contamination of the 3 ppm GABA peak. That may be overcome by using macromolecule-nulled MEGA-editing, though frequency navigators may be necessary to overcome the increased sensitivity to frequency drift.

      The authors achieved their aims of showing that imaging glucose metabolites was possible using standard proton-only MRI systems, without the need for additional multinuclear coils, transmitters and receivers. The evidence if very compelling for the mouse scans but only incomplete for the human scans.

      The ability to quantify metabolites involved in E/I balance has the potential to revolutionise studies into disorders where changes in E/I balance are implicated. This is especially the case for preclinical models. Such studies may be less feasible in clinical studies due to the high cost of universally 13C-labelled glucose, but this proof-of-concept is a promising start.

    3. Reviewer #2 (Public review):

      Summary:

      The main aim of the presented manuscript was to test and validate proton observed proton edited 13C MRS and track the 13C label from uniformly labeled U-13C-glucose into glutamine, glutamate, GABA and lactate in rodent and human brain in vivo.

      Strengths:

      In contrast to already established methods of 13C and 2H MRSI this method applies only proton RF and thus can potentially be implemented on a standard clinical scanner.

      Weaknesses:

      The validation of the method in a human setting is rudimentary. First, even at ultra-high field strength of 7T, the authors did not reach sufficient SNR to detect and quantify GABA with good CV, and further, the experiment needs optimisation to reach metabolic and fractional enrichment steady state and/or for additional conditions to show the possibility of lactate detection.

    4. Reviewer #3 (Public review):

      Summary:

      In their paper, the authors propose using dynamic MEGA-sLASER acquisitions to track the incorporation of ¹³C from uniformly ¹³C-labeled glucose into glutamate(+glutamine) and GABA pools. This approach enables the direct and simultaneous investigation of excitatory and inhibitory neurotransmission and metabolism without requiring a ¹³C radiofrequency coil. While the authors' goals and efforts are commendable, the work has several significant limitations.

      Strengths:

      Use of excellent hardware (¹H cryoprobe, ultra-high-field MRI scanners); Ambitious objectives.

      Weaknesses:

      MRS data analysis requires improvement; Small cohort sizes; No demonstration that the proposed acquisition scheme is superior in terms of robustness, accuracy, or performance.

    1. eLife Assessment

      There is a significant need for improved understanding of the neural circuit mechanisms for learning and memory dysfunction in Alzheimers disease. In this valuable study, Zheng and colleagues compare neural coding between wild-type and App knock-in rats experiencing different contexts. Solid evidence of a difference emerges when "spatial" and "temporal" neural coding dynamics are considered separately. However, the lack of a link between neural coding and behaviorally measured learning, and minor issues with the analyses, are limiting factors of the evidence and the significance of the work.

    2. Reviewer #1 (Public review):

      Summary:

      Remapping is clearly degraded in amyloid models, but people and animals with a lot of pathology often hold onto more function than their spatial maps would predict. The authors' idea is that CA1 carries two things at once: an explicit code where rate maps onto position, and an implicit one in the temporal relationships between cells, and that AD hits the first much harder. They recorded CA1 with tetrodes in App(NL-G-F) and WT rats running an A-B-B-A sequence of open field sessions, repeated daily for six days, with rest in between. They then compared rate map measures against a cofiring measure (pairwise Kendall's tau) and looked at SWR reactivation during rest.

      Strengths:

      (1) The design is right for the question. Alternating back to the familiar arena separates "can the network register a new context" from "can it get back to the old one," and the finding that App rats look OK on the first A to B transition but fall apart on the return is the most striking thing here. The confused cell result, 24% vs 4.5%, is easy to read and hard to dismiss.

      (2) Six consecutive days is also worth something. Most work on this is cross-sectional, and looking at how the coding changes with accumulated experience is the right way to ask about plasticity.

      (3) The PIR analysis in Figure 5 is the bit I liked best. Subtracting each cell's position-predicted rate before computing tau is a reasonable check that the cofiring effects aren't place fields in disguise, and it's more than most papers making this argument do. The theta index control is a good instinct too, though see below on how it's analysed.

      Weaknesses:

      (1) No behaviour. This is the main problem and everything else is secondary. The title, abstract, intro and discussion all turn on preserved learning and memory, and the rats were never tested on anything. They foraged for popcorn in an open field. No discrimination measure, no novelty preference, no probe, nothing. So "learning" ends up being defined as "decoder accuracy went up across days," which makes the central claim circular. Either add a behavioural readout in these animals or take the learning language out of the title and abstract and say what was actually measured, which is experience-dependent change in neural coding.

      (2) Four animals per group is fine for this kind of work, but the statistics don't respect it. Degrees of freedom in the thousands and tens of thousands (t(8616), t(36293), F(1,2674)) treat cells and cell pairs as independent, which they aren't. The mixed models in Figure 1 are the right approach, and I couldn't see why they were dropped everywhere else. The theta result is the clearest casualty: a null across 36,293 cell pairs from four rats isn't evidence that theta coordination is preserved; it's an untested question with n=4.

      (3) Also, the reported df do not always match the stated n. The Methods say 4 per genotype, but several animal-level tests give t(4), which implies 3. Figure 4C gives t(32), and Figure 5C gives t(31) for what look like per-day measures. I couldn't work out what the sampling unit was in each case.

      (3) Some statistics can't be right. I noticed three, without looking hard. For example. Figure 1C, rate overlap: t(471) = 2.3 with p = 2.0e-7. Fig 1G: t(163) = 5.3 with p = 0.48. Fig 3F: F(1,483) = 36.5 with partial eta squared = 0.7, when the almost identical test in the previous sentence gives 0.07. These are likely all typos, but there are enough that the whole set needs going through.

      (4) The dissociation isn't tested with matched methods. Explicit coding gets rate map correlations, PVC, rate overlap, and field size. Implicit coding gets an SVM across six days. The claim is that one improves with experience and the other doesn't, but they're never put through the same analysis. The authors should run the identical decoder on rate vectors and on tau vectors, same cross-validation, day by day, and show the slopes diverging. That would be a real dissociation. Figure 1G does run a rate decoder but only pooled, not across days. As it stands, the difference in learning slopes could partly be the two analyses having different sensitivity.

      (5) The PIR residual may not be as clean as it looks. PIR is observed rate minus rate predicted by the cell's own spatial map. If the spatial map is a worse model of firing in App rats, which is the paper's own claim, then less gets subtracted and more is left in the residual. So a group difference in residual tau structure could be partly downstream of the group difference in place coding quality rather than something independent. This is worth some kind of check, e.g. matching cells on spatial information, or at least reporting how much variance the spatial model explains in each group.

      (6) Reactivation consistency. This carries a lot of the interpretation, and it's the thinnest evidence in the paper. r = 0.2, p = 0.04 in App against r = -0.2, p = 0.06 in WT. That's a difference in significance, not a tested difference between groups, and with both p-values sitting on either side of 0.05, I wouldn't build a mechanism on it. Needs a group x day interaction. Separately, mean pairwise correlation across SWR population vectors depends on how many cells are active, how many events there are (panels show 189 to 491) and how sparse the firing is. SWR rate, duration, participating cells and firing rates would let the reader judge whether "more consistent reactivation" means what's claimed.

      (7) The hyperexcitability to excessive replay to Hebbian consolidation story on pp 21 to 22 runs about a page on the back of one marginal correlation. This section should be cut down and flagged as speculation.

      (8) Missing controls. Things I expected and didn't find: histology confirming tetrode placement in CA1, any pathology verification in this cohort rather than a citation to Pang 2022, and A1-A2 spatial correlation shown next to B2-A2. That last one matters. If the App representation of A just drifts across the day, that's a different story from a specific failure to reinstate A, and the confused cell analysis as built can't tell them apart. Also, with the threshold set at the 95th percentile of the A1B1 baseline, the WT value of 4.5% is basically the chance floor by construction, so the number that carries information is the App one.

      (9) The issue of males only should be mentioned.

    3. Reviewer #2 (Public review):

      This study by Wang et al. longitudinally tracks hippocampal CA1 population dynamics in Alzheimer model rats during repeated exposure to different environments. The authors dissociate two levels of neural coding. "Explicit" spatial coding, assessed by rate maps and population vector correlations, is severely impaired in AD rats and does not improve with experience. In contrast, "implicit" temporal cofiring structure, quantified by pairwise Kendall's tau and population cofiring correlations, becomes progressively more context-specific over days, mirroring behavioral learning. This preserved temporal coding is not merely a byproduct of spatial overlap, as position-independent rate analysis confirms that learning-dependent discrimination arises from intrinsic temporal dynamics rather than from residual spatial tuning. Moreover, offline sharp-wave ripple reactivation shows increasing consistency across days specifically in AD rats. These findings reveal a dissociation in the AD hippocampus and propose that temporally structured population dynamics, rather than spatially selective firing, may support residual cognitive function.

      Overall, the findings are novel and thought-provoking; the analyses are comprehensive and well-controlled, and the proposed re-registration framework offers a compelling new lens for understanding cognitive resilience in Alzheimer's disease, with clear potential to guide future neuromodulation interventions.

      I have only a few minor comments.

      (1) Justification of terminology ("explicit" vs. "implicit").

      The manuscript should clearly define the two terms early in the Introduction, ideally in a dedicated paragraph. In my opinion, the current use of "explicit" and "implicit" is not intuitive and may even be misleading.

      (2) Does the dissociation reflect differential impairment between rest and running states in AD?

      The authors could elaborate on this.

      (3) Effect size in Figure 1D - the difference does not appear very large.

      The statistical significance in Figure 1D is accompanied by relatively modest effect sizes. The authors may tone down the claim about "failure to distinguish" in the manuscript ("failed to distinguish different contexts during the second transition" on Page 8).

      (4) Definition of "confused cells" - why not a fixed correlation threshold?

      The authors define "confused cells" as those whose B2‑A2 spatial correlation exceeds the 95th percentile of the A1‑B1 baseline distribution within the same animal. This seems to be unnecessary. A fixed correlation threshold (e.g., r > 0.5 or r > 0.6) would have a direct biological interpretation: it would identify cells that maintain similar firing fields across two putatively distinct environments, i.e., cells that truly fail to remap. In contrast, a relative percentile threshold defines "confusion" not against an absolute standard of similarity, but against the degree of remapping observed during the first A‑B transition.

      (5) SVM decoding on spike‑train temporal structure - missing justification.

      The authors state that "the temporal structure of spike trains contained distinct contextual information" (P9) and then directly apply an SVM decoder to the data, but the logical bridge is missing. Why is a decoder necessary here, and is it useful for such a task?

      (6) Figure 6 - learning occurs during reactivation but does not transfer to the online state (theta state), even after days of training.

      What does this mean in terms of different phases of memory? Is the consolidation phase affected more? The authors may provide more discussion along these lines.

    4. Reviewer #3 (Public review):

      Summary:

      Determining the ways in which Alzheimer's disease (AD) impacts the neural instantiations of memory is a fundamental aim of neuroscience and likely to be critical to understanding and treating disease progression. Previous work has highlighted how the hippocampal spatial code - typically context-specific - fails to discriminate between different environments in rodent models of AD. Here, Wang et al. leverage new analyses in a rat model of AD to test whether this deficit in 'remapping' reflects an impairment in intrinsic temporal coding. The authors find that the intrinsic temporal code of AD rats does come to effectively discriminate between environments (albeit more weakly than their wild-type counterparts), despite persistent impairments in remapping.

      Strengths:

      One of the major strengths of this work is its focus on distinguishing between two different types of neural impairments in AD. While previous work has highlighted impairments in remapping, these impairments could be due to: (a) impairments in intrinsic neural computations, or (b) impairments in the way these intrinsic neural computations are anchored to the world. The author's evidence supports the latter, with important implications for the nature of these impairments.

      Weaknesses:

      A handful of weaknesses could be addressed to strengthen this work. Firstly, a number of key measures of place code quality and behavioral quality between groups are omitted. Given that the author's interpretations rely on comparisons between groups and often use decoding analyses for central conclusions, indicating whether the recordings are comparable between groups in terms of behaviors and cell counts would be helpful for the reader. Even better, matching cell counts between groups for decoding analyses could ensure that outcomes are not driven by this potential confound.

      Another weakness that is worth addressing involves the pivotal comparison between time-averaged remapping (Figure 4) and intrinsic temporal code discrimination (Figure 3). For temporal code discrimination, the analysis relies on comparisons between pairs of epochs (e.g. from A1B1 to A2B2), while for remapping, the comparison averages together epochs (A1+A2 and B1+B2). As a result, these comparisons are characterizing two slightly different things. Given that this is a key comparison for this paper (cofiring evolves to discriminate contexts while the spatial code does not), I think it would be important to demonstrate that remapping between pairs of epochs also stagnates to more closely mirror the cofiring analysis.

      A final weakness is the sharp wave ripple (SWR) ensemble analysis. Here, the authors extract population vectors during SWR (SWR-PVs) during rest in both WT and AD rats. Next, they compare the extent to which SWR-PVs on AVERAGE resemble other SWR-PVs for that epoch. The authors report that this measure increases with experience in AD rats but not wild-type (WT) rats. While the authors interpret this to mean that SWR-PVs come to be more reliable with experience in AD rats, it could alternatively mean that SWR-PVs become less diverse or 'muddier' with experience, while SWR-PVs in WT rats continue to represent diverse trajectories. To make this analysis compelling, a better measure might be something that quantifies the diversity or content of SWR-PVs and something that quantifies similarity between each SWR-PV and the most similar other SWR-PVs.

    1. eLife Assessment

      This study provides valuable findings on computational measures of learning in human subjects with current depression, individuals remitted from depression, individuals at familial risk, and healthy controls using reinforcement learning and risky decision-making tasks. The evidence is solid but could benefit from clear interpretation of model parameters and psychiatric symptoms. The paper will be of interest to scientists interested in learning, reward processing, value-based decision-making, and psychopathology.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates state-, trait-, and recovery-related computational phenotypes in Major Depressive Disorder (MDD) using an unmedicated, four-group cross-sectional design comprising currently depressed participants, remitted individuals, first-degree relatives at familial risk, and healthy controls. Utilizing a volatile four-armed bandit task and an explicit risk gambling task alongside hierarchical Bayesian modeling, the authors report that remitted participants uniquely display a lower punishment learning rate relative to all other groups. The authors conclude that recovery from MDD does not represent a simple normalization to healthy baseline levels, but rather involves a protective computational recalibration that dampens reactivity to negative outcomes to sustain remission.

      Strengths:

      (1) Highly Valuable Clinical Sample: Evaluates a rare, unmedicated sample across four distinct clinical stages (MDD, REM, REL, CTR), providing an exceptionally controlled framework to disentangle state, trait, and recovery markers.

      (2) Combines dynamic reinforcement learning under environmental volatility with prospect-theoretic decision-making under explicit risk to capture multiple dimensions of value-based choice.

      (3) Challenges the conventional assumption of clinical "normalization," offering a compelling hypothesis that psychiatric recovery may depend on active, compensatory recalibrations of cognitive parameters.

      (4) Utilizes hierarchical Bayesian parameter estimation and leverages Bayesian Model Averaging (BMA) to mitigate single-model selection bias.

      Weaknesses:

      (1) The Discussion characterizes MDD and familial-risk groups as exhibiting "noisier" choice behavior in the gambling task, which directly contradicts the reported higher inverse temperature values that mathematically denote more deterministic choices.

      (2) Framing a reduced punishment learning rate as a "recovery mechanism" overinterprets single-timepoint data, which cannot differentiate an acquired post-episode adaptation from a pre-existing resilience trait.

      (3) The theoretical claim that lower punishment learning is protective in remission directly conflicts with the authors' dimensional findings, where lower punishment learning correlates with worse subclinical apathy and anhedonia in non-depressed participants.

      (4) Model-agnostic choice repetition yielded no significant group effects, contrasting sharply with the robust group differences in model-derived parameters and necessitating posterior predictive checks.

      (5) Fails to provide parameter recovery analyses to demonstrate that punishment learning rates can be reliably disentangled from lapse rates and outcome sensitivities across a 200-trial task structure.

      (6) Selects a lower-ranked model under LOOIC without sufficient quantitative justification, and lacks sensitivity analyses to confirm that group-specific hierarchical priors did not skew estimates given unequal group sizes.

      (7) Relies on several marginal p-values bordering across multiple parameters and symptom correlations without establishing a clear family-wise error or FDR correction strategy.

    3. Reviewer #2 (Public review):

      This manuscript reports on reinforcement learning in participants with current depression, remitted depression (without current depression), in people without depression but with first-degree relatives with depression, as well as healthy controls. Participants completed two common tasks measuring reinforcement learning and risk aversion, and their behavior was fit to computational models assessing processes on these tasks.

      Participants with remitted depression showed a lower punishment learning rate and more value-concordant decisions on the reinforcement learning task. Relatives of depressed participants, as well as people who are currently depressed, had higher inverse temperature, indicating more value-driven choices. Within the non-depressed participants, the punishment learning rate was negatively associated with symptoms of anhedonia and apathy.

      I have reviewed this paper at a previous journal. This revised version is responsive to most of my concerns, particularly in terms of placing the manuscript more in the context of other related literature and providing more details on methods. There are some remaining concerns about sample size and the appropriateness of some of the methods (e.g., interpreting participant-level parameter estimates from hierarchical models estimated using BMA), but the latter has been adequately addressed with sensitivity analyses.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript describes an interesting study (preceded by a pilot study) that combined computational modeling with two tasks to try to tease apart state, trait, and vulnerability-related effects of depression. Specifically, the authors administered a bandit task and a gambling task to healthy controls, currently depressed adults, formerly depressed adults, and adults at familial risk of depression, and then used a suite of computational models to identify group differences in key model parameters. Key results included the detection of lower punishment learning rates (during the bandit task) in the remitted depressed group, which the authors hypothesize may be a compensatory mechanism to counteract the over-reaction to negative feedback that often characterizes depression (and indeed, punishment learning rates were elevated in currently depressed adults). Among non-depressed adults, lower punishment learning rates were associated with more apathy and anhedonia. The remitted depressed group also showed lower lapse rates in the bandit task. In the gambling task, the inverse temperature parameter was elevated in currently depressed adults and in adults at high familial risk for depression; exactly how to interpret this last result seems unclear. This is a revised manuscript, and from what I can see it appears that the authors were responsive to earlier comments.

      Strengths (and summary of weaknesses):

      I think the study has several noteworthy strengths. Testing four groups is a strength, as there is great interest in teasing apart risk factors of depression vs. "scars" of the illness, and the use of unmedicated individuals removes a common confound. The work is hypothesis-driven, the modeling is sophisticated, and the paper is well-written. Moreover, as the authors note, modeling allowed the authors to identify group differences that were not evident in raw behavioral analyses.

      However, I think the manuscript could be further improved because some aspects of the methodology are a bit confusing and/or do not seem optimal. I list these issues in the comments below, but to briefly summarize: (a) I did not understand why the authors estimated separate reward and punishment values for each bandit; (b) the rationale for Bayesian Model Averaging could be strengthened; (c) examining relationships between model parameters and symptom scores only in the control group is suboptimal given the goal to better understand depression; and (d) the group differences in inverse temperature seem like they may depend on potential outliers. I think addressing these concerns would improve the paper and ensure that the study has a strong impact on the field.

      Details regarding weaknesses/concerns:

      (1) I found a basic aspect of the 4-arm bandit modeling confusing - namely, the use of separate reward and punishment value estimates (see equations 1 and 2 in the supplement). The participants are choosing among the bandits (presumably) based on value estimates for each bandit. Clearly, delivery of rewards and punishments affect those estimates, but it is not clear to me why or how there are separate value terms for rewards and punishments; typically, rewards and punishments influence one overall value estimate per bandit. Can the authors clarify? (I see that similar models were used in references 29 and 30, but additional clarification for readers of the current manuscript would be helpful)

      (2) Bayesian Model Averaging (BMA) is new to me and may be new to many readers, and it would be helpful to provide a stronger rationale for the approach. The manuscript argues that model selection amounts to a "winner-takes-all" approach that may introduce bias; maybe, but typically the goal is to figure out which mechanism(s) best explain behavior, and so winner-takes-all is often appropriate. My limited understanding is that BMA is often used when prediction-rather than mechanistic understanding-is the goal. Why is BMA the right choice here?

      (3) The authors performed a confirmatory factor analysis on questionnaire data from healthy controls out of concern that extreme scores in the other groups might bias the factor structure; consequently, they can only relate model parameters to their latent factors in the healthy controls. This does not seem optimal. While the negative relationships between punishment learning rate and both anhedonia and apathy in controls are interesting, the controls are not struggling with anhedonia or apathy. It would be valuable to know if similar relationships obtain in the other groups, particularly because this would speak to the study's goal of distinguishing between risk for depression and state/trait aspects of depression.

      (4) Figure 5 gives the impression that group differences in inverse temperature may depend on potential outliers in the relative and MDD groups. Is that correct?

      (5) The paper notes a group difference in IQ as estimated from the WTAR, and looking at Table 1 it appears that the group difference is driven by a lower WTAR score in the healthy volunteers (HV) from the pilot study. First things first: the score in the HV group is too low to be a standardized WTAR score. Like IQ scores, standardized WTAR scores typically have a mean around 100; scores for the four Study 2 groups look alright, but the Study 1 mean score of 40 is much too low. Can the authors clarify?

    1. eLife Assessment

      This important study provides evidence that Clock and fat body proteasome subunits contribute to dietary restriction-mediated lifespan extension in Drosophila. The evidence supporting these conclusions is solid, yet whether dietary restriction-induced daily rhythmicity of proteasomal genes is CLK-dependent remains incompletely demonstrated, and alternative explanations involving feeding differences and non-circadian functions of CLK cannot be, with the current dataset, excluded. The work will definitely be of interest to researchers studying biological rhythms, nutrition, and aging.

    2. Reviewer #1 (Public review):

      The studies by Hwangbo et al. diligently attempt to account for many of the typically neglected dietary and non-dietary factors.

      Strengths:

      • Work addresses many potential artifacts of dietary (e.g., dehydration stress, macronutrient ratios, and protein source) and non-dietary (e.g., leaky expression of S106-GAL4) manipulations-important factors that are too often overlooked.

      • Balanced and complementary behavioral, molecular, and bioinformatic experiments

      • Show necessity of proteostatic subunits in the fat body for DR-mediated longevity. The findings in the current manuscript lay the ground for future studies that test sufficiency of fat body prosβ3 and rpn7, or necessity of other proteostatic genes in other tissues.

      Comments on revised version:

      The revised manuscript is substantially improved and addresses many of the prior concerns. I have only a few minor recommendations and remaining issues:

      Clarify the interpretation of the Con‑Ex feeding data. The authors describe the ~70% higher intake on 1SY in Clk^Jrk as modest, and note a ~40% higher mean intake on 5SY that is not statistically significant. However, lack of significance can reflect limited power, and these differences are potentially biologically meaningful, given that relatively small changes in nutrient ingestion can substantially affect lifespan. If the average effects are real, the Clk^Jrk flies would be ingesting an effective diet closer to ~1.7SY and ~7SY relative to controls. A shift of the diet-lifespan response curve in Clk^Jrk therefore cannot be fully excluded, particularly given the absence of intermediate diets between 1SY and 5SY and the observation that Clk^Jrk is sometimes shorter‑ and sometimes longer‑lived than controls across different trials and diets.

      Although the core finding is strengthened by using several diet formulations, most additional experiments continue to rely on whole‑food dilution, even as the field is moving toward more defined DR regimens (e.g., yeast‑only or yeast‑extract-based protocols). There remains considerable variability and, in some cases, a lack of clear DR‑mediated lifespan extension in control cohorts (for example, in some GeneSwitch experiments using whole‑food dilution). It would be helpful if the authors briefly commented on this variability and justified their continued use of whole‑food dilution in these experiments.

      Please add a clear Methods description of the feeding assay (Con‑Ex), including fly age, assay duration, dye or tracer conditions, sample processing, quantification, and statistical analysis.

      Please ensure that the survival data shown in Figure 5 and associated supplements are explicitly linked to Cox proportional hazards analyses in the text or figure legends, with clear indication of the models used (e.g., gene, diet, and gene×diet interaction terms). The Methods state that diet is used as a continuous variable; given the non‑linear (U‑shaped) lifespan-diet reaction norm (reduced survival at both 1SY and higher yeast), it would be important to clarify whether 1SY was excluded from these Cox models, or alternatively, to model diet categorically, restrict the continuous analysis to 5-20SY, or apply an appropriate non‑linear transformation (e.g., splines). As written, it is not clear how the Cox model accommodates the non‑linear diet response.

    3. Reviewer #2 (Public review):

      Summary:

      Dietary restriction (DR) increases lifespan, an effect that has been consistently observed in several organisms, but we still lack a clear mechanism to explain this phenomenon. In this work, Hwangbo et al. revisited the role of the circadian clock in DR-mediated lifespan effects. They found that the increase in lifespan produced by DR is missing on a clock mutant, a clock dependency that is also observed at the level of nutrient-dependent egg laying. By conducting RNA-seq with an impressive temporal resolution, they showed that DR triggers an increment in the number of cycling genes expressed in the fat body, the fly functional analog of the mammalian liver. Interestingly, from these genes, a group of them are de novo daily expressed genes, meaning that their expression was not rhythmic under the control diet but appear rhythmically expressed under DR. Among those, genes encoding proteasome subunits are enriched. The authors finally showed that adult-specific knockdown of these genes in the fat body prevents the increase in lifespan under DR, further supporting a role of the proteasome in this process. Overall, the conclusions are mostly supported by the evidence presented, and the authors' discussion nicely frame their results with other research in the field.

      Strengths:

      - Many studies have limited their observations of DR on lifespan to a few dietary conditions which makes the reach of some previous conclusions somewhat limited. The dilution strategy that the authors used in this work provides a strong indication that the effect of DR on lifespan relies on clock expression regardless of the conditions used. Furthermore, the inclusion of the egg-laying assay is a good addition to support this hypothesis.

      - Because the strength of the rhythmicity statistics relies heavily on the number of data points collected, the temporal resolution used for the RNA-seq experiments (every 2 hrs per 48hrs) is remarkable. This allows exquisite dissection of the phase of rhythmic genes in different conditions. The dataset produced in this work might be of use to other groups interested in weighting the role of other represented gene clusters in DR.

      Weaknesses:

      I see only minor flaws in this work, that if addressed, might strengthen the authors' conclusions, particularly:

      - The results of the lifespan assays are quite variable and in some instances contradictory (Fig. S8) across trials, possibly because there are other unaccounted variables we still do not understand. The fecundity assay, in contrast, seems to be a better readout (Fig. 2). Confirming at least the two genes picked for the study (Fig. 5) would be good support for the claim that the proteasome mediates the effects of DR.

      - According to the model, the acute effect of DR on gene expression is related to CLOCK protein function. However, I am not sure how this link was established. It is tempting to assume that CLOCK upstream is the reason for having an increase in rhythmic genes under DR, but the experiments did not test this. The tests conducted either assessed the role of clk or the effect of an impaired proteasome on DR-dependent extension of lifespan. Thus, it is difficult to assert the authors' claims on the link between CLK and the changes in cycling genes and to the proteasome upon DR.

      Comments on revised version:

      In this new version, Hwangbo and colleagues add new data to support a role of the clock in the effects of DR on lifespan. While adding this new data helps to alleviate some of the concerns previously raised, I think there are still some gaps. Below are my main concerns:

      ClkJrk transcriptomic data: Adding this data supports a role of the clock in daily rhythmicity of genes in the fat body, likely due to a circadian role. However, it does not show that de novo rhythmicity of proteasomal genes is clock-related since there is no ClkJrk DR dataset. Thus, there is still a possibility this is a pleiotropic effect. While redoing an entire RNA-seq dataset might not be feasible, a possible way to support the circadian claim would be to use proxy genes observed in the Ctrl vs DR conditions and compare them by qPCR in ClkJrk Ctrl vs DR.

      Feeding data: Are the flies reared in DR conditions, or do they just start the DR at the beginning of the experiment? If so, is it possible the flies will show a different feeding pattern after consecutive days of DR affecting overall (5-10 days) food consumption?

      Clk expression is important for non-circadian roles in the ovaries (Wang et al., Cell Mol Life Sci, 2025). Therefore, it is possible that the fecundity effect is at the low level of the ovary/egg development instead of integration and processing of DR. This might be a confounding effect when interpreting data in Fig 2 in ClkJrk as solely the effect of DR.

      Line 215: Considering the discussion above, I'd rather change "circadian-dependent change" to "daily"<br /> Considering that the ClkJrk RNA-seq transcriptomic was generated, presumably, at a different date/time than the original transcriptomic data from Control vs DR, comparative metrics (seq depth, mapping rate, etc.) between these are needed.

      How is the feeding analysis conducted? I believe the method information was not updated.

    4. Author response:

      The following is the authors’ response to the original reviews

      Correction: In the process of revising the preprint, we discovered that for the 15SY dataset that a single time point (ZT2) out of the 12 timepoint series was inadvertently combined with temporally adjacent time samples (ZT20, 22, 24). We corrected the accompanying GEO submission (Series GSE145509). With this update, we repeated the rhythm analysis with an updated RAIN algorithm as the original Boot-eJTK could not be run as it was outdated with dependent packages no longer maintained or supported, and some are no longer available through standard package managers. The analysis with the corrected ZT2 sample and did not find any significant changes in the major claims of the papers. One minor change is that we no longer observe significant DR-dependent increases in proteasome gene levels. Nonetheless, we still find DR-dependent cycling of proteasome genes and proteasome module network connectivity consistent with the DR-sensitivity of the proteasome pathway. Figure 4 has been updated to reflect this change. After correcting this issue, we revised the manuscript in response to the reviewer comments.

      eLife Assessment

      This study describes important findings on how a core component of the circadian clock impacts the effect of dietary restriction (DR) on longevity and fecundity in Drosophila, which lead the authors to postulate rhythmic control of proteostasis in the fat body as a critical aspect of DR effects. The evidence presented is still incomplete, not fully supporting the conclusions of the study, as alternative hypotheses/explanations have not yet been systematically explored. The work will nevertheless be of substantial interest to researchers working in circadian and cell biology, metabolism, and aging, with an interesting hypothesis to be explored further.

      We sincerely thank eLife for considering our manuscript and express our gratitude to all three reviewers for their time and constructive comments. While acknowledging that there are alternative hypotheses for some of our findings, which make our evidence incomplete, we appreciate that our manuscript was recognized as important and of substantial interest. We have revised the manuscript to acknowledge that the major findings under light-dark conditions could be attributable to being driven by light rather than the circadian clock. We add new data demonstrating that the far majority of cycling genes in LD in control flies are disrupted in Clk<sup>Jrk</sup> consistent with circadian regulation (see also below). Future investigations are needed to test the alternative hypotheses and explanations. Nevertheless, we believe that the manuscript still represents a meaningful advancement in understanding how molecular circadian clocks in peripheral tissues interact with diet to influence systemic lifespan and aging.

      Public Reviews:

      Reviewer #1 (Public Review):

      The studies by Hwangbo et al. diligently attempt to account for many of the typically neglected dietary and non-dietary factors.

      Strengths:

      - Work addresses many potential artifacts of dietary (e.g., dehydration stress, macronutrient ratios, and protein source) and non-dietary (e.g., leaky expression of S106-GAL4) manipulations-important factors that are too often overlooked.

      - Balanced and complementary behavioral, molecular, and bioinformatic experiments

      - Show necessity of proteostatic subunits in the fat body for DR-mediated longevity. The findings in the current manuscript lay the ground for future studies that test sufficiency of fat body prosβ3 and rpn7, or necessity of other proteostatic genes in other tissues.

      Weaknesses:

      - Could the lack of DR response in clock mutants across dietary concentrations be simply because the clock mutants are better at compensatory feeding adjustments to dietary dilutions? If this were the case, there are two major implications to the authors' conclusions:

      a) The Clk mutants are differently responding to dietary dilutions, not to dietary restriction, per se.

      b) Nutritional intake was unaffected by the dietary manipulations. If the changes in fat body proteostasis and lifespan were due to nourishment, it would be expected that the physiology and lifespan do not change.

      Accurate measurements of food consumption and the resulting protein intake could potentially clarify this critical question.

      We thank the reviewer for their positive feedback and also appreciate their raising the important issue of whether Clk^Jrk mutants may be more effective at compensatory feeding. Xu et al. (2008) reported that overall food consumption in Clk^Jrk flies was indistinguishable from control flies.

      We also directly assessed food intake in ~1 week old flies on three diets (1% SY, 5% SY, 15% SY) over two days (48 hours) using the Con-Ex method (Shell et al. 2018). We observed either no significant changes or relatively modest changes in food consumption between iso31 controls and Clk^Jrk flies that are limited compared to the large differences in caloric content between the diets. While we cannot rule out changes in feeding patterns throughout the lifespan, we believe that minimal differential compensatory feeding in the Clk^Jrk mutants are not sufficient to be the primary cause of the lifespan differences observed across diets. These data are added as new Figure 1-figure supplement 3.

      Reviewer #1 (Recommendations For The Authors):

      Hard to find information:

      - Type of yeast used. Should be addressed in the Methods section at least, instead of having to dig through several paragraphs into the Results section.

      Missing information:

      - Agar type and concentration

      - Mifepristone diets: pipetted on top or mixed into food?

      The relevant information has been updated in the Materials and Methods

      Wrong information:

      - Line #159-160: "Fig. 1 and Fig. S3" should be "Fig. 1 and Fig. S2" and then "(Fig. S3)" added to the end of the sentence.

      This information has been corrected in the revision

      Presentation:

      - Paragraphs are very long.

      - Interactions should be denoted by ×, not * or x.

      - Remove markers from mortality graphs. Having bulky markers reduces perceived differences between curves.

      These changes have been updated in the revised version.

      Reviewer #2 (Public Review):

      Dietary restriction (DR) increases lifespan, an effect that has been consistently observed in several organisms, but we still lack a clear mechanism to explain this phenomenon. In this work, Hwangbo et al. revisited the role of the circadian clock in DR-mediated lifespan effects. They found that the increase in lifespan produced by DR is missing on a clock mutant, a clock dependency that is also observed at the level of nutrient-dependent egg laying. By conducting RNA-seq with an impressive temporal resolution, they showed that DR triggers an increment in the number of cycling genes expressed in the fat body, the fly functional analog of the mammalian liver. Interestingly, from these genes, a group of them are de novo daily expressed genes, meaning that their expression was not rhythmic under the control diet but appear rhythmically expressed under DR. Among those, genes encoding proteasome subunits are enriched. The authors finally showed that adult-specific knockdown of these genes in the fat body prevents the increase in lifespan under DR, further supporting a role of the proteasome in this process. Overall, the conclusions are mostly supported by the evidence presented, and the authors' discussion nicely frame their results with other research in the field.

      Strengths:

      - Many studies have limited their observations of DR on lifespan to a few dietary conditions which makes the reach of some previous conclusions somewhat limited. The dilution strategy that the authors used in this work provides a strong indication that the effect of DR on lifespan relies on clock expression regardless of the conditions used. Furthermore, the inclusion of the egg-laying assay is a good addition to support this hypothesis.

      - Because the strength of the rhythmicity statistics relies heavily on the number of data points collected, the temporal resolution used for the RNA-seq experiments (every 2 hrs per 48hrs) is remarkable. This allows exquisite dissection of the phase of rhythmic genes in different conditions. The dataset produced in this work might be of use to other groups interested in weighting the role of other represented gene clusters in DR.

      We are grateful for the reviewer’s positive feedback regarding the robust experimental design in the manuscript.

      Weaknesses:

      I see only minor flaws in this work, that if addressed, might strengthen the authors' conclusions, particularly:

      - The results of the lifespan assays are quite variable and in some instances contradictory (Fig. S8) across trials, possibly because there are other unaccounted variables we still do not understand. The fecundity assay, in contrast, seems to be a better readout (Fig. 2). Confirming at least the two genes picked for the study (Fig. 5) would be good support for the claim that the proteasome mediates the effects of DR.

      We appreciate the reviewer’s comment of seeing “only minor flaws”. We reiterate that we focused on those results which were replicated across trials, providing confidence in the overall conclusions. Nonetheless, we agree that exploring the proteasome role on the DR effect on fecundity would be intriguing and may complement the lifespan data. We now acknowledge this point in our discussion.

      - According to the model, the acute effect of DR on gene expression is related to CLOCK protein function. However, I am not sure how this link was established. It is tempting to assume that CLOCK upstream is the reason for having an increase in rhythmic genes under DR, but the experiments did not test this. The tests conducted either assessed the role of clk or the effect of an impaired proteasome on DR-dependent extension of lifespan. Thus, it is difficult to assert the authors' claims on the link between CLK and the changes in cycling genes and to the proteasome upon DR.

      We have updated our text and model figure suggesting a direct CLK role in the Discussion.

      Reviewer #2 (Recommendations For The Authors):

      As mentioned before, the experiments, in particular the RNA-seq datasets are excellent. Additionally, the discussion provides a good overview of other relevant papers on DR, and the conclusions are mostly supported by the data. Here I provide a couple of suggestions that I believe might improve this work:…

      - Although the model is simple and understandable (Fig. 6), the inclusion of an overall summary or explanation in the figure legend would be appreciated, especially for readers that are not familiar with the terminology.

      - It might be the formatting while parsing the files but some of the in-text citations are between curly brackets (e.g., lines 80, 92, 93).

      - By definition, and unlike Canton-S or Oregon-R strains, w1118 flies are not wild-type but a genetic control. I believe that reference to this on the figures and text may need correction.

      We updated and/or our corrected each of these in the revised version. For wild-type, we more explicitly define this as wild-type for the relevant genetic locus.

      Reviewer #3 (Public Review):

      In this study, Hwangbo and co-workers investigate the extent to which the well-established life extending effects of DR rely on the molecular circadian clock and how the landscape of clock-controlled gene expression changes in the face of DR within the fat body of the fly, a tissue that performs the functions associate with both the liver and adipose tissue of mammals. The authors evidence that DR extends lifespan in a manner that depends on only one of the two major limbs of the fly's molecular circadian clock, namely the positive limb, that DR produces major changes in the identities of cycling clock output genes, and that genes related to the proteosome represent a major component of DR-induced transcript cycling. Though interesting, these conclusions are not strongly supported by the data and there are two major reasons for this. First, the authors rely on only one loss of function genotype each for the loss of positive and negative limb clock gene function. Second, though they wish to address the "circadian transcriptome" under normal and DR conditions, the authors conduct all their work under strong Light/Dark cycles, making it impossible to address circadian phenomena. These shortcomings are problematic in the extreme, as they leave open obvious alternative explanations for the results and fail to directly determine if the rhythmic expression, they observe are clock controlled or merely driven by the light/dark cycles, which themselves produce major effects on activity, feeding, etc., that may be responsible for differentially driving rhythmic transcripts under normal and DR conditions in the fat bodies.

      Major Weakness One: The use of only genotype each for the loss of positive (Clk^JRK) and negative (Per^01) limb of the circadian represents a major challenge for a central conclusion of the study. Phenotypes caused by the loss of a single clock gene may be due to the loss of circadian timekeeping, or they may represent a pleiotropic effect of the loss of function mutant being used. There are multiple precedents for pleiotropic (non-circadian) effects of clock gene mutants. It is, therefore, possible that the differences in the extent of DR mediated life extension between Clk^JRK and Per^01 may not represent a difference between breaking the positive and negative limbs of the clock but may simply reflect a pleiotropic effect of the dominant negative Clk^JRK. This possibility is acknowledged by the authors (lines 343-344). This could be addressed quite easily by extending the analysis to other loss of function mutants, for example, tim01 for the negative limb and cyc01 for the positive. Given the central focus here on the "circadian transcriptome," leaving open this alternative explanation for Clk's role in DR induced life extension represents a major weakness of the study. Furthermore, given the fact that Clk^JRK appears to be short lived on most of the media tested in the study, is it really surprising or informative that they would display lower life extension under DR?

      We confirmed that the large majority of LD oscillating genes in wild-type controls are disrupted in ClkJrk consistent with circadian clock regulation (Figure 3-figure supplement 1). As noted, we formally acknowledged that the circadian clock mutant alleles used here, and in fact any circadian clock alleles, can have pleiotropic, i.e., non-circadian, clock effects. This would only be partially mitigated by adding more (but also potentially pleiotropic) clock mutant alleles. Very challenging circadian resonance experiments (see Xu et al, 2019) are the gold standard for resolving circadian clock v. non-clock effects which are beyond the scope of this study which we now add to our discussion.

      We also note that foxo mutants are both short-lived and exhibit a robust lifespan extension to dietary restriction and thus the ClkJrk mutant is distinct in this regard. We have added this point to the Discussion.

      Major Weakness Two: The authors have not established that any of cycling transcripts they have detected in the fat body under normal and DR conditions are driven by the circadian clock. This is because: 1.) they have conducted their transcriptomic analysis on cells taken from flies entrained to light dark cycles, which can themselves drive daily changes in expression levels and 2.) they have not shown that the cycling measured on normal diet or DR conditions depends on a functional circadian clock. The "significant reorganization of the circadian transcriptome" is presented as a major conclusion of this study, but the authors have not addressed circadian control of transcription at all here, either by an examination of transcription under free-running conditions and/or in loss of function clock mutants.

      In addition, there is a logical gap in this study. The authors have shown that DR produces less life extension in Clk^JRK mutants than Per^01 or wild-type controls. They then show that DR produces changes in the rhythmic transcriptome when flies are place on DR. The central model presented in Fig. 6 shows/concludes that CLK drives increases in proteome-related transcript rhythms under DR. This conclusion could have been directly tested by asking if the changes in rhythmic gene expression induced by DR are gone the loss of function Clk mutants, or if the transcriptomic landscapes fail to differ between feeding conditions in these mutants.

      In conclusion, the study falls far short of directly testing the ideas it puts forth, greatly limiting its impact and interest.

      As noted above, we also examined the diurnal transcriptome in ClkJrk (at 4 hour resolution) and found that of the 290 genes that were detectably rhythmic in wild-type just 13 were rhythmic in ClkJrk consistent with the notion that oscillations depend on Clk (Figure 3-figure supplement 1). We now add this analysis to the manuscript. Nonetheless, we cannot exclude a role for light and thus have opted to use “diurnal” in place of “circadian” where appropriate for observed rhythms under LD conditions.

      Reviewer #3 (Recommendations For The Authors):

      Line 140 "showed an almost identical response" was a little hard to understand at first. Consider clarifying.

      This has been rephrased for clarity

      The authors claim that Clk mutants are "much longer lived" than wild-type controls on two of the relatively low calorie diets. Figure S3C certainly argues otherwise, and it's not clear how the data in 1C and S2C and warrant the use of "much" here.

      This wording has been rephrased and corrected in the revised version. The low-calorie diet shown in Figure 1-figure supplement 3 contains a higher sucrose concentration (5%) than those used in Figure 1 and Figure 1-figure supplement 2 (1%). This observation suggests that sucrose may play an independent role in the survival of ClkJrk mutants under malnutrition conditions.

      The authors should provide the rationale for the use of a dominant negative form of Clk for their experiments. Would the available amorphic allele be a better choice?

      As ClkJrk is the first described Clk allele and it is probably the most well characterized. As a dominant negative version which is still capable of dimerizing and binding DNA it is less susceptible to compensation by redundant bHLH transcription factors as has been observed for between mouse Clock and NPAS2 (Debruyne et al, 2006).

      It is not clear why the authors have chosen to examine transcriptomes so soon after transfer to DR. Why not wait longer. The authors provide context that changes are already taking place at the early time-point used, but would waiting a bit provide a more robust indication of how DR is changing the fat body?

      We noted in the manuscript that the effects of DR on survival are evident relatively soon (~2d) after a diet shift. We were interested in identifying those changes in daily transcription that would be occurring during that early time span and potentially be a cause rather than an effect of survival changes.

    1. eLife Assessment

      This valuable study provides a detailed three-dimensional characterization of primary cilia organization in the postnatal mouse growth plate, revealing reproducible spatial patterns in ciliation, ciliary length, and orientation. The evidence supporting these descriptive findings is solid, based on high-quality quantitative imaging and complementary genetic, mechanical, and transcriptomic approaches. However, evidence for the broader mechanistic conclusions is incomplete, particularly regarding whether ciliary orientation is uncoupled from basal-body and cell orientation and whether its stability under altered mechanical loading reflects a cell-intrinsic program. The study provides a helpful foundation for understanding primary cilia organization and mechanobiology in the growing skeleton, while the mechanisms underlying these observations remain to be established.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript provides fundamental insight into ciliary biology, specifically, how ciliary axoneme orientation is governed by microenvironmental bending or intrinsic cytoskeletal steering rather than strictly basal body docking coordinates.

      Strengths:

      There are three major strengths in this manuscript. First, combining high-resolution imaging with deep-tissue sectioning yields impressive lateral resolution, enabling robust separation of dual centrioles within the crowded chondrocyte extracellular matrix. Second, the authors established an automated pipeline that evaluates thousands of individual cells across multiple anatomical regions and differentiation zones, lending strong statistical weight to positional and volumetric measurements. Lastly, they demonstrate that ciliation peaks in the peripheral resting zone and ciliary length peaks in hypertrophic cells, providing a compelling cellular explanation for why Ift88 deletion impacts peripheral growth plate geometry and hypertrophic expansion.

      Weaknesses:

      There are three major weaknesses in this manuscript. First, the paper lacks explicit descriptions of data mentioned in the Abstract and Methods, including the RNA-seq differential expression, WGCNA modules, and immobilization/ciliary alignment data. Second, while the Methods section mentions correcting for "Z-blur" / point-spread function distortion in 3D spherical coordinate calculations, further detail is required on how orientations are disambiguated from optical sectioning depth artifacts. Lastly, the RNA-seq analysis demonstrates that ambulatory unloading alters hedgehog and primary cilia gene signatures, yet axoneme orientation itself remains static. The narrative requires a clearer mechanistic synthesis regarding how mechanical loading modulates ciliary signaling if physical alignment of the cilia is refractory to mechanical force in chondrocytes.

    3. Reviewer #2 (Public review):

      Summary:

      The aim of this work was to characterise in detail the cellular organisation of the growth plate in the growing mouse limb, the effect of mechanical loading due to physical activity, and the role of primary cilia in mediating this effect as putative mechanosensors. Primary cilia are generally thought to function as mechanosensors in a range of different organs and tissues, e.g. in the kidney. Exposure to mechanical loading is a normal part of post-natal limb growth, and the central hypothesis underlying this work was that primary cilia will play an important role in transducing the effects of mechanical loading into cellular responses during this process. To investigate this, the authors used a mouse model encoding fluorescent markers for primary cilia, and also allowing conditional knockout of IFT88, a protein that plays a key role in primary cilium biogenesis. The authors compared the effects of mechanical loading by comparing tissue from mice with normal or surgically immobilised limbs. To investigate the role of primary cilia, the authors performed the same experiments with mice treated with tamoxifen to induce IFT88 knockout.

      Strengths:

      A major strength of this work is that it studied primary cilia in a fully in vivo system. The authors employed cutting-edge imaging approaches and image analysis pipelines to rigorously study cellular organisation, centriole positioning, cilium length, and orientation across thousands of cells in limbs from 18 animals. The imaging results presented are of an exceptionally high-quality. The analysis of imaging data is extremely quantitative and robust and utilised appropriate statistical analyses, which were clearly stated throughout. The imaging data were complemented with transcriptomic data and analyses, which provided an orthogonal dimension for understanding cellular responses. An interesting and unexpected outcome from this is that the primary cilia are oriented in the same direction, and inclined at a roughly 45{degree sign} angle to the mediolateral and proximal-distal axes of the limb. The authors speculate as to how this might arise and the role it may play in sensing.

      Weaknesses:

      Overall, the work reported in this study was of a very high quality, and I could not find any significant shortcomings. However, I did feel that the paper was not very well written in many places, which made it difficult to read.

      Overall, I think that the authors did achieve their aims in this study. It will be interesting to unpick the cellular mechanisms that lead to alignment of the cilia, the role of this alignment in mechanosensing, and the molecular mechanisms by which the cilia sense mechanical strain. Thus, this work provides a fertile ground for future studies, which will have important consequences for the study of primary cilia in vivo.

    4. Reviewer #3 (Public review):

      Summary:

      This study aims to characterize the three-dimensional organization of primary cilia in the postnatal mouse growth plate and to determine how ciliary prevalence, length, position, and orientation vary across anatomical regions and stages of chondrocyte differentiation. The authors combine volumetric imaging and automated image analysis with conditional disruption of IFT88, limb immobilization, and bulk transcriptomics. The work provides a valuable anatomical dataset and identifies several interesting spatial patterns, including increased ciliation in the lateral resting zone, longer cilia in hypertrophic chondrocytes, and, most notably, a non-random orientation of ciliary projections despite a much broader distribution of centriole positions.

      The descriptive evidence is largely solid, and the imaging dataset should be useful to researchers studying primary cilia, skeletal development, and tissue mechanobiology. However, the evidence is incomplete for several of the central mechanistic conclusions. In particular, the current analyses do not establish that ciliary orientation is uncoupled from basal-body position, and they leave unresolved how ciliary orientation relates to cell orientation. The immobilization experiment also supports a narrower conclusion than the proposed cell-intrinsic orientation program, while the final model linking ciliary angle to multidirectional signal integration remains speculative. Overall, the authors succeed in identifying an interesting and reproducible anatomical pattern, but the mechanism underlying that pattern remains largely open.

      Strengths:

      A major strength is the scale and anatomical context of the imaging. Quantifying thousands of cilia and tens of thousands of centrioles in three dimensions while preserving information about growth-plate zone and position across the limb is technically demanding. This allows the authors to identify regional differences that would be lost in dissociated cells or bulk tissue measurements.

      The study also provides several potentially useful observations. Ciliation is higher in the lateral periphery, particularly in the resting zone, cilia are longer in hypertrophic chondrocytes, and ciliary projections show a reproducible non-random orientation. The latter is the most interesting result of the study and provides a useful foundation for asking how organelle orientation is established within a developing tissue.

      I also appreciated that the authors place these observations in several biological contexts rather than stopping at a descriptive atlas. IFT88 disruption alters ciliation and growth-plate cell organization, while immobilization changes growth-plate dimensions, cell morphology and orientation, and the transcriptome even though the measured ciliary properties remain comparatively stable. These perturbations give the anatomical observations useful biological context.

      Weaknesses

      The main concern is that the claim that ciliary orientation is uncoupled from basal-body position is not directly demonstrated. The manuscript shows that centriole positions are broadly distributed and, separately, that ciliary projections have a preferred tissue-level orientation. Different population-level distributions, however, do not establish independence within individual cells. For example, basal bodies could be broadly distributed while their position determines which of two opposite directions along a common tissue axis the cilium adopts. This possibility is particularly relevant because Centrin-2 labels both centrioles, whereas only one serves as the basal body. The current data therefore support a difference between the population distributions of position and orientation, but not yet the stronger claim that the two are uncoupled.

      A related gap is the relationship between cell orientation and ciliary orientation. The manuscript measures both, and cell orientation changes with growth-plate region, IFT88 deletion, and immobilization, while ciliary orientation appears relatively stable. Yet the two measurements are never directly related within the same cells. It is therefore unclear whether cilia adopt a reproducible angle relative to the major axis of their own cell, whether this relationship changes between the center and periphery, or whether changes in cell organization can occur independently of local ciliary alignment. This seems important for interpreting the tissue-level orientation pattern.

      The statistical treatment of orientation also deserves caution. These are circular or spherical measurements, yet much of the analysis relies on linear distributions and Kolmogorov-Smirnov tests. This is particularly problematic around the 0{degree sign}/360{degree sign} boundary, where values near 350{degree sign} and 30{degree sign} are geometrically close but appear separated in a linear representation. The manuscript also does not clearly distinguish between a preferred axis, where opposite directions are equivalent, and a preferred polarity, where one direction is favored. In addition, thousands of cilia are nested within a much smaller number of mice, so the apparent statistical power should not be driven primarily by pooled object counts. Given that regional differences are a central theme of the paper, it would also be useful to know more clearly whether the preferred axis or the strength of the orientation bias differs between the middle and lateral growth plate at the animal level. I also could not identify a formal comparison of the centriole or ciliary distributions with an appropriate uniform circular or spherical null. The reported tests mainly compare zones and regions, so the claims that centriole position is non-preferential and ciliary orientation is non-random are not yet statistically established in the form presented.

      The IFT88 conditional knockout is not carried through to the principal orientation question. The authors examine cell size, cell-axis organization, ciliation, and cilium length, but do not report whether the remaining cilia retain the preferred orientation or whether centriole positioning changes. Given the central role of this genetic perturbation in the manuscript, this leaves the genetic and orientation arms of the study somewhat disconnected.

      Finally, the immobilization experiment and the mechanistic interpretation should be separated more carefully. Only four animals were analyzed in the offloaded and contralateral conditions, and medial and lateral regions were averaged because of the small sample size. The experiment shows that the established ciliary orientation remains relatively stable over a two-week postnatal interval despite clear changes elsewhere in the tissue. It does not exclude a role for mechanical forces earlier in establishing the axis, nor does a nonsignificant difference with four animals demonstrate equivalence. Similarly, the proposed cell-intrinsic orientation program and the final "single-axis blindness" model are interesting hypotheses, but the study does not yet identify what establishes the axis or how the observed ciliary angle would alter sensitivity to a force or biochemical gradient. Those ideas are worth discussing, but they should remain clearly separated from the observations directly supported by the data.

    5. Author response:

      In response to the valuable reviewers’ comments and suggested changes, we are finalising changes to the manuscript, in order to resubmit a revised version, and a document with full author responses, that reflects all the review comments.

      These changes include adding points of clarity, improving accuracy on wording of key messages, adding additional interpretation of the data, including additional data and analyses that reflect open questions raised, and more discussion concerning these unanswered questions, which are subjects of future work.

      In response to specific points we were asked to provisionally address (actions in italics):

      Reviewer #1. We are pleased the reviewer sees the insight these data bring. We indeed think it likely that cilia axoneme orientation is governed by the immediate microenvironment and/or changes to the cytoskeleton and are actively looking to explore this.

      To address the 3 areas of concern:

      (1) Our re-writing of the abstract and the results concerning transcriptomic data seeks to overcome weaknesses in descriptions of these data.

      (2) Further detail is being added on how z-distortion is corrected for, so accuracy is the same in all axis and orientation measurements are robust.

      (3) We will add to the discussion to add our thoughts as to why ciliary and cilia signalling genes are regulated by immobilisation, but that immobilisation does not apparently affect cilia structure.

      Reviewer #2. Thank you for such broadly positive comments, we are pleased the scale and depth of the quantitative analyses comes across, but will make sure that revisions throughout improve the quality of the writing describing these. We are actively exploring means to test ideas for how cilia axoneme become orientated in this way and what the function is. These preliminary ideas will be reflected in the discussion.

      Reviewer #3. Thank you for such a detailed and thoughtful review. They will ensure the data are presented to their very best and we will address the concerns raised. This descriptive study had a hypothesis, and made discoveries which surprised us, we have tried, as you say, to put this in some context of the role of cilia and the role of mechanical forces in GP biology. Most notably we are considering that uncoupled is not the correct term here. To address areas of concern;

      (1) We agree ‘uncoupled’ is not the correct word here. We cannot find a pattern of correlation between centriole position and orientation. However, the two can’t be ‘uncoupled’ and there is no proven independence on a single-cell level (that cilia position and cilia orientation are not in any way mechanistically linked). We will make changes and add more details on what we have considered in this area.

      (2) Similarly, we have not correlated in each cell, cellular orientation and ciliary orientation, though we have made attempts and not yet found a relationship. However, again, this is not the same as one being absent. We might have expected ciliary orientation to change as cell orientation does (through zones or with pertubations) as we have seen in vitro but this remains to be fully explored and is one subject of follow-up work. We are considering column populations and per animal considerations of the data.

      (3) We did take a cautious approach to statistics and specialist advice, but advice was not to overcomplicate things when there are 2 main messages related to centriole position and cilia orientation. Firstly, centriolar position appears random or without preference thus distribution of position on cell, is homogenous. Second, ciliary orientation angle is not a homogenous distribution, as would be expected if random with this number of measurements. We do, and will add comments to this effect, have to mindful of large dataset, but do not think this means we are looking at false discoveries due to number of comparisons. We are considering how better to reflect this. To address these important points we will add a section to the methods and discussion and will endeavour to change results to this end as it is a central point of the manuscript.

      (4) We will add new data concerning IFT88cKO and centriole position and orientation.

      (5) We will add a critique of the immobilisation experiments to ensure the relatively diminished power is clear. We agree force may have set things up initially, we will ensure this is discussed and we will ensure our proposal for why cilia orientation is this way is framed as a hypothesis. We have preliminary data, but this is the subject of an entire new project thus not yet supported by robust experimental evidence so is speculative at this stage and we will ensure this is clear.

    1. eLife Assessment

      This Review Article addresses a significant and timely topic that concerns the role of cryo-EM in transforming RNA structural biology from static, individual conformations toward the reconstruction of dynamic conformational ensembles and energy landscapes. The authors describe eight case studies spanning different RNA classes and argue that conformational heterogeneity is a source of mechanistic information, rather than a limitation.

    2. Reviewer #1 (Public review):

      Summary:

      The review addresses an important and timely topic that concerns the role of cryo-EM in transforming RNA structural biology from a "static" discipline to one increasingly concerned with conformational ensembles and molecular dynamics. The scope is well within the eLife standards, and the style and general architecture do fit eLife.

      Strengths:

      The review is extremely well written, well-conceived and clear. The main strengths are in the breadth of coverage, the clear theme, and the inclusion of practical examples that explain in detail the construct design, sample preparation, vitrification, and data analysis. The manuscript will certainly be impactful and valuable, especially for readers who are not specialists in cryo-EM, as it provides an accessible overview of recent advances across a wide range of RNA systems.

      Weaknesses:

      My only reservation is that, currently, the review reads too much like a list of examples. The authors should make an effort, and I am sure they are well up to it, to try to synthesise the message, provide more critical insights and amalgamate the text better, to really reach a wider audience.

      If revised along the lines detailed below, I am sure that the review will become an authoritative and influential resource for the RNA structural biology community.

    3. Reviewer #2 (Public review):

      Summary:

      In this review, the authors set out to synthesize how cryo-EM is reshaping RNA structural biology, moving the field from the determination of static, individual conformations toward the reconstruction of dynamic conformational ensembles and energy landscapes. Through eight case studies spanning ribozymes, riboswitches, viral RNAs, and synthetic RNA assemblies, they aim to show how cryo-EM has revealed mechanisms of RNA motion, including folding, ligand-dependent switching, and cooperative assembly, and to provide a practical account of the experimental and computational challenges (construct design, sample preparation, vitrification, data analysis) that are specific to dynamic RNA targets.

      Strengths:

      The manuscript succeeds in bringing together a wide and genuinely current range of case studies illustrating the field's shift toward dynamics-focused cryo-EM, several published within the last one to two years. The dedicated "Challenges" section, which walks through construct engineering, buffer and vitrification optimization, grid screening, and heterogeneity-resolving computational approaches, is a particularly useful and practical contribution; it goes beyond simply cataloguing structures and gives readers new to the area a genuine methodological roadmap. The figures are detailed and well matched to the quantitative claims made in the text (helix rotations, distances, RMSDs), which strengthens the paper's value as a reference resource.

      Weaknesses:

      The coverage of two areas in particular, the SL5 viral RNA element and RNA quaternary/multimeric assemblies, would benefit from incorporating additional recent primary literature that is directly relevant but currently omitted. This does not undermine the manuscript's core narrative, but it means the review is presently less complete than it could be as a field synthesis, particularly for readers using it to identify the full body of recent work on these specific RNA classes.

      More concerning is that one specific structural claim, describing conformer heterogeneity in the cobalamin riboswitch (Case study 3, holo dimer 4), appears to invert the finding reported in its own source paper (Ding, Deme et al., 2023). As written, the manuscript states that P2 and the distal half of P6 are structured in dimer 4, whereas the source paper reports that these are the regions that could not be modeled. This is worth flagging prominently because it is a factual claim about a specific structure, not an interpretive point, and readers relying on this review as a secondary source could come away with an inverted understanding of that structure's flexibility.

      The manuscript also contains several minor internal inconsistencies. None of these individually threatens the paper's core arguments, but together they suggest the manuscript would benefit from a careful proofreading and reference-list audit pass.

      Overall, the authors largely achieve their stated aim. The case studies convincingly illustrate that cryo-EM can now resolve discrete and continuous conformational states of RNA at near-atomic resolution, and the Challenges section substantiates the claim that construct design, sample preparation, and computational innovations have been jointly responsible for this progress. The gaps in coverage of SL5 and multimeric RNA literature, and the inverted claim in Case study 3, are the main respects in which the manuscript falls short of being a fully comprehensive and accurate synthesis at this stage, but these are correctable issues rather than flaws in the overall argument or framework.

      This review is likely to be a useful entry point and practical reference for researchers moving into RNA cryo-EM, particularly given the level of methodological detail in the Challenges section. Its impact would be strengthened by two additions. First, a short discussion of recent cryo-EM advances in tRNA would round out the manuscript's coverage of classical RNA structural targets alongside the ribozyme, riboswitch, and viral RNA case studies already included. Second, the raiA non-coding RNA, currently mentioned only briefly, has in the last two years become a genuine model system for cryo-EM-based ncRNA structure determination, progressing from a single novel-fold discovery to a comparative structural framework spanning multiple raiA subtypes and candidate protein partners. Expanding this into a full case study would let the manuscript showcase, in a single worked example, exactly the kind of field-level progression (from novel fold to scaffold-based strategy and further to comparative structural framework) that the review's own framing describes as the trajectory of the field as a whole.

    4. Reviewer #3 (Public review):

      This review describes how cryo-EM is moving RNA structural biology from the determination of static structures toward the characterization of conformational ensembles. Eight case studies spanning self-splicing introns, riboswitches, viral RNA elements and synthetic assemblies show how cryo-EM has captured folding intermediates, hinge-mediated domain motions and ligand-dependent switching, and the authors pair these examples with practical guidance on construct design, sample preparation, vitrification and heterogeneity analysis. The argument that conformational heterogeneity is a source of mechanistic information, rather than a limitation, is well supported, and Table 1 will be a useful reference for laboratories entering the field. The manuscript is timely and of broad interest, and I recommend publication after minor revision.

      (1) Scaffold-based structure determination (page 14, lines 5-9). The section on chimeric RNAs presents the scaffold strategies as a single group and cites Haack et al. (2025) and Langeberg & Kieft (2023) in one parenthetical, so individual results are not attributed to their sources. The distinction between these studies is substantive. Earlier scaffolds based on the Tetrahymena group I intron (Langeberg & Kieft, 2023) or on RNA origami (Sampedro Vallina et al., Nucleic Acids Res. 51:4613-4624, 2023) resolved the appended RNAs at approximately 4.4-5 Å. Haack et al. (2025) reported the first RNA scaffold to yield a high-resolution structure of the target itself, resolving the ligand-binding pocket of the thiamine pyrophosphate (TPP) riboswitch at 2.5 Å and extending nucleotide-level cryo-EM analysis to small RNAs that had previously been intractable. I ask the authors to attribute each result to its source and to state explicitly that Haack et al. (2025) achieved the first high-resolution structure of a scaffolded target RNA. Because the same study captured the ligand-free TPP riboswitch in an open, Y-shaped conformation, a direct example of the ligand-dependent switching that is central to this review, the TPP riboswitch should also appear in the riboswitch section (page 8, lines 15-36), together with the fluoride riboswitch of Langeberg & Kieft (2023).

      (2) Page 6, lines 8-14. Self-splicing introns are described collectively as evolutionary ancestors of the spliceosome that reside in pre-mRNA transcripts. The proposed ancestral relationship to spliceosomal introns and snRNAs applies to group II introns. Group I introns, including the Tetrahymena intron, which interrupts a pre-rRNA, initiate splicing with an exogenous guanosine and are not considered spliceosomal precursors.

    1. eLife Assessment

      This study provides an important insight into how the medial and lateral entorhinal cortices interact through distinct excitatory and inhibitory pathways. Using anatomical tracing, optogenetics, and electrophysiology, the authors show that glutamatergic medial entorhinal neurons provide broad excitatory input to lateral entorhinal, while long-range SST+ interneurons deliver selective inhibition to layer I. These findings reveal a novel layer- and cell-type-specific organization of medial to lateral entorhinal connectivity with implications for spatial and episodic memory. The work is convincing, but validation of injection specificity and viral spread is needed to fully confirm the anatomical interpretations; with these clarifications, this will be a significant contribution to understanding entorhinal-hippocampal circuit organization.

    2. Reviewer #1 (Public review):

      The study addresses the organisation of synaptic connections from medial to lateral entorhinal cortex. Classic anatomical work has suggested these connections exist but very little is known about their identity or functional impact. The manuscript argues that these projections are mediated by glutamatergic neurons, providing excitatory input from MEC to all layers of LEC, and by SST+ve interneurons sending inhibitory projections to L1 of LEC. This appears the most likely interpretation of the data. Potential concerns about confounds due to spread of virus/tracer from the injection site are addressed in the supplemental figures. My view is that the weight of evidence favours the authors' interpretation although the evidence isn't quite compelling.

      Knowing the configuration of projections from MEC to LEC is important for thinking about circuit mechanisms for spatial cognition and episodic memory. This study adds to an emerging view that MEC and LEC can interact directly, indicating that the cell-type level organisation of these interactions is asymmetric and identifying an intriguing long range inhibitory pathway.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Nilssen et al. presents a comprehensive study of the circuitry linking the medial and lateral entorhinal cortices (MEC and LEC). Using a combination of anatomical tracing, optogenetics, and in vitro electrophysiology, the authors convincingly demonstrate that the MEC sends both glutamatergic and long-range inhibitory SST+ GABAergic projections to the LEC, with distinct laminar and cell-type-specific targeting. Notably, they reveal that SST+ inhibitory projections selectively suppress the activity of layer IIa neurons, whereas excitatory inputs preferentially engage neurons in layers IIb and III, thereby differentially modulating hippocampal-projecting populations.

      Strengths:

      The experiments are carefully executed, the results are compelling, and the conclusions are well supported by the data. This work will be of broad interest to researchers studying memory circuits, cortical inhibition, and the organization of long-range connectivity.

      Weaknesses:

      Although the in vivo relevance of these connections remains to be determined, this is an important and timely contribution to our understanding of entorhinal-hippocampal interactions.

      Comments on revised version.

      The authors have addressed my comments satisfactorily, and I am satisfied with the changes made in the revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study addresses the organisation of synaptic connections from the medial to the lateral entorhinal cortex. Classic anatomical work has suggested these connections exist, but very little is known about their identity or functional impact. The manuscript argues that these projections are mediated by glutamatergic neurons, providing excitatory input from MEC to all layers of LEC, and by SST+ve interneurons sending inhibitory projections to L1 of LEC. This appears to be the most likely interpretation of the data, although in my opinion, more could be done to rule out the possible impact of the spread of the virus/tracer from the injection site.

      While this concern might seem overly picky, the importance of this level of detail is nicely shown by the authors' previous work clarifying connectivity from postrhinal to entorhinal cortices through careful analysis of similar types of data (Doan et al. 2019). If additional analyses/data can address the concern here, then I think this will be an important set of fundamental results that will influence thinking about circuit mechanisms for spatial cognition and episodic memory. In particular, it will nicely add to an emerging view that MEC and LEC can interact directly, showing that the organisation of these interactions is asymmetric and identifying a potentially interesting long-range inhibitory pathway.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Nilssen et al. presents a comprehensive study of the circuitry linking the medial and lateral entorhinal cortices (MEC and LEC). Using a combination of anatomical tracing, optogenetics, and in vitro electrophysiology, the authors convincingly demonstrate that the MEC sends both glutamatergic and long-range inhibitory SST+ GABAergic projections to the LEC, with distinct laminar and cell-type-specific targeting. Notably, they reveal that SST+ inhibitory projections selectively suppress the activity of layer IIa neurons, whereas excitatory inputs preferentially engage neurons in layers IIb and III, thereby differentially modulating hippocampal-projecting populations.

      Strengths:

      The experiments are carefully executed, the results are compelling, and the conclusions are well supported by the data. This work will be of broad interest to researchers studying memory circuits, cortical inhibition, and the organization of long-range connectivity.

      Weaknesses:

      Although the in vivo relevance of these connections remains to be determined, this is an important and timely contribution to our understanding of entorhinal-hippocampal interactions.

      The request for validation of injection specificity and viral spread, as detailed in the comments and suggestions of the two reviewers has been provided in the revised version. We added supplementary figures 1,2 and 6 as well as an extra insert into the old supplementary figure 5, now supplementary figure 9.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Interpretation of the retrograde labelling experiment in Figure 1A-C is challenging, as the spread of the tracer at the injection site is not shown. It's important to see the full dorsal-ventral extent of the injection site in order to establish that the labelling of neurons in MEC results from projections to LEC and not adjacent areas (including MEC).

      As mentioned in our initial reply, we are fully aware of the risks associated with an incomplete assessment of injection sites and viral spread, so we provide a new Supplementary Fig. 1 showing 6 dorsoventral levels of the case shown in Fig. 1A. The injection site in case of FG often shows a core of damaged tissue with a halo of substantial unspecific fluorescence. Outside of the injection side, one only sees retrogradely labeled somata (often recognizable by FG signal clustered in lysosomes) and dendritic elements. As indicated in the legend, we report some tracer leakage along the needle track in temporal and perirhinal cortex, areas that receive only sparse MEC projections, but there is no apparent spread of the injection into MEC.

      Numbers are also quite low. E.g., Figure 1A is N=1/2 for retrograde labelling experiments.

      The reviewer would be correct if the experiments were meant to analyze MEC projections to LEC in full anatomical detail. This was not our intention (several papers addressed this pathway in detail), we merely aimed to distinguish between glutamatergic and potential GABAergic contributions to this pathway and to establish optimal coordinates in slices to prepare for the electrophysiological recording experiments. Two animals suffice for this purpose and using more would be against the aim of reducing the use of experimental animals as much as possible.

      The rationale here for the use of AAV2-CAG-tdTomato as a retrograde tracer is unclear. My understanding is that this is more effective as an anterograde tracer. Some clarification and validation would be important.

      The reviewer is correct that AAV2 is generally considered an effective anterograde tracer, but tracing the connectivity of entorhinal cortex with AAVs has been proven to be notoriously difficult, in particular retrograde tracing of inputs to layer II. In a neighboring lab in the centre, headed by Edvard and May-Britt Moser, it was established that AAV2 types show very efficient retrograde transport and that is why we decided to use the virus. Also, in our hands the virus showed excellent retrograde transport that served our purpose

      (2) Interpretation of the anterograde experiments in Figure 1D-E would also benefit from showing evidence that the injection sites are restricted to MEC. It should be straightforward to make a supplemental figure showing labelling at all dorsoventral levels.

      More careful analysis of the axon labelling in the dentate gyrus could also help make a case for the selectivity of the injection site for eGFP. In this case, only the intermediate portion of the molecular layer of the DG should be labelled. In the image shown, the labelled band is quite wide, but it's hard to tell if this reflects the plane of section or is because it also includes labelling in the outer molecular layer (which would be indicative of LEC expression).

      We thank the reviewer for these two suggestions, and we have prepared a new Supplementary Fig. 2 in line with this.

      Numbers are also on the low side for these experiments.

      See our response above

      (3) For optogenetic experiments in Figure 2, the selectivity of targeting of AAV to MEC is assessed through the specificity of labelling in the DG. This is great, but it's important to show that this specificity is maintained at all dorsoventral levels.

      Higher resolution images of labelling in LEC could also be helpful. It's hard to tell from the images in 2A if labelling is axonal or is in the soma adjacent to the nuclear NeuN signal (which would indicate a lack of selectivity for MEC).

      We thank the reviewer for these two suggestions and provide a new Supplementary Fig. 6, showing both the details of AAV1 being present only in neuropil in MEC not in somata as well as the specific labeling in the middle molecular layer of DG in detail. Including all dorsoventral levels would not provide additional information in view of the very well-established topographical organization of the entorhinal to dentate projection, reaching approximately 20 -25 % of the full long axis of DG (Van Groen et al., 2003)

      In addition, we have again carefully screened all tissue from the electrophysiological experiments for possible leakage of virus from MEC to LEC. We decided to exclude recordings from one mouse, which had labelling in MEC that was close to the border with LEC. Neuron counts have therefore been adjusted (pages 7-9) and the example recording showing responses to TTX/4-AP exposure in Figure 2B has been exchanged.

      (4) The analysis of excitatory and inhibitory opto-responses in Figure 2 is nice. It may be helpful to report quantification of the rise and decay kinetics of the synaptic currents. They appear much slower for the inhibitory input, which may be functionally important.

      This would indeed be nice to add, but it would not significantly impact or change the main message of our study. Since the lab of the senior author (MPW) has been discontinued and the resources for conducting these analyses are not readily available anymore, we have found it difficult to comply with the reviewer’s request

      (5) More direct evidence for SST axons projecting from MEC to LEC would strengthen the conclusions made. E.g., in experiments where the SST neurons are labelled, is it possible to follow the axons? Do they project as expected from the MEC to the LEC?

      In our view the tracing data provide convincing evidence in support of a direct projection from MEC to LEC by SST neurons, as shown in horizontal brain sections where SST axons labelled in MEC of an SST<sup>Cre</sup> mouse projects within Layer I from the site of origin in MEC to Layer I of MEC (Supplementary Figure 3). Similar visualizations were not possible to obtain in our electrophysiological experiments where semicoronal slices were used. This cutting angle has been shown to be optimal to preserve most of the axon and the dendritic tree of LEC neurons (Tahvildari and Alonso, 2005; Canto and Witter 2012), but does not maintain the projection from MEC to LEC.

      Minor Points:

      (1) "These layers are heavily innervated by medial entorhinal axons (Figure 1F...". I don't see a 1F.

      This has been corrected; should have been Figure 1E.

      (2) Methods should report series resistance values for patch-clamp experiments (range and mean).

      Fully agree and this information has now been added on page 22 of the manuscript:

      Under Voltage clamp: ‘Recordings with series resistance ≤ 25 MΩ were accepted, with an average of 16.2 MΩ for voltage clamp recorded neurons (range, 4.9 – 24.9 MΩ).’

      Under Current clamp: ‘All recordings (series resistance: 18.9 MΩ, 5.0 – 66.7 MΩ; mean, range) were included for analysis.’

      Reviewer #2 (Recommendations for the authors):

      (1) Please specify in the figure or, alternatively, in the figure legend which virus was used in each group shown in Figures 2H and 2I. This is somewhat confusing, since Figure 2E illustrates a specific combination of viruses and mouse lines that only corresponds to part of Figure 2H. While this information is provided in the text, including it directly in the figure would help the reader.

      We thank the reviewer for this excellent suggestion, and we have implemented this in the new version of figure 2.

      (2) In Figure 4, regarding the inputs from PIR, cLEC, and PER to LEC, the inhibitory components recruited by each input were not examined as thoroughly as for the MEC inputs. In fact, some inhibitory interneurons were double-labeled in the GAD67 mice (Figure 4B), which could also influence the responses of LEC neurons, especially for PER inputs. Recordings in Figure 4C appear to have been obtained near the reversal potential for inhibition, which may have prevented the observation of inhibitory effects. The authors could discuss this point in the Results.

      The reviewer is correct and this issue is now addressed in the relevant section in the results (page 11):

      ‘It should be noted, however, that it is possible that inhibitory effects could have been masked in some recordings, due to the resting membrane potential in our recordings being close to the theoretical chloride equilibrium potential. This could be particularly relevant for the inputs from PER, an area where we found LEC-projecting GABAergic neurons (Fig. 4B) and which is known to provide long-distance inhibition to LEC (Pinto et al., 2006; Apergis- Schoute et al., 2007).’

      (2) A diagram summarizing the known connections among MEC, LEC, and the hippocampal formation, highlighting the relevant cell types, layers, and the new connections identified in this study, would be a valuable addition, perhaps as a supplementary figure.

      We appreciate the suggestion, though find a full summary of known connectivity a bit overdone. Instead, we included a new figure 6 that summarizes the main new findings of the paper in the context of LEC projections to the hippocampal formation.

      (3) Although the main focus is on MEC-LEC connectivity, the experiments examining interactions with other cortical areas and converging inputs would benefit from a discussion of how MEC-driven inhibition of LEC might influence those inputs and shape the resulting output to the hippocampus. Including a short paragraph addressing this in the Discussion section would strengthen the manuscript.

      Excellent suggestion although we did speculate briefly in the result section on the possible effect. We have added a short paragraph in the discussion (page 15), reiterating the part in the results (page 13, last paragraph of results). We also briefly discussed the potential functional relevance of the suppression of the pathway from layer IIa to DG-CA3/CA2 versus the facilitation of activity in the pathway from layers IIb/III to CA1 and subiculum (last section of the discussion).

      (4) Lastly, it would be interesting to know what the main source of activation is for the SST long-range LEC projecting neurons. Are these neurons recruited in a feedback manner by the activity of MEC excitatory cells? I realize this question is beyond the scope of the present study, but if the authors have any data or insights related to this point, including a brief discussion would be valuable.

      This is an interesting thought, and we have included a new supplementary figure (supplementary Figure 5) showing data from experiments mapping monosynaptic inputs to MEC SST neurons using rabies virus. Although it was not possible to target only MEC SST neurons that project to LEC, the data show which are the main extrinsic inputs to the population of MEC SST neurons, most likely including those that project to LEC.

    1. eLife Assessment

      This is an important study of critical period plasticity, focused on temperature manipulations, and how different parts of the Drosophila larval motor circuit adapt or maladapt. The work convincingly demonstrates that components of the motor network respond in distinct ways to heat shocks, and the combination of functional, structural, and electrophysiological approaches makes the study of significant interest. The work points to central interneurons as primary drivers of maladaptive changes, while motoneurons and neuromuscular junctions show compensatory or homeostatic adjustments. The authors have thoughtfully and substantially addressed the concerns raised in the original reviews through new temporal-window experiments, additional receptor and behavioral analyses, clearer discussion of the study's limitations, and an integrative model.

    2. Reviewer #1 (Public review):

      Summary:

      The authors examine the impact of heat stress during an embryonic CP in Drosophila, focusing on the larval locomotor network. They show that elevated temperature increases neuronal activity and, when applied during the CP, results in long-term instability of the network which manifests in prolonged seizure recovery times. At the neuromuscular junction, substantial structural changes occur, including terminal overgrowth and altered receptor composition, yet synaptic transmission remains preserved due to homeostatic regulation. Motoneurons display reduced excitability but receive increased synaptic input from premotor interneurons. These findings suggest that maladaptive instability originates within the central circuitry rather than at the neuromuscular junction, where changes seem to be homeostatically compensated. The study concludes that different network components exhibit distinct and hierarchical responses to CP perturbations, with premotor interneurons setting the tone for downstream adjustments in motoneurons.

      Strengths:

      The work takes advantage of the unique accessibility of the Drosophila system. A major strength of the study is the integration of structural, physiological, and behavioral analyses, which allows the authors to draw a comprehensive picture of how CP perturbations shape the locomotor network. The choice of an ecologically relevant stimulus (heat stress) is particularly convincing, as it links experimental manipulations more closely to natural environmental conditions. The experiments are carefully designed, and the results are robust and consistent with previous findings in the field, while also extending them in new directions. Importantly the work clarifies how temperature perturbations within distinct developmental time windows affect different properties of motor circuit formation.

      Weaknesses:

      A small limitation of the study is that it remains difficult to integrate maladaptive (seizure recovery) and adaptive/homeostatic phenotypes within a single mechanistic framework, leaving some space for interpretation.

      Comments on revised version.

      I think the authors did a great job at revising the manuscript and they addressed all my comments.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a thoughtful and well-executed study of critical period plasticity in the Drosophila larval motor circuit. The authors examined how transient heat, 32C, during embryonic stage, altered network properties, showing that premotor interneurons A27h increase excitatory drive onto motoneurons, which respond with a reduction in excitability. At the NMJ, synaptic terminals expand and GluRIIA distribution shifts, yet synaptic transmission remains largely unaffected. Despite these local compensations, the treated larvae display slower crawling and prolonged recovery from seizures, indicating that the network is functionally compromised.

      Strengths:

      (1) One of the major strengths of this study is the elegant dissection of a defined circuit, tracking changes from premotor interneurons through motoneurons to the NMJ. The multimodal approach provides a comprehensive view of how connected elements respond to CP perturbations.

      (2) An interesting finding is that NMJ morphology changes dramatically without corresponding deficits in synaptic transmission, challenging the common assumption that larger boutons necessarily indicate stronger synapses.

      (3) Another intriguing result is that even with two layers of homeostatic compensation, locomotor behavior is still impaired, highlighting the limits of compensation and underscoring the critical role of CP timing.

      (4) Beyond these scientific insights, the study benefits from a well-defined, tractable system and simple experimental manipulations, which together make the results highly interpretable and reproducible.

      Comments on revised version.

      The authors have carefully considered my comments and recommendations and have made substantial efforts to improve the clarity and validity of the study. Although additional electrophysiology experiments using shorter heat stress windows were not feasible, the authors performed additional analyses of postsynaptic GluRs and provided a clearer discussion of the study's limitations. Overall, this is a strong and well-written paper that establishes a valuable foundational framework for addressing interesting and important questions about adaptive responses in developing neural circuits.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This is an important study of critical period plasticity, focused on temperature manipulations, and how different parts of the Drosophila larval motor circuit adapt or maladapt. The work convincingly demonstrates that components of the motor network respond in distinct ways to the heat shock, and the combination of functional, structural, and electrophysiological approaches makes the study of significant interest. The work points to central interneurons as primary drivers of maladaptive changes, while motoneurons and neuromuscular junctions show compensatory or homeostatic adjustments. The study is methodologically rigorous, contributing important insights into critical period biology using a tractable invertebrate model.

      We thank the reviewers for their thoughtful critique and suggestions. We agree with these and, where possible, we have attempted to address these, improving this study. As outlined below, we have revised the manuscript substantively and included additional data and figures.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors examine the impact of heat stress during an embryonic CP in Drosophila, focusing on the larval locomotor network. They show that elevated temperature increases neuronal activity and, when applied during the CP, results in long-term instability of the network, which manifests in prolonged seizure recovery times. At the neuromuscular junction, substantial structural changes occur, including terminal overgrowth and altered receptor composition, yet synaptic transmission remains preserved due to homeostatic regulation. Motoneurons display reduced excitability but receive increased synaptic input from premotor interneurons. These findings suggest that maladaptive instability originates within the central circuitry rather than at the neuromuscular junction, where changes seem to be homeostatically compensated. The study concludes that different network components exhibit distinct and hierarchical responses to CP perturbations, with premotor interneurons setting the tone for downstream adjustments in motoneurons.

      Strengths:

      The work takes advantage of the unique accessibility of the Drosophila system. A major strength of the study is the integration of structural, physiological, and behavioral analyses, which allows the authors to draw a comprehensive picture of how CP perturbations shape the locomotor network. The choice of an ecologically relevant stimulus (heat stress) is particularly convincing, as it links experimental manipulations more closely to natural environmental conditions. The experiments are carefully designed, and the results are robust and consistent with previous findings in the field, while also extending them in new directions.

      Weaknesses:

      The study leaves some uncertainty regarding the experimental design and interpretation. The change from short to prolonged heat shock manipulations raises the possibility that the effects observed may not be confined to the critical period alone - this could be experimentally addressed or simply rephrased in the text.

      We agree that clarity about the experimental paradigm is important and have addressed this as suggested, within text, figures and figure legends: the duration of embryo exposure to 32˚C heat stress is now unambiguously stated and each figure has a graphical illustrations of the heat stress paradigm. For example, experiments represented in Figures 1, 3 (new data) and 8 (new data) used short, defined periods of a few hours of heat stress, aimed to identify specific windows of development that are sensitive to 32˚C heat stress. These also show that behavioural changes result from heat stress experienced during the specific 2-hour window that defines the critical period of the developing central locomotor circuitry, namely from 17-19 hours after egg laying - previously identified by Giachello & Baines (2015). Longer exposure to 32˚C heat stress during embryogenesis result in the same phenotypes when this 2-hour window is included, causing the same level of reduced larval crawling speed and lowered network stability, which manifests in increased seizure recovery times. This is as one might expect from a critical period of nervous system development.

      Where the neuromuscular junction is concerned, where we identified embryonic heat stress causing phenotypes that are evident at late larval stages, a more complex model has emerged. Following suggestions from both reviewers to explore shorter heat stress exposures during embryogenesis, additional experiments (see Figure 3) we identified what might be a critical period for the body wall muscles. This is an earlier window of development, within 13-16 hours after egg laying, which is sensitive to heat stress in terms of the levels of the GluRIIA glutamate receptor subunit that will be expressed in the late larva. This developmental period is characterised by muscles acquiring their electrical properties (Broadie & Bate, 1993), i.e. comparable to the central locomotor network transitioning through its critical period at the time that it becomes active. The neuromuscular junction is composed of both presynaptic motoneurons and postsynaptic muscles, and therefore this composite structure is subject to multiple, sequential critical periods. The characterisation of changes to neuromuscular junction synaptic physiology was carried out using heat stress throughout most of embryogenesis (Figure 4). We think this appropriate from the perspectives of having included all relevant critical periods (muscle and CNS) to explore how this composite structure responds to environmental heat stress; also based on our observations that for each critical period phenotypes are defined by the experience during the critical period and not exacerbated by prolonged heat stress either side.

      In addition, the maladaptive (seizure recovery) and adaptive/homeostatic phenotypes are not always clearly distinguished or highlighted, which makes it harder to appreciate how the different levels of the network plasticity fit together into a single mechanistic framework.

      Following the suggestion, we have tried to clarify the mechanistic framework in a new figure that aims to summarise the model in Figure 9.

      The question of whether phenotypes that result from an embryonic heat stress manipulation are adaptive or maladaptive is difficult to resolve. This is partly due to the nature of critical periods, since perturbations during these developmental windows can cause significant, long-lasting maladaptations that are challenging to reconcile from a perspective of adaptive plasticity. Secondly, in light of the nature of this animal, which has evolved a particularly rapid development and large brood sizes, any deviation from the evolved optimum developmental temperature of 25˚C could constitute a reduction in fitness. We interpret the phenotypes we see along those lines: network instability that results from critical-period perturbations is a manifestation of a sub-optimally tuned network, as is a reduction in larval crawling speed.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a thoughtful and well-executed study of critical period plasticity in the Drosophila larval motor circuit. The authors examined how transient heat, 32 {degree sign}C, during the embryonic stage, altered network properties, showing that premotor interneurons A27h increase excitatory drive onto motoneurons, which respond with a reduction in excitability. At the NMJ, synaptic terminals expand and GluRIIA distribution shifts, yet synaptic transmission remains largely unaffected. Despite these local compensations, the treated larvae display slower crawling and prolonged recovery from seizures, indicating that the network is functionally compromised.

      Strengths:

      (1) One of the major strengths of this study is the elegant dissection of a defined circuit, tracking changes from premotor interneurons through motoneurons to the NMJ. The multimodal approach provides a comprehensive view of how connected elements respond to CP perturbations.

      (2) An interesting finding is that NMJ morphology changes dramatically without corresponding deficits in synaptic transmission, challenging the common assumption that larger boutons necessarily indicate stronger synapses.

      (3) Another intriguing result is that even with two layers of homeostatic compensation, locomotor behavior is still impaired, highlighting the limits of compensation and underscoring the critical role of CP timing.

      (4) Beyond these scientific insights, the study benefits from a well-defined, tractable system and simple experimental manipulations, which together make the results highly interpretable and reproducible.

      Weaknesses:

      There are a few areas where the manuscript could be strengthened.

      (1) Although A27h premotor neurons are well characterized, the claim that they are the causal driver of downstream changes would be strengthened by additional experiments or a clearer discussion of the temporal hierarchy.

      We have tried to clarify the model of the temporal hierarchy (new Figure 9). This is a model and as such will hopefully help us collectively to think about this system and underlying processes, while also inviting this perspective to be challenged. The model we propose is compatible with our observations, namely that the premotor circuitry might change in response to a critical period heat stress (e.g. increasing their synaptic drive onto motoneurons), followed by homeostatic adjustment by the postsynaptic motoneurons (e.g. by reduction of their excitability), thus serving to maintain overall normal motoneuron firing patterns (see Figure 6).

      However, synaptic communication is commonly regulated in both antero- and retrograde directions. Therefore, while compatible with the observations we have made, bi-directional information flow could also be instructive during the CNS critical period.

      (2) While 32 {degree sign}C heat stress is presented as ecologically relevant, it produces maladaptive behavioral outcomes, raising questions about the ecological and mechanistic interpretation of the model. In particular, most experiments, with the exception of Figure 1, used prolonged (24h) heat treatments, which could introduce developmental effects beyond the CP itself. Comparing shorter and longer heat exposures would help clarify the specificity of the CP response.

      We agree. For a detailed response on this point, please response to Reviewer #1 above.

      (3) While there are schematics for experimental procedures, a circuit diagram tracing information flow and indicating where structural and functional changes occur would help readers better understand the findings.

      We have created Figure 9 as a working model.

      (4) Finally, the main paradox of the study, that robust homeostatic compensations occur yet behavior remains impaired, could be explored in more depth in the Discussion.

      We have tried to address this in the discussion.

      Reviewer #3 (Public review):

      Summary:

      During development, neural circuits undergo brief windows of heightened neuronal plasticity (e.g., critical periods) that are thought to set the lifelong functional properties of underlying circuits. These authors, in addition to others within the Drosophila community, previously characterized a critical period in late fly embryonic development, during which alterations to neuronal activity impact late-stage larval crawling behavior. In the current study, the authors use an ethologically-relevant activation paradigm (increased temperature) to boost motor activity during embryogenesis, followed by a series of electrophysiology and imaging-based experiments to explore how 3 distinct levels of the circuit remodel in response to increases in embryonic motor activity. Specifically, they find that each level of the circuit responds differently, with increased excitatory drive from excitatory pre-motor neurons, reduced excitability in motor neurons, and no physiological changes at the NMJ despite dramatic morphological differences. Together, these data suggest that early life experience in the motor neuron drives compensatory changes at each level of the circuit to stabilize overall network output.

      Strengths:

      The study was well-written, and the data presented were clear and an important contribution to the field.

      Weaknesses:

      The sample sizes and what they referred to throughout the distinct studies were unclear. In the legends, the authors should clearly state for each experiment N=X, and if N refers to an NMJ, for example, instead of an individual animal, they should state N=X NMJs per N=X animals. This will help readers better understand the statistical impact of the study.

      This is a good point. For the majority, each data point is derived from a unique specimen, unless explicitly stated otherwise, for NMJ size on muscle DA1 (Figure 3) and for larval crawling data, where each larva was measured up to three time, once per unique 5-minute crawling interval.

      Recommendations for the authors:

      Reviewing Editor Comments:

      In addition to revising the text and making interpretive changes as suggested by the reviewers, we invite you to consider the following:

      (1) Either rephrase the conclusions on the role of the 2h CP and discuss the effects of temperature during embryonic development. Alternatively, to validate the idea of a longer CP window, directly compare the results of a few key experiments using the 2h and 24h heat treatment.

      We have addressed this within the text, as suggested. In the text and figure legends, we have made clear distinctions between exposure to 32˚C heat stress during most of embryogenesis vs a specific developmental window of a few hours. In figures, we have provided diagrams that graphically illustrate the period of heat stress exposure.

      Our ability to experimentally test differences between precise vs broader heat stress periods during embryonic development have been constrained due to the departure of scientists, who were able to carry out electrophysiological recordings (as also explained below). As the next best alternative, we focus on imaging and behavioural analyses. These demonstrated that the developing body wall muscles are sensitive to heat stress during an earlier phase, from 13-16 hours after egg laying, when the body wall muscles become electrically active. It precedes the critical period of the central locomotor circuitry (17-19 hours after egg laying), when neurons in the CNS become electrically and synaptically active (16 hours after egg laying). We think this an exciting additional insight, demonstrating sequential critical periods as different parts of the locomotor network become active: first the body wall muscles, followed by the central circuitry.

      (2) Clarifying the homeostatic responses and shedding light on how they engage with the maladaptive changes described would benefit the study. Furthermore, adding more information about the anatomical and structural changes and how they relate to the intrinsic and synaptic changes would also benefit the study.

      We have tried to address this within the text and with a summary diagram (Figure 8), as suggested.

      Reviewer #1 (Recommendations for the authors):

      (1) It remains unclear whether the authors want to conclude that reduced network stability is not due to changes at the motoneuron level, but rather at the premotor level. Although this idea is mentioned in the results and discussion, it does not appear in the abstract or introduction, which leaves the different findings disconnected. Clarifying and highlighting this conclusion throughout the manuscript would strengthen the narrative.

      We have added additional experiments and changed the manuscript to address this point. These showed that there are distinct phases of embryonic development during which heat stress causes changes to NMJ structure vs to larval behaviour (seizure recovery times/network stability and crawling speed) - additional data in suppl. Fig. 2 and Fig. 7). In the Drosophila embryo, the body wall muscles develop and acquire their electrical properties before central neurons do, and these phases correlate with sensitivity to heat stress.

      (2) In the results section related to Figure 1, the logical link between the CP protocol and the functional assessment of the locomotor network at different temperatures is not sufficiently explained. It is not clear what this assessment is meant to test or demonstrate. A more explicit statement of the rationale and correction of what seems to be a typographical error in the final sentence of the paragraph would help to clarify the authors' intent.

      We have tried to rectify this by changes in the manuscript and to Figure 1, to make the sequence of panels more intuitive.

      (3) In the second results section, the experimental strategy shifts from using a short 2-hour heat shock to a 24-hour manipulation. The reasoning - that short manipulations in different windows yield no phenotype - is understandable, but a 24-hour perturbation may have broader consequences beyond the CP, simply by virtue of its longer duration. Moreover, 24h is roughly the duration of embryonic development at 25C. When at 32C, embryos should develop faster; therefore, is the 24h heat shock extending to L1?

      Yes, the 24 hour heat stress extends into the first few hours of the L1 larval stage.

      In order to validate the use of a longer window, the author should show how it affects the developmental time. Moreover, one should test that a few important observations remain the same with 2h and 24h heat perturbation. Alternatively, one cannot conclude that the phenotype is due to the rather narrow previously defined CP rather than to other effects associated with the overall embryonic developmental time and coordination. This would not make the results less interesting, but it would be important to assess whether the effects can be solely attributed to the 2h CP.

      We have compared the impact of heat stress experience during the majority of embryogenesis, including the CP that had been defined for the central locomotor network (17-19 hours after egg laying) with shorter heat stress manipulations during consecutive phases of embryogenesis until larval hatching. As outlined above in response to point (1) by the Reviewing Editor, reduced stability of the central network and associated reduction in larval crawling occurs when heat stress is experienced during the CP of the central locomotor network (17-19 hours after egg laying). Prolonged heat stress experience for 24 hours leads to indistinguishable outcomes, as long as this 2-hour CP window is included (see Fig. 1 and Fig. 8).

      However, this suggestion by Reviewer #1 led us to identify a second CP for the body wall muscles (see Fig. 3). NMJ overgrowth and changes to the postsynaptic glutamate receptor composition result from earlier heat stress experiences, and those are comparable to the effects caused by 24-hour heat stress exposure when this earlier muscle CP is included.

      Therefore, NMJ development is affected by consecutive CPs, an earlier one linked to body wall muscle development, followed by a later one that impacts the presynaptic motoneurons and their upstream circuitry. Nevertheless, the larval NMJ and behavioural phenotypes that we have identified appear to result from sensitivity to heat stress during these respective CP windows, with no clear evidence of cumulative effects on these phenotypes resulting from longer heat stress exposure during embryogenesis.

      (4) In session 3, the authors note that GluRIIA reductions were most pronounced in proximal regions of the NMJ. However, this is not explicitly quantified in the figures or methods. Including such quantification, or clarifying where it can be found, would make this observation more convincing.

      We have analysed anti-GluRIIA signal intensities in proximal vs distal boutons, comparing different ROI selection processes (e.g. thresholding to a full NMJ/anti-HRP mask and to an anti-GluRIIB mask, which is more selective to postsynaptic sites). Analysis of multiple data sets did not show statistical significance, but instead confirmed that comparable reductions in anti-GluRIIA signal manifest in both proximal and distal boutons, following an embryonic 32C heat stress, relative to controls. We have therefore removed relevant speculative statements.

      (5) In session 4, the authors conclude that motoneurons undergo a decrease in excitability to adjust to greater premotor drive. Is there anatomical evidence for this, such as an increase in input synapses?

      We previously quantified change in excitatory presynaptic synaptic contact number onto aCC motoneuron dendrites in third instar larvae following an embryonic pharmacological activity manipulation: overexcitation of the developing network following introduction of PTX via feeding to gravid females*. No significant structural changes were seen. Although this is a different manipulation of the developing network, all our data to date suggest that heat stress manipulations during the embryonic critical period signal via the same pathways, at least in part due to temperature increases leading to activity increases. Because such a quantification is technically challenging and extremely time-consuming due to the low level of marking individual motoneurons, we did not think it informative or in scope for this project.

      *See Figure 5 in this publication: Hunter I, Coulson B, Pettini T, Davies JJ, Parkin J, Landgraf M, Baines RA. Balance of activity during a critical period tunes a developing network. Elife. 2024 Jan 9;12:RP91599. doi: 10.7554/eLife.91599. PMID: 38193543; PMCID: PMC10945558.

      The interpretation of the optogenetic experiments would also benefit from clarification. If motoneurons are less excitable yet receive more drive, one might expect no net change, rather than the differences observed. Alternatively, could the excitability of the premotor neuron itself have changed, either intrinsically or in relation to Chronos expression? Measuring premotor activity directly during optogenetic activation could help to resolve this ambiguity.

      These are good suggestions. Yes, we think that motoneuron excitability has changed as a result of heat stress - see paper submitted in parallel and published since: Sobrido-Cameán et al., 2025, PLoS Biology. Unfortunately, the team members, who could have carried out this type of analysis had moved on by submission of the manuscript. Therefore, we were no able to experimentally pursue these questions further.

      (6) In session 6, the authors report slower propagation of premotor activity waves after CP heat stress, but the logic of the experiment is not sufficiently explained. How does this finding relate to the enhanced premotor drive described earlier? Only timing is quantified; information about amplitude and wave dynamics would strengthen the interpretation. These results could also be discussed in relation to Figure 1B, where acute heat stress increased motoneuron activity. One interesting possibility could be that CP manipulations might adaptively prepare the larva to function at different temperatures. Experiments testing wave propagation at 32 {degree sign}C (Figure 6) or, conversely, motoneuron activity after CP manipulations (Figure 1B paradigm) would provide valuable evidence for such an adaptive role.

      As per above, unfortunately, the team member who could have carried out this type of analysis had moved on by submission of the manuscript. However, we have tried to address the question of whether there is an adaptive element to the adjustments that result from embryonic heat stress experience. Specifically, we carried out behavioural tests on how larvae respond with changes in crawling speed to acute changes in ambient temperature (new Fig. 7).

      We interpret our findings as follows: that heat stress during embryonic development leads to sub-optimal outcomes with regard to network stability as well as default and maximum crawling speed. The precise causes for this will be difficult to unpick. Behaviourally, when challenging larvae with an acute change in ambient temperature, we saw that slow crawling larvae do respond comparatively normally to a relative increase in ambient temperature by speeding up (in effect an escape response). This demonstrates that embryonic heat stress causes a change in the default crawling speed, while principally maintaining behavioural responses to changes in ambient temperature. It appears that animals that had experienced heat stress during embryonic development, by adjusting their default speed downward, maintain a dynamic response range into a higher temperature range than controls (35C vs 29C, respectively). Potentially, this could be an adaptive outcome to living at higher temperatures, though such an interpretation would require a body of work. 

      Nevertheless, every aspect we have assayed suggests that heat stress experience during the CP leads to sub-optimal outcomes: of network instability, slower default and slower maximal crawling speeds.

      Minor points:

      (1) Abstract: "has suboptimal outcomes;" should be corrected to "has suboptimal outcomes,".

      Corrected

      (2) Abstract: "we find that transient embryonic..." would improve readability with a capitalized "We".

      We are unsure about the sentence this refers to. If this sentence, then we suggest that this could remain as was.

      "Within the central nervous system, we find transient embryonic CP perturbation leads to increased synaptic drive from premotor interneurons to motoneurons..."

      The intention here is to differentiate between changes within the CNS vs at the NMJ.

      (3) Abstract: The sentence "Present the larva ... as an experimental model system..." overstates novelty, as the system has already been established in prior work. Instead, this study could highlight temperature manipulation as an ecologically relevant way to probe CPs.

      We have adopted this suggestion.

      Reviewer #2 (Recommendations for the authors):

      I would recommend:

      (1) Perform additional experimental support and dissuasions for the causal role of the premotor neuron and the network disability.

      Unfortunately, it has not been possible to carry out additional e-phys experimental work due to key people having moved on and now unable to carry out such experiments, and no replacements in sight to do so. Instead, we have focused on other work that we could do, namely to test the effect of different heat stress windows during embryonic development on GluRIIA vs GluRIIB expression at postsynaptic sites. This shows that indeed the effect seen following a 24-hour heat stress is replicated by a much shorter window of heat stress. For NMJs GluRII composition the critical period is different from the critical period of the CNS. This replicates the different developmental timings of maturation: the body wall muscles express ion channels and attain their electrical properties several hours before central neurons. We have provided these additional data as a new supplementary figure to Fig. 2. We have changed the main text to note the caveat of longer heat stress manipulations potentially leading to additional or more exacerbated phenotypes.

      (2) I would suggest that the authors expand the discussion on why two layers of homeostatic adjustment fail to preserve behavior. Is this simply a limit of plasticity?

      This is a difficult aspect to address well, beyond the purely speculative and potentially confusing. We would like to suggest that to do so requires a basic understanding on what pressures neurons/networks respond to (heat-caused over-activation and/or metabolic); and from that perspective to gauge how they adjust to those pressures, i.e. what the adjustments are trying to "achieve". This in turn should inform on whether such adjustments are homeostatic or anti-homeostatic in nature, and whether there are limits beyond which we consider a system "breaking".

      (3) 32C heat was described as "ecologically relevant". However, it produces maladaptive outcomes. The author should consider reframing it as a stressor that reveals CP sensitivity, rather than an adaptive signal.

      This is a good suggestion, and we have implemented changes accordingly.

      (4) It is important to distinguish the effects of transient (2h) vs prolonged heat exposure to confirm the manipulation targeted CP specifically.

      We have tried to make these distinctions clearer within text and figure legends. As per response to (1) above, we generated and analysed additional data to show that changes in GluRIIA are induced during a defined shorter developmental time window, not exacerbated by prolonged heat stress exposure during embryogenesis.

      (5) I would definitely recommend adding a diagram tracking information flow and showing where structural and functional changes occur.

      This is a helpful suggestion, which we have tried to implement with a new figure (Fig. 8).

      Overall, this is a strong and well-written paper that produced some unexpected results, and added a solid model circuit to study CP plasticity at the circuit level.

      Reviewer #3 (Recommendations for the authors):

      I identified one typo: "activity manipulations during the embryonic CP are artificial, We asked to what extent" The "W" of "We" should be lower case.

      Now corrected.

    1. eLife Assessment

      This paper represents a valuable contribution to our understanding of how local field potential (LFP) oscillations and beta band coordination between the hippocampus and prefrontal cortex of rats may relate to learning. Through a set of appropriate analyses, including oscillation-oscillation coupling, oscillation-spike modulation, and behavioral dependant measurements, the study presents convincing evidence for uncoupled beta activity between the two regions. This work will be of interest to researchers working on interaction of brain regions during spatial learning in rodents.

    2. Reviewer #1 (Public review):

      Wang, Zhou et al. investigated coordination between prefrontal cortex (PFC), and hippocampus (Hp), during reward delivery via analyzing beta oscillation. Beta oscillations are associated with various cognitive functions but their role in coordinating brain networks during learning is still not thoroughly studied. Authors focused on the changes in power, peak frequencies and coherence of beta oscillations in two regions when rats learn a spatial task thru days. Contradicting with authors hypothesis, beta oscillations in those two regions during reward delivery were not coupled in spectral or temporal aspects. They were, however, able to show reverse changes in beta oscillations in PFC and Hp as the animal's performance got better. Authors were also able to show a small subset of cell population in PFC that are modulated by both beta oscillations in PFC and sharp wave ripples in Hp. A similarly modulated cell population was not observed in Hp. These results are valuable in pointing out distinct periods during a spatial task when two regions modulate their activity independent from each other.

      Authors made a detailed analysis of the data to support their conclusions. Few more points of discussion would clarify the results of the paper.

      (1) One of the big conclusions of the paper is how the beta burst power is changing after learning the task (Figure 3). Authors have also showed in Figure 6-1, how the SWR power and rate are changing thru the training days. Did they observe a change in coordination of Beta bursts and SWR between the days, which would also reflect how experience changes the coordination?

      (2) Authors have shown in detail the opposite relationship between Beta phase locking and SWR modulation in Hippocampus in Figure 7I. This might require a different analysis, but is it possible to make a discussion on predicting a cell firing in a SWR after it fires in a beta burst.

      Other than these two points, authors have addressed previous comments and made a convincing analysis of their data.

    3. Reviewer #2 (Public review):

      Using electrophysiological recordings in freely moving rats during a spatial navigation task, this study investigated the role of beta oscillations in the hippocampal-prefrontal network. Through a set of appropriate analyses-including oscillation-oscillation coupling, oscillation-spike modulation, and behavioral dependant measurements-the study presents convincing evidence for uncoupled beta activity between the two regions. These findings offer important insights into the network mechanisms of spatial navigation and may have significant implications for related neurological disorders.

      Comments on revised version.

      The authors have carried out additional analyses and made corresponding revisions to the manuscript in response to the earlier review comments, which have made the conclusions more convincing. However, please note that the last question about coexistence of beta and SWR has not been fully answered. Please supplement your response to address this point completely.

    4. Reviewer #3 (Public review):

      This paper explored the role of beta rhythms in the context of spatial learning and mPFC-hippocampal dynamics. The authors characterized mPFC and hippocampal beta oscillations, examining how their coordination and their spectral profiles related to learning and prefrontal neuronal firing. Rats performed two tasks, a Y-maze and F-maze, with the F-maze task being more cognitively demanding. Across learning, prefrontal beta oscillation power increased while beta frequency decreased. In contrast, hippocampal beta power and beta frequency decreased. This was particularly for the well-performed and well-learned Y-maze paradigm. The authors identified the timing of beta oscillations, revealing an interesting shift in beta burst timing relative to reward entry as learning progressed. They also discovered an interesting population of prefrontal neurons that were tuned to both prefrontal beta and hippocampal sharp-wave ripple events, revealing a spectrum of SWR-excited and SWR-inhibited neurons that were differentially phase locked to prefrontal beta rhythms.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Wang, Zhou et al. investigated coordination between the prefrontal cortex (PFC) and the hippocampus (Hp), during reward delivery, by analyzing beta oscillations. Beta oscillations are associated with various cognitive functions, but their role in coordinating brain networks during learning is still not thoroughly understood. The authors focused on the changes in power, peak frequencies, and coherence of beta oscillations in two regions when rats learn a spatial task over days. Inconsistent with the authors' hypothesis, beta oscillations in those two regions during reward delivery were not coupled in spectral or temporal aspects. They were, however, able to show reverse changes in beta oscillations in PFC and Hp as the animal's performance got better. The authors were also able to show a small subset of cell populations in PFC that are modulated by both beta oscillations in PFC and sharp wave ripples in Hp. A similarly modulated cell population was not observed in Hp. These results are valuable in pointing out distinct periods during a spatial task when two regions modulate their activity independently from each other.

      The authors included a detailed analysis of the data to support their conclusions. However, some clarifications would help their presentation, as well as help readers to have a clear understanding.

      (1) The crucial time point of the analysis is the goal entry. However, it needs a better explanation in the methods or in figures of what a goal entry in their behavioral task means.

      We appreciate Reviewer 1 pointing out this shortcoming and will clarify the description in the revised manuscript. Each goal is located at the end of the arm, and is equipped with a reward delivery unit. The unit has an infrared sensor. The rat breaks the infrared beam when it enters the goal. Figures 1 and 2 have been updated to clearly indicate the time of goal entry. The main text and methods have been updated with the explanation.

      (2) Regarding Figure 2, the authors have mentioned in the methods that PFC tetrodes have targeted both hemispheres. It might be trivial, but a supplementary graph or a paragraph about differences or similarities between contralateral and ipsilateral tetrodes to Hp might help readers.

      We appreciate this suggestion, which has led to an interesting finding. The coherence and burst coordination were similar for ipsi- and contralateral PFC and hippocampus. Interestingly, we found PFC beta activity was more coherent within each PFC hemisphere compared with across hemispheres. This was observed for coherence and burst time. This suggests there is hemispheric localization of beta oscillations. These results are shown in Fig. 2-1.

      (3) The authors have looked at changes in burst properties over days of training. For the coincidence of beta bursts between PFC and Hp, is there a change in the coincidence of bursts depending on the day or performance of the animal?

      This is now reported in Fig. 3-3. After quantifying the proportion of independent and coincident bursts as function of experiment day or performance, we found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (4) Regarding the changes in performance through days as well as variance of the beta burst frequency variance (Figures 3C and 4C); was there a change in the number of the beta bursts as animals learn the task, which might affect variance indirectly?

      The difference in the burst count across days did not explain the results. We performed a permutation test (Fig. 4-4), where we randomly shuffled the day identity to control for the count difference across days. The change in variance remains significant.

      (5) In the behavioral task, within a session, animals needed to alternate between two wells, but the central arm (1) was in the same location. Did the authors alternate the location of well number 1 between days to different arms? It is possible that having well number 1 in the same location through days might have an effect on beta bursts, as they would get more rewards in well number 1?

      The central arm remained the same across days since we needed the animals to learn the alternation task. In our experience, the animal needs a few days to learn the alternation rule when we switch the central arm location. For this experiment, we were interested in the initial learning process, and we kept the central arm constant. Switching the central arm location is a great suggestion for a follow-up experiment where we can understand the effects of reward contingency change on beta bursts.

      (6) The animals did not increase their performance in the F maze as much as they increased it in the Y maze. It would be more helpful to see a comparison between mazes in Figure 5 in terms of beta burst timing. It seems like in Y maze, unrewarded trials have earlier beta bursts in Y maze compared to F maze. Also, is there a difference in beta burst frequencies of rewarded and unrewarded trials?

      We performed the analysis and found burst timing was similar between the two mazes (Fig. 4-2). Bursts on rewarded trials occurred later than those on unrewarded trials (Fig. 4). Interestingly, PFC bursts on rewarded trials were lower in frequency compared with unrewarded trials. CA1 bursts during rewarded and unrewarded trials had similar frequencies (Fig. 4-3).

      (7) For individual cell analysis, the authors recorded from Hp and the behavioral task involved spatial learning. It would be helpful to readers if authors mention about place field properties of the cells they have recorded from. It is known that reward cells firing near reward locations have a higher rate to participate in a sharp wave ripple. Factoring in the place field properties of the cells into the analysis might give a clearer picture of the lack of modulation of HP cells by beta and sharp wave ripples.

      As recommended, we quantified the mean speed, mean distance to goal locations, and spatial information for CA1 cells (Fig. 7 J-L). We found SWR reactivated CA1 cells had higher speed and were spiking further away from the goals compared with non-reactivated CA1 cells. This is consistent with prior work that shows SWR-associated reactivation in dorsal CA1 can correspond to trajectories taken as the animal moves towards goals. In intermediate CA1, the content of reactivations is biased toward place representations closer to goals (Jin et al., 2024). CA1 cells with or without phase locking to beta oscillations had similar spatial firing properties.

      Reviewer #1 (Recommendations for the authors):

      (1) Please make a figure representing what the goal entry means in Figure 1.

      We have updated Fig. 1 to clearly show the definition of goal entry. We also edited the main text and methods to better explain the definition of goal entry.

      (2) For Figure 1-1, please either change the contrast of the pictures, or define the lesioned areas, as it is a bit difficult to see the lesioned parts, especially in PFC.

      We have increased the contrast for Fig. 1-1 for the histology to better show the lesions.

      (3) For Figure 1-2, is it possible to show a beta burst from Hp?

      Yes, we provided two examples of beta bursts from the hippocampus alongside burst examples from PFC in Fig. 1-2.

      (4) Is it possible to make a supplementary table showing the number of tetrodes recorded from each animal per day, plus the number of isolated single cells?

      Yes, we have included tetrode counts in Table 1-2 and cell counts in Tables 5-2 to 5-4.

      Reviewer #2 (Public review):

      (1) When presenting the power spectra for the representative example (Figure 1), it would be appropriate to display a broader frequency band-including delta, theta, and gamma (up to ~100 Hz), rather than only the beta band.

      We agree the extended frequency range provides a better overview of the spectral characteristics during the goal period. We have now included example spectrograms up to 100 Hz to show the spectral content for a wider range of frequencies (Fig. 1-2). Further, we have included additional analyses to compare the spectral characteristics between periods when the animal was moving on the maze or immobile at the goal, for frequencies up to 100 Hz (Fig. 1C-H, Fig. 1-3 and 1-4). We used both Welch’s periodogram (Fig. 1-3) and continuous wavelet transform (Fig. 1-4) to demonstrate our findings on beta oscillations in both regions are robust and consistent.

      What was the rat's locomotor state (e.g., running speed) after entering the reward location, during which the LFPs were recorded?

      Because goal entry is defined as the time the animals break the infrared beam at the goal (response to Reviewer 1), the rat would have come to a stop. We have added the time-aligned speed profile to the spectra and raw data examples in the manuscript (Fig. 1B, Fig. 1-4, Fig. 2A, and Fig. 6A). In addition, we added the quantification of the animal’s speed at the time of beta bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1D).

      If the rats stopped at the goal but still consumed the reward (i.e., exhibited very low running speed), theta rhythms might still occasionally occur, and sharp-wave ripples (SWRs) could be observed during rest.

      We typically find low theta power in the hippocampus after the animal reaches the goal location and as it consumes reward. Reviewer 2 is correct about occasional theta power at the goal. To compare differences in LFP characteristics between maze running and goal locations, we added additional analyses in Fig. 1-3 and 1-4. We did find SWRs during goal periods (Fig. 6) and we quantified SWR properties in an additional analysis in Fig. 6-1.

      Do beta bursts also occur during navigation prior to goal entry? It would be beneficial to display these rhythmic activities continuously across both the navigation and goal entry phases.

      We did not find consistent beta bursts in PFC or CA1 on approach to goal entry. We generated an additional goal entry-aligned spectrogram, and quantification (Fig. 1-4) to show that beta oscillations in both regions increased after goal entry. This was also supported by the Welch’s periodogram method (Fig. 1C-H, Fig. 1-3). Beta oscillations in the hippocampus during locomotion or exploration have been reported (Ahmed & Mehta, 2012; Berke et al., 2008; França et al., 2014; França et al., 2021; Iwasaki et al., 2021; Lansink et al., 2016; Rangel et al., 2015).

      Additionally, given that the hippocampal theta rhythm is typically around 7-8 Hz, while a peak at approximately 15-16 Hz is visible in the power spectra in Figure 1C, the authors should clarify whether the 22 Hz beta activity represents a genuine oscillation rather than a harmonic of the theta rhythm.

      We performed further spectral analysis comparing times when the animal is moving on the maze with times when the animal is immobile at the goal (Fig. 1-3 and 1-4). The results point to the beta frequency oscillations in both regions are unlikely to be harmonics of theta. The beta frequency bands in the spectrogram are independent of the theta band. We were initially concerned about the possibility that the 22 Hz power in CA1 may be a harmonic rather than a standalone oscillation band. If these are harmonics of theta, we should expect to find coincident theta at the time of bursts in the beta frequency. In Fig. 1B, Fig. 1-5, and Fig. 2A, we show examples of the raw LFP traces from CA1. Here, the detected bursts are not accompanied by visible theta-frequency activity. For PFC, we do not always see persistent theta-frequency oscillations like CA1. In PFC, we found beta bursts were frequent and visually identifiable when examining the LFP. We provided examples of the PFC LFP (Fig. 1B, Fig. 1-5, and Fig. 2A). In these cases, we see clear beta frequency oscillations lasting several cycles and these are not accompanied by any visible oscillations in the theta frequency in the LFP trace.

      (2) The authors claim that beta activity is independent between CA1 and PFC, based on the low coherence between these regions. However, it is challenging to discern beta-specific coherence in CA1; instead, coherence appears elevated across a broader frequency band (Figure 2 and Figure 2-1D). An alternative explanation could be that the uncoupled beta between CA1 and PFC results from low local beta coherence within CA1 itself.

      This is a legitimate concern, and we used three methods to characterize coherence and coordination between the two regions. First, we calculated coherence for tetrode pairs for times when the animal was at goals (Fig. 2B), which provides a general estimation of coherence across frequencies but lack any temporal resolution. Second, we calculated burst-aligned coherence (Fig. 2-2), which provides temporal resolution relative to the burst, but the multi-taper method is constrained by the time-frequency resolution trade-off. Third, we quantified the timing between the burst peaks (Fig. 2D), which described the timing differences but the peaks for the bursts may not be symmetric. Each method has its own caveats, but we drew our conclusion from the combination of results from these three analyses, which pointed to similar conclusions.

      Reviewer 2 is correct in pointing out the uniformly high coherence within CA1 across the frequency range we examined. When we inspected the raw LFP across multiple tetrodes in CA1, they were similar to each other (Fig. 2A). This likely reflects the uniformity in the LFP across recording sites in CA1, which is what we saw with coherence values across the frequency range (Fig. 2B). We found that CA1 coherence between tetrode pairs within CA1 was statistically higher than tetrode pairs in PFC across the frequency range (Fig. 2B and C), thus our results are unlikely to be explained by low beta coherence within CA1 itself. The burst-aligned coherence using a multi-taper method also supports this. The coherence values within CA1 at the time of CA1 bursts were ~0.8-0.9.

      (3) In Figure 2-1E-F, visual inspection of the box plots reveals minimal differences between PFC-Ind and PFC-Coin/CA1-Coin conditions, despite reported statistical significance. It may be necessary to verify whether the significance arises from a large sample size.

      We will include the sample sizes in Table 2-1. We repeated the analyses based on average values per day for each animal (Fig. 2-2 E-F). The pattern of significance remains consistent.

      (4) In Figure 3 and Figure 4, although differences in power and frequency appear to change significantly across days, these changes are not easily discernible by visual inspection. It is worth considering whether these variations are related to increased task familiarity over days, potentially accompanied by higher running speeds.

      We agree with Reviewer 2 that familiarity increases across days, and the animal is likely running faster. The analysis for Fig. 3 (previously Fig. 3 and 4) includes only data from periods when the animal was at the goal and was not moving. We added a supplemental figure (Fig. 3-1), which shows that the speed of the animal was below 0.5 cm/s at the time of the analyzed bursts. We used linear mixed-effects models to quantify the relationship between power, frequency and day or behavioral quintile, which accounts for repeated measurements across animals.

      (5) The stronger spiking modulation by local beta oscillations shown in Figure 6 could also be interpreted in the context of uncoupled beta between CA1 and PFC. In this analysis, only spikes occurring during beta bursts should be included, rather than all spikes within a trial. The authors should verify the dataset used and consider including a representative example illustrating beta modulation of single-unit spiking.

      We agree with Reviewer 2 that the stronger modulation to local beta is another piece of evidence indicating uncoupled beta between the two regions. We appreciate this suggestion and have revised Fig. 5 (previously Fig. 6) to include examples illustrating beta modulation for single units. These are spike-phase raster and histograms. We want to clarify that in the revised manuscript, the spikes were only from periods when the animal was at the goal location (5 s after entry) and did not include the running period between goals. Although beta power fluctuates in bursts, our data show there are ongoing beta oscillations throughout this period (Fig. 1-3 and Fig. 1-4), which prompted us to examine the entire period.

      (6) As observed in Figure 7D, CA1 beta bursts continue to occur even after 2.5 seconds following goal entry, when SWRs begin to emerge. Do these oscillations alternate over time, or do they coexist with some form of cross-frequency coupling?

      This is a very helpful suggestion and led to some interesting findings. We performed two additional analyses: 1) burst/SWR cross-correlation on a shorter timescale and 2) SWR-aligned spectrogram for the beta frequency range. PFC beta burst timing and power appear to be anti-correlated with SWRs detected in CA1 (Fig. 6G and K). In contrast CA1 beta bursts were more likely to occur with SWRs (Fig. 6H and L). These results suggest there is temporal coordination between ongoing beta oscillations in PFC and SWRs in the hippocampus during waking. Cortical beta oscillations are reduced during hippocampal SWRs, perhaps to support the transient switch in global cortical states that accompanies SWRs.

      To examine potential cross-frequency coupling between SWRs and beta oscillations, we computed the mean SWR band power (150-250).

      Reviewer #3 (Public review):

      Summary:

      This paper explored the role of beta rhythms in the context of spatial learning and mPFC-hippocampal dynamics. The authors characterized mPFC and hippocampal beta oscillations, examining how their coordination and their spectral profiles related to learning and prefrontal neuronal firing. Rats performed two tasks, a Y-maze and an F-maze, with the F-maze task being more cognitively demanding. Across learning, prefrontal beta oscillation power increased while beta frequency decreased. In contrast, hippocampal beta power and beta frequency decreased. This was particularly the case for the well-performed and well-learned Y-maze paradigm. The authors identified the timing of beta oscillations, revealing an interesting shift in beta burst timing relative to reward entry as learning progressed. They also discovered an interesting population of prefrontal neurons that were tuned to both prefrontal beta and hippocampal sharp-wave ripple events, revealing a spectrum of SWR-excited and SWR-inhibited neurons that were differentially phase locked to prefrontal beta rhythms.

      In sum, the authors set out to examine how beta rhythms and their coordination were related to learning and goal occupancy. The authors identified a set of learning and goal-related correlates at the level of LFP and spike-LFP interactions, but did not report on spike-behavioral correlates.

      Strengths:

      Pairing dual recordings of medial prefrontal cortex (mPFC) and CA1 with learning of spatial memory tasks is a strength of this paper. The authors also discovered an interesting population of prefrontal neurons modulated by both beta and CA1 sharpwave ripple (SWR) events, showing a relationship between SWR-excited and SWR-inhibited neurons and beta oscillation phase.

      Weaknesses:

      Moreover, there is little detail provided about sample sizes and how data sampling is being performed (e.g., rats, sessions, or trials), raising generalizability concerns.

      We appreciate Reviewer 3’s thoughtful suggestions for making our claims convincing. We have included information about sample sizes in the revised manuscript.

      The authors report on a task where rats were performing sub-optimally (F-maze), weakening claims.

      Our experiment was designed to create a scenario in which one task was learned (Y-maze) and another was not (F-maze). This contrast allows us to determine differences in neural correlates of learning versus familiarity. The design produced a learned and not learned task with similar levels of familiarity over 5 days, within the same animal.

      Likewise, it is questionable as to whether mPFC and hippocampus are dually required to perform a no-delay Y-maze task at day 5, where rats are performing near 100%.

      We agree with Reviewer 3 that the mPFC and hippocampus may not be required when the animal reaches stable performance on day 5 (Deceuninck & Kloosterman, 2024). The data we collected spans the full range of early learning (day 1) to proficiency (day 5). We wanted to understand the dynamics of beta across these learning stages, which have not been reported previously.

      Recent studies suggest mPFC and hippocampus are likely to be needed, in some capacity, for learning continuous spatial alternation tasks on a range of maze geometries. Lesions, inactivation or waking activity perturbation of hippocampus or hippocampus and mPFC on the W maze alternation task slowed learning (Jadhav et al., 2012; Kim & Frank, 2009; Maharjan et al., 2018). More recently, optogenetic silencing of mPFC after sharp wave ripples on the Y-maze alternation affected performance when the center arm was switched (den Bakker et al., 2023). The Y and F-mazes in our study both share the continuous alternation rule, where the animal needed to avoid visiting a previously visited location on the outbound choice relative to the center, and always return to the center location.

      Further, the performance characteristics on the outbound and inbound components of our Y task are similar to the W task. We have analyzed the “inbound” and “outbound” performance of the animals on the Y-maze alternation task, and they are similar to the W maze alternation task. The “inbound” or reference location component is learned quickly whereas the “outbound”, alternation component is learned slowly.

      There would be little reason to suspect strong oscillatory coupling when task performance is poor and/or independent of mPFC-HPC communication (Jones and Wilson, 2005) potentially weakening conclusions about independent beta rhythms.

      Although many studies have examined the oscillatory coupling properties at the theta frequency between mPFC-HPC (Hyman et al., 2005; Jones & Wilson, 2005; Siapas et al., 2005), our understanding of beta frequency coordination between the two regions is less established, especially at goal locations. Our work suggests beta frequency coordination at goal locations does not share properties with those of theta frequency coupling between mPFC and HPC, which occurs primarily during movement on the maze. Our first novel finding is that beta oscillations occur at goal locations in mPFC and HPC; our second is that the beta frequency dynamics in these regions are surprisingly distinct. We are not aware of prior work describing these properties at goal locations in spatial navigation tasks, especially their temporal coordination.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusions from this article would be made much stronger if the authors (1) record from rats performing a task known to be dependent on mPFC-HPC communication (e.g. a spatial working memory task) or (2) record from the F-maze in well-trained rats (75-80% performance is common), or (3) show that Y-maze task performance is dependent on mPFC-hippocampal communication (see Maharjan et al., 2018, which used a W-track). It is possible that learning the Y-maze depends on mPFC-hippocampal communication, but that with asymptotic performance, this changes. This would put beta oscillation coupling findings into a nuanced perspective.

      We appreciate the recommendation. The objective of this current manuscript is to present previously unknown properties of beta oscillations in hippocampal-prefrontal cortical networks. We agree that further investigation is required to fully dissect the functional contribution of beta dynamics in these networks. We are in the process of doing that.

      The rule on the Y-maze in our experiment is identical to the continuous alternation rule on the W maze in Maharjan et al., 2018 and Kim and Frank, 2009 (Author response image 1), which showed the PFC and hippocampus are required for normal learning, respectively. The Y and W mazes share the same topology; there is one junction connecting three arms. The W maze has two 90-degree-angle turns which are not choice points. Thus, both maze tasks have one choice point and involve learning an alternation rule. Further the learning properties share similarities. For the Y-maze, the inbound portion (return to center) (Author response image 1) was quicker to learn than the outbound portion (alternation) (Author response image 1). The same pattern is observed in the W maze learning task (Maharjan et al., 2018, Fig. 3A and D, Kim and Frank, Fig. 4C). We agree an experiment is needed to show that Y-maze task learning is dependent on the function of hippocampal-prefrontal cortical networks. The shared topology, rule definition, and behavior profile suggest that the Y and W maze tasks engage similar learning processes.

      Author response image 1.

      W and Y maze alternation tasks share the same rule. Schematics illustrate a comparison of W and Y maze alternation task rules. The W maze has been inverted for visual comparison with the Y maze. Performance grouped by in- or outbound trials. Inbound trials originate from the side goals (2 or 3). Outbound trials originate from the center goal (1). Performance on the inbound trials was higher than that on the outbound trials, which is comparable to previously published alternation tasks on W-shaped mazes.

      (2) Typically, when analyzing LFP profiles, experimenters include running velocity/speed. It should be ruled out whether spectral changes are confounded in any way by speed or time spent in the goal zones.

      We appreciate this suggestion for better conveying our definition of goal period. This was raised by other reviewers. We have now included the speed profiles for the goal-entry-aligned spectrograms (Fig. 1B, 1-4, 2A, and 6A), as well as the speed quantification at the times of bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1 D). We strictly define the goal period after entry into the sensor on the reward delivery device at the end of the arms. This is to ensure the animal is immobile for the period to avoid confounds related to speed. We also ensured the time periods are comparable between the trials we analyzed.

      (3) The authors should describe in each analysis how data are sampled, and for extracellular electrophysiology experiments, cell counts from each rat reported. A table would be ideal. For example, it is unknown if entrainment analysis is performed on data collected from one rat or from all rats.

      We have now added the missing information (Table 1-2, 5-1 to 5-3). We performed analyses using data from all animals.

      (4) There was no profiling of mPFC neurons in terms of their behavioral correlates. The authors should strongly consider examining how individual neurons encode task variables (e.g., trial correctness, reward location...) and can do so using a generalized linear model. Adding an analysis of behavioral correlates could nicely tie into the beta-SWR analyses. For example, are SWR-beta rhythm-modulated neurons also behaviorally modulated?

      This is a very helpful suggestion. We have added analyses on the behavioral correlates of PFC and CA1 neurons. We quantified three metrics: 1) whether PFC and CA1 spiking activity can distinguish goal location based on firing rate or phase preference, 2) firing distance to goal and 3) spatial information (Fig. 5 and 7). We also performed the same analysis based on SWR and beta modulation status. We found PFC cells that were both SWR- and beta-modulated showed the strongest task firing relationship (Fig. 7).

      (5) Figure 2:

      Are these data analyzed from well-trained rats? What is your N (rats/sessions/trials/epochs)?

      These results in Fig. 2 are from all days. We added the breakdown in Table 2-1.

      (6) Figure 3:

      (a) The authors show that on the Y-maze, performance, beta oscillations power, and beta oscillation frequency change over days. However, for the F-maze, performance improves but appears to taper off at 60% and PFC beta frequencies do not change with learning. Do you have rats performing this task well above chance (e.g., 75-80%?), and if so, do beta oscillation frequencies in the mPFC gradually change?

      We did not observe rats performing above chance on the F-maze over 5 days. They do not appear to learn the alternation rule on this task. This is why we used the F-maze as the “non-learner” control. The F-maze task design for this study appears to be difficult to learn and would likely require a much longer training period for the performance to exceed chance. We agree an important follow-up question is whether the effects reported here are generally observed across learning in different tasks. This is a future direction we are actively pursuing.

      One existing data point may partially address the relevance of beta power change with learning. We analyzed the power as a function of performance quintile on the Y-maze (learned) or F-maze (not learned). The power changed with performance quintile on the Y-maze (learned) but not the F-maze (not learned) (Fig. 3D), pointing to the change in power being associated with learning status.

      (b) Does the proportion of beta bursts change with learning? What about the proportion of coherent events?

      We performed this analysis and the results are shown in Fig. 3-3. We found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (c) Is the reduction in beta oscillation frequency in the mPFC related to running behavior? This should be ruled out. Same for the hippocampus.

      The reduction is not due to running because we only included bursts when the animal was at the goal and immobile. We added a figure to show speed at the time of bursts (Fig. 3-1).

      (d) Why don't you also show coherence as a function of learning?

      This is a great suggestion. Coherence as a function of day or performance is now shown in Fig. 3-2. Overall, the trends were weak suggesting there was no strong change in coherence over days or as a function of learning. Although some of the linear mixed effects models were statistically significant, the marginal R<sup>2</sup> values (R<sup>2</sup><sub>m</sub>) were very low, indicating the effects of performance or day on coherence were small. Beta frequency coherence within each brain region either remained the same or slightly decreased across days (Fig. 3-2 D and F) or with performance (Fig. 3-2 J). For coherence across regions, we only found a weak but significant increase for the well-performed Y-maze task across days (Fig. 3-2 B).

      (e) The statistics in the caption are great, but maybe consider using a table and a supplemental figure

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (7) Figure 4

      (a) This figure is better suited as supplemental to Figure 3.

      We have combined Fig. 4 with Fig. 3.

      (b) Statistics might be better suited in a table and a supplemental figure.

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (8) Figure 5

      (a) The shift of beta rhythms being linked to learning is interesting, albeit confusing, given that error trials were accompanied by even more 'precise' beta rhythm timing. Is it possible that beta rhythms are accompanied by an error signal? Or maybe instead related to running behavior?

      We verified this observation was not due to speed since we only included periods when the animal was at the goal and immobile. We added a new supplemental figure (Fig. 4-1) to report this analysis.

      (b) Changes in beta rhythm timing in the F-maze could be entirely explained by the data in the Y-maze, given that F-maze performance was, in general, very low. To test if beta rhythm shifting is a consequence of learning, the authors should show performance on the F-maze without the Y-maze, and vice versa.

      As suggested, we also analyzed frequency variance data from each maze separately and found learning-related changes were more prominent on the Y-maze (learned) than on the F-maze (not learned) (Fig. 4-4 C). The F-maze data serves as a valuable control within the same animal for the learning-related effects we are reporting.

      We agree that the Y-maze data is consistent with performance-related changes to beta timing. The addition of the F-maze provides a within-animal comparison for a task in which the animal performed less well. We agree an additional control group would provide further evidence to support learning-related changes in beta timing. However, the F-maze will likely take much longer to learn, therefore the additional days of exposure on the task will be a different confound. Thus, we wanted to ensure that we compared a “learned” with a “not learned” task within the same animal, with a comparable exposure period.

      (9) Figure 6

      (a) How many cells are analyzed here? Are they all pyramidal neurons? The authors should add the number of cells analyzed from each rat. A table would suffice.

      We added tables (Table 5-2 to 5-4) to show this information. For our analysis, we included all units.

      (b) The authors should include statistics about recorded units. Peak-to-trough timing and interspike interval are commonly used. Even demonstration of action potentials from individual cells is valuable for proof of concept.

      We have added more information on the cell type classification (Table 5-2) and how we classified the cells (Methods) and also added example waveforms in Fig. 5 (previously Fig. 6).

      (c) Why did the authors stop comparing the F-maze and Y-maze for entrainment analysis?

      The data for each maze was comparable. Per the reviewer’s suggestion, we added the entrainment analysis for each maze separately (Fig. 5-1).

      (10) Figure 7

      (a) The authors should minimally include instantaneous velocity to show that during putative ripple events, the animal is in fact, quiescent.

      We added the speed at the time of ripples in Fig. 6-1 D.

      (b) Figure 7C: The authors should show how time spent in the reward zones varies over days. Is SWR power simply changing due to behavioral occupancy differences (e.g., a statistical sampling problem)? An analysis of the SWR rate over days would be valuable.

      We agree and have added additional SWR properties across days in Fig. 6-1. This includes SWR power (Fig. 6-1 A), rate (Fig. 6-1 B), duration (Fig. 6-1 C) and speed (Fig. 6-1 D). To ensure we are comparing the equivalent time spent at goals, the SWR properties were calculated from the 10 s after goal entry (Fig. 6-1) and we only included goal visits that lasted at least 10 s. This controls for the potential confounds from occupancy differences across days.

      (c) Why did the authors stop comparing the F-maze and Y-maze for SWR analysis?

      The SWR properties were similar in both tasks. We have now added separate analyses for each maze in Fig. 6-1 and 6-2.

      (11) Figure 8

      (a) The authors discovered that mPFC neurons more strongly entrained to SWRs compared to their own beta rhythms, potentially indicating coordination of mPFC and hippocampus (lines 275:276). What about mPFC spiking to CA1 beta?

      We quantified mPFC spiking entrainment to CA1 beta in Fig. 5 O and P. We found mPFC cells were more strongly entrained to the local beta within mPFC compared with CA1 beta.

      (b) Lines 277:278: How did the authors arise at 8.6% being the expected proportion of neurons entrained by both SWRs and beta rhythms?

      We have added a better explanation of how we arrived at the expected proportion. The calculation is based on the joint probability between the proportion of neurons modulated by SWRs (30%) and the proportion of neurons with phase locking to beta rhythms (22%). The expected proportion (6.6%) is calculated by multiplying the two values (30% ´ 22%) under the assumption that the two classes are independent. Deviations from the expected proportions are determined using a Fisher’s exact test.

      We note that in the revised manuscript we reanalyzed the data with more stringent criteria. We now include SWRs with durations greater than 50 ms, rather than 15 ms, and we quantify beta spike phase-locking using only the spikes within the first 5 s after goal entry, for goal visits lasting at least 5 s.

      With these more conservative criteria, the proportion of SWR- and beta-modulated cells (8%) is no longer significantly different from the expected proportion (6.6%, Fisher’s exact test p=0.09). Although this differs from our original significance test, the trend, and importantly, the distinct task correlates for this population (Fig. 7 C-D) still hold. We have updated the revised manuscript with these findings.

      Our original result was: “The subset of PFC cells that are modulated by both SWR and beta (11%) is greater than the expected proportion (8.6%) under the assumption that SWR and beta can modulate the population independently (Fisher exact test, p=0.021), although the size of the difference is small.”

      (c) The discovery that neurons modulated by both beta and SWRs are unique from those simply modulated by beta is really interesting. The authors discovered an interesting relationship among dually entrained mPFC neurons whereby SWR-excited mPFC neurons were entrained to the peak of beta, whereas SWR-inhibited mPFC neurons were entrained to the trough of beta. Then, in lines285:286, the authors write:

      (12) "This relationship was not observed for PFC cells that are modulated by beta but not modulated by SWRs (Fig. 8C, right, Fig. 8-1A)."

      Why would the authors expect this relationship to exist when the mPFC neurons were not SWR modulated? Were the authors referring to something else?

      We should have phrased this more clearly. We edited the text to better convey the expected SWR and beta modulation patterns for the control populations. We wanted to ensure the relationship between SWR modulation direction and beta phase preference was not observed for the cells that were not SWR-modulated.

      Furthermore, what about mPFC neurons that are SWR modulated but not beta modulated? For completeness, the authors should examine these.

      We agree this is an important comparison and have now included this population in the new Fig. 7. As expected, we do not find any relationship between SWR modulation direction and spike preference to beta for the SWR-modulated and not beta-modulated population.

      (13) Lines 294:296: The authors should elaborate on how they obtained an expected proportion of CA1 beta modulated neurons.

      We have now added a description for calculating the expected number of beta-modulated CA1 neurons. These are the neurons that have a Rayleigh test p-value less than 0.05.

      (14) What are these neurons doing to predict task information? Do they at all? Do they differ from beta-only neurons or non-phase-locked neurons?

      This is a great suggestion, and we have added analyses to show the task correlates for the cells. We found brain region-specific differences for task correlates that depended on how the cell was modulated by SWRs or local beta oscillations. This is shown in the updated Fig. 7 and the accompanying supplemental Figs. 7-1 to 7-2.

      For PFC cells, SWR and beta modulation status defined a subpopulation with a strong task structure correlation. There was a positive correlation between the direction of SWR modulation and the spiking distance relative to goals. SWR-excited cells were more active further away from goals, whereas the SWR-inhibited cells were active closer to goals. This correlation was not found for SWR-modulated PFC cells that were not beta-modulated. For PFC cells, the direction of SWR modulation is known to be correlated with movement speed, consistent with the hypothesis that movement-active cells become reactivated during SWRs, and immobility-active cells become suppressed (Jadhav et al., 2016; Yu et al., 2017). We found the same relationship, the direction of SWR modulation was positively correlated with mean spiking speed. Our results show beta modulation marks a subpopulation of PFC cells with stronger task-structure correlates.

      For CA1 cells, we found the expected relationship between SWR modulation and task-structure correlates. SWR-excited CA1 cells spiked further away from the goal and when the animal was moving. This is consistent with the reactivation of trajectory-related spatial firing patterns during movement on the maze. This pattern was observed irrespective of beta modulation status.

      Ahmed, O. J., & Mehta, M. R. (2012). Running speed alters the frequency of hippocampal gamma oscillations. J Neurosci, 32(21), 7373-7383. https://doi.org/10.1523/JNEUROSCI.5110-11.2012

      Berke, J. D., Hetrick, V., Breck, J., & Greene, R. W. (2008). Transient 23-30 Hz oscillations in mouse hippocampus during exploration of novel environments. Hippocampus, 18(5), 519-529. https://doi.org/10.1002/hipo.20435

      Deceuninck, L., & Kloosterman, F. (2024). Disruption of awake sharp-wave ripples does not affect memorization of locations in repeated-acquisition spatial memory tasks. Elife, 13. https://doi.org/10.7554/eLife.84004

      den Bakker, H., Van Dijck, M., Sun, J. J., & Kloosterman, F. (2023). Sharp-wave-ripple associated activity in the medial prefrontal cortex supports spatial rule switching. Cell Rep, 42(8), 112959. https://doi.org/10.1016/j.celrep.2023.112959

      França, A. S., do Nascimento, G. C., Lopes-dos-Santos, V., Muratori, L., Ribeiro, S., Lobão-Soares, B., & Tort, A. B. (2014). Beta2 oscillations (23-30 Hz) in the mouse hippocampus during novel object recognition. Eur J Neurosci, 40(11), 3693-3703. https://doi.org/10.1111/ejn.12739

      França, A. S. C., Borgesius, N. Z., Souza, B. C., & Cohen, M. X. (2021). Beta2 Oscillations in Hippocampal-Cortical Circuits During Novelty Detection. Front Syst Neurosci, 15, 617388. https://doi.org/10.3389/fnsys.2021.617388

      Hyman, J. M., Zilli, E. A., Paley, A. M., & Hasselmo, M. E. (2005). Medial prefrontal cortex cells show dynamic modulation with the hippocampal theta rhythm dependent on behavior. Hippocampus, 15(6), 739-749. https://doi.org/10.1002/hipo.20106

      Iwasaki, S., Sasaki, T., & Ikegaya, Y. (2021). Hippocampal beta oscillations predict mouse object-location associative memory performance. Hippocampus, 31(5), 503-511. https://doi.org/10.1002/hipo.23311

      Jadhav, S. P., Kemere, C., German, P. W., & Frank, L. M. (2012). Awake hippocampal sharp-wave ripples support spatial memory. Science (New York, N.Y.), 336(6087), 1454-1458. https://doi.org/10.1126/science.1217230

      Jadhav, S. P., Rothschild, G., Roumis, D. K., & Frank, L. M. (2016). Coordinated Excitation and Inhibition of Prefrontal Ensembles during Awake Hippocampal Sharp-Wave Ripple Events. Neuron, 90(1), 113-127. https://doi.org/10.1016/j.neuron.2016.02.010

      Jin, S. W., Ha, H. S., & Lee, I. (2024). Selective reactivation of value- and place-dependent information during sharp-wave ripples in the intermediate and dorsal hippocampus. Sci Adv, 10(32), eadn0416. https://doi.org/10.1126/sciadv.adn0416

      Jones, M. W., & Wilson, M. A. (2005). Theta Rhythms Coordinate Hippocampal–Prefrontal Interactions in a Spatial Memory Task. PLoS Biology, 3(12). https://doi.org/10.1371/journal.pbio.0030402

      Kim, S. M., & Frank, L. M. (2009). Hippocampal Lesions Impair Rapid Learning of a Continuous Spatial Alternation Task. PLoS ONE, 4(5). https://doi.org/10.1371/journal.pone.0005494

      Lansink, C. S., Meijer, G. T., Lankelma, J. V., Vinck, M. A., Jackson, J. C., & Pennartz, C. M. (2016). Reward Expectancy Strengthens CA1 Theta and Beta Band Synchronization and Hippocampal-Ventral Striatal Coupling. J Neurosci, 36(41), 10598-10610. https://doi.org/10.1523/JNEUROSCI.0682-16.2016

      Maharjan, D. M., Dai, Y. Y., Glantz, E. H., & Jadhav, S. P. (2018). Disruption of dorsal hippocampal-prefrontal interactions using chemogenetic inactivation impairs spatial learning. Neurobiol Learn Mem, 155, 351-360. https://doi.org/10.1016/j.nlm.2018.08.023

      Rangel, L. M., Chiba, A. A., & Quinn, L. K. (2015). Theta and beta oscillatory dynamics in the dentate gyrus reveal a shift in network processing state during cue encounters. Front Syst Neurosci, 9, 96. https://doi.org/10.3389/fnsys.2015.00096

      Siapas, A. G., Lubenov, E. V., & Wilson, M. A. (2005). Prefrontal Phase Locking to Hippocampal Theta Oscillations. Neuron, 46(1), 141-151. https://doi.org/10.1016/j.neuron.2005.02.028

      Yu, J. Y., Kay, K., Liu, D. F., Grossrubatscher, I., Loback, A., Sosa, M.,…Frank, L. M. (2017). Distinct hippocampal-cortical memory representations for experiences associated with movement versus immobility. Elife, 6, e27621. https://doi.org/10.7554/eLife.27621

    1. eLife Assessment

      This study provides convincing evidence that the genetics of local adaptation in Arabidopsis is shaped by fluctuations in the environment and interactions with genotype and location. This is an important dataset contributing to the developing understanding of non-linear selection in plants and beyond.